Virtual avatar generation method and device based on large model, agent, electronic device, and storage medium

By processing images and 3D object shapes using large-scale models and texture-generating large-scale models, a 3D virtual image matching the user's needs is generated, solving the problems of long generation time and low matching degree in existing technologies, and realizing efficient virtual image display.

CN119444977BActive Publication Date: 2025-11-07BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411336634.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-11-07
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

In the fields of e-commerce and film and animation, generating 3D virtual avatars that match user needs takes a lot of time and the degree of matching is low, resulting in poor display effects.

Method used

The target image is processed using a large model to obtain object description information. The image and the 3D object shape are processed by a texture-generating large model to generate a 3D object with target texture information, and a virtual image is generated based on this.

Benefits of technology

It improves the matching degree between virtual images and target images, realizes the automatic generation of three-dimensional virtual images that meet user needs, and enhances the display effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119444977B_ABST
    Figure CN119444977B_ABST
Patent Text Reader

Abstract

The present disclosure provides a large model-based virtual image generation method and device, an agent, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical field of computer vision, deep learning, large model, etc., and can be applied to the scene of AIGC, digital person, intelligent e-commerce, etc. The method comprises: processing a target image comprising a target object by using a large model to obtain object description information, the target object having texture information; processing the target image and a to-be-processed image representing the shape of a three-dimensional object by using a texture generation type large model to obtain a target three-dimensional object having target texture information, the three-dimensional object being determined based on the object description information, and the target texture information matching the texture information; and generating a virtual image based on the target three-dimensional object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of computer vision, deep learning, large model, etc., and can be applied to the scenarios of AIGC (Artificial Intelligence Generative Content), digital human, intelligent e-commerce, etc. BACKGROUND

[0002] With the development of the Internet e-commerce, animation games, video production, etc., virtual images can be designed to realize interaction with users. For example, in the field of Internet e-commerce, three-dimensional virtual images can be used to introduce the functions of goods, so as to improve the display effect of goods. SUMMARY

[0003] The present disclosure provides a large model-based virtual image generation method, device, agent, electronic device, storage medium and program product.

[0004] According to an aspect of the present disclosure, a large model-based virtual image generation method is provided, including: processing a target image including a target object by using a large model to obtain object description information, the target object having texture information; processing the target image and a to-be-processed image representing the object form of a three-dimensional object by using a texture generation type large model to obtain a target three-dimensional object having target texture information, the three-dimensional object being determined based on the object description information, and the target texture information matching the texture information; and generating a virtual image based on the target three-dimensional object.

[0005] According to another aspect of the present disclosure, an artificial intelligence agent is provided, configured to execute the method provided by the embodiments of the present disclosure.

[0006] According to another aspect of the present disclosure, a large model-based virtual image generation device is provided, including: an object description information obtaining module, configured to process a target image including a target object by using a large model to obtain object description information, the target object having texture information; a target three-dimensional object obtaining module, configured to process the target image and a to-be-processed image representing the object form of a three-dimensional object by using a texture generation type large model to obtain a target three-dimensional object having target texture information, the three-dimensional object being determined based on the object description information, and the target texture information matching the texture information; and a virtual image generation module, configured to generate a virtual image based on the target three-dimensional object.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the embodiments of the present disclosure.

[0008] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to make the computer perform the method provided by the embodiments of the present disclosure.

[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method provided by the embodiments of the present disclosure.

[0010] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0012] Figure 1 An exemplary system architecture to which the large model-based virtual avatar generation method and device according to the embodiments of the present disclosure can be applied is schematically shown;

[0013] Figure 2 A flowchart of the large model-based virtual avatar generation method according to the embodiments of the present disclosure is schematically shown;

[0014] Figure 3 A schematic diagram of a texture generation-based large model according to the embodiments of the present disclosure is schematically shown;

[0015] Figure 4 A schematic diagram of a texture generation-based large model according to another embodiment of the present disclosure is schematically shown;

[0016] Figure 5 A schematic diagram of a texture generation-based large model according to still another embodiment of the present disclosure is schematically shown;

[0017] Figure 6 A schematic diagram of a texture generation-based large model according to still another embodiment of the present disclosure is schematically shown;

[0018] Figure 7 An application scenario diagram of the large model-based virtual avatar generation method according to the embodiments of the present disclosure is schematically shown;

[0019] Figure 8 A block diagram of a large model-based virtual image generation apparatus according to an embodiment of the present disclosure is schematically shown;

[0020] Figure 9 A structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown; and

[0021] Figure 10 A block diagram of an electronic device suitable for implementing a large model-based virtual image generation method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0022] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered in a descriptive sense only. It will thus be appreciated that various modifications and changes can be made to the embodiments described herein without departing from the scope and spirit of the disclosure. Likewise, the description and the illustrations are given solely for the purpose of illustrating the general principles of the present disclosure and should not be taken as limiting. The descriptions and illustrations are not intended to include all features and aspects of the present disclosure.

[0023] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information involved comply with relevant legal regulations, necessary security measures are taken, and do not violate public order and good customs.

[0024] The inventors have found that in the fields of Internet e-commerce, film and animation, etc., a three-dimensional virtual image can be driven to perform a specified action task to display goods and video plots. In addition, users can also make three-dimensional virtual images based on personal needs. However, it usually takes a lot of time to make a virtual image that matches the user's needs, and the matching degree between the generated virtual image and the user's actual needs is low, thereby reducing the display effect of the virtual image.

[0025] Embodiments of the present disclosure provide a large model-based virtual image generation method, apparatus, agent, electronic device, storage medium and program product. The large model-based virtual image generation method comprises: processing a target image comprising a target object by using a large model to obtain object description information, the target object having texture information; processing the target image and a to-be-processed image representing the shape of a three-dimensional object by using a texture generation-based large model to obtain a target three-dimensional object having target texture information, the three-dimensional object being determined based on the object description information, and the target texture information matching the texture information; and generating a virtual image based on the target three-dimensional object.

[0026] According to an embodiment of the present disclosure, by processing the target image by using the large model, the object attribute information such as the contour, shape, style, etc. of the target object in the target image can be learned based on the powerful visual understanding capability of the large model, so that the output object description information can accurately represent the object attributes of the target object. The three-dimensional object determined based on the object description information can match the object attribute information represented by the object description information, so that the to-be-processed image representing the object shape of the three-dimensional object can accurately represent the object shape of the target object. By processing the target image and the to-be-processed image by using the texture generation large model, the texture information of the target object in the target image and the object shape represented by the to-be-processed image can be accurately fused, so that the generated target three-dimensional object can accurately represent the matching relationship between the object shape and the texture information of the target object in the three-dimensional space, so that the virtual avatar generated according to the target three-dimensional object can accurately represent the shape and texture of each target object in the target image in the three-dimensional space, improve the matching degree between the virtual avatar and the target image, and further accurately generate a three-dimensional virtual avatar that matches the user's demand in an automated manner.

[0027] Figure 1 An example system architecture to which the large model-based virtual avatar generation method and device according to an embodiment of the present disclosure can be applied is schematically shown.

[0028] It should be noted that, Figure 1 The example shown is only an example of a system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the example system architecture to which the large model-based virtual avatar generation method and device can be applied can include terminal devices, but the terminal devices can not need to interact with the server to implement the large model-based virtual avatar generation method and device provided by the embodiments of the present disclosure.

[0029] As Figure 1 shown, the system architecture 100 according to this embodiment can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.

[0030] The user can use the terminal devices 101, 102, and 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0031] The terminal devices 101, 102, and 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, etc.

[0032] The server 105 can be a server providing various services, such as a background management server providing support for the content browsed by the user using the terminal devices 101, 102, and 103 (only as an example). The background management server can analyze and process the received user requests and other data, and feed back the processing results (such as web pages, information, or data, etc. obtained or generated according to the user requests) to the terminal devices.

[0033] The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0034] It should be noted that the large model-based virtual image generation method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the large model-based virtual image generation apparatus provided by the embodiments of the present disclosure can generally be arranged in the server 105. The large model-based virtual image generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105. Correspondingly, the large model-based virtual image generation apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105.

[0035] For example, any one of the terminal devices 101, 102, and 103 can acquire a target image input by a user, and then send the acquired target image to the server 105, so that the server 105 processes the target image by using a large model to obtain object description information, processes the target image and a to-be-processed image representing the object form of a three-dimensional object by using a texture generation large model to obtain a target three-dimensional object having target texture information, and generates a virtual avatar based on the target three-dimensional object. Or the target image and the to-be-processed image are processed by a server or a server cluster capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105, and finally a virtual avatar is generated.

[0036] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above description is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks, and servers.

[0037] Figure 2 A flowchart of a large model-based virtual avatar generation method according to an embodiment of the present disclosure is schematically shown.

[0038] As shown in Figure 2 The large model-based virtual avatar generation method includes operations S210-S230.

[0039] In operation S210, a target image including a target object having texture information is processed by using a large model to obtain object description information.

[0040] In operation S220, the target image and a to-be-processed image representing the object form of a three-dimensional object are processed by using a texture generation large model to obtain a target three-dimensional object having target texture information.

[0041] In operation S230, a virtual avatar is generated based on the target three-dimensional object.

[0042] According to an embodiment of the present disclosure, the target object in the target image can be any type of object such as a person, clothing, accessories, tools, buildings, vehicles, etc. in the target image. The texture information of the target object can be represented based on the pixels of the image region corresponding to the target object in the target image. For example, the texture information can be represented based on the RGB (Red Green Blue) information of the pixels in the target image. It should be noted that the target object can represent a real object such as a person or an animal, or the target object can also represent a virtual object such as an animated character.

[0043] According to an embodiment of the present disclosure, a large model can refer to a deep learning model with large-scale model parameters. A large model generally contains hundreds of millions, billions, tens of billions, hundreds of billions, or even more than ten trillion model parameters. The large model can include a VideoChat model, a Video-LlaMA model, and other multi-modal large models with visual understanding capabilities. The large model involved in the embodiments of the present disclosure can be a general large model, or can also be a specialized large model obtained by fine tuning based on sample object description information and sample images. The embodiments of the present disclosure do not limit this. The large model can be used to process information of any modality such as images, texts, videos, and audios.

[0044] According to an embodiment of the present disclosure, by processing the target image using the large model, the object attribute information of the target object in the target image such as the contour, shape, and style of the target object can be learned based on the powerful visual understanding capability of the large model, so that the output object description information can accurately represent the object attribute of the target object. Determining the three-dimensional object based on the object description information can include retrieving a three-dimensional object based on the object description information. The three-dimensional object shape of the three-dimensional object can be adapted to the object attribute represented by the object description information.

[0045] It should be understood that the target object can be represented based on a two-dimensional image region in the target image. The three-dimensional object can be represented based on a three-dimensional object element (for example, a three-dimensional space point, a three-dimensional grid element) in a three-dimensional space.

[0046] According to an embodiment of the present disclosure, the three-dimensional object is determined based on the object description information.

[0047] In one example, the three-dimensional object can be obtained by retrieving from a preset three-dimensional object library based on the object description information. The preset three-dimensional object in the preset three-dimensional object library can be associated with preset description information. The three-dimensional object can be determined from the preset three-dimensional object library based on the similarity between the object description information and the preset description information.

[0048] It should be understood that, under the condition of obtaining relevant authorization, the three-dimensional object matching the object description information can also be obtained by retrieving in any database based on the object description information. For example, the three-dimensional object can be obtained by retrieving in an open-source three-dimensional model database. The embodiments of the present disclosure do not limit the specific way of determining the three-dimensional object, as long as the three-dimensional object matches the object attribute information represented by the object description information.

[0049] According to an embodiment of the present disclosure, the to-be-processed image can include a two-dimensional image, for example, a two-dimensional image that can be a grayscale image, a binary image, a UV map, etc. The pixels of the to-be-processed image can have a spatial position mapping relationship with the object elements of the three-dimensional object. The to-be-processed image can be obtained by performing two-dimensional spatial mapping on the object pixels of the three-dimensional object. Alternatively, the to-be-processed image can also be determined based on a two-dimensional UV map used to render the three-dimensional object.

[0050] According to an embodiment of the present disclosure, the target texture information is matched with the texture information. The target texture information of the target three-dimensional object can have a mapping relationship with the texture information of the target object in space.

[0051] According to an embodiment of the present disclosure, the to-be-processed image and the target image are processed by using the texture generation large model. Based on the powerful image semantic understanding ability of the texture generation large model, the semantic attribute relationship between the object form represented by the to-be-processed image and the texture information of the specified position of the target object in the target image can be learned, and then the texture semantic attribute of the object form represented by the to-be-processed image can be migrated based on the texture information of the target object. In this way, the generated target three-dimensional object can accurately represent the texture semantic attribute of the target object, and the target texture information of the target three-dimensional object can be matched with the texture information of the target object, thereby improving the representation accuracy of the target three-dimensional object representing the target object.

[0052] According to an embodiment of the present disclosure, generating a virtual avatar based on the target three-dimensional object can include fusing a plurality of target three-dimensional objects according to the positional relationship represented by the target image to obtain a virtual avatar. Alternatively, the virtual avatar that matches the user's demand can also be obtained by fusing the target three-dimensional object with other preset three-dimensional virtual objects. The specific way of generating a virtual avatar is not limited in the embodiments of the present disclosure.

[0053] According to an embodiment of the present disclosure, the target object includes at least one of the following: a clothing object, a body part object, a vehicle object, and a building object.

[0054] According to an embodiment of the present disclosure, the clothing object can represent the clothing displayed in the target image. For example, a suit, a long skirt, and other clothing worn by a person. The target three-dimensional object corresponding to the clothing object can be a three-dimensional clothing model matched with the texture information of the clothing object.

[0055] According to an embodiment of the present disclosure, the body part object can represent any body part such as a head, a hand, an arm, etc. The target three-dimensional object corresponding to the body part object can be a three-dimensional body part model representing the body part. By fusing the three-dimensional body part model with the target texture information according to the posture and position of the person in the target image, the generated avatar can more accurately represent the posture and texture of the person in the target image, and the avatar can more accurately simulate the target image, thereby improving the matching degree between the avatar and the user's needs.

[0056] According to an embodiment of the present disclosure, the vehicle object can represent any type of movable vehicle such as a car, a ship, an airplane, etc. The target three-dimensional object corresponding to the vehicle object can be a three-dimensional vehicle model with target texture information.

[0057] According to an embodiment of the present disclosure, the building object can represent any type of building such as a house, a bridge, a warehouse, etc. The target three-dimensional object corresponding to the building object can be a three-dimensional house model, a three-dimensional bridge model, etc.

[0058] According to an embodiment of the present disclosure, the texture generation model includes a texture generation network. The texture generation network can be constructed based on a generative model algorithm. For example, the texture generation network can be constructed based on any type of generative model algorithm such as a GAN (Generative Adversaria Networks) model, a VAE (Variational auto-encoder) model, a Flow-based Models, a Diffusion Models, etc. However, it is not limited thereto, and the texture generation network can also be constructed based on other types of deep learning algorithms. For example, the texture generation network can be constructed based on a convolutional neural network algorithm.

[0059] According to an embodiment of the present disclosure, the image to be processed includes a position mapping image. The position pixel of the position mapping image represents the three-dimensional coordinates of the object element in the three-dimensional object. For example, the position pixel can store the three-dimensional space coordinates (x, y, z) of the object mesh element of the three-dimensional object.

[0060] According to an embodiment of the present disclosure, the position pixel representing the three-dimensional coordinates of the object element in the three-dimensional object can be understood as that the position pixel in the position mapping image can store the three-dimensional coordinates of the object element. The position mapping image can be determined based on the three-dimensional coordinates of the object element in the three-dimensional object.

[0061] It should be noted that the corresponding position mapping image can be constructed for the preset three-dimensional object in the preset three-dimensional object library. For example, the mapping pixel information of the preset UV map used to render the preset three-dimensional object can be updated, and the preset texture information stored in the mapping pixels of the preset UV map is updated to the three-dimensional coordinates of the object elements of the preset three-dimensional object, so as to obtain the position mapping image.

[0062] According to an embodiment of the present disclosure, processing the target image and the image representing the object form of the three-dimensional object by using the texture generation model can include: based on the texture generation network, performing feature fusion on the object style feature and the position mapping image to obtain a target texture map matched with the texture information; and updating the three-dimensional object based on the target texture information of the target texture map to obtain a target three-dimensional object.

[0063] According to an embodiment of the present disclosure, the object style feature is determined based on the target image. For example, the object style feature can be obtained by performing feature extraction on the target image based on an image feature extraction algorithm. The image feature extraction algorithm can include any type of deep learning algorithm such as a convolutional neural network algorithm, and the specific type of the image feature extraction algorithm is not limited by the embodiments of the present disclosure.

[0064] According to an embodiment of the present disclosure, the object style feature represents the texture semantic attribute, shape detail semantic attribute, pattern semantic attribute, etc. of the target object and the image semantic attribute. Based on the texture generation network, the object style feature and the position mapping image are fused, and based on the three-dimensional coordinates stored in the position pixels as prior information, the texture generation network is used to learn the corresponding relationship between the three-dimensional coordinates of the object elements and the object style feature by using a large-scale model parameter, so that the image style related to the target object in the target image is mapped to the three-dimensional coordinates corresponding to the object elements in the three-dimensional space, so that the generated target texture map can more accurately store the target texture information according to the position of the object elements of the three-dimensional object in the three-dimensional space. This can make the target texture map more accurately store the position of the texture information of the target object in the three-dimensional space, so that the generated target three-dimensional object can more accurately match the texture information of the target object, and the representation accuracy for the target object is improved.

[0065] It should be understood that the mapping pixels of the target texture map can have a spatial mapping relationship with the object elements of the three-dimensional object, and the target three-dimensional object with the target texture information can be obtained by unfolding the target texture map in the three-dimensional space.

[0066] In one example, the target texture map can be processed based on a rendering engine to obtain a target three-dimensional object with target texture information.

[0067] According to an embodiment of the present disclosure, the feature fusion based on the texture generation network and according to the object style feature and the position mapping image can include: performing feature fusion on the position mapping image and a shape mask image in the to-be-processed image to obtain initial fusion features; and performing feature fusion based on the initial fusion features and the object style feature to obtain target fusion features.

[0068] According to an embodiment of the present disclosure, the mask pixels of the shape mask image represent whether the object elements of the three-dimensional object can store texture information, and the mask pixels have a position mapping relationship with the object elements.

[0069] According to an embodiment of the present disclosure, the shape mask image can be determined based on a UV map corresponding to the three-dimensional object. For example, the initial texture information of the map pixels in the UV map can be updated as mask information, and the mask information can represent whether the object elements can store texture information based on a preset code, numerical value or representation.

[0070] In one example, the pixel value of the mask pixel of the shape mask image can be represented as 1 or 0. The mask pixel being 1 represents that the object pixel corresponding to the mask pixel needs to have target texture information, and the mask pixel being 0 represents that the object pixel corresponding to the mask pixel does not have target texture information.

[0071] According to an embodiment of the present disclosure, the feature fusion on the position mapping image and the shape mask image in the to-be-processed image can include processing the position mapping image and the shape mask image based on an attention network algorithm to obtain initial fusion features.

[0072] In one embodiment, the position mapping image and the shape mask image can be processed based on a cross-attention network algorithm, so that the three-dimensional coordinates of the object pixels represented by the position pixels and the semantic attributes of whether the texture information needs to be stored represented by the mask pixels are fully fused. In this way, the initial fusion features can be used for feature fusion with the object style feature, so as to control the style semantic attributes of the object style feature to be accurately fused into the object elements of the three-dimensional object according to the correspondence between the pixel positions of the target object and the three-dimensional spatial positions of the object elements of the three-dimensional object, so as to generate target fusion features that can more accurately represent the texture information of the target object.

[0073] According to an embodiment of the present disclosure, based on the initial fusion feature obtained by fusing the mask pixels and the position pixels of the shape mask image, the texture generation network can be controlled to update the texture semantic attributes of the object elements of the three-dimensional object to which the target texture information needs to be added, and the texture generation network can be controlled to avoid updating the texture semantic attributes of the object elements of the three-dimensional object to which the target texture information does not need to be added, so that the target fusion feature can more accurately control the target texture information of the object elements of the target three-dimensional object, avoid style feature migration errors caused by errors in the position mapping relationship between the object elements of the three-dimensional object and the pixels of the target object, and improve the representation accuracy of the texture semantic attributes of the target object for the target three-dimensional object.

[0074] According to an embodiment of the present disclosure, the target texture map can be determined based on the target fusion feature. For example, the target fusion feature can be processed based on any type of machine learning algorithm such as a fully connected layer to obtain the target texture map.

[0075] According to an embodiment of the present disclosure, the texture generation model further includes a first style feature extraction network, and the first style feature extraction network includes a down-sampling layer and an up-sampling layer having a U-shaped network structure. The object style feature can include at least one level of down-sampled style feature and at least one level of up-sampled style feature obtained by processing the target image by the down-sampling layer and the up-sampling layer. The at least one level of down-sampled style feature can be output by the down-sampling layer, and the at least one level of up-sampled style feature can be output by the up-sampling layer.

[0076] According to an embodiment of the present disclosure, the down-sampling layer or the up-sampling layer can be constructed based on a convolutional neural network algorithm. However, it is not limited thereto, and the down-sampling layer or the up-sampling layer can also be constructed based on other types of neural network algorithms such as an attention network algorithm. Embodiments of the present disclosure do not limit the specific algorithm type for constructing the down-sampling layer or the up-sampling layer.

[0077] In one example, the down-sampling layer or the up-sampling layer can be constructed based on a convolutional neural network algorithm and a pooling algorithm.

[0078] According to an embodiment of the present disclosure, the feature fusion based on the initial fusion feature and the object style feature to obtain the target fusion feature includes: performing a feature encoding operation on the initial fusion feature and the at least one level of down-sampled style feature by using a texture encoder of the texture generation network to obtain a first intermediate fusion feature; and performing a feature decoding operation on the first intermediate fusion feature and the at least one level of up-sampled style feature by using a texture decoder of the texture generation network to obtain the target fusion feature.

[0079] According to an embodiment of the present disclosure, the texture generation network can include a texture encoder and a texture decoder. The texture encoder and the texture decoder can be constructed based on a network structure of an encoder and a decoder in a diffusion model.

[0080] According to an embodiment of the present disclosure, the feature encoding operation can be understood as a feature encoding process performed based on the initial fusion feature and the down-sampled style feature. The feature encoding operation can perform feature fusion on the initial fusion feature and the down-sampled style feature by introducing noise information, so that the first intermediate fusion feature can sufficiently fuse high-level style semantic attributes in the target image according to the spatial position of the three-dimensional object represented by the initial fusion feature. The feature decoding operation can be understood as performing feature fusion and feature decoding processing based on the first intermediate fusion feature and the up-sampled style feature, so that the target fusion feature can fuse style semantic attributes according to the spatial position of the object elements of the three-dimensional object and restore the resolution. In this way, the target fusion feature can more accurately represent the style details of the three-dimensional object, and improve the matching degree between the target texture information of the target three-dimensional object and the texture information of the target object.

[0081] In one example, preset noise information can be introduced in the process of the feature encoding operation and the feature decoding operation, and the initial fusion feature and the object style feature are fused through the feature encoding operation and the feature decoding operation, and feature denoising is performed to obtain the target fusion feature.

[0082] According to an embodiment of the present disclosure, the texture encoder can include a plurality of encoder layers, and the texture decoder can include a plurality of decoder layers. The encoder layers and the decoder layers can be used to perform the feature encoding operation and the feature decoding operation.

[0083] In one example, the encoder layer or the decoder layer can be constructed based on an attention network algorithm, for example, can be constructed based on a cross-attention network algorithm or a self-attention network algorithm.

[0084] In one example, performing the feature encoding operation includes: fusing at least one level of the first encoding feature and at least one level of the down-sampled style feature to obtain a next level of the first encoding feature. The first level of the first encoding feature can be determined by performing the feature encoding operation on the initial fusion feature and the first level of the down-sampled style feature.

[0085] In one example, performing the feature decoding operation includes: fusing at least one level of the first decoding feature and at least one level of the up-sampled style feature to obtain a next level of the first decoding feature. The first level of the first decoding feature is determined by performing the feature decoding operation on the first intermediate fusion feature and the first level of the up-sampled style feature. The target fusion feature is determined based on the decoding feature of the last level.

[0086] Figure 3 A schematic diagram of a texture generation-based large model according to an embodiment of the present disclosure is shown.

[0087] As shown in Figure 3 , the texture generation-based large model 300 can include a first style feature extraction network 310 and a texture generation network including a texture encoder 321 and a texture decoder 322. The texture generation network can be constructed based on a diffusion model algorithm. The first style feature extraction network 310 can include down-sampling layers and up-sampling layers determined based on a U-shaped network structure. The down-sampling layers and the up-sampling layers can include a first-level down-sampling layer 3111 and a third-level up-sampling layer 3112, a second-level down-sampling layer 3121 and a second-level up-sampling layer 3122, and a third-level down-sampling layer 3131 and a first-level up-sampling layer 3132. The down-sampling layers or the up-sampling layers can be constructed based on a convolutional neural network algorithm. The first U-shaped substructure composed of the first-level down-sampling layer 3111 and the third-level up-sampling layer 3112, the second U-shaped substructure composed of the second-level down-sampling layer 3121 and the second-level up-sampling layer 3122, and the third U-shaped substructure composed of the third-level down-sampling layer 3131 and the first-level up-sampling layer 3132 can be used to extract style features of different scales, respectively. The texture encoder 321 includes a first encoder layer 3211, a second encoder layer 3212, and a third encoder layer 3213 constructed based on an attention network algorithm. The texture decoder 322 includes a first decoder layer 3221, a second decoder layer 3222, and a third decoder layer 3223 constructed based on an attention network algorithm.

[0088] As shown in Figure 3 , the target image 301 is input into the first-level down-sampling layer 3111 to obtain first-level down-sampling style features. The first-level down-sampling style features are input into the second-level down-sampling layer 3121 to obtain second-level down-sampling style features. The second-level down-sampling style features are input into the third-level down-sampling layer 3131 to obtain third-level down-sampling style features. The third-level down-sampling style features and feature maps determined based on the third-level down-sampling style features are input into the first-level up-sampling layer 3132 to output first-level up-sampling style features. The first-level up-sampling style features and feature maps determined based on the second-level down-sampling style features are input into the second-level up-sampling layer 3122 to output second-level up-sampling style features. The second-level up-sampling style features and feature maps determined based on the first-level down-sampling style features are input into the third-level up-sampling layer 3112 to output third-level up-sampling style features.

[0089] The initial fused feature 302 obtained by feature fusion based on the position mapping image and the shape mask image, and the first-level downsampled style feature are input into the first encoder layer 3211. The first encoder layer 3211 fuses the initial fused feature 302 and the first-level downsampled style feature using an attention network algorithm, and performs a first-level feature encoding operation based on noise information to obtain the first-level coded feature. The first-level coded feature and the second-level downsampled style feature are input into the second encoder layer 3212. The second encoder layer 3212 fuses the first-level coded feature and the second-level downsampled style feature using an attention network algorithm, and performs a second-level feature encoding operation based on noise information to obtain the second-level coded feature. The first-level coded feature and the third-level downsampled style feature are input into the third encoder layer 3213. The third encoder layer 3213 fuses the first-level coded feature and the third-level downsampled style feature using an attention network algorithm, and performs a third-level feature encoding operation based on noise information to obtain the third-level coded feature. The first coding feature at level 3 is determined as the first intermediate fusion feature.

[0090] like Figure 3 As shown, the first encoded feature of level 3 is used as the first intermediate fusion feature. The first intermediate fusion feature and the first upsampled style feature of level 1 are input into the first decoder layer 3221. The first decoder layer 3221 can process the first intermediate fusion feature and the first upsampled style feature of level 1 based on an attention mechanism to perform feature decoding operation and obtain the first decoded feature of level 1. The first decoded feature of level 1 and the second upsampled style feature of level 2 are input into the second decoder layer 3222. The second decoder layer 3222 can process the first decoded feature of level 1 and the second upsampled style feature of level 2 based on an attention mechanism to perform feature decoding operation and obtain the first decoded feature of level 2. The first decoded feature of level 2 and the second upsampled style feature of level 3 are input into the third decoder layer 3223. The third decoder layer 3223 can process the first decoded feature of level 2 and the second upsampled style feature of level 3 based on an attention mechanism to perform feature decoding operation and obtain the first decoded feature of level 3. The target fusion feature 303 can be determined based on the first decoded feature of level 3. Image rendering based on target fusion feature 303 can yield target texture map 304.

[0091] According to embodiments of this disclosure, performing feature decoding operations on a first intermediate fusion feature and at least one level upsampled style feature using a texture decoder of a texture generation network includes: performing feature decoding operations on the first intermediate fusion feature, at least one level upsampled style feature, and at least one level position feature using a texture decoder based on an attention mechanism.

[0092] According to an embodiment of the present disclosure, the position feature is determined according to the position mapping image. The position feature can be obtained by performing feature extraction on the position mapping image, for example, the position mapping image can be extracted based on any type of algorithm such as a convolutional neural network. Alternatively, the position feature can also be obtained by performing feature fusion on the initial fusion feature and the position mapping image. For example, the initial fusion feature and the position mapping image can be fused based on any type of deep learning algorithm such as a convolutional neural network algorithm, an attention network algorithm, etc. to obtain one or more levels of position features. The embodiments of the present disclosure do not limit the specific manner of determining the position feature.

[0093] In one example, the initial fusion feature and the position mapping image can be processed based on a plurality of convolutional neural network layers in cascade to obtain a plurality of levels of position features. The texture encoder can be used to process the first intermediate fusion feature, at least one level of up-sampled style feature, and at least one level of position feature to perform a feature decoding operation to obtain the target fusion feature.

[0094] According to an embodiment of the present disclosure, the texture generation large model can further include a position feature extraction network, which can include a plurality of levels of position feature extraction layers in cascade. The position feature extraction layer can be constructed based on a convolutional neural network, or the position feature extraction layer can also be constructed based on other types of neural network algorithms. The position feature includes a plurality of levels, and the plurality of levels of position features are determined by processing the target image using the plurality of levels of position feature extraction layers.

[0095] According to an embodiment of the present disclosure, the plurality of levels of position feature extraction layers in cascade can perform feature extraction on the position mapping image from a plurality of scales, so that the plurality of levels of position features can represent the position mapping relationship between the three-dimensional spatial position of the object element and the two-dimensional pixel based on the position mapping semantic attribute of the position mapping image. By using the texture encoder to fuse the first intermediate fusion feature, the plurality of levels of up-sampled style features, and the plurality of levels of position features during the feature decoding operation, the position mapping relationship and the style semantic attribute can be accurately and finely fused to improve the representation accuracy of the mapping relationship between the target fusion feature and the coordinates of the object element, thereby improving the matching degree of the target three-dimensional object and the target object in the target image.

[0096] Figure 4 The principle diagram of the texture generation large model according to another embodiment of the present disclosure is schematically shown.

[0097] As Figure 4As shown, the texture generation type large model 400 can include a first style feature extraction network 410, a texture generation network including a texture encoder 421 and a texture decoder 422, and a position feature extraction network 430. The texture generation network can be constructed based on a diffusion model algorithm. The first style feature extraction network 410 can include down-sampling layers and up-sampling layers determined based on a U-shaped network structure. The down-sampling layers and the up-sampling layers can include a first-level down-sampling layer 4111 and a third-level up-sampling layer 4112, a second-level down-sampling layer 4121 and a second-level up-sampling layer 4122, and a third-level down-sampling layer 4131 and a first-level up-sampling layer 4132. The down-sampling layers or the up-sampling layers can be constructed based on a convolutional neural network algorithm. The first-level down-sampling layer 4111 and the third-level up-sampling layer 4112 constitute a U-shaped substructure, the second-level down-sampling layer 4121 and the second-level up-sampling layer 4122 constitute another U-shaped substructure, and the third-level down-sampling layer 4131 and the first-level up-sampling layer 4132 constitute a third U-shaped substructure. The three nested U-shaped substructures can be used to extract style features of different scales, respectively. The texture encoder 421 includes a first encoder layer 4211, a second encoder layer 4212, and a third encoder layer 4213 constructed based on an attention network algorithm. The texture decoder 422 includes a first decoder layer 4221, a second decoder layer 4222, and a third decoder layer 4223 constructed based on an attention network algorithm. The position feature extraction network 430 includes a first position feature extraction layer 431, a second position feature extraction layer 432, a third position feature extraction layer 433, and a fourth position feature extraction layer 434 connected in cascade. The first position feature extraction layer 431, the second position feature extraction layer 432, the third position feature extraction layer 433, and the fourth position feature extraction layer 434 are respectively constructed based on different scale convolutional neural network algorithms.

[0098] As shown, Figure 4 The target image 401 is input into the first-level down-sampling layer 4111 to obtain first-level down-sampling style features. The first-level down-sampling style features are input into the second-level down-sampling layer 4121 to obtain second-level down-sampling style features. The second-level down-sampling style features are input into the third-level down-sampling layer 4131 to obtain third-level down-sampling style features. The third-level down-sampling style features and feature maps determined based on the third-level down-sampling style features are input into the first-level up-sampling layer 4132 to output first-level up-sampling style features. The first-level up-sampling style features and feature maps determined based on the second-level down-sampling style features are input into the second-level up-sampling layer 4122 to output second-level up-sampling style features. The second-level up-sampling style features and feature maps determined based on the first-level down-sampling style features are input into the third-level up-sampling layer 4112 to output third-level up-sampling style features.

[0099] As shown, Figure 4As shown, the position mapping image 403 is added to the initial fusion feature 402 and then input into the first position feature extraction layer 431 to obtain the first-level position feature. The first-level position feature is input into the second position feature extraction layer 432 to obtain the second-level position feature. The second-level position feature is input into the third-level position feature extraction layer 433 to obtain the third-level position feature. The third-level position feature is input into the fourth-level position feature extraction layer 434 to obtain the fourth-level position feature.

[0100] like Figure 4 As shown, the initial fused feature 402 obtained by feature fusion based on the position mapping image and the shape mask image, and the first-level downsampled style feature are input to the first encoder layer 4211. The first encoder layer 4211 fuses the initial fused feature 402 and the first-level downsampled style feature using an attention network algorithm, and performs a first-level feature encoding operation based on noise information to obtain the first-level coded feature. The first-level coded feature and the second-level downsampled style feature are input to the second encoder layer 4212. The second encoder layer 4212 fuses the first-level coded feature and the second-level downsampled style feature using an attention network algorithm, and performs a second-level feature encoding operation based on noise information to obtain the second-level coded feature. The first-level coded feature and the third-level downsampled style feature are input to the third encoder layer 4213. The third encoder layer 4213 fuses the first-level coded feature and the third-level downsampled style feature using an attention network algorithm, and performs a third-level feature encoding operation based on noise information to obtain the third-level coded feature. The first coding feature at level 3 is determined as the first intermediate fusion feature.

[0101] like Figure 4As shown, the first coding feature at the third level is taken as a first intermediate fusion feature, and the first intermediate fusion feature, the first up-sampling style feature at the first level, and the position feature at the fourth level are input into the first decoder layer 4221. The first decoder layer 4221 can perform a feature decoding operation on the first intermediate fusion feature, the first up-sampling style feature at the first level, and the position feature at the fourth level based on an attention mechanism to obtain a first decoding feature at the first level. The first decoding feature at the first level, the second up-sampling style feature at the second level, and the position feature at the third level are input into the second decoder layer 4222. The second decoder layer 4222 can perform a feature decoding operation on the first decoding feature at the first level, the second up-sampling style feature at the second level, and the position feature at the third level based on an attention mechanism to obtain a first decoding feature at the second level. The first decoding feature at the second level, the position feature at the second level, and the third up-sampling style feature at the third level are input into the third decoder layer 4223. The third decoder layer 4223 can perform a feature decoding operation on the first decoding feature at the second level, the position feature at the second level, and the third up-sampling style feature at the third level based on an attention mechanism to obtain a first decoding feature at the third level. Feature fusion based on the first decoding feature at the third level and the position feature at the first level can determine the target fusion feature 404. Image rendering based on the target fusion feature 404 can obtain the target texture map 405.

[0102] According to an embodiment of the present disclosure, the texture generation model further comprises a second style feature extraction network, and the second style feature extraction network comprises a plurality of cascaded style feature extraction layers. The style feature extraction layer can be constructed based on any type of neural network algorithm such as a convolutional neural network algorithm or an attention network algorithm.

[0103] According to an embodiment of the present disclosure, the object style feature comprises a plurality of style features obtained by processing the object style feature by a plurality of style extraction layers. The plurality of style features can represent style semantic attributes of the target image at different scales.

[0104] According to an embodiment of the present disclosure, the feature fusion based on the initial fusion feature and the object style feature to obtain the target fusion feature comprises: performing a feature encoding operation on the initial fusion feature by a texture encoder of the texture generation network to obtain a second intermediate fusion feature; and performing a feature decoding operation on the second intermediate fusion feature and at least one style feature by a texture decoder of the texture generation network to obtain the target fusion feature.

[0105] According to an embodiment of the present disclosure, the texture encoder can include cascaded multi-stage encoder layers, and the encoder layers can be constructed based on an attention network algorithm. The first-stage encoder layer can process the initial fusion feature and the noise information based on the attention network algorithm to implement a feature encoding operation on the initial fusion feature to obtain a first-stage encoded feature. The second-stage encoder layer can process the first-stage encoded feature and the noise information based on the attention network algorithm to implement a feature encoding operation on the first-stage encoded feature to obtain a second-stage encoded feature. The second intermediate fusion feature can be determined based on a last-stage encoded feature output by a last-stage encoder layer. For example, the last-stage encoded feature can be autonomously fused to obtain the second intermediate fusion feature.

[0106] According to an embodiment of the present disclosure, a feature decoding operation can be performed on the second intermediate fusion feature and the at least one level of style feature based on multi-stage decoder layers of the texture decoder. For example, based on an attention network algorithm, the first-stage decoder layer is used to process the second intermediate fusion feature and the at least one level of style feature to implement a feature decoding operation on the second intermediate fusion feature to obtain a first-stage decoded feature. Based on the attention network algorithm, the second-stage decoder layer is used to process the first-stage decoded feature and the at least one level of style feature to implement a feature decoding operation on the first-stage decoded feature to obtain a second-stage decoded feature. The target fusion feature is determined based on a last-stage decoded feature obtained by performing a feature decoding operation by the last-stage decoder layer.

[0107] According to an embodiment of the present disclosure, by fusing the second intermediate fusion feature and the at least one level of style feature in the process of performing the feature decoding operation by the texture decoder, the association between the positional semantic attribute of the object element and the style semantic attribute can be enhanced in the stage of restoring the semantic attribute of the target map, so as to improve the association precision between the target fusion feature and the three-dimensional spatial coordinates of the object element with respect to the style semantic attribute, and also reduce the data processing amount of the texture generative large model for generating the target fusion feature, save the computing overhead, and improve the representation precision of the target fusion feature representing the target texture information of the three-dimensional object.

[0108] Figure 5 A schematic diagram of a texture generative large model according to yet another embodiment of the present disclosure is shown.

[0109] As Figure 5As shown, the texture generation model 500 may include a second style feature extraction network 510 and a texture generation network. The texture generation network includes a texture encoder 521 and a texture decoder 522. The texture generation network can be constructed based on a diffusion model algorithm. The second style feature extraction network 510 may include multiple style feature extraction layers determined based on a cascaded network structure. The multiple style feature extraction layers may include a first style feature extraction layer 511, a second style feature extraction layer 512, a third style feature extraction layer 513, and a fourth style feature extraction layer 514. The style feature extraction layers can be constructed based on a convolutional neural network algorithm, and the multiple style feature extraction layers can be used to extract style feature maps at different scales of the target image. The texture encoder 521 includes a first encoder layer 5211, a second encoder layer 5212, and a third encoder layer 5213 constructed based on an attention network algorithm. The texture decoder 522 includes a first decoder layer 5221, a second decoder layer 5222, and a third decoder layer 5223 constructed based on an attention network algorithm.

[0110] like Figure 5 As shown, the initial style features are obtained by processing the target image 501 through the initial style feature extraction layer. The initial style features are then added to the initial fused features 502 obtained by feature fusion based on the position mapping image and the shape mask image to achieve feature fusion, resulting in the target initial style features. The target initial style features are input into the first style feature extraction layer 511 to extract style semantic features at the first scale, resulting in the first-level style features. The first-level style features are input into the second style feature extraction layer 512 to extract style semantic features at the second scale, resulting in the second-level style features. The second-level style features are input into the third style feature extraction layer 513 to extract style semantic features at the third scale, resulting in the third-level style features. The third-level style features are input into the fourth style feature extraction layer 514 to extract style semantic features at the fourth scale, resulting in the fourth-level style features.

[0111] In the feature encoding stage, the texture encoder of the texture generation network performs a feature encoding operation on the initial fusion feature to obtain a second intermediate fusion feature can include: a first encoder layer 5211 as a first-level encoder layer, which can process the initial fusion feature 502 and the noise information based on an attention network algorithm to implement the feature encoding operation on the initial fusion feature to obtain a first-level encoding feature. A second encoder layer 5212 as a second-level encoder layer, which can process the first-level encoding feature and the noise information based on an attention network algorithm to implement the feature encoding operation on the first-level encoding feature to obtain a second-level encoding feature. A third encoder layer 5213 as a third-level encoder layer, which can process the second-level encoding feature and the noise information based on an attention network algorithm to implement the feature encoding operation on the second-level encoding feature to obtain a third-level encoding feature. The third-level encoding feature can be used as a second-level intermediate fusion feature.

[0112] In the feature decoding stage, the first decoder layer 5221 of the texture encoder 522 performs a first-level feature decoding operation by fusing the second intermediate fusion feature and the fourth-level style feature based on an attention mechanism to obtain a first-level decoding feature. The second decoder layer 5222 performs a second-level feature decoding operation by fusing the first-level decoding feature and the third-level style feature based on an attention mechanism to obtain a second-level decoding feature. The third decoder layer 5223 performs a third-level feature decoding operation by fusing the second-level decoding feature and the second-level style feature based on an attention mechanism to obtain a third-level decoding feature. The third-level decoding feature is fused with the first-level style feature to obtain a target fusion feature 503. Based on the target fusion feature 503, image rendering can be performed to obtain a target texture map 504.

[0113] According to an embodiment of the present disclosure, the texture decoder of the texture generation network performs a feature decoding operation on the second intermediate fusion feature and at least one level of style feature includes: based on an attention mechanism, the texture decoder performs a feature decoding operation on the second intermediate fusion feature, at least one level of style feature and at least one level of position feature.

[0114] According to an embodiment of the present disclosure, the feature decoding operation can be performed on the second intermediate fusion feature, the at least one level style feature and the at least one level position feature based on the multi-level decoder layers of the texture decoder. For example, based on the attention network algorithm, the second intermediate fusion feature, the at least one level style feature and the at least one level position feature are processed by using the first level decoder layer to implement the feature decoding operation on the second intermediate fusion feature to obtain the first level decoded feature. Based on the attention network algorithm, the first level decoded feature, the at least one level style feature and the at least one level position feature are processed by using the second level decoder layer to implement the feature decoding operation on the first level decoded feature to obtain the second level decoded feature. The target fusion feature is determined based on the last level decoded feature obtained by performing the feature decoding operation based on the last level decoder layer.

[0115] According to an embodiment of the present disclosure, by fusing the second intermediate fusion feature, the at least one level style feature and the at least one level position feature in the process of performing the feature decoding operation by using the texture decoder, the correlation between the position semantic attribute and the style semantic attribute of the object element can be enhanced based on the coordinates of the object element in the three-dimensional space represented by the position feature in the stage of recovering the semantic attributes of the target map, so as to improve the correlation accuracy between the style semantic attribute and the three-dimensional space coordinates of the object element of the target fusion feature, and also reduce the data processing amount of the texture generative large model for generating the target fusion feature, save the computing overhead, and improve the representation accuracy of the target fusion feature representing the target texture information of the three-dimensional object.

[0116] Figure 6 The schematic diagram of the texture generative large model according to another embodiment of the present disclosure is shown.

[0117] As Figure 6As shown, the texture generation large model 600 may include a second style feature extraction network 610, a texture generation network, and a location feature extraction network 630. The texture generation network includes a texture encoder 621 and a texture decoder 622. The texture generation network can be constructed based on a diffusion model algorithm. The second style feature extraction network 610 may include multiple style feature extraction layers determined based on a cascaded network structure. The multiple style feature extraction layers may include a first style feature extraction layer 611, a second style feature extraction layer 612, a third style feature extraction layer 613, and a fourth style feature extraction layer 614. The style feature extraction layers can be constructed based on a convolutional neural network algorithm, and the multiple style feature extraction layers can be used to extract style feature maps at different scales of the target image. The texture encoder 621 includes a first encoder layer 6211, a second encoder layer 6212, and a third encoder layer 6213 constructed based on an attention network algorithm. The texture decoder 622 includes a first decoder layer 6221, a second decoder layer 6222, and a third decoder layer 6223 constructed based on an attention network algorithm. The location feature extraction network 630 includes a cascaded first location feature extraction layer 631, a second location feature extraction layer 632, a third location feature extraction layer 633, and a fourth location feature extraction layer 634. The first location feature extraction layer 631, the second location feature extraction layer 632, the third location feature extraction layer 633, and the fourth location feature extraction layer 634 are constructed based on convolutional neural network algorithms of different scales.

[0118] like Figure 6 As shown, the initial style features are obtained by processing the target image 501 through the initial style feature extraction layer. The initial style features are then added to the initial fused features 502 obtained by feature fusion based on the position mapping image and the shape mask image to achieve feature fusion, resulting in the target initial style features. The target initial style features are input into the first style feature extraction layer 511 to extract style semantic features at the first scale, resulting in the first-level style features. The first-level style features are input into the second style feature extraction layer 512 to extract style semantic features at the second scale, resulting in the second-level style features. The second-level style features are input into the third style feature extraction layer 513 to extract style semantic features at the third scale, resulting in the third-level style features. The third-level style features are input into the fourth style feature extraction layer 514 to extract style semantic features at the fourth scale, resulting in the fourth-level style features.

[0119] like Figure 6As shown, the position mapping image 603 is added to the initial fusion feature 602 to realize feature fusion, and the fused initial position fusion feature is input into a first position feature extraction layer 631 to obtain a first-level position feature. The first-level position feature is input into a second position feature extraction layer 632 to obtain a second-level position feature. The second-level position feature is input into a third position feature extraction layer 633 to obtain a third-level position feature. The third-level position feature is input into a fourth position feature extraction layer 634 to obtain a fourth-level position feature.

[0120] In the feature encoding stage, the texture encoder of the texture generation network is used to perform feature encoding operation on the initial fusion feature to obtain the second intermediate fusion feature can include: the first encoder layer 6211 as the first-level encoder layer can process the initial fusion feature 602 and the noise information based on the attention network algorithm to realize the feature encoding operation on the initial fusion feature to obtain the first-level encoding feature. The second encoder layer 6212 as the second-level encoder layer can process the first-level encoding feature and the noise information based on the attention network algorithm to realize the feature encoding operation on the first-level encoding feature to obtain the second-level encoding feature. The third encoder layer 6213 as the third-level encoder layer can process the second-level encoding feature and the noise information based on the attention network algorithm to realize the feature encoding operation on the second-level encoding feature to obtain the third-level encoding feature. The third-level encoding feature can be used as the second-level intermediate fusion feature.

[0121] In the feature decoding stage, the first decoder layer 6221 of the texture encoder 622 performs the first-level feature decoding operation by fusing the second intermediate fusion feature, the fourth-level position feature and the fourth-level style feature based on the attention mechanism to obtain the first-level decoding feature. The second decoder layer 6222 performs the second-level feature decoding operation by fusing the first-level decoding feature, the third-level position feature and the third-level style feature based on the attention mechanism to obtain the second-level decoding feature. The third decoder layer 6223 performs the third-level feature decoding operation by fusing the second-level decoding feature, the second-level position feature and the second-level style feature based on the attention mechanism to obtain the third-level decoding feature. The third-level decoding feature is fused with the first-level position feature and the first-level style feature to obtain the target fusion feature 604. Based on the target fusion feature 603, image rendering can be performed to obtain the target texture map 605.

[0122] According to an embodiment of the present disclosure, processing the target image and the image representing the object form of the three-dimensional object using the texture generation large model can include: processing the target image and the object depth map of the image to be processed based on the texture generation large model to obtain a target texture image; and updating the texture attribute of the three-dimensional object based on the target texture image to obtain a target three-dimensional object.

[0123] According to an embodiment of the present disclosure, the object depth map represents an image of the three-dimensional object from a specified perspective. The specified perspective can represent a front perspective, a back perspective, a top perspective, etc. of the three-dimensional object. However, the specified perspective is not limited to this, and can also include any other perspective for observing the three-dimensional object, such as a perspective for observing the three-dimensional object at an angle of 45 degrees from above, and the embodiment of the present disclosure does not limit the specific type of the specified perspective.

[0124] According to an embodiment of the present disclosure, the depth image pixel of the object depth map can include depth information representing an object element of the three-dimensional object from the specified perspective, and the pixel coordinate of the depth image pixel can correspond to coordinate information of a projection plane corresponding to the object element from the specified perspective. The pixel coordinate and the depth information of the object depth map pixel can represent the coordinates of the object element in the three-dimensional space. For example, the pixel coordinate and the depth information of the depth map can be converted in the coordinate system based on the camera parameters corresponding to the specified perspective, to obtain the coordinates of the object element that can be observed from the specified perspective.

[0125] According to an embodiment of the present disclosure, the texture attribute updating of the three-dimensional object according to the target texture image can include updating the texture information of the three-dimensional object using the pixel texture information of the target texture image based on the position mapping relationship between the image pixels of the target texture image and the object elements of the three-dimensional object, so as to obtain the target three-dimensional object with the target texture information.

[0126] According to an embodiment of the present disclosure, the texture generation large model can learn the mapping relationship between the object depth map of the specified perspective and the object element through model fine-tuning. Therefore, by processing the object depth map of the target image and the to-be-processed image through the texture generation large model, the texture generation large model can control the texture semantic attribute fusion of the texture information of the target object in the target image according to the coordinates of the object element represented in the three-dimensional space based on the depth image pixel and the depth information of the object depth map, so as to accurately update the texture semantic attribute of the object element of the three-dimensional object from the specified perspective. Therefore, the obtained target texture image can more accurately represent the result of displaying the three-dimensional object from the specified perspective according to the texture semantic attribute of the target object. In this way, the target three-dimensional object generated based on the target texture image can more accurately represent the texture semantic attribute of the target object in the three-dimensional space, and the matching degree between the target three-dimensional object and the target object in the target image can be improved.

[0127] According to an embodiment of the present disclosure, the texture attribute updating of the three-dimensional object based on the target texture image to obtain the target three-dimensional object can comprise: determining, based on the target texture image and the object depth map, a pixel mapping relationship between a first pixel of the target texture image and a second pixel of the initial object map; updating, based on the pixel mapping relationship, the initial object map according to target texture information of the target texture image to obtain a target object map; and performing object rendering based on the target object map to obtain the target three-dimensional object.

[0128] According to an embodiment of the present disclosure, the initial object map is determined based on the three-dimensional object. The initial object map may, for example, be a UV map used for rendering to obtain the three-dimensional object. The pixel coordinates of the second pixel of the initial object map can have a position mapping relationship with the three-dimensional coordinates of the object element in the three-dimensional space.

[0129] According to an embodiment of the present disclosure, the first pixel and the second pixel having the pixel mapping relationship can represent the object element at the same position in the three-dimensional space. The first pixel of the target texture image and the depth map pixel of the object depth map can be aligned in the coordinate system based on a preset camera parameter, thereby obtaining the aligned target texture image and the object depth map. Based on the aligned first pixel and the depth map pixel, the aligned first pixel can be enabled to determine the second pixel representing the object element at the same position from the initial object map based on the depth information of the associated depth map pixel, thereby the pixel mapping relationship between the first pixel and the second pixel can be determined.

[0130] According to an embodiment of the present disclosure, updating the initial object map according to the target texture information of the target texture image based on the pixel mapping relationship can comprise: updating the initial object map using the target texture information of the target texture image under a plurality of specified viewing angles based on the pixel mapping relationship.

[0131] According to an embodiment of the present disclosure, the target texture image under a plurality of specified viewing angles can represent the target texture information of the object element of the three-dimensional object observed by the plurality of specified viewing angles. The second pixel having the pixel mapping relationship is updated by the target texture information of the first pixel under the plurality of specified viewing angles, respectively, so that the map pixel of the obtained target object map can represent the target texture information of the object element of the plurality of specified viewing angles. Therefore, the target three-dimensional object is rendered based on the target object map, so that the target texture information of the target three-dimensional object can be more accurately matched with the texture information of the target object to improve the matching degree between the target three-dimensional object and the target object.

[0132] Figure 7 An application scenario diagram of a large model-based virtual avatar generation method according to an embodiment of the present disclosure is schematically shown.

[0133] As Figure 7As shown, in this application scenario, the user can input a two-dimensional target image 110 and a text command "Please describe the hairstyle, bangs, face shape, and clothing style of the two-dimensional character in the picture". The target image 701 can represent a virtual character wearing a purple dress. The text command is used as a prompt word of the large model, and the large model is used to process the target image 701 and the text command to obtain object description information 710. The object description information 710 can include character description information of the virtual character and clothing description information of the clothing object. The character description information may, for example, be: 1) gender: female; 2) hairstyle: cute short hair; 3) bang type: straight bangs; 4) glasses shape: ……, etc. The clothing description information may, for example, be: 9) clothing style: short-sleeved dress; 10) clothing length: ……, etc.

[0134] Based on the character description information and the clothing description information, a search is performed in a preset three-dimensional object library to obtain a search result 720, which includes a virtual character three-dimensional object 721 and a clothing three-dimensional object 722. The clothing three-dimensional object 722 can be a three-dimensional model without texture information.

[0135] The clothing three-dimensional object 722 and the target image 701 are input into a texture generation large model to obtain a target three-dimensional object 702. The target three-dimensional object 702 can be a texture clothing three-dimensional object having the same or similar texture information as the clothing in the target image 701. The target three-dimensional object 702 and the virtual character three-dimensional object 721 are rendered in a three-dimensional space to obtain a three-dimensional virtual character matching the virtual character wearing a purple dress in the target image 701. The three-dimensional virtual character can be used in animation production, intelligent customer service, and other scenarios.

[0136] According to an embodiment of the present disclosure, the virtual character generation method based on a large model further includes: in response to a modification request for the displayed virtual character, determining a modification prompt word matching the modification request; processing the target image, the to-be-processed image, and the modification prompt word based on the texture generation large model to obtain an updated target three-dimensional object; and generating an updated virtual character according to the updated target three-dimensional object.

[0137] According to an embodiment of the present disclosure, the modification request can be determined based on a user's modification operation on the currently displayed virtual character. For example, the user can input a text "Please lengthen the sleeve length of the clothes" based on any interactive operation type, and generate a modification request based on the input text.

[0138] According to an embodiment of the present disclosure, the modification prompt word matched with the modification request can be used to represent the modification demand for the currently displayed target three-dimensional object represented by the modification request. For example, the modification prompt word can be: "lengthen the sleeve length of the heroine's clothes". The modification prompt word can be used to control the texture generation type large model to process the target image and the image to be processed, so as to lengthen the sleeve of the current three-dimensional clothes model according to the modification intention represented by the modification prompt word, and generate an updated three-dimensional clothes model as the target three-dimensional object. The updated virtual image can be obtained by fusing the updated target three-dimensional object with other target three-dimensional objects that do not need to be modified.

[0139] According to an embodiment of the present disclosure, the virtual image generation method based on a large model can further include: in response to the target driving instruction, driving the virtual image to perform an action related to the target driving instruction.

[0140] According to an embodiment of the present disclosure, the target driving instruction can be generated based on the operation request of the user, or the target driving instruction can also be generated based on other manners. For example, the target driving instruction can be determined in response to the generation of the target three-dimensional object.

[0141] According to an embodiment of the present disclosure, the target driving instruction can include action parameters representing any action such as mouth shape action, expression change action, posture change action, etc. The virtual image can be controlled to perform the action based on the action parameters, so as to meet the action demand represented by the target driving instruction. In this way, the virtual image generation method based on a large model provided by the embodiment of the present disclosure can be used in any AIGC scene such as game animation, video production, intelligent e-commerce, etc.

[0142] Based on the virtual image generation method based on a large model provided by the present disclosure, the present disclosure further provides a virtual image generation device based on a large model. The following will be described in detail Figure 8 The device will be described in detail.

[0143] Figure 8 The block diagram of the virtual image generation device based on a large model according to an embodiment of the present disclosure is schematically shown.

[0144] As Figure 8 shown, the virtual image generation device based on a large model 800 can include an object description information obtaining module 810, a target three-dimensional object obtaining module 820 and a virtual image generation module 830.

[0145] The object description information obtaining module 810 is configured to process a target image including a target object having texture information by using a large model, to obtain object description information.

[0146] The target three-dimensional object obtaining module 820 is configured to process the target image and the image representing the object form of the three-dimensional object by using the texture generation large model to obtain a target three-dimensional object having target texture information, the three-dimensional object being determined based on the object description information, and the target texture information being matched with the texture information; and

[0147] The virtual image generation module 830 is configured to generate a virtual image based on the target three-dimensional object.

[0148] According to an embodiment of the present disclosure, the texture generation large model comprises a texture generation network, and the image to be processed comprises a position mapping image, and a position pixel of the position mapping image represents a three-dimensional coordinate of an object element in the three-dimensional object.

[0149] According to an embodiment of the present disclosure, the target three-dimensional object obtaining module comprises a target texture map obtaining sub-module and a first obtaining sub-module.

[0150] The target texture map obtaining sub-module is configured to perform feature fusion based on the texture generation network and according to the object style feature and the position mapping image to obtain a target texture map matched with the texture information, wherein the object style feature is determined based on the target image.

[0151] The first obtaining sub-module is configured to update the three-dimensional object based on the target texture information of the target texture map to obtain the target three-dimensional object.

[0152] According to an embodiment of the present disclosure, the target texture map obtaining sub-module comprises an initial fusion feature obtaining unit and a target fusion feature obtaining unit.

[0153] The initial fusion feature obtaining unit is configured to perform feature fusion on the position mapping image and a shape mask image in the image to be processed to obtain an initial fusion feature, wherein a mask pixel of the shape mask image represents whether an object element of the three-dimensional object can store texture information, and the mask pixel has a position mapping relationship with the object element.

[0154] The target fusion feature obtaining unit is configured to perform feature fusion based on the initial fusion feature and the object style feature to obtain a target fusion feature, wherein the target texture map is determined based on the target fusion feature.

[0155] According to an embodiment of the present disclosure, the texture generation large model further comprises a first style feature extraction network, the first style feature extraction network comprises a down-sampling layer and an up-sampling layer having a U-shaped network structure, and the object style feature comprises at least one level of down-sampling style feature and at least one level of up-sampling style feature obtained by processing the target image by using the down-sampling layer and the up-sampling layer.

[0156] According to an embodiment of the present disclosure, the target fusion feature obtaining unit comprises: a first intermediate fusion feature obtaining subunit and a first target fusion feature obtaining subunit.

[0157] The first intermediate fusion feature obtaining subunit is configured to perform a feature encoding operation on the initial fusion feature and the at least one level down-sampled style feature by using a texture encoder of the texture generation network to obtain the first intermediate fusion feature.

[0158] The first target fusion feature obtaining subunit is configured to perform a feature decoding operation on the first intermediate fusion feature and the at least one level up-sampled style feature by using a texture decoder of the texture generation network to obtain the target fusion feature.

[0159] According to an embodiment of the present disclosure, the first target fusion feature obtaining subunit is configured to perform the feature decoding operation on the first intermediate fusion feature, the at least one level up-sampled style feature and at least one level position feature by using the texture decoder based on an attention mechanism, wherein the position feature is determined according to a position mapping image.

[0160] According to an embodiment of the present disclosure, the texture generation large model further comprises a second style feature extraction network, and the second style feature extraction network comprises a cascade of M level style feature extraction layers.

[0161] According to an embodiment of the present disclosure, the target fusion feature obtaining unit comprises: a second intermediate fusion feature obtaining subunit and a second target fusion feature obtaining subunit.

[0162] The second intermediate fusion feature obtaining subunit is configured to perform a feature encoding operation on the initial fusion feature by using a texture encoder of the texture generation network to obtain the second intermediate fusion feature; and

[0163] The second target fusion feature obtaining subunit is configured to perform a feature decoding operation on the second intermediate fusion feature and the at least one level style feature by using a texture decoder of the texture generation network to obtain the target fusion feature.

[0164] According to an embodiment of the present disclosure, the second target fusion feature obtaining subunit is configured to perform the feature decoding operation on the second intermediate fusion feature, the at least one level style feature and at least one level position feature by using the texture decoder based on an attention mechanism, wherein the position feature is determined according to a position mapping image.

[0165] According to an embodiment of the present disclosure, the texture generation large model further comprises a position feature extraction network, and the position feature extraction network comprises a cascade of multiple level position feature extraction layers.

[0166] According to an embodiment of the present disclosure, the position feature comprises a plurality of levels, and the multi-level position feature is determined by processing the target image by using a multi-level position feature extraction layer.

[0167] According to an embodiment of the present disclosure, the target three-dimensional object obtaining module comprises a target texture image obtaining submodule and a second obtaining submodule.

[0168] The target texture image obtaining submodule is configured to process the object depth map of the target image and the to-be-processed image based on the texture generation type large model to obtain a target texture image, wherein the object depth map represents an image of the three-dimensional object at a specified view angle.

[0169] The second obtaining submodule is configured to perform texture attribute updating on the three-dimensional object based on the target texture image to obtain a target three-dimensional object.

[0170] According to an embodiment of the present disclosure, the second obtaining submodule comprises a pixel mapping relationship determining unit, a target object map determining unit, and a target three-dimensional object obtaining unit.

[0171] The pixel mapping relationship determining unit is configured to determine a pixel mapping relationship between a first pixel of the target texture image and a second pixel of an initial object map based on the target texture image and the object depth map, wherein the initial object map is determined based on the three-dimensional object.

[0172] The target object map determining unit is configured to update the initial object map based on the target texture information of the target texture image to obtain a target object map based on the pixel mapping relationship.

[0173] The target three-dimensional object obtaining unit is configured to perform object rendering based on the target object map to obtain a target three-dimensional object.

[0174] According to an embodiment of the present disclosure, the target object map determining unit comprises an updating submodule.

[0175] The updating submodule is configured to update the initial object map based on the pixel mapping relationship by using target texture information of the target texture image at a plurality of specified view angles.

[0176] According to an embodiment of the present disclosure, the avatar generation apparatus based on a large model further comprises a modification prompt word determining module, a target three-dimensional object modifying module, and an avatar modifying module.

[0177] The modification prompt word determining module is configured to determine a modification prompt word matched with a modification request in response to the modification request for the displayed avatar.

[0178] The target three-dimensional object modifying module is configured to process the target image, the to-be-processed image, and the modification prompt word based on the texture generation type large model to obtain an updated target three-dimensional object.

[0179] a virtual avatar modification module configured to generate an updated virtual avatar according to the updated target three-dimensional object.

[0180] According to embodiments of the present disclosure, the target object comprises at least one of the following: a clothing object, a body part object, a vehicle object, a building object.

[0181] According to embodiments of the present disclosure, the large model based virtual avatar generation apparatus further comprises a driving module.

[0182] The driving module is configured to drive the virtual avatar to perform an action related to the target driving instruction in response to the target driving instruction.

[0183] Based on the large model based virtual avatar generation method described above, embodiments of the present disclosure further provide an artificial intelligence agent, which is configured to execute the large model based virtual avatar generation method provided in the above embodiments. The agent will be described in detail below in combination with the figures.

[0184] Figure 9 A structural block diagram of an artificial intelligence agent according to embodiments of the present disclosure is schematically shown.

[0185] In embodiments of the present disclosure, inspired by the Von Neumann structure in modern computer theory, as shown in Figure 9 The AI agent 900 can include five core modules: an input module 910, a control module 920, a storage module 930, a calculation module 940, and an output module 950.

[0186] The input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals or data from the outside world (e.g. users or external environment), and converting them into a format that the AI agent 900 can understand and process. The input module 910 is the first link for the AI agent 900 to interact with the outside world, which enables the AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to it.

[0187] In examples, the input module 910 can input information such as the target image, the image to be processed or the three-dimensional object described above.

[0188] In examples, the control module 920 is the core support for the AI agent 900 to process complex tasks. The control module 920 can execute the large model based virtual avatar generation method described above.

[0189] In an example, the control module 920 will interact with the storage module 930, the operation module 940 and / or the output module 950 constantly during the running process. However, it is noted that in the embodiments of the present disclosure, the control module 920 initiates the communication with the storage module 930, the operation module 940 and / or the output module 950 as a single initiator, and there is no communication coupling between the storage module 930, the operation module 940 and / or the output module 950.

[0190] In an example, the performance of the control module 920 can be closely related to the large model on which the AI agent 900 is based. In order to fully exert the capability of the large model, the internal structure of the control module 920 can be designed to be highly configurable and scalable in order to cope with various different types of tasks and requirements in real scenarios.

[0191] The storage module 930 can be responsible for memorizing historical dialogues, event streams and the like. The prompt information and data resources as described above can be included in the storage module 930.

[0192] In an example, after obtaining the search request, the AI agent 900 can process the target image by using the large model, or can also process the target image by using the texture generation type large model to obtain the target three-dimensional object. The target image can be stored in the storage module 930. The AI agent 900 can determine the three-dimensional object, the image to be processed and the like from the storage module 930 and feed them back to the control module 920. Then, the control module 920 can obtain the target three-dimensional object corresponding to the target image by using the feedback prompt information and data resources, and deliver the target three-dimensional object to the output module 950.

[0193] The operation module 940 can be regarded as a pre-defined tool library. The rendering engine and the like as described above can be included in the operation module 940.

[0194] In an example, the output module 950 can output the target three-dimensional object or the virtual image as described above.

[0195] The AI agent 900 according to the embodiments of the present disclosure can simply and effectively improve the intelligent degree and improve the flexibility and versatility.

[0196] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0197] According to the embodiments of the present disclosure, an electronic device comprises at least one processor and a memory connected in communication with the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0198] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the large model based virtual image generation method described above.

[0199] According to an embodiment of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the large model based virtual image generation method.

[0200] Figure 10 A block diagram schematically illustrates an electronic device suitable for implementing the large model based virtual image generation method according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0201] As shown in Figure 10 The device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0202] Various components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc., an output unit 1007, such as various types of displays, speakers, etc., the storage unit 1008, such as a magnetic disk, an optical disk, etc., and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0203] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The computing unit 1001 performs various methods and processes described above, such as the large model based virtual avatar generation method. For example, in some embodiments, the large model based virtual avatar generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded onto the RAM 1003 and executed by the computing unit 1001, one or more steps of the large model based virtual avatar generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the large model based virtual avatar generation method by any other suitable means, such as by means of firmware.

[0204] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0205] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0206] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0207] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0208] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0209] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0210] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, using the flow. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions of the present disclosure can be achieved, which are not limited herein.

[0211] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A large model-based virtual avatar generation method, comprising: processing a target image including a target object by using a large model to obtain object description information, the target object having texture information; processing the target image and a to-be-processed image representing the object form of a three-dimensional object by using a texture generation large model to obtain a target three-dimensional object having target texture information, the three-dimensional object being determined based on the object description information, and the target texture information matching the texture information; generating a virtual avatar based on the target three-dimensional object; wherein the to-be-processed image includes a position mapping image, and a position pixel of the position mapping image represents a three-dimensional coordinate of an object element in the three-dimensional object; processing the target image and the to-be-processed image representing the object form of a three-dimensional object by using a texture generation large model comprises: using an object style feature and the position mapping image as prior information, and performing feature fusion on the object style feature and the position mapping image by using a texture generation network of the texture generation large model to obtain a target texture map matching the texture information, the object style feature being determined based on the target image, and the prior information being used to make the texture generation network learn the correspondence between the three-dimensional coordinate of the object element and the object style feature; updating the three-dimensional object based on the target texture information of the target texture map to obtain the target three-dimensional object.

2. The method of claim 1, wherein, the feature fusion performed by the texture generation network of the texture generation large model on the object style feature and the position mapping image comprises: performing feature fusion on the position mapping image and a shape mask image in the to-be-processed image to obtain initial fusion features, wherein a mask pixel of the shape mask image represents whether an object element of the three-dimensional object can store texture information, and the mask pixel has a position mapping relationship with the object element; performing feature fusion based on the initial fusion features and the object style feature to obtain target fusion features, wherein the target texture map is determined based on the target fusion features.

3. The method of claim 2, wherein, The texture generation large model further comprises a first style feature extraction network, and the first style feature extraction network comprises a down-sampling layer and an up-sampling layer having a U-shaped network structure, and the object style feature comprises at least one level of down-sampling style feature and at least one level of up-sampling style feature obtained by processing the target image by using the down-sampling layer and the up-sampling layer; wherein the feature fusion based on the initial fusion features and the object style feature to obtain target fusion features comprises: performing feature encoding on the initial fusion features and at least one level of the down-sampling style feature by using a texture encoder of the texture generation network to obtain first intermediate fusion features; and performing feature decoding on the first intermediate fusion features and at least one level of the up-sampling style feature by using a texture decoder of the texture generation network to obtain the target fusion features.

4. The method of claim 3, wherein, the feature decoding performed by the texture decoder of the texture generation network on the first intermediate fusion features and at least one level of the up-sampling style feature comprises: performing, by the texture decoder of the texture generation network, a feature decoding operation on the second intermediate fusion feature and at least one level of the style feature to obtain the target fusion feature.

5. The method according to claim 2, wherein, The texture generation model further comprises a second style feature extraction network, the second style feature extraction network comprising a plurality of cascaded style feature extraction layers, and the object style feature comprises a plurality of style features obtained by processing the object style feature by using the plurality of cascaded style feature extraction layers. The feature fusion based on the initial fusion feature and the object style feature comprises: performing, by the texture encoder of the texture generation network, a feature encoding operation on the initial fusion feature to obtain a second intermediate fusion feature; and performing, by the texture decoder of the texture generation network, a feature decoding operation on the second intermediate fusion feature and at least one level of the style feature to obtain the target fusion feature.

6. The method according to claim 5, wherein, The feature decoding operation performed by the texture decoder of the texture generation network on the second intermediate fusion feature and at least one level of the style feature comprises: performing, by the texture decoder of the texture generation network, a feature decoding operation on the second intermediate fusion feature, at least one level of the style feature, and at least one level of a position feature based on an attention mechanism, the position feature being determined according to the position mapping image.

7. The method according to claim 4 or 6, wherein, The texture generation model further comprises a position feature extraction network, the position feature extraction network comprising a plurality of cascaded position feature extraction layers. The position feature comprises a plurality of levels, and the plurality of levels of the position feature are determined by processing the target image by using the plurality of cascaded position feature extraction layers.

8. The method of claim 1, wherein, The processing of the target image and the image representing the object form of the three-dimensional object by using the texture generation model further comprises: processing, by the texture generation model, an object depth map of the target image and the image to obtain a target texture image, the object depth map representing an image of the three-dimensional object at a specified view angle; and updating a texture attribute of the three-dimensional object based on the target texture image to obtain a target three-dimensional object.

9. The method of claim 8, wherein, The updating of the texture attribute of the three-dimensional object based on the target texture image to obtain the target three-dimensional object comprises: determining, based on the target texture image and the object depth map, a pixel mapping relationship between a first pixel of the target texture image and a second pixel of an initial object map, the initial object map being determined based on the three-dimensional object; updating, based on the pixel mapping relationship, the initial object map according to target texture information of the target texture image to obtain a target object map; and performing object rendering based on the target object map to obtain the target three-dimensional object.

10. The method of claim 9, wherein, The updating of the initial object map according to the target texture information of the target texture image based on the pixel mapping relationship comprises: updating, based on the pixel mapping relationship, the initial object map according to target texture information of a plurality of target texture images at the specified view angle.

11. The method of claim 1 or 8, further comprising: in response to a modification request for the virtual image displayed, determining a modification prompt word matched with the modification request; processing the target image, the image to be processed and the modification prompt word based on the texture generation large model to obtain an updated target three-dimensional object; and generating an updated virtual image according to the updated target three-dimensional object.

12. The method of claim 1, wherein, The target object includes at least one of: clothing objects, body part objects, vehicle objects, building objects.

13. The method of claim 1, further comprising: in response to a target driving instruction, driving the virtual image to perform an action related to the target driving instruction.

14. An artificial intelligence agent system comprising: input module, storage module, control module, operation module and output module; The input module is used to input target image and image to be processed; The storage module is used to store the target image and the image to be processed; the control module is used to execute the method according to any one of claims 1 to 13; The control module initiates communication to the storage module, the operation module and the output module as a single initiator, and the output module is used to output virtual image.

15. A large model based virtual image generation device, comprising: an object description information obtaining module for processing a target image including a target object using a large model to obtain object description information, the target object having texture information; a target three-dimensional object obtaining module for processing the target image and an image to be processed representing the object form of a three-dimensional object using a texture generation large model to obtain a target three-dimensional object having target texture information, the three-dimensional object being determined based on the object description information, the target texture information matching the texture information; and a virtual image generation module for generating a virtual image based on the target three-dimensional object; wherein the texture generation large model includes a texture generation network, and the image to be processed includes a position mapping image, the position pixels of the position mapping image representing the three-dimensional coordinates of object elements in the three-dimensional object; wherein the target three-dimensional object obtaining module includes: a target texture map obtaining submodule for using the texture generation network of the texture generation large model to perform feature fusion on object style features and the position mapping image based on the three-dimensional coordinates represented by the position pixels as prior information, to obtain a target texture map matching the texture information, the object style features being determined based on the target image, and the prior information being used to make the texture generation network learn the correspondence between the three-dimensional coordinates of object elements and object style features; and a first obtaining submodule for updating the three-dimensional object based on the target texture information of the target texture map to obtain the target three-dimensional object.

16. The apparatus of claim 15, wherein, The target texture map obtaining submodule includes: An initial fusion feature obtaining unit is configured to perform feature fusion on a shape mask image in the position mapping image and the to-be-processed image to obtain initial fusion features, wherein a mask pixel of the shape mask image represents whether an object element of the three-dimensional object can store texture information, and the mask pixel and the object element have a position mapping relationship; A target fusion feature obtaining unit is configured to perform feature fusion on the initial fusion features and the object style features to obtain target fusion features, wherein the target texture map is determined based on the target fusion features.

17. The apparatus of claim 16, wherein, The texture generation large model further includes a first style feature extraction network, the first style feature extraction network includes a down-sampling layer and an up-sampling layer having a U-shaped network structure, and the object style features include at least one level of down-sampling style features and at least one level of up-sampling style features obtained by processing the target image by using the down-sampling layer and the up-sampling layer. The target fusion feature obtaining unit includes: A first intermediate fusion feature obtaining subunit is configured to perform feature encoding on the initial fusion features and at least one level of the down-sampling style features by using a texture encoder of the texture generation network to obtain first intermediate fusion features; and A first target fusion feature obtaining subunit is configured to perform feature decoding on the first intermediate fusion features and at least one level of the up-sampling style features by using a texture decoder of the texture generation network to obtain the target fusion features.

18. The apparatus of claim 17, wherein, The first target fusion feature obtaining subunit is configured to: perform feature decoding on the first intermediate fusion features, at least one level of the up-sampling style features, and at least one level of position features by using the texture decoder based on an attention mechanism, wherein the position features are determined according to the position mapping image.

19. The apparatus of claim 16, wherein, The texture generation large model further includes a second style feature extraction network, the second style feature extraction network includes a plurality of cascaded style feature extraction layers, and the object style features include a plurality of style features obtained by processing the object style features by using the plurality of cascaded style feature extraction layers. The target fusion feature obtaining unit includes: A second intermediate fusion feature obtaining subunit is configured to perform feature encoding on the initial fusion features by using a texture encoder of the texture generation network to obtain second intermediate fusion features; and A second target fusion feature obtaining subunit is configured to perform feature decoding on the second intermediate fusion features and at least one level of the style features by using a texture decoder of the texture generation network to obtain the target fusion features.

20. The apparatus of claim 19, wherein, The second target fusion feature obtaining subunit is configured to: perform feature decoding on the second intermediate fusion features, at least one level of the style features, and at least one level of position features by using the texture decoder based on an attention mechanism, wherein the position features are determined according to the position mapping image.

21. The apparatus of claim 18 or 20, wherein, The texture generation large model further includes a position feature extraction network, the position feature extraction network includes a plurality of cascaded position feature extraction layers. The position feature includes multiple levels, and the multiple levels of the position feature are determined by processing the target image using the multiple level position feature extraction layer.

22. The apparatus of claim 15, wherein, The target three-dimensional object obtaining module includes: A target texture image obtaining sub-module is configured to process object depth maps of the target image and the image to be processed based on the texture generative large model to obtain a target texture image, where the object depth map represents an image of the three-dimensional object at a specified view angle; and A second obtaining sub-module is configured to update a texture attribute of the three-dimensional object based on the target texture image to obtain the target three-dimensional object.

23. The apparatus of claim 22, wherein, The second obtaining sub-module includes: A pixel mapping relationship determining unit is configured to determine a pixel mapping relationship between a first pixel of the target texture image and a second pixel of an initial object map based on the target texture image and the object depth map, where the initial object map is determined based on the three-dimensional object; A target object map determining unit is configured to update the initial object map based on the target texture information of the target texture image according to the pixel mapping relationship to obtain a target object map; and A target three-dimensional object obtaining unit is configured to perform object rendering based on the target object map to obtain the target three-dimensional object.

24. The apparatus of claim 23, wherein, The target object map determining unit includes: An updating sub-unit is configured to update the initial object map based on the pixel mapping relationship using target texture information of multiple target texture images at the specified view angle.

25. The apparatus of claim 15 or 22, further comprising: A modification prompt word determining module is configured to determine a modification prompt word matched with a modification request for the virtual figure in response to the modification request; A target three-dimensional object modifying module is configured to process the target image, the image to be processed, and the modification prompt word based on the texture generative large model to obtain an updated target three-dimensional object; and A virtual figure modifying module is configured to generate an updated virtual figure according to the updated target three-dimensional object. The target object includes at least one of the following:

26. The apparatus of claim 15, wherein, A clothing object, a body part object, a vehicle object, and a building object.

27. The apparatus of claim 15, further comprising: A driving module is configured to drive the virtual figure to perform an action related to a target driving instruction in response to the target driving instruction.

28. An electronic device, comprising: At least one processor; And A memory connected in communication with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 13. The computer instructions are used to enable the computer to perform the method of any one of claims 1 to 13.

29. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, 30. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 13. ​

Citation Information

Patent Citations

  • Virtual object image synthesis method and device, electronic equipment and storage medium

    CN112150638A

  • Method and device for generating personalized texture map

    CN115345980A