Virtual image generation method, computer terminal, storage medium and program product

By generating a mesh model and rendering the mesh model with attribute information images from multiple perspectives, the problem of poor quality of virtual images in existing technologies is solved, and high-quality virtual image generation is achieved from different perspectives.

CN121937592APending Publication Date: 2026-04-28ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALIBABA (CHINA) CO LTD
Filing Date
2024-10-18
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The virtual avatars generated from single images in the existing technology are of poor quality, especially when the avatars are moving, they are easily damaged or the viewpoint is limited, and they cannot perform diverse movements.

Method used

By generating a mesh model and multiple attribute information images from different perspectives based on a single image containing a real object, and using these images to render the mesh model, a three-dimensional virtual image is generated, ensuring high visual appeal from different perspectives.

Benefits of technology

It improves the quality of virtual character generation, avoids damage during large-scale movement changes, and achieves diverse movement variations and high-quality visual effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937592A_ABST
    Figure CN121937592A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual image generation method, a computer terminal, a storage medium and a program product. The method relates to the field of artificial intelligence and digital human and comprises the steps that an input instruction acting on an operation interface is responded, a first image corresponding to the input instruction is determined, and the first image comprises an image of a real object in a target posture; a grid model and a plurality of second images are generated based on the first image, and different second images are used for representing attribute information of the real object at different visual angles; rendering the grid model based on the plurality of second images to generate a virtual image corresponding to the real object; and displaying the virtual image on the operation interface. The technical problem that the quality of a virtual image generated according to a single image is poor in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and digital humans, and more specifically, to a method for generating virtual avatars, a computer terminal, a storage medium, and a program product. Background Technology

[0002] With the development of video generation technology, the secondary creation of some popular materials and film and television materials has gradually become a popular form of entertainment. Currently, the common methods for generating virtual characters from images include 2D (Dimensional) rendering and 3D+2D rendering. However, virtual characters generated by these two methods usually have obvious drawbacks. For example, virtual characters generated by 2D rendering are usually limited by the image display direction and cannot generate virtual characters in other directions that are different from the image display direction. When making large-scale changes in movement, the virtual character may be damaged or the result of the movement change may be poor. On the other hand, virtual characters generated by 3D+2D rendering are usually limited by fixed action templates and cannot make diverse movement changes, resulting in poor quality of the actual generated virtual character. Summary of the Invention

[0003] This application provides a method for generating virtual avatars, a computer terminal, a storage medium, and a program product, to at least solve the technical problem of poor quality of virtual avatars generated from a single image in related technologies.

[0004] According to one aspect of the embodiments of this application, a method for generating a virtual avatar is provided, comprising: responding to an input command applied to an operation interface, determining a first image corresponding to the input command; generating a mesh model and a plurality of second images based on the first image; rendering the mesh model based on the plurality of second images to generate a virtual avatar corresponding to a real object; and displaying the virtual avatar on the operation interface.

[0005] According to another aspect of the embodiments of this application, a method for generating a virtual image is also provided, comprising: acquiring a first image, wherein the first image contains an image of a real object in a target pose; generating a mesh model and a plurality of second images based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; and rendering the mesh model based on the plurality of second images to generate a virtual image corresponding to the real object.

[0006] According to another aspect of the embodiments of this application, a method for generating a virtual avatar is also provided, comprising: obtaining a first image by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes a first image, and the first image contains an image of a real object in a target pose; generating a mesh model and a plurality of second images based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; rendering the mesh model based on the plurality of second images to generate a virtual avatar corresponding to the real object; and outputting the virtual avatar by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the virtual avatar.

[0007] According to one aspect of the embodiments of this application, a virtual avatar generation apparatus is provided, comprising: an image determination module, configured to determine a first image corresponding to an input command applied to an operation interface, wherein the first image includes an image of a real object in a target pose; an information generation module, configured to generate a mesh model and multiple second images based on the first image, wherein different second images are used to characterize attribute information of the real object from different perspectives; a model rendering module, configured to render the mesh model based on the multiple second images to generate a virtual avatar corresponding to the real object; and an avatar display module, configured to display the virtual avatar on the operation interface.

[0008] According to another aspect of the embodiments of this application, a virtual avatar generation apparatus is also provided, comprising: an image acquisition module for acquiring a first image, wherein the first image contains an image of a real object in a target pose; a first generation module for generating a mesh model and a plurality of second images based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; and a second generation module for rendering the mesh model based on the plurality of second images to generate a virtual avatar corresponding to the real object.

[0009] According to another aspect of the embodiments of this application, a virtual avatar generation apparatus is also provided, comprising: an image calling module, configured to obtain a first image by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the first image, and the first image contains an image of a real object in a target pose; a third generation module, configured to generate a mesh model and multiple second images based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; a fourth generation module, configured to render the mesh model based on the multiple second images to generate a virtual avatar corresponding to the real object; and an avatar output module, configured to output the virtual avatar by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the virtual avatar.

[0010] According to another aspect of the embodiments of this application, a computer terminal is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0011] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0012] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0013] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the methods in various embodiments of this application.

[0014] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

[0015] In this embodiment, a first image corresponding to an input command is determined in response to an input command applied to the operation interface; a mesh model and multiple second images are generated based on the first image; the mesh model is rendered based on the multiple second images to generate a virtual image corresponding to the real object; and the virtual image is displayed on the operation interface. By generating a mesh model matching the real object and multiple attribute information from different perspectives (i.e., second images) based on a single first image containing the real object, and using the second images to render the mesh model to obtain a three-dimensional virtual image, it is possible to ensure that the virtual image has high visibility from different perspectives and will not be damaged when performing large-scale actions, thereby improving the quality of the generated virtual image and solving the technical problem of poor quality of virtual images generated from a single image in related technologies.

[0016] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 This is a schematic diagram illustrating a virtual avatar generation scenario according to this application;

[0019] Figure 2 This is a structural block diagram of a computing environment for generating virtual images, as shown in this application.

[0020] Figure 3 This is a flowchart illustrating a method for generating a virtual avatar according to this application;

[0021] Figure 4 This is a schematic diagram illustrating a mesh model generation process according to an embodiment of this application;

[0022] Figure 5 This is a schematic diagram illustrating a second image generation process according to an embodiment of this application;

[0023] Figure 6 This is a schematic diagram illustrating a skeletal binding result according to an embodiment of this application;

[0024] Figure 7 This is a schematic diagram illustrating a virtual avatar video generation process according to an embodiment of this application;

[0025] Figure 8 This is another method for generating a virtual image according to embodiments of this application;

[0026] Figure 9 This is another method for generating a virtual image according to embodiments of this application;

[0027] Figure 10 This is a structural block diagram of a virtual image generation device according to an embodiment of this application;

[0028] Figure 11 This is a structural block diagram of another virtual image generation device according to an embodiment of this application;

[0029] Figure 12 This is a structural block diagram of another virtual image generation device according to an embodiment of this application;

[0030] Figure 13 This is a structural block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0034] GaussianSplatting: An advanced surface reconstruction and rendering technology.

[0035] Rigging: A skeletal binding technique that attaches a skeleton to a 3D human model for driving.

[0036] DiT: Diffusion Transformers, an advanced sequence modeling neural network.

[0037] 3D-VAE: 3D Variational AutoEncoder is a neural network that compresses three-dimensional representations.

[0038] Triplanes: An advanced 3D compressed representation.

[0039] T-Pose: A human posture with arms outstretched in a T-shape.

[0040] Skinning Weight: The weight of skinning on bones.

[0041] LBS: Linear Blending Skinning technology.

[0042] According to an embodiment of this application, a method for generating a virtual avatar is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0043] Considering that machine learning models consume significant computing resources on mobile terminals, the methods described above in this application can be applied to, for example... Figure 1 The application scenarios shown are not limited to these. Figure 1 This is a schematic diagram of a virtual avatar generation scenario shown in this application, in which... Figure 1 In the application scenario shown, the machine learning model is deployed on server 10. Server 10 can connect to one or more clients 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. Clients 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Clients 20 can interact with users through a graphical user interface to invoke the large model, thereby implementing the method provided in this application embodiment.

[0044] It should be noted that, provided that the client device's operating resources can meet the deployment and operation conditions of the large model, the embodiments of this application can be performed on the client device.

[0045] Figure 2 This is a structural block diagram of a computing environment for generating virtual images, as shown in this application. Figure 2 As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (represented as 210-1, 210-2, ... in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.

[0046] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).

[0047] The services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.

[0048] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers within a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.

[0049] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201, and executing one or more functions of one service may require calling one or more functions of another service. For example... Figure 2 As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.

[0050] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.

[0051] Under the aforementioned operating environment, this application provides the following: Figure 3 The method for generating the virtual avatar shown is illustrated. It should be noted that the method for generating the virtual avatar in this embodiment can be derived from... Figure 1The server execution of the illustrated embodiment is also considered to be a user-interactive process when generating a virtual avatar. Therefore, the virtual avatar generation method of this embodiment can also be executed on common terminal devices, such as mobile phones and computers. Figure 3 This is a flowchart illustrating a method for generating a virtual avatar according to this application. Figure 3 As shown, the method may include the following steps:

[0052] Step S302: In response to the input command applied to the operation interface, determine the first image corresponding to the input command.

[0053] The first image contains an image of a real object in the target pose.

[0054] The aforementioned operation interface refers to an interface through which users can input operation commands according to their needs. The aforementioned input commands can be commands for inputting the first image. The aforementioned first image can be a reference image used to generate the corresponding virtual avatar. The first image can contain at least one object from which the virtual avatar needs to be generated, i.e., the aforementioned real object. The real object can include, but is not limited to, objects existing in real life such as people, buildings, and vehicles. Considering that when people use video generation technology for entertainment, the generated virtual avatars are usually of different people, this application mainly uses people as the aforementioned real object for description. The aforementioned target pose can be the pose currently presented by the real object in the first image. For example, if the first image contains a person standing with their hands on their hips, then the aforementioned target pose can refer to standing with hands on hips. Generally, in order to better generate a virtual avatar that matches the real object, the aforementioned real object presented in the first image can be a front view of the real object.

[0055] In one optional embodiment, when a user input command is received on the operation interface, such as an instruction to upload an image triggered by the user in the interface provided by the video generation application, the virtual avatar generation system (hereinafter referred to as the generation system) can determine the first image corresponding to the input command, such as an image of a real object in the target pose selected by the user. Based on the first image, the generation system can determine the real object for which a virtual avatar needs to be generated.

[0056] Step S304: Generate a mesh model and multiple second images based on the first image.

[0057] Different second images are used to represent the attribute information of real objects from different perspectives.

[0058] In one optional embodiment, to generate a virtual avatar capable of diverse action changes without image distortion or poor perspective during these changes, after acquiring the first image, the generation system can generate a mesh model and multiple second images for constructing the virtual avatar, based on the real object contained in the first image. The generated mesh model can represent the shape, structure, surface details, and other information of the virtual avatar, while the multiple second images can represent the texture, color, and other information of the virtual avatar from different perspectives. Based on the generated mesh model and multiple second images, the generation system can generate a corresponding virtual avatar. To ensure the accuracy of the generated virtual avatar, when generating the mesh model and multiple second images, the generation system can first identify the first image, determining the target pose, surface texture, and other features of the real object. Then, based on the target pose of the real object, a mesh model in three-dimensional space is generated using mesh generation techniques, such as monocular depth estimation, pre-trained generative adversarial networks, and variational autoencoders. Simultaneously, based on the surface texture of the real object, the attribute information corresponding to the virtual avatar at the perspective corresponding to the first image is determined. Considering that the perspective of the real object in the first image is fixed and there is usually only one perspective, this means that the attribute information of the real object determined from the first image usually corresponds to only one perspective. Directly using the attribute information under this perspective to render the mesh model may result in an incomplete virtual image, which seriously affects the visual experience of the virtual image. Therefore, in order to improve the quality of the generated virtual image, after extracting the attribute information corresponding to the virtual image under the perspective of the first image, the generation system can further deduce the attribute information corresponding to the virtual image under other perspectives based on this attribute information, so as to generate the corresponding texture image based on the attribute information under different perspectives, i.e., the second image mentioned above.

[0059] In deriving attribute information from other perspectives, the generation system can first identify the attribute information from the current perspective to determine the object type of the object containing that attribute information, such as the style and type of clothing worn by a person. Then, based on the object type, it determines the overall attribute information of the object. Finally, it uses the overall attribute information to determine the attribute information from other perspectives, ensuring the accuracy of the determined attribute information. Alternatively, the generation system can directly generate attribute information from other perspectives through texture synthesis, mapping, or other methods based on the style, texture change trends, and other characteristics of the attribute information from the current perspective, thereby improving the efficiency of determining attribute information from other perspectives. It should be noted that the process of deriving attribute information from other perspectives mentioned here is merely an illustrative example and is not intended to be limiting.

[0060] Step S306: Render the mesh model based on multiple second images to generate a virtual image corresponding to the real object.

[0061] In one optional embodiment, after generating the mesh model and multiple second images, the generation system can use the multiple second images to render the mesh model from different perspectives to generate a virtual image corresponding to the real object. During model rendering, it is sufficient to ensure that the perspective presented by the mesh model matches the perspective corresponding to the second image. The specific rendering process can be found in related technologies and is not limited here.

[0062] Step S308: Display the virtual avatar on the operation interface.

[0063] After generating the virtual avatar, the generation system can then display the virtual avatar to the user on the user interface for easy viewing.

[0064] In this embodiment, a first image corresponding to an input command is determined in response to an input command applied to the operation interface; a mesh model and multiple second images are generated based on the first image; the mesh model is rendered based on the multiple second images to generate a virtual image corresponding to the real object; and the virtual image is displayed on the operation interface. By generating a mesh model matching the real object and multiple attribute information from different perspectives (i.e., second images) based on a single first image containing the real object, and using the second images to render the mesh model to obtain a three-dimensional virtual image, it is possible to ensure that the virtual image has high visibility from different perspectives and will not be damaged when performing large-scale actions, thereby improving the quality of the generated virtual image and solving the technical problem of poor quality of virtual images generated from a single image in related technologies.

[0065] In this embodiment of the application, generating a mesh model based on a first image includes: using a pose transformation model to transform a real object in the first image from a target pose to a preset pose to obtain a target image, wherein the pose transformation model is a machine learning model; and generating a mesh model based on the target image using an object generation model, wherein the accuracy of the mesh model is less than the preset accuracy, and the object generation model is a machine learning model.

[0066] The aforementioned preset pose can refer to a standard pose commonly used when utilizing a pose conversion model, such as a T-pose. The accuracy of the mesh model generated under the preset pose, such as the matching degree between the shape of the mesh model and the real object in the first image, is usually greater than the accuracy of the mesh model generated under other poses. The aforementioned object generation model can be a general generative model trained on general object data, which is a type of machine learning model. The aforementioned mesh model accuracy can refer to the number of meshes presented on the model. Generally, the more meshes on the mesh model, the higher the accuracy of the mesh model. The aforementioned preset accuracy can refer to the accuracy used to measure the generation efficiency of the mesh model. Generally, the efficiency of generating mesh models with different accuracies is usually different. When users use virtual avatars for entertainment, the generation system usually needs to control the generation time of the virtual avatar to avoid negative emotional impact on the user. Therefore, the preset accuracy can be the accuracy set in the object generation model for quickly generating mesh models. Since the aforementioned general object data does not have a specific data direction, the accuracy of the mesh model generated by the object generation model is usually low, for example, lower than the aforementioned preset accuracy, and the corresponding efficiency of generating the mesh model is higher.

[0067] In one optional embodiment, to ensure the accuracy of the generated mesh model, before generating the mesh model, the generation system can first identify the target pose of the real object in the first image to determine whether the target pose of the real object is a preset pose. For example, the target pose can be matched with the preset pose, and the matching degree can be used to determine whether the two are the same. If the target pose of the real object is the preset pose, the generation system can directly use the pre-trained object generation model to generate the mesh model corresponding to the real object based on the first image. If the target pose of the real object is not the preset pose, the generation system can first use the pre-trained pose transformation model to transform the target pose of the real object in the first image into the preset pose to obtain the transformed target image. Then, the object generation model is used to generate the mesh model corresponding to the real object based on the target image, thereby ensuring the accuracy of the generated mesh model.

[0068] In one optional embodiment, in order to improve the efficiency of generating mesh models, the pose conversion model and object generation model used above can both be machine learning models. The accuracy of the mesh model generated by the machine learning model can also be set to be less than the preset accuracy mentioned above, so as to reduce the generation time of surface details of the mesh model and thus improve the generation efficiency. When generating virtual images using the mesh model in the future, the accuracy of the mesh model can be further optimized, for example, by increasing the number of meshes in the mesh model, to ensure the viewpoint effect of the generated virtual image.

[0069] In this embodiment, the object generation model includes a feature extraction module, a sequence modeling module, and a decoding module. The process of generating a mesh model based on a target image using the object generation model includes: using the feature extraction module to extract features from the target image to obtain semantic and pixel information; using the sequence modeling module to perform attention processing on the semantic information, pixel information, and time steps to obtain preset spatial features, wherein the time steps are used to identify the steps for sampling the semantic and pixel information; and using the decoding module to perform model recovery from the preset spatial features to generate the mesh model.

[0070] In this embodiment, the preset spatial features are used to characterize the features obtained by projecting the mesh model onto multiple planes, and can be a representation of Triplanes.

[0071] In one optional embodiment of this application, in order to improve the accuracy of the generated mesh model, at least the aforementioned feature extraction module, sequence modeling module, and decoding module can be set in the object generation model. Correspondingly, when the object generation model is used to process the target image to generate a mesh model, the generation system can first use the aforementioned feature extraction module to extract features from the target image to obtain semantic information and pixel information of the target image, such as spatial features, depth features, shape features, edge features, category, attributes, etc. Then, the semantic information and pixel information used to identify the extracted semantic information and pixel information, as well as the step of sampling these two types of information in the feature extraction module, i.e., the aforementioned time step, are input into the aforementioned sequence modeling module for attention processing to obtain the corresponding preset spatial features, i.e., the features that can be obtained by projecting the mesh model onto multiple planes, thereby determining the spatial features that can be obtained when observing the mesh model from different perspectives. Finally, the decoding module is used to restore the model from the preset spatial features, and the mesh model corresponding to the real object in the first image can be generated.

[0072] In this embodiment of the application, the feature extraction module includes a semantic information extraction network and a pixel information extraction network. The feature extraction module performs feature extraction on the target image to obtain semantic information and pixel information of the target image, including: using the semantic information extraction network to extract semantic information from the target image to obtain semantic information; and using the pixel information extraction network to extract pixel information from the target image to obtain pixel information.

[0073] In one optional embodiment of this application, in order to improve the accuracy of the obtained semantic information and pixel information, at least the above-mentioned semantic information extraction network and pixel information extraction network can be configured in the feature extraction module. Correspondingly, the semantic information extraction network can be used to extract semantic information from the target image to obtain the semantic information of the target image, and the pixel information extraction network can be used to extract pixel information from the target image to obtain the pixel information of the target image.

[0074] In this embodiment of the application, the decoding module includes a decoder and a geometric mapping network. The decoding module performs model recovery on preset spatial features to generate a mesh model, which includes: decoding the preset spatial features using the decoder to obtain the model features of the mesh model; and performing model recovery on the model features using the geometric mapping network to obtain the mesh model.

[0075] In one optional embodiment of this application, to ensure the accuracy of the decoded mesh model, at least the aforementioned decoder and geometric mapping network can be configured in the decoding module. Correspondingly, the decoder can first be used to decode the obtained preset spatial features to obtain the model features of the mesh model, such as size features, shape features, depth features, edge features, etc. Then, the geometric mapping network is used to recover the model based on the decoded model features to obtain the corresponding mesh model. Considering that the preset spatial features generated through semantic information, pixel information, and corresponding time steps are typically used to describe the two-dimensional features presented by the mesh model in a two-dimensional plane, the decoder used when decoding the preset spatial features can be a 2D convolutional decoder. The process of decoding the preset spatial features using the decoder can be seen as the process of recovering two-dimensional sparse point cloud data into three-dimensional dense point cloud data.

[0076] In this embodiment, the decoder is trained on the decoding module in the pre-trained 3D reconstruction module using the first training data. The first training data includes: a real mesh model, training point cloud data, and multiple preset features. The training point cloud data is used to represent the point cloud data obtained by sampling the real mesh model. The multiple preset features are used to represent the bias information of the real mesh model relative to the preset mesh model. The pre-trained 3D reconstruction module is trained based on the preset mesh model.

[0077] The aforementioned 3D reconstruction module can refer to a pre-trained 3D-VAE module used to reconstruct 3D models. It can be trained based on a preset mesh model. Considering that the types of real objects used to generate virtual images may differ, the 3D reconstruction module used here can be a reconstruction module matching the type of the real object in the first image. For example, if the real object is a person, the aforementioned preset mesh model can be a general human body model, and the corresponding 3D reconstruction module can be a module for 3D reconstruction of the human body. If the real object is a vehicle, the aforementioned preset mesh model can be a general vehicle model, and the corresponding 3D reconstruction module can be a module for 3D reconstruction of the vehicle. The 3D reconstruction module may include a decoding module, which is used to analyze the features of the mesh model. The aforementioned real mesh model can refer to a mesh model actively constructed for the training object. The aforementioned training point cloud data can refer to point cloud data obtained by sampling the real mesh model, which can be considered as dense point cloud data. The aforementioned preset mesh model can refer to a mesh model generated for the image corresponding to the training object, which can be used to train the 3D reconstruction model. The point cloud data sampled from the preset mesh model can be considered as sparse point cloud data. The aforementioned preset features can refer to the offset information between the real mesh model and the preset mesh model, such as the spatial relationship, geometric features, topological relationship, density information, scale information, etc. between the corresponding point cloud data of the two.

[0078] In one optional embodiment, to improve the accuracy of the decoder, the generation system can actively construct a corresponding real mesh model for a preset object and sample the real mesh model to obtain training point cloud data corresponding to the real mesh model. At the same time, the real mesh model is compared with a pre-constructed preset mesh model to obtain multiple preset features corresponding to the real mesh model, i.e., the bias information between the two. Then, the real mesh model, training point cloud data and multiple preset features are used as the first training data to train the decoding module in the 3D reconstruction module, thereby obtaining a decoder with higher accuracy.

[0079] In this embodiment, the 3D reconstruction module includes a decoding module and an encoding module. The encoding module is used to encode the point cloud data based on multiple preset features to obtain the training space features corresponding to the point cloud data. The decoding module is used to decode the training space features to obtain the training model features of the real mesh model. The network parameters of the decoding module are adjusted based on the loss function between the real mesh model and the recovered mesh model. The recovered mesh model is obtained by using a geometric mapping network to recover the training model features.

[0080] In one optional embodiment, to ensure the accuracy of the trained 3D reconstruction module, an encoding module corresponding to the decoding module can be configured in the 3D reconstruction module. When training the 3D reconstruction module, the encoding module can first encode the point cloud data based on multiple preset features to obtain the training space features corresponding to the point cloud data. Then, the decoding module can decode the obtained training space features to obtain the training model features of the real mesh model. Based on the obtained training model features, the decoding module can use a geometric mapping model to recover the model to obtain the recovered mesh model corresponding to the current training model features. Finally, the recovered mesh model can be matched with the corresponding real mesh model to construct a loss function between the two. The network parameters of the decoding module can be adjusted using this loss function to improve the accuracy of the decoding module, thereby improving the accuracy of the 3D reconstruction module.

[0081] To facilitate understanding of the above process of generating the mesh model, Figure 4 This is a schematic diagram illustrating a mesh model generation process according to an embodiment of this application, such as... Figure 4 As shown, 401 represents multiple preset features, 402 represents training point cloud data, 403 represents the first image, 404 represents the feature generation module, 405 represents multiple generated spatial features, 406 represents the decoder, 407 represents the geometric mapping network, 408 represents the mesh model, 409 represents the semantic information extraction network, 410 represents the pixel information extraction network, 411 represents semantic information, 412 represents pixel information, 413 represents the time step, and 414 represents the sequence modeling module DiT.

[0082] like Figure 4 As shown, when generating the mesh model, multiple spatial features can be generated in the feature generation module using the training point cloud data corresponding to the real object model and multiple preset features. These spatial features are then used to train the decoding module in the 3D reconstruction module to obtain a decoder and geometric mapping network with high accuracy. Then, when generating the mesh model, semantic information extraction network and pixel information extraction network are used to extract semantic information and pixel information from the first image and determine the time step when extracting these two pieces of information. The sequence modeling module is then used to fuse the semantic information, pixel information, and time step to obtain preset spatial features. The preset spatial features are then input into the trained decoder for decoding to obtain the model features of the mesh model. Finally, the geometric mapping network is used to restore the model features to obtain the mesh model corresponding to the real object in the first image.

[0083] In this embodiment of the application, generating multiple second images based on a first image includes: inputting the first image into an image generation model, using the image generation model to perform texture inference on different perspectives of the mesh model, and obtaining multiple second images.

[0084] In one alternative embodiment, in order to improve the accuracy of the derived second image, a pre-trained image generation model can be used to infer the texture that the mesh model may present under different viewpoints based on the first image, thereby obtaining multiple second images that match the first image and the mesh model.

[0085] In this embodiment, the image generation model includes a feature extraction module, a semantic extraction module, a feature processing module, and a decoder. The image generation model performs texture reasoning on different perspectives of a mesh model to obtain multiple second images. This includes: using the feature extraction module to extract features from a first image to obtain first image features; using the semantic extraction module to extract features from the first image to obtain semantic information of the first image; using the feature processing module to fuse the first image features and semantic information according to different perspectives to obtain second image features of multiple second images, wherein multiple noisy images correspond one-to-one with multiple second images; and using the decoder to decode the second image features to obtain multiple second images.

[0086] In one optional embodiment, to improve the accuracy of the generated second image, at least the aforementioned feature extraction module, semantic extraction module, feature processing module, and decoder can be configured in the image generation module. Correspondingly, when generating multiple second images, the feature extraction module can first extract features from the acquired first image to obtain the first image features. Simultaneously, the semantic extraction module can extract features from the first image to obtain its semantic information. Then, the feature processing module can fuse the extracted first image features and semantic information to obtain the image features of multiple second images. To ensure the rationality of generating second images based on the obtained second image features, different perspectives of the aforementioned mesh model can be introduced during the fusion of the first image and semantic information, thereby generating second image features of multiple second images from different perspectives. After generating the second image features of multiple second images, the generation system can further decode the second images using the aforementioned decoder to obtain the aforementioned multiple second images.

[0087] In this embodiment of the application, the image generation model further includes: an attitude encoder, wherein the method further includes: determining multiple normal images based on a mesh model, wherein different normal images are used to characterize the structural information of the mesh model under different viewpoints; encoding the multiple normal images using the attitude encoder to obtain third image features of the multiple normal images; fusing the first image features and semantic information according to different viewpoints using a feature processing module to obtain second image features of multiple second images, including: fusing the first image features, semantic information, multiple noise images and third image features using a feature processing module to obtain second image features, wherein the second image features are aligned with the third image features, and the multiple noise images correspond one-to-one with the multiple second images.

[0088] In one optional embodiment, to improve the consistency of the generated second image with the normals observed from different viewpoints when viewing the mesh model, the aforementioned attitude encoder can be configured in the image generation model. Before generating the second image features of the second image, the generation system can first identify the structural information of the generated mesh model from different viewpoints to obtain multiple normal images corresponding to multiple viewpoints. Then, the attitude encoder is used to encode these multiple normal images to obtain the third image features of these multiple normal images. Correspondingly, when generating the second image features, the third image features can also be introduced as a reference into the fusion process of the first image features, semantic information, and multiple noise images corresponding to the multiple second images to obtain the aforementioned second image features. Thus, while ensuring the robustness of the obtained second image features, the second image features can be aligned with the third image features. For example, it is ensured that the normal information contained in the second image features is the same as the normal information contained in the third image features. This ensures that the viewpoint presented by the second image generated based on the second image features is the same as the viewpoint presented by the aforementioned normal images, thereby ensuring that the second image decoded based on the second image features can fit the generated mesh model.

[0089] In this embodiment, the image generation model is obtained by fine-tuning a pre-trained generation model using training images, and the pre-trained generation model is trained using 3D model data.

[0090] The aforementioned pre-trained generative model can refer to a general image generation model. Using this model, the second image corresponding to different types of objects can be determined. The aforementioned training image can refer to an image corresponding to an object of the same type as the real object in the first image. After fine-tuning the pre-trained generative model using the training image, the fine-tuned model can be made to be more focused on generating the second image corresponding to the first image.

[0091] In one optional embodiment, the pre-trained generative model can be fine-tuned using training images to obtain an image generation model with higher accuracy. The corresponding pre-trained generative model can be trained using 3D model data. The specific training process can be referred to in related technologies and is not limited here.

[0092] To facilitate understanding the process of generating second images from different perspectives Figure 5 This is a schematic diagram illustrating a second image generation process according to an embodiment of this application. In this diagram, 401 represents the first image, 501 represents the feature extraction module, 502 represents the semantic extraction module, 503 represents the feature processing module, 504 represents the decoder, 505 represents the normal images corresponding to the mesh model from different viewpoints, 506 represents the attitude encoder, 507 represents the noisy image, and 508 represents the second image. When generating the second image, the feature extraction module and the semantic extraction module are first used to extract features from the first image to obtain first image features and semantic information. Simultaneously, the attitude encoder is used to encode the normal images corresponding to the mesh model from different viewpoints to obtain third image features of multiple normal images. Then, the feature processing module is used to fuse the first image features, semantic information, third image features, and multiple noisy images to obtain the corresponding second image features. Finally, the decoder is used to decode the second image features to obtain multiple second images from different viewpoints.

[0093] In this embodiment of the application, the above method further includes: constructing a Gaussian sputtering model based on a mesh model; optimizing the Gaussian sputtering model based on multiple second images and a preset sphere radius to obtain a virtual image.

[0094] The aforementioned Gaussian sputtering model can refer to a model constructed based on a mesh model using Gaussian Splatting technology.

[0095] In one optional embodiment, in order to improve the quality of the virtual image obtained by rendering the mesh model using the second image, the generation system can first use Gaussian Splatting technology to reconstruct the surface of the mesh model before rendering to obtain a Gaussian sputtering model with higher accuracy, and then use multiple second images to render the Gaussian sputtering model. At the same time, the sphere radius of the Gaussian sputtering model is constrained by a preset sphere radius, thereby obtaining a virtual image with higher quality.

[0096] In this embodiment of the application, the above method further includes: responding to an interactive command applied to the operation interface, displaying a target video containing a virtual image on the operation interface, wherein the target video is obtained by driving the virtual image based on the offset data of the surface vertices of the virtual image in the voxel space, the offset data is obtained based on the target coefficient and the bone skinning weight of the real three-dimensional model, and the target coefficient is used to characterize the linear bone skinning coefficient of the virtual image in the voxel space.

[0097] The aforementioned target video can refer to a video generated by controlling the virtual avatar's movements according to a user-provided action template. The aforementioned real 3D model can refer to a 3D model corresponding to the object used as a reference in the video, which belongs to the same type as the real object.

[0098] In one optional embodiment, if an interactive command is received from the user on the operation interface, the generation system can also control the virtual avatar's actions according to the action template selected by the user to generate a template video containing the virtual avatar, and display the template video to the user on the operation interface for easy viewing. Correspondingly, in order to ensure that the virtual avatar can stably run according to the action template, it is common practice to first use Rigging technology to skeletally bind the virtual avatar after its generation, and then have the user or application device move the bound skeletal points to generate a series of actions. However, this method is relatively complex. Therefore, in order to reduce the difficulty of generating the target video and improve the efficiency of generating the target video, the generation system can directly control the virtual avatar's actions based on the offset data of the virtual avatar's surface vertices in voxel space, such as the displacement distance and curve of the surface vertices, to generate the corresponding target video. The offset data used to drive the virtual avatar's actions can be determined by the linear skeleton skinning coefficient (LBS coefficient) of the virtual avatar in voxel space, i.e., the aforementioned target coefficient, and the skeleton skinning weight of the real 3D model.

[0099] In this embodiment of the application, the above method further includes: constructing a voxel space based on a preset resolution and a real 3D model, wherein the voxel space surrounds the surface vertices of the real 3D model; determining the distance between any voxel in the voxel space and the surface vertex of the real 3D model; and determining the bone skinning weight of any voxel based on the distance and the bone skinning weight corresponding to the surface vertex, thereby obtaining the bone skinning weight of the real 3D model.

[0100] In one optional embodiment, to improve the accuracy of the determined offset data, the generation system can first construct a voxel space surrounding the surface vertices of the real 3D model based on a preset resolution and the real 3D model. Then, it determines the distance between any voxel in the voxel space and a surface vertex of the real 3D model, such as the Euclidean distance. Finally, based on this distance and the bone skinning weights corresponding to different surface vertices, it determines the bone skinning weight of any voxel, thereby obtaining the bone skinning weight of the real 3D model. Wherein, if the surface vertex is p and the voxel center point is c, the corresponding Euclidean distance can be expressed as d(c, p). The bone skinning weights of different voxels can be determined using the k-nearest neighbor algorithm, and the corresponding formula for determining the bone skinning weight can be expressed as:

[0101]

[0102] Among them, w p This represents the bone skinning weight corresponding to the surface vertex p, and C represents a constant.

[0103] Figure 6 This is a schematic diagram of a skeleton binding result according to an embodiment of this application, wherein 601 represents the voxel space surrounding the real three-dimensional model, 602 represents a voxel, 603 represents a surface vertex, and each voxel can take the 5 surface vertices closest to that voxel.

[0104] To facilitate understanding of the process of generating the target video, Figure 7 This is a schematic diagram illustrating a virtual avatar video generation process according to an embodiment of this application, wherein 401 represents the first image, 701 represents the pose encoder, 702 represents the target image, 703 represents the object generation model, 408 represents the mesh model, 704 represents the image generation model, 705 represents the virtual avatar, 706 represents the skeleton binding module, and 707 represents the target video. Figure 7 As shown, when generating the target video, the generation system can first adjust the pose of the real object in the first image to obtain the target image. Then, it can use the object generation model and the image generation model to construct the corresponding mesh model and virtual image respectively. The system can then use the skeleton binding module to bind the skeleton of the mesh model. Finally, the system can control the motion of the skeleton-bound virtual image to generate the corresponding target video.

[0105] According to an embodiment of this application, a method for generating a virtual avatar is also provided. Figure 8 This is another method for generating virtual images according to embodiments of this application, such as... Figure 8 As shown, the method may include the following steps:

[0106] Step S802: Obtain the first image. The first image contains an image of the real object in the target pose.

[0107] Step S804: Generate a mesh model and multiple second images based on the first image. Different second images are used to represent the attribute information of the real object from different viewpoints.

[0108] Step S806: Render the mesh model based on multiple second images to generate a virtual image corresponding to the real object.

[0109] In one optional embodiment, when generating a virtual avatar, the generation system can first acquire a first image containing the real object from which the virtual avatar needs to be generated, then generate a corresponding mesh model and multiple second images based on the first image, and finally render the mesh model using the second images to obtain the virtual avatar corresponding to the real object. For details, please refer to the preceding text; further elaboration is not provided here.

[0110] According to an embodiment of this application, a method for generating a virtual avatar is also provided. Figure 9 This is another method for generating virtual images according to embodiments of this application, such as... Figure 9 As shown, the method may include the following steps:

[0111] Step S902: Obtain the first image by calling the first interface. The first interface includes a first parameter, the value of which includes the first image, which contains an image of a real object in the target pose.

[0112] Step S904: Generate a mesh model and multiple second images based on the first image. The different second images are used to represent the attribute information of the real object from different viewpoints.

[0113] Step S906: Render the mesh model based on multiple second images to generate a virtual image corresponding to the real object.

[0114] Step S908: Output the virtual avatar by calling the second interface. The second interface includes a second parameter, the value of which includes the virtual avatar.

[0115] In one optional embodiment, when generating a virtual avatar, the generation system can first obtain the parameter value of the first parameter by calling the first interface, i.e., obtain a first image containing a real object in the target pose. Then, based on the first image, it generates a mesh model of the real object and multiple second images, and uses the multiple second images to render the mesh model, thereby obtaining a virtual avatar of the real object. Finally, it can call the second interface to output the parameter value of the second parameter, i.e., output the virtual avatar, for user viewing. Specific details can be found in the preceding text and will not be repeated here.

[0116] It should be noted that the preferred embodiments involved in the above embodiments of this application are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.

[0117] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0118] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0120] According to an embodiment of this application, a virtual avatar generation apparatus for implementing the above-described virtual avatar generation method is also provided. Figure 10 This is a structural block diagram of a virtual image generation device according to an embodiment of this application, such as... Figure 10 As shown, the device includes: an image determination module 1002, an information generation module 1004, a model rendering module 1006, and an image display module 1008.

[0121] The image determination module 1002 is used to respond to input commands applied to the operation interface and determine the first image corresponding to the input command, wherein the first image contains an image of the real object in the target pose; the information generation module 1004 is used to generate a mesh model and multiple second images based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; the model rendering module 1006 is used to render the mesh model based on the multiple second images to generate a virtual image corresponding to the real object; and the image display module 1008 is used to display the virtual image on the operation interface.

[0122] It should be noted that the image determination module 1002, information generation module 1004, model rendering module 1006, and image display module 1008 mentioned above correspond to steps S302 to S308 in the above embodiments. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of a device and run in the AR / VR device provided in the above embodiments.

[0123] According to an embodiment of this application, a virtual avatar generation apparatus for implementing the above-described virtual avatar generation method is also provided. Figure 11 This is a structural block diagram of another virtual image generation device according to an embodiment of this application, such as... Figure 11 As shown, the device includes: an image acquisition module 1102, a first generation module 1104, and a second generation module 1106.

[0124] The image acquisition module 1102 is used to acquire a first image, which contains an image of a real object in a target pose; the first generation module 1104 is used to generate a mesh model and multiple second images based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; the second generation module 1106 is used to render the mesh model based on the multiple second images to generate a virtual image corresponding to the real object.

[0125] It should be noted that the image acquisition module 1102, the first generation module 1104, and the second generation module 1106 mentioned above correspond to steps S802 to S806 in the above embodiments. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of a device and run in the AR / VR device provided in the above embodiments.

[0126] According to an embodiment of this application, a virtual avatar generation apparatus for implementing the above-described virtual avatar generation method is also provided. Figure 12 This is a structural block diagram of another virtual image generation device according to an embodiment of this application, such as... Figure 12 As shown, the device includes: an image retrieval module 1202, a third generation module 1204, a fourth generation module 1206, and an image output module 1208.

[0127] The image retrieval module 1202 is used to obtain a first image by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the first image, and the first image contains an image of a real object in a target pose; the third generation module 1204 is used to generate a mesh model and multiple second images based on the first image, wherein different second images are used to represent the attribute information of the real object from different perspectives; the fourth generation module 1206 is used to render the mesh model based on multiple second images to generate a virtual image corresponding to the real object; and the image output module 1208 is used to output the virtual image by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the virtual image.

[0128] It should be noted that the image retrieval module 1202, the third generation module 1204, the fourth generation module 1206, and the image output module 1208 correspond to steps S902 to S908 in the above embodiments. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should also be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. These modules can also run as part of a device in the AR / VR device provided in the above embodiments.

[0129] It should be noted that the preferred embodiments involved in the above embodiments of this application are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.

[0130] Embodiments of this application may provide a computer terminal, which may include a server and a client. The server may be any one of the servers in a server device group or a cloud server.

[0131] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0132] In this embodiment, the computer terminal described above can execute the program code in the method.

[0133] Optionally, Figure 13 This is a structural block diagram of a computer terminal according to an embodiment of this application. As shown in the figure, the computer terminal A may include: one or more (only one is shown in the figure) processors 1302, memory 1304, memory controller, and peripheral interfaces, wherein the peripheral interfaces are connected to a radio frequency module, an audio module, and a display.

[0134] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to computer terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0135] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: responding to an input command applied to the operation interface, determining a first image corresponding to the input command, wherein the first image contains an image of a real object in a target pose; responding to a generation command applied to the operation interface, displaying a virtual image corresponding to the real object on the operation interface, wherein the virtual image is generated by rendering a mesh model based on multiple second images, the mesh model and multiple second images are generated based on the first image, and different second images are used to characterize the attribute information of the real object from different perspectives.

[0136] Those skilled in the art will understand that Figure 13 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, PDA, mobile Internet device (MID), PAD, and other terminal devices. Figure 13 This does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include components that are more complex than those described above. Figure 13 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 13 The different configurations shown.

[0137] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0138] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.

[0139] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a computer terminal cluster, or in any mobile terminal in a mobile terminal cluster.

[0140] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: in response to an input command applied to the operation interface, determining a first image corresponding to the input command, wherein the first image contains an image of a real object in a target pose; in response to a generation command applied to the operation interface, displaying a virtual image corresponding to the real object on the operation interface, wherein the virtual image is generated by rendering a mesh model based on multiple second images, the mesh model and the multiple second images are generated based on the first image, and different second images are used to characterize the attribute information of the real object from different perspectives.

[0141] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.

[0142] Embodiments of this application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the method provided in the above embodiments.

[0143] Embodiments of this application also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.

[0144] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0145] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0147] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0148] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0149] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating a virtual avatar, characterized in that, include: In response to an input command applied to the user interface, a first image corresponding to the input command is determined, wherein the first image contains an image of a real object in a target pose; A mesh model and multiple second images are generated based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; The mesh model is rendered based on the multiple second images to generate a virtual image corresponding to the real object; The virtual image is displayed on the user interface.

2. The method according to claim 1, characterized in that, The generation of the mesh model based on the first image includes: The pose transformation model is used to transform the real object in the first image from the target pose to a preset pose to obtain the target image, wherein the pose transformation model is a machine learning model; Based on the target image, the mesh model is generated using an object generation model, wherein the accuracy of the mesh model is less than a preset accuracy, and the object generation model is a machine learning model.

3. The method according to claim 2, characterized in that, The object generation model includes a feature extraction module, a sequence modeling module, and a decoding module, wherein generating the mesh model based on the target image using the object generation model includes: The feature extraction module is used to extract features from the target image to obtain the semantic information and pixel information of the target image; The sequence modeling module is used to perform attention processing on the semantic information, the pixel information, and the time step to obtain preset spatial features, wherein the time step is used to identify the steps of sampling the semantic information and the pixel information; The decoding module is used to recover the model of the preset spatial features to generate the mesh model.

4. The method according to claim 3, characterized in that, The feature extraction module includes a semantic information extraction network and a pixel information extraction network. The step of using the feature extraction module to extract features from the target image to obtain the semantic information and pixel information of the target image includes: The semantic information is extracted from the target image using the semantic information extraction network to obtain the semantic information; The pixel information is obtained by extracting pixel information from the target image using the pixel information extraction network.

5. The method according to claim 3, characterized in that, The decoding module includes a decoder and a geometric mapping network, wherein the step of using the decoding module to perform model recovery on the preset spatial features and generate the mesh model includes: The preset spatial features are decoded using the decoder to obtain the model features of the mesh model; The model features are recovered using the geometric mapping network to obtain the mesh model.

6. The method according to claim 5, characterized in that, The decoder is obtained by training the decoding module in the pre-trained 3D reconstruction module using the first training data. The first training data includes: a real mesh model, training point cloud data, and multiple preset features. The training point cloud data is used to represent the point cloud data obtained by sampling the real mesh model. The multiple preset features are used to represent the bias information of the real mesh model relative to the preset mesh model. The pre-trained 3D reconstruction module is trained based on the preset mesh model.

7. The method according to claim 6, characterized in that, The 3D reconstruction module includes a decoding module and an encoding module. The encoding module is used to encode the point cloud data based on the multiple preset features to obtain the training space features corresponding to the point cloud data. The decoding module is used to decode the training space features to obtain the training model features of the real mesh model. The network parameters of the decoding module are adjusted based on the loss function between the real mesh model and the recovered mesh model. The recovered mesh model is obtained by recovering the training model features using the geometric mapping network.

8. The method according to claim 3, characterized in that, The preset spatial features are used to characterize the features obtained by projecting the mesh model onto multiple planes.

9. The method according to claim 1, characterized in that, The generation of multiple second images based on the first image includes: The first image is input into the image generation model, and the image generation model is used to perform texture inference on different perspectives of the mesh model to obtain the plurality of second images.

10. The method according to claim 9, characterized in that, The image generation model includes: a feature extraction module, a semantic extraction module, a feature processing module, and a decoder. The step of using the image generation model to perform texture inference on different perspectives of the mesh model to obtain the plurality of second images includes: The feature extraction module is used to extract features from the first image to obtain the first image features of the first image; The semantic extraction module is used to extract features from the first image to obtain the semantic information of the first image; The feature processing module fuses the first image features and the semantic information according to the different perspectives to obtain the second image features of the plurality of second images; The decoder is used to decode the features of the second image to obtain the plurality of second images.

11. The method according to claim 10, characterized in that, The image generation model further includes a pose encoder, wherein the method further includes: Based on the mesh model, multiple normal images are determined, wherein different normal images are used to characterize the structural information of the mesh model from different viewpoints; The attitude encoder is used to encode multiple normal images to obtain the third image features of the multiple normal images; The step of fusing the first image features and the semantic information according to the different perspectives using the feature processing module to obtain the second image features of the plurality of second images includes: fusing the first image features, the semantic information, the plurality of noisy images and the third image features using the feature processing module to obtain the second image features, wherein the second image features are aligned with the third image features, and the plurality of noisy images correspond one-to-one with the plurality of second images.

12. The method according to claim 8, characterized in that, The image generation model is obtained by fine-tuning a pre-trained generation model using training images, and the pre-trained generation model is obtained by training with 3D model data.

13. The method according to claim 1, characterized in that, The method further includes: Based on the aforementioned mesh model, a Gaussian sputtering model is constructed; The Gaussian sputtering model is optimized based on the multiple second images and the preset sphere radius to obtain the virtual image.

14. The method according to any one of claims 1 to 13, characterized in that, The method further includes: In response to an interactive command applied to the operation interface, a target video containing the virtual image is displayed on the operation interface. The target video is obtained by driving the virtual image based on the offset data of the surface vertices of the virtual image in voxel space. The offset data is obtained based on the target coefficients and the bone skinning weights of the real 3D model. The target coefficients are used to characterize the linear bone skinning coefficients of the virtual image in the voxel space.

15. The method according to claim 14, characterized in that, The method further includes: The voxel space is constructed based on a preset resolution and the real 3D model, wherein the voxel space surrounds the surface vertices of the real 3D model; The distance between any voxel in the voxel space and a surface vertex of the real 3D model is determined. Based on the distance and the bone skinning weights corresponding to the surface vertices, the bone skinning weights of any voxel are determined, thus obtaining the bone skinning weights of the real 3D model.

16. A method for generating a virtual avatar, characterized in that, include: Acquire a first image, wherein the first image contains an image of a real object in a target pose; A mesh model and multiple second images are generated based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; The mesh model is rendered based on the multiple second images to generate a virtual image corresponding to the real object.

17. A method for generating a virtual avatar, characterized in that, include: A first image is obtained by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the first image, and the first image contains an image of a real object in the target pose; A mesh model and multiple second images are generated based on the first image, wherein different second images are used to characterize the attribute information of the real object from different perspectives; The mesh model is rendered based on the multiple second images to generate a virtual image corresponding to the real object; The virtual image is output by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the virtual image.

18. A computer terminal, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 17.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 17.

20. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 17.