Virtual human generation method and device, computer equipment and storage medium

By processing a single original perspective image and/or text description, combining multi-perspective generation and three-dimensional Gaussian function reconstruction, dividing the Mesh and performing iterative optimization, the problem of balancing the speed and quality of virtual human generation is solved, and efficient and detailed virtual human model construction is achieved, which is suitable for virtual human generation and animation production in the financial field.

CN120655859APending Publication Date: 2025-09-16PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510724268.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively balance the contradiction between 3D virtual human generation speed and generation quality. Especially in financial business scenarios, existing methods find it difficult to achieve fast and high-fidelity virtual human model construction.

Method used

By performing image processing on a single original perspective image and/or text description as input, a new perspective image is generated. A multi-perspective diffusion model and a three-dimensional Gaussian function are used to reconstruct a coarse model of the virtual human, which is divided into a mesh. Random perspective two-dimensional image rendering, noise addition, and stable diffusion processing are then performed, and the mesh is iteratively optimized to generate a fine model.

Benefits of technology

It achieves efficient and detailed construction of virtual human models, improves the realism and detail expression of the model, and significantly improves rendering efficiency. It is suitable for virtual human generation and animation production in the financial field and has a certain degree of versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655859A_ABST
    Figure CN120655859A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI of financial scenes, and discloses a virtual human generation method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the image processing of an input single original view angle image and / or text description, and obtaining a new view angle image; performing multi-view-angle generation processing on the new-view-angle image to obtain a multi-view-angle image; performing reconstruction processing on the multi-view image through a three-dimensional Gaussian function to obtain a virtual human rough model; extracting basic structure data of the virtual human rough model, and dividing the basic structure data into Mesh grids; performing random view angle two-dimensional image rendering on the Mesh grid to obtain a corresponding rendered image, and performing noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image; and iteratively calculating the loss of the rendered image and the corresponding optimized image so as to optimize the Mesh grid and obtain a virtual human fine model. The method has the advantages that the virtual human model generation process is simplified, and the high-fidelity model generation effect is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI technology in financial scenarios, and in particular to a method, apparatus, computer equipment, and storage medium for generating a virtual human. Background Art

[0002] As AI-generated content (AIGC) technology continues to mature, extracting sufficient information from a single viewpoint to enable rapid and high-fidelity 3D virtual human modeling has become an extremely challenging task. The generated 3D virtual humans can be redesigned, accelerating innovation in fields such as finance. Currently, a two-stage approach to generating 3D virtual humans has gained significant traction. The first stage generates multi-view images using a diffusion model and constructs a rough 3D model using SDS technology. The second stage then refines this rough model.

[0003] In some financial business scenarios, to achieve diversified marketing scenario design, more and more institutions are turning to AI technology to streamline the 3D virtual human generation process. However, current optimization-based methods struggle to effectively balance the speed and quality of 3D virtual human generation. Summary of the Invention

[0004] The present invention provides a virtual human generation method, apparatus, computer equipment and storage medium to solve the problem of difficulty in effectively balancing the generation speed and generation quality of 3D virtual humans.

[0005] In a first aspect, a method for generating a virtual human is provided, comprising:

[0006] Perform image processing on the input single original perspective image and / or text description to obtain a new perspective image;

[0007] Performing multi-perspective generation processing on the new-perspective image to obtain a multi-perspective image;

[0008] Reconstructing the multi-view images using a three-dimensional Gaussian function to obtain a rough model of a virtual human;

[0009] Extracting basic structural data of the virtual human coarse model and dividing it into Mesh grids;

[0010] Performing random perspective two-dimensional image rendering on the Mesh grid to obtain a corresponding rendered image, and performing noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image;

[0011] The loss of the rendered image and the corresponding optimized image is iteratively calculated to optimize the Mesh grid and obtain a virtual human model.

[0012] In a second aspect, a virtual human generation device is provided, comprising:

[0013] A first image generation module is used to perform image processing on a single input original perspective image and / or text description to obtain a new perspective image;

[0014] A second image generation module is used to perform multi-perspective generation processing on the new-perspective image to obtain a multi-perspective image;

[0015] A coarse generation module, configured to reconstruct the multi-view images using a three-dimensional Gaussian function to obtain a coarse model of a virtual human;

[0016] A model segmentation module is used to extract the basic structural data of the virtual human coarse model and divide it into Mesh grids;

[0017] An image optimization module is used to perform random perspective two-dimensional image rendering on the Mesh grid to obtain a corresponding rendered image, and to perform noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image;

[0018] A fine generation module is used to iteratively calculate the loss of the rendered image and the corresponding optimized image to optimize the Mesh grid and obtain a virtual human fine model.

[0019] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned virtual human generation method when executing the computer program.

[0020] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned virtual human generation method are implemented.

[0021] In the scheme implemented by the above-mentioned virtual human generation method, device, computer equipment and storage medium, a new perspective image is obtained by performing image processing on a single original perspective image and / or text description as input; a multi-perspective generation process is performed on the new perspective image to obtain a multi-perspective image; the multi-perspective image is reconstructed using a three-dimensional Gaussian function to obtain a coarse model of the virtual human; the basic structural data of the coarse model of the virtual human is extracted and divided into a mesh grid; the mesh grid is rendered with a random perspective two-dimensional image to obtain a corresponding rendered image, and the rendered image is subjected to noise and stable diffusion processing to obtain a corresponding optimized image; the loss of the rendered image and the corresponding optimized image is iteratively calculated to optimize the mesh grid and obtain a fine model of the virtual human. The present invention is aimed at virtual human model construction scenarios in the financial field. By combining technical means such as mesh grid division, random perspective two-dimensional image rendering, noise and stable diffusion processing, and iterative loss optimization, it successfully implements an efficient and precise virtual human model construction method. This method effectively optimizes the virtual human model generation process and maintains a high-fidelity model generation effect. This method not only enhances the realism and detail of virtual human models but also significantly improves rendering efficiency, providing strong technical support for virtual human generation, animation, and real-time rendering in applications such as finance. Furthermore, the solution is versatile and can be extended to the construction of virtual human models in other fields, opening up new avenues for the development and application of virtual human technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0023] Figure 1 This is a schematic diagram of an application environment of a method for generating a virtual human in one embodiment of the present invention;

[0024] Figure 2 This is a flow chart of a method for generating a virtual human according to an embodiment of the present invention;

[0025] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S201;

[0026] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S202;

[0027] Figure 5 yes Figure 2A schematic flow chart of a specific implementation of step S203;

[0028] Figure 6 yes Figure 2 A schematic flow chart of a specific implementation of step S204;

[0029] Figure 7 yes Figure 2 A schematic flow chart of a specific implementation of step S205;

[0030] Figure 8 yes Figure 2 A schematic flow chart of a specific implementation of step S206;

[0031] Figure 9 is a structural diagram of a virtual human generation device in one embodiment of the present invention;

[0032] Figure 10 is a structural diagram of a computer device in one embodiment of the present invention;

[0033] Figure 11 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.

[0035] The virtual human generation method provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the client communicates with the server through a network. The server can receive a single original perspective image and / or text description input by the user through the client to perform image processing to obtain a new perspective image; perform multi-perspective generation processing on the new perspective image to obtain a multi-perspective image; reconstruct the multi-perspective image through a three-dimensional Gaussian function to obtain a coarse model of a virtual human; extract the basic structural data of the coarse model of the virtual human and divide it into a Mesh grid; perform random perspective two-dimensional image rendering on the Mesh grid to obtain a corresponding rendered image, and perform noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image; iteratively calculate the loss of the rendered image and the corresponding optimized image to optimize the Mesh grid and obtain a fine model of the virtual human. In the present invention, the generation process of the virtual human model is effectively optimized, and a high-fidelity model generation effect is maintained. It also effectively improves the customization demand for three-dimensional virtual humans in the financial business line to meet specific marketing scenarios or metaverse asset generation. In addition, the method of the present invention uses only a single image as input to automatically complete the subsequent generation of the virtual human model, and both professionals and non-professionals can quickly get started. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0036] See also Figure 2 As shown, Figure 2 A flowchart of a method for generating a virtual human provided by an embodiment of the present invention includes the following steps:

[0037] S201, performing image processing on a single original perspective image and / or text description input to obtain a new perspective image;

[0038] In step S201, to ensure the quality of the new perspective image, a single original perspective image and its corresponding text description can be input as a data pair. For example, in a scene with a virtual person representing a financial mascot, a frontal photo of the financial mascot and text describing its characteristics, such as "wearing a red Tang suit, holding a gold ingot, and smiling," can be input as a data pair.

[0039] like Figure 3 As shown, in one embodiment, step S201 includes:

[0040] S301, receiving an original viewing angle image and a text description including target viewing angle parameters input by a user;

[0041] S302, performing multimodal feature encoding and fusion on the original view image and text description to generate a conditional control signal, wherein the view parameters in the text description are parsed into spatial transformation parameters;

[0042] S303, inputting the conditional control signal and the spatial transformation descriptor into a pre-trained single-view generative diffusion model, generating a latent feature representation of the target view through an iterative denoising process of the single-view generative diffusion model, and achieving view consistency control by dynamically adjusting the spatial transformation parameters during the iterative denoising process;

[0043] S304: Decode the latent feature representation and output a new perspective image that is consistent with the input image scene semantics and satisfies the target perspective constraint.

[0044] Steps S301-S304 can be implemented using a deep learning framework. First, through multimodal feature encoding and fusion, the information from the image and text description is effectively combined, improving the accuracy and reliability of the perspective conversion. Second, using a pre-trained single-perspective generative diffusion model, an iterative denoising process is used to generate a latent feature representation of the target perspective, ensuring the high quality and richness of detail in the new perspective image. Finally, perspective consistency control is achieved by dynamically adjusting the spatial transformation parameters, ensuring the semantic consistency of the new perspective image with the input image scene while satisfying the constraints of the target perspective.

[0045] In some preferred embodiments, the single-view generation diffusion model may adopt the open source Karlo diffusion model. The Karlo diffusion model can efficiently and accurately complete the image perspective conversion task, greatly enhancing the diversity and flexibility of the image perspective.

[0046] In some scenes of virtual people of financial mascots, users can convert the financial mascot from the original perspective image to a new perspective image of any target perspective, such as from a front perspective to a side perspective, or from a top-down perspective to an upward perspective. This flexible perspective conversion capability enriches the mascot's expression form.

[0047] S202, performing multi-perspective generation processing on the new perspective image to obtain a multi-perspective image;

[0048] In step S202, multi-view images refer to multiple images from different perspectives, providing a rich data foundation for the subsequent construction of the virtual human model. This multi-angle image data helps to fully capture the shape and surface structure of the financial mascot to be constructed, such as the example above, which is crucial for building a high-fidelity virtual human model.

[0049] like Figure 4 As shown, in one embodiment, step S202 includes:

[0050] S401, inputting the new view image into a pre-trained multi-view generative diffusion model, where the multi-view generative diffusion model includes a basic visual diffusion architecture and a parameter efficient adaptation layer, where the parameter efficient adaptation layer dynamically adjusts the model parameters through low-rank matrix decomposition;

[0051] S402, propagating the geometric features and texture features of the new view image across views through a multi-stage denoising process of a multi-view generative diffusion model, and iteratively generating a view-conditioned feature vector in a latent space;

[0052] S403: Decode the view-conditioned feature vector into a multi-view image containing spatial continuity.

[0053] In steps S401-S403, the multi-perspective generation diffusion model can make full use of the information in the new perspective image, and accurately capture and process the geometric features and texture features of the new perspective image through its internal neural network structure. Specifically, step S401 uses the basic visual diffusion architecture to perform preliminary processing on the new perspective image, and realizes dynamic adjustment of model parameters through the parameter efficient adaptation layer to adapt to changes in images from different perspectives. Step S402 is the denoising process. The model gradually eliminates the noise in the new perspective image through a multi-stage denoising strategy, and at the same time propagates the geometric features and texture features of the new perspective image across perspectives, so that these features can be effectively integrated and optimized in the latent space. Finally, in step S403, the model decodes the optimized perspective-conditioned feature vector into a multi-perspective image containing spatial continuity, thereby achieving the goal of generating high-quality images from different perspectives.

[0054] In steps S401 to S403, the multi-view diffusion model can be generated using the open-source Zero123 diffusion model, which has a built-in LoRA layer. Because the weights of the LoRA and Zero123 diffusion models can be directly superimposed across all network layers, there's no need for a separate integration mechanism. The output is a multi-view image from any perspective.

[0055] During the construction of virtual human models of some financial mascots as shown in the examples above, users can use this multi-perspective generation process to further convert the new perspective image of the financial mascot into multiple images with different perspectives. These multi-perspective images not only enrich the visual expression of the mascot, but also provide a solid foundation for the subsequent construction of more realistic and three-dimensional virtual human models. For example, in the virtual live broadcast or virtual interactive scene of the financial mascot, by showing images of the mascot from different perspectives, the audience's sense of immersion and participation can be enhanced, making the mascot's image more vivid and three-dimensional. In addition, these multi-perspective images can also be used for aspects such as motion capture and expression simulation of the mascot, further enhancing the interactive experience and realism of the virtual human.

[0056] S203, reconstructing the multi-view images using a three-dimensional Gaussian function to obtain a rough model of the virtual human;

[0057] In step S203, the geometric shape features of the virtual human model are accurately captured from the multi-view images through a three-dimensional Gaussian function, and a relatively rough virtual human model that basically conforms to the actual proportion is generated for subsequent refinement processing.

[0058] like Figure 5 As shown, in one embodiment, step S203 includes:

[0059] S501, obtaining a multi-view two-dimensional image set of multi-view images;

[0060] S502: Input the multi-view two-dimensional image set into the three-dimensional Gaussian function, perform gradient back propagation between the image space and the three-dimensional space through the differentiable rendering engine, iteratively optimize the Gaussian basis element attribute parameters to generate a three-dimensional editable virtual human rough model.

[0061] Steps S501-S502 enable the rapid generation of a 3D editable coarse virtual human model from multi-view images. This process not only improves model generation efficiency but also ensures the accuracy and authenticity of the generated virtual human model in 3D space. By utilizing a 3D Gaussian function and a differentiable rendering engine, key information from the multi-view images can be accurately captured and converted into Gaussian basis element attribute parameters in 3D space, thereby generating a high-quality coarse virtual human model.

[0062] Taking the construction of the aforementioned financial mascot as an example, a set of 2D images of the mascot from multiple perspectives was obtained. These 2D images were then fed into a 3D Gaussian function, and the gradient backpropagation between the image space and 3D space was performed using a differentiable rendering engine. During this process, the attribute parameters of the Gaussian basis elements, such as position, shape, and size, were continuously iteratively optimized until a preliminary 3D editable rough model of the financial mascot was generated. While this rough model was somewhat crude, it captured the main features and morphology of the financial mascot, laying a solid foundation for subsequent fine-tuning and optimization.

[0063] S204, extracting basic structural data of the virtual human coarse model and dividing it into Mesh grids;

[0064] In step S204, the purpose of dividing the avatar coarse model into meshes is to facilitate subsequent fine-tuning and optimization of the coarse avatar model. Meshes can be used to divide the coarse avatar model into multiple small regions, each of which can be independently deformed and adjusted, thereby achieving more detailed avatar form shaping. Furthermore, mesh division helps improve avatar rendering efficiency and realism, making the avatar appear more natural and lifelike in various scenarios.

[0065] like Figure 6 As shown, in one embodiment, step S204 includes:

[0066] S601, extracting a rough model of a virtual human using a polygonal network to obtain a shape and surface structure described by fixed points, edges, and faces;

[0067] S602: Set the three-dimensional space division degree to divide the shape and surface structure of the virtual human rough model to obtain multiple sub-blocks, and use the three-dimensional space grid to search and query, and generate corresponding Mesh grids.

[0068] In steps S601 to S602, the mesh extraction process is first performed on the virtual human rough model. This process involves visual extraction using polygon mesh technology to describe the morphology and surface structure of the three-dimensional object in detail through vertices, edges and faces. 3 The virtual human coarse model is subdivided into several sub-blocks. Then, an 8*8*8 grid structure is used for search and query to generate the corresponding high-precision mesh.

[0069] Taking the rough model of the financial mascot in the example above as an example, after extracting its shape and surface structure using polygonal mesh technology, the rough model was subdivided into multiple sub-blocks, each representing a portion of the mascot model. An 8x8x8 grid structure was then used to search and query in three-dimensional space, ensuring that each sub-block was accurately divided and a corresponding mesh was generated. This successfully divided the rough model into meshes, laying the foundation for the subsequent creation of more refined virtual human models and improving rendering efficiency.

[0070] S205, performing random perspective two-dimensional image rendering on the Mesh to obtain a corresponding rendered image, and performing noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image;

[0071] In step S205, image rendering ensures that the mesh presents a realistic two-dimensional image from different viewing angles. Noise processing simulates real-world image noise and enhances image robustness. Stable diffusion further smooths image noise while preserving key image features, resulting in an optimized image that is both realistic and clear.

[0072] like Figure 7 As shown, in one embodiment, step S205 includes:

[0073] S701, performing two-dimensional image rendering on the Mesh by randomly sampling the view angle to obtain a rendered image;

[0074] S702 , adding noise to the rendered image and inputting the noise into a pre-trained stable diffusion model for denoising optimization to generate an optimized image.

[0075] In steps S701-S702, by randomly selecting different sampling perspectives, different sides and details of the Mesh can be captured, thereby generating diverse two-dimensional rendering images. These two-dimensional rendering images can more comprehensively reflect the three-dimensional structure of the Mesh.

[0076] Next, noise is added to the rendered image and then input into a pre-trained stable diffusion model. The stable diffusion model here can adopt a generative model based on deep learning (such as a stable diffusion model). The stable diffusion model can gradually add noise to convert the data distribution into a simple noise distribution, and then learn the reverse process to generate a high-quality image from the noise to output the corresponding optimized image.

[0077] Taking the rough model of the financial mascot in the example above as an example, we rendered the mascot's mesh from multiple randomly selected viewpoints, generating a series of 2D rendered images. These rendered images showcase the mascot's appearance from different perspectives, increasing the diversity and realism of the rendered images. We then added a certain amount of noise to these rendered images to simulate the image noise that can occur in real-world photography. These noisy rendered images were then fed into a pre-trained stable diffusion model. This model leverages deep learning techniques to gradually remove noise from the images while preserving key features such as the mascot's shape, color, and texture. After optimization using the stable diffusion model, we obtained realistic and clear optimized images that not only showcase the mascot's 3D structure but also provide strong support for subsequent fine-tuning and rendering efficiency improvements.

[0078] S206, iteratively calculating the loss of the rendered image and the corresponding optimized image to optimize the Mesh grid and obtain a virtual human model;

[0079] In step S206, by comparing the rendered image with the optimized image, a deep learning algorithm is used to iteratively adjust the mesh's vertex positions and shapes to minimize the loss function. This process gradually brings the rendered image closer to the target optimized image, thereby finely depicting the virtual human's appearance, such as facial contours and facial expressions, to obtain a high-quality virtual human model.

[0080] like Figure 8 As shown, in one embodiment, step S206 includes:

[0081] S801, calculating the mean square error loss between the rendered image and the optimized image, and back-propagating the mean square error loss to the optimizer based on fractional distillation sampling to update the mesh vertex coordinates;

[0082] S802 , repeatedly performing the mean square error loss calculation and back propagation process between the rendered image and the optimized image until a preset number of iterations or loss convergence condition is reached, thereby obtaining a virtual human fine model.

[0083] In steps S801-S802, the mesh vertex coordinates are continuously optimized iteratively, gradually reducing the difference between the rendered and optimized images, allowing the virtual human model to gradually approach the real-world effect. Once a predetermined number of iterations or loss convergence condition is reached, the virtual human model is considered optimized and can be used in subsequent applications such as virtual human generation, animation, or real-time rendering.

[0084] More specifically, the mean squared error (MSE) loss is used as an optimization metric. This loss function quantifies the degree of difference between the rendered image and the optimized image. In each iteration, the mean squared error loss between the current rendered image and the optimized image is calculated and back-propagated to an optimizer based on fractional distillation sampling. Based on the loss gradient information, the optimizer intelligently adjusts the vertex coordinates of the mesh to minimize the difference between the rendered and optimized images. This process is iterative, with the steps of calculating the loss, back-propagating, and updating the mesh vertex coordinates repeatedly until a preset number of iterations or loss convergence condition is reached. After convergence, the resulting virtual human model closely approximates the target optimized image, exhibiting detailed appearance features and realistic visual effects. This virtual human model not only retains the unique shape and color texture of the financial mascot, but also, through an iterative optimization process, further enhances its realism and detail, providing a solid foundation for subsequent applications.

[0085] As can be seen, in the above solution, for virtual human model construction scenarios in the financial sector, an efficient and sophisticated virtual human model construction method was successfully implemented by combining technical means such as mesh partitioning, random perspective 2D image rendering, noise addition and stable diffusion processing, and iterative loss optimization. This method effectively optimizes the virtual human model generation process while maintaining high-fidelity model generation results. This method not only improves the realism and detail expression of the virtual human model, but also significantly improves rendering efficiency, providing strong technical support for virtual human generation, animation production, or real-time rendering in application scenarios such as the financial sector. In addition, the solution is also universal and can be extended to the construction of virtual human models in other fields, opening up new avenues for the development and application of virtual human technology.

[0086] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0087] In one embodiment, a virtual human generation device is provided, which corresponds to the virtual human generation method in the above embodiment. Figure 9 As shown, the virtual human generation device includes a first image generation module 901, a second image generation module 902, a coarse generation module 903, a model segmentation module 904, an image optimization module 905 and a fine generation module 906. The functional modules are described in detail as follows:

[0088] The first image generation module 901 is configured to perform image processing on a single input original perspective image and / or text description to obtain a new perspective image;

[0089] The second image generation module 902 is configured to perform multi-perspective generation processing on the new-perspective image to obtain a multi-perspective image;

[0090] A coarse generation module 903 is used to reconstruct the multi-view images using a three-dimensional Gaussian function to obtain a coarse model of the virtual human;

[0091] The model segmentation module 904 is used to extract the basic structural data of the virtual human rough model and divide it into meshes;

[0092] The image optimization module 905 is used to perform random perspective two-dimensional image rendering on the Mesh grid to obtain a corresponding rendered image, and to perform noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image;

[0093] The fine generation module 906 is used to iteratively calculate the loss of the rendered image and the corresponding optimized image to optimize the Mesh grid and obtain a virtual human fine model.

[0094] In one embodiment, the first image generation module 901 is specifically configured to:

[0095] Receive the original view image and the text description containing the target view parameters input by the user;

[0096] Multimodal feature encoding and fusion of the original view image and text description are performed to generate conditional control signals, where the view parameters in the text description are parsed into spatial transformation parameters;

[0097] The conditional control signal and the spatial transformation descriptor are input into a pre-trained single-view generative diffusion model. The latent feature representation of the target view is generated through the iterative denoising process of the single-view generative diffusion model. During the iterative denoising process, the spatial transformation parameters are dynamically adjusted to achieve view consistency control.

[0098] The latent feature representation is decoded and a new view image is output that is semantically consistent with the input image scene and satisfies the target view constraint.

[0099] In one embodiment, the second image generation module 902 is specifically configured to:

[0100] The new view image is input into the pre-trained multi-view generative diffusion model, which includes a basic visual diffusion architecture and a parameter-efficient adaptation layer. The parameter-efficient adaptation layer dynamically adjusts the model parameters through low-rank matrix decomposition.

[0101] The geometric and texture features of the new-view image are propagated across views through a multi-stage denoising process of a multi-view generative diffusion model, and a view-conditioned feature vector is iteratively generated in the latent space.

[0102] The view-conditioned feature vector is decoded into a multi-view image containing spatial continuity.

[0103] In one embodiment, the coarse generation module 903 is specifically configured to:

[0104] Acquire a multi-view two-dimensional image set of multi-view images;

[0105] A multi-view two-dimensional image set is input into a three-dimensional Gaussian function, and the gradient backpropagation between the image space and the three-dimensional space is performed through a differentiable rendering engine. The attribute parameters of the Gaussian basis elements are iteratively optimized to generate a three-dimensional editable virtual human rough model.

[0106] In one embodiment, the model segmentation module 904 is specifically configured to:

[0107] The rough model of the virtual human is extracted using polygonal networks to obtain the shape and surface structure described by fixed points, edges and faces;

[0108] The three-dimensional space division degree is set to divide the shape and surface structure of the virtual human rough model to obtain multiple sub-blocks, and the three-dimensional space grid is used for search and query, and the corresponding Mesh grid is generated.

[0109] In one embodiment, the image optimization module 905 is specifically configured to:

[0110] Render the Mesh grid into a two-dimensional image by randomly sampling the perspective;

[0111] After adding noise to the rendered image, the image is input into the pre-trained stable diffusion model for denoising optimization to generate an optimized image.

[0112] In one embodiment, the detail generation module 906 is specifically configured to:

[0113] Calculate the mean squared error loss between the rendered image and the optimized image, and backpropagate the mean squared error loss to the optimizer based on fractional distillation sampling to update the mesh vertex coordinates;

[0114] The mean square error loss calculation and back propagation process between the rendered image and the optimized image are repeated until the preset number of iterations or the loss convergence condition is reached to obtain a virtual human model.

[0115] The present invention provides a virtual human generation device, which successfully implements an efficient and sophisticated virtual human model construction method in the virtual human model construction scenario in the financial field by combining Mesh mesh division, random perspective two-dimensional image rendering, noise addition and stable diffusion processing, and iterative loss optimization and other technical means. This method effectively optimizes the generation process of the virtual human model and maintains a high-fidelity model generation effect. This method not only improves the realism and detail expression of the virtual human model, but also significantly improves the rendering efficiency, providing strong technical support for virtual human generation, animation production or real-time rendering in application scenarios such as the financial field. In addition, the solution also has a certain degree of versatility and can be extended to the construction of virtual human models in other fields, opening up new paths for the development and application of virtual human technology.

[0116] For the specific definition of the virtual human generation device, please refer to the definition of the virtual human generation method above and will not be repeated here. The various modules in the above-mentioned virtual human generation device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0117] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 10As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a virtual human generation method.

[0118] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 11 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. The computer program is executed by the processor to implement functions or steps on the client side of a virtual human generation method.

[0119] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0120] Perform image processing on the input single original perspective image and / or text description to obtain a new perspective image;

[0121] Performing multi-perspective generation processing on the new perspective image to obtain a multi-perspective image;

[0122] The multi-view images are reconstructed using a three-dimensional Gaussian function to obtain a rough model of the virtual human;

[0123] Extract the basic structural data of the virtual human coarse model and divide it into Mesh grids;

[0124] Perform random perspective 2D image rendering on the Mesh grid to obtain the corresponding rendered image, and perform noise addition and stable diffusion processing on the rendered image to obtain the corresponding optimized image;

[0125] The loss of the rendered image and the corresponding optimized image is iteratively calculated to optimize the Mesh mesh and obtain a virtual human model.

[0126] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0127] Perform image processing on the input single original perspective image and / or text description to obtain a new perspective image;

[0128] Performing multi-perspective generation processing on the new perspective image to obtain a multi-perspective image;

[0129] The multi-view images are reconstructed using a three-dimensional Gaussian function to obtain a rough model of the virtual human;

[0130] Extract the basic structural data of the virtual human coarse model and divide it into Mesh grids;

[0131] Perform random perspective 2D image rendering on the Mesh grid to obtain the corresponding rendered image, and perform noise addition and stable diffusion processing on the rendered image to obtain the corresponding optimized image;

[0132] The loss of the rendered image and the corresponding optimized image is iteratively calculated to optimize the Mesh mesh and obtain a virtual human model.

[0133] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0134] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).

[0135] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0136] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for generating a virtual human, characterized in that: include: Perform image processing on the input single original perspective image and / or text description to obtain a new perspective image; Performing multi-perspective generation processing on the new-perspective image to obtain a multi-perspective image; Reconstructing the multi-view images using a three-dimensional Gaussian function to obtain a rough model of a virtual human; Extracting basic structural data of the virtual human coarse model and dividing it into Mesh grids; Performing random perspective two-dimensional image rendering on the Mesh grid to obtain a corresponding rendered image, and performing noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image; The loss of the rendered image and the corresponding optimized image is iteratively calculated to optimize the Mesh grid and obtain a virtual human model.

2. The method for generating a virtual human according to claim 1, wherein: The step of performing image processing on the input single original perspective image and / or text description to obtain a new perspective image includes: Receiving the original viewing angle image and a text description including target viewing angle parameters input by a user; Performing multimodal feature encoding and fusion on the original perspective image and text description to generate a conditional control signal, wherein the perspective parameters in the text description are parsed into spatial transformation parameters; Inputting the conditional control signal and the spatial transformation descriptor into a pre-trained single-view generative diffusion model, generating a latent feature representation of the target view through an iterative denoising process of the single-view generative diffusion model, and achieving view consistency control by dynamically adjusting spatial transformation parameters during the iterative denoising process; The latent feature representation is decoded to output a new perspective image that is semantically consistent with the input image scene and satisfies the target perspective constraint.

3. The method for generating a virtual human according to claim 1, wherein: The performing multi-perspective generation processing on the new-perspective image to obtain the multi-perspective image includes: Inputting the new-view image into a pre-trained multi-view generative diffusion model, wherein the multi-view generative diffusion model includes a basic visual diffusion architecture and a parameter efficient adaptation layer, wherein the parameter efficient adaptation layer dynamically adjusts the model parameters through low-rank matrix decomposition; Propagate the geometric features and texture features of the new-view image across views through a multi-stage denoising process of the multi-view generative diffusion model, and iteratively generate a view-conditioned feature vector in a latent space; The view-conditioned feature vector is decoded into a multi-view image containing spatial continuity.

4. The method for generating a virtual human according to claim 1, wherein: The reconstructing process of the multi-view images by using a three-dimensional Gaussian function to obtain a coarse model of a virtual human includes: Acquire a multi-view two-dimensional image set of the multi-view image; The multi-view two-dimensional image set is input into a three-dimensional Gaussian function, and gradient backpropagation between the image space and the three-dimensional space is performed through a differentiable rendering engine, and Gaussian basis element attribute parameters are iteratively optimized to generate a three-dimensional editable virtual human rough model.

5. The method for generating a virtual human according to claim 1, wherein: The extracting of the basic structural data of the virtual human coarse model and dividing it into Mesh grids includes: Extracting the rough model of the virtual human using a polygonal network to obtain a shape and surface structure described by fixed points, edges, and faces; The three-dimensional space division degree is set to divide the shape and surface structure of the virtual human rough model to obtain multiple sub-blocks, and the three-dimensional space grid is used for search and query, and a corresponding Mesh grid is generated.

6. The method for generating a virtual human according to claim 1, wherein: The performing random perspective two-dimensional image rendering on the Mesh to obtain a corresponding rendered image, and performing noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image, including: Performing two-dimensional image rendering on the Mesh by randomly sampling perspectives to obtain a rendered image; After adding noise to the rendered image, the image is input into a pre-trained stable diffusion model for denoising optimization to generate an optimized image.

7. The method for generating a virtual human according to claim 1, wherein: The iterative calculation of the loss of the rendered image and the corresponding optimized image to optimize the Mesh grid and obtain a virtual human model includes: Calculating the mean square error loss between the rendered image and the optimized image, and back-propagating the mean square error loss to the optimizer based on fractional distillation sampling to update the mesh vertex coordinates; The mean square error loss calculation and back propagation process between the rendered image and the optimized image are repeatedly performed until a preset number of iterations or loss convergence condition is reached, thereby obtaining a virtual human fine model.

8. A virtual human generation device, characterized in that: include: A first image generation module is used to perform image processing on a single input original perspective image and / or text description to obtain a new perspective image; A second image generation module is used to perform multi-perspective generation processing on the new-perspective image to obtain a multi-perspective image; A coarse generation module, configured to reconstruct the multi-view images using a three-dimensional Gaussian function to obtain a coarse model of a virtual human; A model segmentation module is used to extract the basic structural data of the virtual human coarse model and divide it into Mesh grids; An image optimization module is used to perform random perspective two-dimensional image rendering on the Mesh grid to obtain a corresponding rendered image, and to perform noise addition and stable diffusion processing on the rendered image to obtain a corresponding optimized image; A fine generation module is used to iteratively calculate the loss of the rendered image and the corresponding optimized image to optimize the Mesh grid and obtain a virtual human fine model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the virtual human generation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the virtual human generation method according to any one of claims 1 to 7 are implemented.