Three-dimensional virtual object generation method and device, medium and computer equipment
Through parameter estimation and prior model training technology, combined with three-dimensional Gaussian splashing parameters, the problem of traditional 3D digital life generation requires a lot of professional skills and time, and achieves fast and high-quality 3D model generation.
Patent Information
- Application Number
- CN202411855585.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional 3D digital life generation requires a lot of professional skills and time, and the final product is inconsistent due to the differences in artists, making it difficult to meet the needs of fast iterative game development.
Parameter estimation and prior model training technology are used to obtain the target object image for parameter estimation, train the virtual object prior model, and generate a high-quality 3D model with three-dimensional Gaussian splattering parameters.
It reduces manual intervention and reliance on high-level professional skills, reduces inconsistency problems caused by personnel differences, and can generate high-quality 3D models faster to adapt to the development needs of rapid iteration.
Smart Images

Figure CN119991891A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of three-dimensional animation technology, and in particular to a method, device, medium, and computer equipment for generating a three-dimensional virtual object. Background Art
[0002] During the game development process, in order to realize the NPC's 3D modeling and its action redirection, in order to achieve the purpose of diversified NPC movement and high-quality CG effects.
[0003] Traditional 3D digital human generation requires manual modeling using professional 3D reconstruction software (such as Autodesk Maya, 3Ds Max, etc.), which is a classic method used in game development for a long time. This method requires the art team to have a high level of professional skills, including the original artist drawing 2D multi-perspective concept drawings, and the 3D modeler building a detailed 3D model based on these design drawings. However, this method has significant manpower and time cost issues, and due to differences in communication and collaboration between different artists, the final product may not meet expectations, affecting the overall consistency of the project.
[0004] In recent years, with the development of artificial intelligence technology, deep learning-based methods have gradually become a new option, especially 3D generative adversarial networks (GANs). Through the adversarial training mechanism, 3D GAN can learn the shape and texture features of 3D digital humans from two-dimensional images to generate realistic 3D models. The advantage of this method is that it can reduce dependence on manual intervention and may accelerate the creation process. However, it also faces its own technical bottlenecks: high-quality 3D model generation requires a large amount of well-annotated training data, which not only increases the cost of preliminary preparation, but may also be limited by copyright or privacy issues. Moreover, the training of deep learning models usually requires strong computing power support, especially when dealing with complex scenes, and the long training cycle may be difficult to meet the needs of fast-iteration game development. Summary of the invention
[0005] In view of this, the embodiments of the present application provide a method, device, medium, and computer equipment for generating a three-dimensional virtual object, which utilize parameter estimation and prior model training technology to reduce the degree of human intervention, reduce the dependence on a high-level professional art team, reduce the inconsistency problem of the final product caused by personnel differences, and can generate high-quality 3D models more quickly.
[0006] According to one aspect of the present application, a method for generating a three-dimensional virtual object is provided, the method comprising:
[0007] Acquire a target object image, wherein the target object image includes a target object having a face portion and a body portion;
[0008] Performing parameter estimation on the target object image to obtain target object facial parameters and target object body parameters, and training a virtual object prior model based on the target object facial parameters and the target object body parameters to obtain target object prior parameters corresponding to the trained virtual object prior model;
[0009] Generate three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image;
[0010] A three-dimensional virtual object is rendered according to the target object priori parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object.
[0011] In an optional implementation, the step of performing parameter estimation on the target object image to obtain target object facial parameters and target object body parameters includes:
[0012] Performing FLAME prior parameter estimation on the target object image to obtain first prior parameters, wherein the first prior parameters include first facial prior parameters; and performing SMPLX prior parameter estimation on the target object image to obtain second prior parameters, wherein the second prior parameters include second facial prior parameters and joint point parameters;
[0013] The second facial priori parameters in the second priori parameters are replaced with the first facial priori parameters, and the facial parameters and the body parameters of the target object are determined based on the replaced second priori parameters.
[0014] In an optional implementation, the training of the virtual object prior model based on the target object facial parameters and the target object body parameters to obtain the target object prior parameters corresponding to the trained virtual object prior model includes:
[0015] Initializing the prior parameters of the virtual object prior model, wherein the prior parameters include body prior parameters and face prior parameters;
[0016] Inputting the target object facial parameters and the target object body parameters into the virtual object prior model to obtain the virtual object three-dimensional feature information;
[0017] Performing a two-dimensional projection on the three-dimensional feature information of the virtual object to obtain virtual object feature projection information, performing loss calculation based on the virtual object feature projection information and target object feature point information corresponding to the target object image, and optimizing the prior parameters based on the obtained first loss calculation result until the model training condition is met and the training is completed;
[0018] The prior parameters in the trained virtual object prior model are obtained as the target object prior parameters.
[0019] In an optional implementation, the performing loss calculation based on the virtual object feature projection information and the target object feature point information corresponding to the target object image, and optimizing the priori parameters based on the obtained first loss calculation result, includes:
[0020] Extracting the facial key point coordinates and joint bone point coordinates corresponding to the target object image;
[0021] Performing facial loss calculation based on the facial key point coordinates and the facial key point projection coordinates in the virtual object feature projection information, and performing joint loss calculation based on the joint bone point coordinates and the joint bone point projection coordinates in the virtual object feature projection information;
[0022] The facial prior parameters are optimized based on the facial loss, and the body prior parameters are optimized based on the joint loss.
[0023] In an optional implementation, the generating of three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image includes:
[0024] Initializing three-dimensional structured feature parameters of the preset three-dimensional virtual object model, sampling the three-dimensional structured feature parameters according to a grid based on vertex position coding corresponding to the preset three-dimensional virtual object model, and obtaining vertex features of virtual object vertices corresponding to the preset three-dimensional virtual object model;
[0025] A three-dimensional Gaussian splash parameter generation model is used to generate three-dimensional Gaussian splash parameters corresponding to the vertices of the virtual object according to the vertex features and the target object image.
[0026] In an optional embodiment, the body prior parameters in the prior parameters include body shape parameters, body posture parameters and joint offset parameters, and the face prior parameters in the prior parameters include face shape parameters and face offset parameters;
[0027] The performing three-dimensional virtual object rendering according to the target object prior parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object includes:
[0028] Rendering the initialized virtual object prior model using the body shape parameters, the facial shape parameters, the facial offset parameters and the three-dimensional Gaussian splash parameters to obtain an initial three-dimensional virtual object model, and skinning the initial three-dimensional virtual object model based on the body posture parameters and the joint offset parameters to obtain an intermediate three-dimensional virtual object model;
[0029] Perform two-dimensional projection on the intermediate three-dimensional virtual object model to obtain full projection information of the virtual object, perform loss calculation based on the full projection information of the virtual object and the target object image, and optimize the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model based on the obtained second loss calculation result until the three-dimensional virtual object model generated based on the optimized three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model meets the three-dimensional target object rendering conditions.
[0030] In an optional implementation, the loss calculation is performed based on the full projection information of the virtual object and the target object image, and the three-dimensional structural feature parameters and the three-dimensional Gaussian splash parameter generation model are optimized based on the obtained second loss calculation result, including:
[0031] extracting a target object mask from the target object image;
[0032] Based on the full projection information of the virtual object and the target object mask, performing pixel value loss calculation, structural similarity loss calculation, and perceptual difference loss calculation;
[0033] The three-dimensional Gaussian splash parameter generation model is optimized based on pixel value loss, and the three-dimensional structured feature parameters of the preset three-dimensional virtual object model are optimized based on structural similarity loss and perceptual difference loss.
[0034] According to another aspect of the present application, a method for redirecting a motion of a three-dimensional virtual object is provided, comprising:
[0035] Acquire a target object video, and extract a target object image sequence from the target object video;
[0036] Perform parameter estimation on each target object image sequence respectively to obtain a body parameter sequence corresponding to each target object image sequence;
[0037] According to the body parameter sequence, the target object prior parameters corresponding to the target object and the three-dimensional Gaussian splash parameters, the target object is redirected in motion, wherein the target object prior parameters and the three-dimensional Gaussian splash parameters are obtained by the three-dimensional virtual object generation method mentioned above.
[0038] In an optional embodiment, the body parameter sequence includes a body posture parameter sequence, and the target object prior parameters include target object body shape parameters, target object joint offset parameters, target object facial shape parameters, and target object facial offset parameters;
[0039] The step of redirecting the target object's motion according to the body parameter sequence, the target object's priori parameters corresponding to the target object, and the three-dimensional Gaussian splash parameters includes:
[0040] A preset virtual object prior model is rendered according to the target object body shape parameters, the target object facial shape parameters, the target object facial offset parameters and the three-dimensional Gaussian splash parameters, and the rendered virtual object prior model is skinned based on the body posture parameter sequence and the target object joint offset parameters to obtain a motion redirection model corresponding to the target object.
[0041] According to another aspect of the present application, a device for generating a three-dimensional virtual object is provided, the device comprising:
[0042] An image acquisition module, used for acquiring a target object image, wherein the target object image includes a target object having a face part and a body part;
[0043] a priori parameter determination module, used to perform parameter estimation on the target object image to obtain target object facial parameters and target object body parameters, and train a virtual object priori model based on the target object facial parameters and the target object body parameters to obtain target object priori parameters corresponding to the trained virtual object prior model;
[0044] A Gaussian splash parameter determination module, configured to generate three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image;
[0045] A virtual object generation module is used to perform three-dimensional virtual object rendering according to the target object prior parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object.
[0046] In an optional implementation manner, the priori parameter determination module is further used to:
[0047] Performing FLAME prior parameter estimation on the target object image to obtain first prior parameters, wherein the first prior parameters include first facial prior parameters; and performing SMPLX prior parameter estimation on the target object image to obtain second prior parameters, wherein the second prior parameters include second facial prior parameters and joint point parameters;
[0048] The second facial priori parameters in the second priori parameters are replaced with the first facial priori parameters, and the facial parameters and the body parameters of the target object are determined based on the replaced second priori parameters.
[0049] In an optional implementation manner, the priori parameter determination module is further used to:
[0050] Initializing the prior parameters of the virtual object prior model, wherein the prior parameters include body prior parameters and face prior parameters;
[0051] Inputting the target object facial parameters and the target object body parameters into the virtual object prior model to obtain the virtual object three-dimensional feature information;
[0052] Performing a two-dimensional projection on the three-dimensional feature information of the virtual object to obtain virtual object feature projection information, performing loss calculation based on the virtual object feature projection information and target object feature point information corresponding to the target object image, and optimizing the prior parameters based on the obtained first loss calculation result until the model training condition is met and the training is completed;
[0053] The prior parameters in the trained virtual object prior model are obtained as the target object prior parameters.
[0054] In an optional implementation manner, the priori parameter determination module is further used to:
[0055] Extracting the facial key point coordinates and joint bone point coordinates corresponding to the target object image;
[0056] Performing facial loss calculation based on the facial key point coordinates and the facial key point projection coordinates in the virtual object feature projection information, and performing joint loss calculation based on the joint bone point coordinates and the joint bone point projection coordinates in the virtual object feature projection information;
[0057] The facial prior parameters are optimized based on the facial loss, and the body prior parameters are optimized based on the joint loss.
[0058] In an optional implementation manner, the Gaussian splash parameter determination module is further used to:
[0059] Initializing three-dimensional structured feature parameters of the preset three-dimensional virtual object model, sampling the three-dimensional structured feature parameters according to a grid based on vertex position coding corresponding to the preset three-dimensional virtual object model, and obtaining vertex features of virtual object vertices corresponding to the preset three-dimensional virtual object model;
[0060] A three-dimensional Gaussian splash parameter generation model is used to generate three-dimensional Gaussian splash parameters corresponding to the vertices of the virtual object according to the vertex features and the target object image.
[0061] In an optional embodiment, the body prior parameters in the prior parameters include body shape parameters, body posture parameters and joint offset parameters, and the face prior parameters in the prior parameters include face shape parameters and face offset parameters;
[0062] The virtual object generation module is further used for:
[0063] Rendering the initialized virtual object prior model using the body shape parameters, the facial shape parameters, the facial offset parameters and the three-dimensional Gaussian splash parameters to obtain an initial three-dimensional virtual object model, and skinning the initial three-dimensional virtual object model based on the body posture parameters and the joint offset parameters to obtain an intermediate three-dimensional virtual object model;
[0064] Perform two-dimensional projection on the intermediate three-dimensional virtual object model to obtain full projection information of the virtual object, perform loss calculation based on the full projection information of the virtual object and the target object image, and optimize the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model based on the obtained second loss calculation result until the three-dimensional virtual object model generated based on the optimized three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model meets the three-dimensional target object rendering conditions.
[0065] In an optional implementation, the virtual object generation module is further used to:
[0066] extracting a target object mask from the target object image;
[0067] Based on the full projection information of the virtual object and the target object mask, performing pixel value loss calculation, structural similarity loss calculation, and perceptual difference loss calculation;
[0068] The three-dimensional Gaussian splash parameter generation model is optimized based on pixel value loss, and the three-dimensional structured feature parameters of the preset three-dimensional virtual object model are optimized based on structural similarity loss and perceptual difference loss.
[0069] According to another aspect of the present application, a motion redirection device for a three-dimensional virtual object is provided, comprising:
[0070] A video acquisition module, used to acquire a target object video and extract a target object image sequence from the target object video;
[0071] A sequence determination module is used to perform parameter estimation on each target object image sequence to obtain a body parameter sequence corresponding to each target object image sequence;
[0072] A redirection module is used to redirect the action of the target object according to the body parameter sequence, the target object prior parameters corresponding to the target object and the three-dimensional Gaussian splash parameters, wherein the target object prior parameters and the three-dimensional Gaussian splash parameters are obtained by the above-mentioned three-dimensional virtual object generation device.
[0073] In an optional embodiment, the body parameter sequence includes a body posture parameter sequence, and the target object prior parameters include target object body shape parameters, target object joint offset parameters, target object facial shape parameters and target object facial offset parameters; the redirection module is also used to: render a preset virtual object prior model according to the target object body shape parameters, the target object facial shape parameters, the target object facial offset parameters and the three-dimensional Gaussian splash parameters, and skin the rendered virtual object prior model based on the body posture parameter sequence and the target object joint offset parameters to obtain an action redirection model corresponding to the target object.
[0074] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method for generating a three-dimensional virtual object or the method for redirecting the action of a three-dimensional virtual object is implemented.
[0075] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein when the processor executes the program, the method for generating a three-dimensional virtual object or the method for redirecting the action of a three-dimensional virtual object is implemented.
[0076] By means of the above technical solution, a method, device, medium, and computer equipment for generating a three-dimensional virtual object provided in an embodiment of the present application obtains a target object image, first performs parameter estimation and virtual object prior model training, obtains target object prior parameters, and then combines the three-dimensional Gaussian splash parameter generation, and finally realizes the rendering of the three-dimensional virtual object based on the target object prior parameters and the three-dimensional Gaussian splash parameters. Compared with the traditional manual modeling method, the embodiment of the present application utilizes parameter estimation and prior model training technology to reduce the degree of manual intervention, reduce the dependence on a high-level professional art team, reduce the inconsistency problem of the final product caused by personnel differences, and can generate high-quality 3D models faster. In addition, by combining the target object prior parameters and the three-dimensional Gaussian splash parameters, it helps to reduce the dependence on labeled data and further improve development efficiency.
[0077] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0079] Figure 1 A schematic diagram of a process for generating a three-dimensional virtual object provided in an embodiment of the present application is shown;
[0080] Figure 2 A schematic diagram showing a flow chart of another method for generating a three-dimensional virtual object provided in an embodiment of the present application;
[0081] Figure 3 A schematic diagram showing a flow chart of another method for generating a three-dimensional virtual object provided in an embodiment of the present application;
[0082] Figure 4 A schematic diagram showing a flow chart of another method for generating a three-dimensional virtual object provided in an embodiment of the present application;
[0083] Figure 5 A schematic diagram showing a flow chart of another method for generating a three-dimensional virtual object provided in an embodiment of the present application;
[0084] Figure 6 A schematic diagram of a process of redirecting a three-dimensional virtual object provided in an embodiment of the present application is shown;
[0085] Figure 7 A schematic diagram showing a flow chart of another method for redirecting a three-dimensional virtual object provided in an embodiment of the present application is shown;
[0086] Figure 8 A schematic diagram of the structure of a device for generating a three-dimensional virtual object provided in an embodiment of the present application is shown;
[0087] Fig. 9 A schematic structural diagram of a three-dimensional virtual object redirection device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0088] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.
[0089] In this embodiment, a method for generating a three-dimensional virtual object is provided. Figure 1 As shown, the method includes:
[0090] Step 101: Acquire a target object image, wherein the target object image includes a target object having a face part and a body part.
[0091] The embodiment of the present application provides a method for generating a three-dimensional virtual object, which can be used to implement 3D modeling of NPC (non-player character) and other characters in game development. First, a target object image is obtained, and the image contains the target object (i.e., the prototype of the virtual character such as NPC to be modeled). In order to ensure the three-dimensional modeling effect, these images should contain the face and body parts of the target object. The target object can be a human body, an animal, etc.
[0092] Step 102: perform parameter estimation on the target object image to obtain target object facial parameters and target object body parameters, and train a virtual object prior model based on the target object facial parameters and the target object body parameters to obtain target object prior parameters corresponding to the trained virtual object prior model.
[0093] In the embodiment of the present application, after obtaining the target object image, the target object image is parameter estimated to obtain specific parameters of the face and body. Specifically, the SMPLX human body prior parameter estimation and other methods can be selected for parameter estimation. Then, these parameters are used to train a virtual object prior model, such as the SMPLX human body prior model. The virtual object prior model contains initialization parameters that express the characteristics of the character. The characteristics of the target object will be learned by training the model, and the parameter values of these initialization parameters will also change accordingly. After training, the parameters of the virtual object prior model are extracted as the target object prior parameters to obtain the target object prior parameters that can express the facial and body characteristics of the target object, so as to facilitate the subsequent three-dimensional modeling of the virtual object based on these parameters.
[0094] Step 103 : generating three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image.
[0095] In the embodiment of the present application, in order to more delicately express the detailed features of the target object, such as color features, etc., the target object image can also be used to generate the three-dimensional Gaussian splash parameters of the target object, so that the above-mentioned target object prior parameters and the three-dimensional Gaussian splash parameters can be combined to achieve more expressive and realistic three-dimensional modeling of virtual objects. Specifically, the three-dimensional Gaussian splash parameters can be generated for each virtual object vertex in the preset three-dimensional virtual object model based on the target object image. This process can be regarded as a process of adding texture or details, so that the subsequently generated 3D model is more realistic and delicate.
[0096] Step 104 : Rendering a three-dimensional virtual object according to the target object priori parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object.
[0097] In the embodiment of the present application, finally, the three-dimensional virtual object is rendered in combination with the target object prior parameters and the three-dimensional Gaussian splash parameters, thereby generating a realistic 3D virtual object based on the target object, and realizing virtual object modeling in scenarios such as game development or animation development.
[0098] By applying the technical solution of this embodiment, by acquiring the target object image, first performing parameter estimation and virtual object prior model training, obtaining the target object prior parameters, and then combining the three-dimensional Gaussian splash parameter generation, finally achieving the rendering of the three-dimensional virtual object based on the target object prior parameters and the three-dimensional Gaussian splash parameters. Compared with the traditional manual modeling method, the embodiment of the present application utilizes parameter estimation and prior model training technology, reduces the degree of manual intervention, reduces the dependence on a high-level professional art team, reduces the inconsistency of the final product caused by personnel differences, and can generate high-quality 3D models faster. In addition, by combining the target object prior parameters and the three-dimensional Gaussian splash parameters, it helps to reduce the dependence on labeled data and further improve development efficiency.
[0099] In an embodiment of the present application, optionally, the parameter estimation of the target object image to obtain the target object facial parameters and the target object body parameters includes: performing FLAME prior parameter estimation on the target object image to obtain first prior parameters, wherein the first prior parameters include first facial prior parameters; and performing SMPLX prior parameter estimation on the target object image to obtain second prior parameters, wherein the second prior parameters include second facial prior parameters and joint point parameters; replacing the second facial prior parameters in the second prior parameters with the first facial prior parameters, and determining the target object facial parameters and the target object body parameters based on the replaced second prior parameters.
[0100] In this embodiment, two prior parameters, FLAME and SMPLX, are estimated for the target object image. By estimating the FLAME prior parameters for the target object image, a first prior parameter including a first facial prior parameter is obtained, and by estimating the SMPLX prior parameters for the target object image, a second prior parameter including a second facial prior parameter and a joint point parameter is obtained. Figure 2 As shown, taking the target object as a human body as an example, the FLAME prior parameter estimation is the FLAME facial parameter estimation, and the SMPLX prior parameter estimation is the SMPLX human body parameter estimation. Subsequently, the facial part in the second prior parameter is replaced with a more accurate FLAME facial prior parameter, thereby obtaining an optimized parameter set. Based on this optimized parameter set, the facial parameters and body parameters of the target object are finally determined. The embodiment of the present application combines the advantages of the two prior parameter estimations of FLAME and SMPLX, improves the accuracy and precision of the facial and body parameter estimation, especially captures facial features more accurately, and helps to generate a more realistic virtual image that is more in line with the characteristics of the target object. Of course, the facial prior parameters and joint point parameters obtained by the SMPLX prior parameter estimation can also be directly used as the body parameters of the target object.
[0101] In the embodiment of the present application, optionally, Figure 3 As shown, the virtual object prior model is trained based on the target object facial parameters and the target object body parameters to obtain the target object prior parameters corresponding to the trained virtual object prior model, including:
[0102] Step 301, initializing the prior parameters of the virtual object prior model, wherein the prior parameters include body prior parameters and face prior parameters;
[0103] Step 302: input the target object facial parameters and the target object body parameters into the virtual object prior model to obtain the virtual object three-dimensional feature information;
[0104] Step 303: perform two-dimensional projection on the three-dimensional feature information of the virtual object to obtain virtual object feature projection information, perform loss calculation based on the virtual object feature projection information and target object feature point information corresponding to the target object image, and optimize the prior parameters based on the obtained first loss calculation result until the model training condition is met and the training is completed;
[0105] Step 304: Acquire the prior parameters in the trained virtual object prior model as the target object prior parameters.
[0106] In this embodiment, if Figure 2As shown, first, a virtual object prior model including body prior parameters and facial prior parameters is initialized, such as a SMPLX human body prior model, and the initialized body prior parameters and facial prior parameters are used as learnable bias parameters to be optimized in the subsequent model training process. For example, the body prior parameters include body shape parameters SMPLX_Shape_params, body posture parameters SMPLX_Pose_params, and joint offset parameters SMPLX_Joint_offset, and the facial prior parameters include facial shape parameters FLAME_Shape_params and facial offset parameters FLAME_face_offset. Subsequently, the target object facial parameters and the target object body parameters are input into the initialized virtual object prior model, and the three-dimensional feature information of the virtual object, i.e., SMPLX human body 3D information, is generated through the model. Then, these three-dimensional feature information are projected into a two-dimensional space to obtain feature projection information of the virtual object. By comparing these projection information with the actual feature point information in the target object image, a first loss value is calculated, and this loss value is used to optimize the prior parameters of the model. This training process continues until the preset model training conditions are met (for example, the loss is less than the preset value) and the training is completed. Finally, the prior parameters are extracted from the trained model as the prior parameters of the target object. The embodiment of the present application trains a virtual object prior model and uses the facial and body parameters of the target object for training, which can refine the model's performance in facial and body features, improve the accuracy and details of the generated virtual object, and enable it to better adapt to and reflect the characteristics of the target object, thereby generating a more realistic virtual object that is more in line with the image of the target object.
[0107] In an embodiment of the present application, optionally, the loss calculation described in step 303 is performed based on the virtual object feature projection information and the target object feature point information corresponding to the target object image, and the prior parameters are optimized based on the first loss calculation result obtained, including: extracting the facial key point coordinates and joint bone point coordinates corresponding to the target object image; performing facial loss calculation based on the facial key point coordinates and the facial key point projection coordinates in the virtual object feature projection information, and performing joint loss calculation based on the joint bone point coordinates and the joint bone point projection coordinates in the virtual object feature projection information; optimizing the facial prior parameters based on the facial loss, and optimizing the body prior parameters based on the joint loss.
[0108] In this embodiment, if Figure 2As shown, the coordinate information of the facial key points (partial facial key points in the figure) and joint bone points (2D bone points of the human body joints in the figure) are extracted from the target object image. Then, these coordinate information are compared with the facial key point projection coordinates (SMPLX facial 2D key points in the figure) and joint bone point projection coordinates (SMPLX human body joint 2D bone points in the figure) in the virtual object feature projection information, and the facial loss L1_loss and joint loss L2_loss are calculated respectively. The facial prior parameters are optimized based on the facial loss, and the body prior parameters are optimized based on the joint loss. By calculating the facial loss and joint loss separately, the optimization process of the facial prior parameters and the body prior parameters can be more accurately guided, and the performance accuracy of the model on facial and body features can be improved.
[0109] In an embodiment of the present application, optionally, the three-dimensional Gaussian splash parameter generation for the virtual object vertices of a preset three-dimensional virtual object model based on the target object image includes: initializing the three-dimensional structured feature parameters of the preset three-dimensional virtual object model, sampling the three-dimensional structured feature parameters according to a grid based on vertex position encoding corresponding to the preset three-dimensional virtual object model, and obtaining vertex features of the virtual object vertices corresponding to the preset three-dimensional virtual object model; generating a model through a three-dimensional Gaussian splash parameter generation model, and generating three-dimensional Gaussian splash parameters corresponding to the virtual object vertices according to the vertex features and the target object image.
[0110] In this embodiment, if Figure 4 As shown, when generating three-dimensional Gaussian splash parameters, first initialize the three-dimensional structured feature parameters of the preset three-dimensional virtual object model, such as Triplane and Triplane_face. These parameters describe the basic shape and structure of the body and face of the virtual object model. Then, the three-dimensional structured feature parameters are sampled using a grid based on vertex position encoding, so as to obtain the vertex features of each virtual object vertex of the preset three-dimensional virtual object model, which may specifically include information such as the position, shape, and relationship with surrounding vertices of the vertex. Then, a model such as MLPs is generated by three-dimensional Gaussian splash parameters, and corresponding three-dimensional Gaussian splash parameters (opacity, scale, center position, RGB value, etc.) are generated for each vertex of the virtual object model in combination with the vertex features and the target object image. These parameters describe the deformation, displacement, color, and other information of the vertex in three-dimensional space, which can make the virtual object model closer to the target object in shape and details. By generating three-dimensional Gaussian splash parameters for each vertex of the virtual object model, the model can be made more refined and realistic in shape and details, thereby better reflecting the characteristics of the target object.
[0111] In the embodiment of the present application, optionally, the body prior parameters in the prior parameters include body shape parameters, body posture parameters and joint offset parameters, and the face prior parameters in the prior parameters include face shape parameters and face offset parameters; Figure 5 As shown, the three-dimensional virtual object rendering is performed according to the target object prior parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object, including:
[0112] Step 501: Render the initialized virtual object prior model using the body shape parameters, the facial shape parameters, the facial offset parameters and the three-dimensional Gaussian splash parameters to obtain an initial three-dimensional virtual object model, and skin the initial three-dimensional virtual object model based on the body posture parameters and the joint offset parameters to obtain an intermediate three-dimensional virtual object model.
[0113] Step 502: perform two-dimensional projection on the intermediate three-dimensional virtual object model to obtain full projection information of the virtual object, perform loss calculation based on the full projection information of the virtual object and the target object image, and optimize the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model based on the second loss calculation result, until the three-dimensional virtual object model generated based on the optimized three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model meets the three-dimensional target object rendering conditions.
[0114] In this embodiment, the body prior parameters include body shape parameters SMPLX_Shape_params, body posture parameters SMPLX_Pose_params, and joint offset parameters SMPLX_Joint_offset, which together describe the body shape, posture, and joint position of the virtual object. The facial prior parameters include facial shape parameters FLAME_Shape_params and facial offset parameters FLAME_face_offset, which are used to describe the facial shape and details of the virtual object. During the rendering process, the initialized virtual object prior model is first rendered using the body shape parameters SMPLX_Shape_params, FLAME_Shape_params, facial offset parameters FLAME_face_offset, and three-dimensional Gaussian splash parameters to obtain an initial three-dimensional virtual object model. Then, based on the body posture parameters
[0115] SMPLX_Pose_params and joint offset parameters SMPLX_Joint_offset are used to skin the initial 3D virtual object model, that is, to adjust the vertex positions on the model surface to reflect the changes in body posture and joint positions, thereby obtaining an intermediate 3D virtual object model. Next, the intermediate 3D virtual object model is projected in two dimensions to obtain the full projection information of the virtual object, which contains all the features of the virtual object in two-dimensional space, such as shape, texture, color, etc. Then, a loss calculation is performed based on the full projection information of the virtual object and the target object image to obtain a second loss calculation result. This loss value reflects the degree of difference between the virtual object and the target object in two-dimensional space. Finally, the 3D structured feature parameters and the 3D Gaussian splash parameter generation model are optimized based on the second loss calculation result. This process is iterative until the 3D virtual object model generated based on the optimized 3D structured feature parameters and the 3D Gaussian splash parameter generation model meets the 3D target object rendering conditions, such as the similarity between the virtual object and the target object in terms of shape, texture, color, etc. reaches a preset threshold. The embodiments of the present application can generate a more realistic three-dimensional virtual object that conforms to the characteristics of the target object by finely controlling parameters such as body shape, posture, joint position, facial shape and details. By iteratively optimizing the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model, the generated three-dimensional virtual object is closer to the target object in shape, texture, color, etc., thereby improving the authenticity of the modeling.
[0116] In an embodiment of the present application, optionally, the loss calculation described in step 402 is performed based on the full projection information of the virtual object and the target object image, and the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model are optimized based on the second loss calculation result obtained, including: extracting the target object mask in the target object image; performing pixel value loss calculation, structural similarity loss calculation, and perceptual difference loss calculation based on the full projection information of the virtual object and the target object mask; optimizing the three-dimensional Gaussian splash parameter generation model based on pixel value loss, and optimizing the three-dimensional structured feature parameters of the preset three-dimensional virtual object model based on structural similarity loss and perceptual difference loss.
[0117] In this embodiment, during the loss calculation and optimization process of the preset three-dimensional virtual object model, the mask of the target object is first extracted from the target object image to identify the position and range of the target object in the image. Then, based on the full projection information of the virtual object and the target object mask, three types of loss calculations are performed: pixel value loss RGB_Loss calculation, structural similarity loss SSIM_Loss calculation, and perceptual difference loss lpip_Loss calculation. The pixel value loss calculation directly compares the difference in pixel value between the virtual object projection and the target object image; the structural similarity loss calculation evaluates the structural similarity between the two; the perceptual difference loss calculation takes into account the difference in human visual perception, which can be specifically evaluated by a deep learning model. Then, the model is optimized according to these loss values. Specifically, the three-dimensional Gaussian splash parameter generation model is optimized based on the pixel value loss to adjust the details and texture of the virtual object surface to make it closer to the target object. At the same time, the three-dimensional structured feature parameters of the preset three-dimensional virtual object model are optimized based on the structural similarity loss and the perceptual difference loss to adjust the overall shape and structure of the virtual object to make it more consistent with the characteristics of the target object. The embodiment of the present application can generate a more realistic three-dimensional virtual object that conforms to the characteristics of the target object by comprehensively considering multiple factors such as pixel values, structural similarities and perceptual differences.
[0118] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, a method for redirecting the action of a three-dimensional virtual object is provided, such as Figure 6 As shown, the method includes:
[0119] Step 601: Acquire a target object video, and extract a target object image sequence from the target object video.
[0120] Step 602: perform parameter estimation on each target object image sequence to obtain a body parameter sequence corresponding to each target object image sequence.
[0121] Step 603: redirect the target object's motion according to the body parameter sequence, the target object's prior parameters and the three-dimensional Gaussian splash parameters corresponding to the target object, wherein the target object prior parameters and the three-dimensional Gaussian splash parameters are obtained by the three-dimensional virtual object generation method described above.
[0122] In the above embodiment, based on the above three-dimensional virtual object generation method, as Figure 7As shown, the action of the virtual object model based on the target object can also be redirected based on the target object video (i.e., the guided motion video). Specifically, a video containing the target object action is obtained, and an image sequence of the target object is extracted. The image sequence is a series of continuous frames in the video, which together record the action process of the target object. Then, the parameters of the target object in each frame image are estimated, and at least the body posture parameters SMPLX Pose_params are obtained and a body parameter sequence is formed, which accurately describes the action changes of the target object in the video. Further, the three-dimensional virtual object generation method described above is used to generate corresponding target object prior parameters and three-dimensional Gaussian splash parameters for the target object. According to the body parameter sequence and these prior parameters, the action of the target object is redirected, and the action of the target object in the video is mapped to the three-dimensional virtual object, so that the three-dimensional virtual object can simulate the action of the target object. In this process, the three-dimensional Gaussian splash parameters can enable the three-dimensional virtual object to maintain details and texture changes similar to the target object during the action simulation process. Furthermore, since motion redirection usually involves the conversion between different body shapes and motion styles, the target object prior parameters can better reflect the target object's body characteristics and motion patterns, thus achieving motion redirection more accurately.
[0123] In an embodiment of the present application, optionally, the body parameter sequence includes a body posture parameter sequence, and the target object prior parameters include target object body shape parameters, target object joint offset parameters, target object facial shape parameters and target object facial offset parameters; step 603 includes: rendering a preset virtual object prior model according to the target object body shape parameters, the target object facial shape parameters, the target object facial offset parameters and the three-dimensional Gaussian splash parameters, and skinning the rendered virtual object prior model based on the body posture parameter sequence and the target object joint offset parameters to obtain a motion redirection model corresponding to the target object.
[0124] In this embodiment, if Figure 7As shown, the body parameter sequence includes a body posture parameter sequence, which describes the movement changes of the target object in the video. The target object prior parameters include the target object body shape parameters SMPLX_Shape_params, which describe the body shape characteristics of the target object, such as height, body shape, etc., and also include the target object joint offset parameters SMPLX_Joint_offset, which describe the small offset of the target object joint relative to the standard position, and also include the target object facial shape parameters FLAME_Shape_params, which describe the facial shape characteristics of the target object, such as face shape, distribution of facial features, etc., and also include the target object facial offset parameters FLAME_face_offset, which describe the small offset of the target object facial features relative to the standard position. Using the target object body shape parameters SMPLX_Shape_params, the target object facial shape parameters FLAME_Shape_params and the target object facial offset parameters FLAME_face_offset, as well as the three-dimensional Gaussian splash parameters 3DGS, the preset virtual object prior model is rendered to generate a preliminary three-dimensional virtual object, whose shape, texture and details are similar to the target object. Further, based on the body posture parameter sequence and the target object joint offset parameter SMPLX_Joint_offset, the rendered virtual object prior model is skinned, thereby associating the surface vertices of the three-dimensional model with the bones or joints, so that when the bones or joints move, the surface vertices will move accordingly, thereby simulating the real action effect. In this process, the body posture parameter sequence describes the action changes of the target object, and the target object joint offset parameter is used to adjust the joint position to ensure that the action redirected model is consistent with the target object in both action and form. After rendering and skinning, a three-dimensional virtual object that is highly similar to the target object in shape, texture, action and details is finally obtained, namely the action redirection model.
[0125] In a specific application scenario, the redirection process of a three-dimensional virtual object includes: S1. Obtaining the prior parameters corresponding to the target object, including the target object body shape parameter SMPLX_Shape_params, the target object joint offset parameter SMPLX_Joint_offset, the target object facial shape parameter FLAME_Shape_params, and the target object facial offset parameter FLAME_face_offset. S2. Obtaining the three-dimensional Gaussian splash parameters 3DGS corresponding to the target object. S3. Obtaining a guided motion video containing the target object, extracting the image sequence corresponding to the video, performing parameter estimation on each image in the image sequence, obtaining the body posture parameter SMPLX Pose_params of each image, and forming a body posture parameter sequence. S4. Render the preset virtual object prior model according to the target object body shape parameters SMPLX_Shape_params, the target object facial shape parameters FLAME_Shape_params, the target object facial offset parameters FLAME_face_offset and the three-dimensional Gaussian splash parameters 3DGS, and skin the rendered virtual object prior model based on the body posture parameter SMPLX Pose_params sequence and the target object joint offset parameter SMPLX_Joint_offset to obtain the action redirection model corresponding to the target object.
[0126] Specifically, in S1, first, a target object image containing a target object is obtained, FLAME facial prior parameter estimation is performed on the target object image, and SMPLX body prior parameter estimation is performed on the target object image, and the facial parameters in SMPLX are replaced with the FLAME facial parameter estimation results, that is, the FALME facial parameter estimation results are used as the target object facial parameters, and the SMPLX body parameter estimation results are used as the target object body parameters. Then, initialize the prior parameters of the virtual object prior model, the prior parameters including body prior parameters (including body shape parameters SMPLX_Shape_params, body posture parameters SMPLX_Pose_params, joint offset parameters SMPLX_Joint_offset) and face prior parameters (including face shape parameters FLAME_Shape_params, face offset parameters FLAME_face_offset); input the target object face parameters and the target object body parameters into the virtual object prior model to obtain the three-dimensional feature information of the virtual object; perform two-dimensional projection on the three-dimensional feature information of the virtual object to obtain the virtual object feature projection information; extract the facial key point coordinates and joint bone point coordinates corresponding to the target object image; calculate the facial loss based on the facial key point coordinates and the facial key point projection coordinates in the virtual object feature projection information, and calculate the joint loss based on the joint bone point coordinates and the joint bone point projection coordinates in the virtual object feature projection information; optimize the facial prior parameters based on the facial loss, and optimize the body prior parameters based on the joint loss. Finally, the prior parameters in the trained virtual object prior model are obtained as the target object prior parameters (including the optimized target object body shape parameters SMPLX_Shape_params, target object body posture parameters SMPLX_Pose_params, target object joint offset parameters SMPLX_Joint_offset, target object facial shape parameters FLAME_Shape_params, target object facial offset parameters FLAME_face_offset).
[0127] In S2, first, the three-dimensional structured feature parameters of the preset three-dimensional virtual object model are initialized, and the three-dimensional structured feature parameters are sampled according to the grid based on vertex position encoding corresponding to the preset three-dimensional virtual object model to obtain the vertex features of the virtual object vertices corresponding to the preset three-dimensional virtual object model; through the three-dimensional Gaussian splash parameter generation model, the three-dimensional Gaussian splash parameters 3DGS corresponding to the virtual object vertices are generated according to the vertex features and the target object image. Secondly, the initialized virtual object prior model is rendered using the target object body shape parameters SMPLX_Shape_params, the target object facial shape parameters FLAME_Shape_params, the target object facial offset parameters FLAME_face_offset and the three-dimensional Gaussian splash parameters 3DGS to obtain an initial three-dimensional virtual object model, and the initial three-dimensional virtual object model is skinned based on the target object body posture parameters SMPLX_Pose_params and the target object joint offset parameters SMPLX_Joint_offset to obtain an intermediate three-dimensional virtual object model; the intermediate three-dimensional virtual object model is two-dimensionally projected to obtain full projection information of the virtual object, a loss calculation is performed based on the full projection information of the virtual object and the target object image, and the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model are optimized based on the second loss calculation result, until the three-dimensional virtual object model generated based on the optimized three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model meets the three-dimensional target object rendering conditions. Among them, during the optimization process, the target object mask in the target object image is extracted; based on the full projection information of the virtual object and the target object mask, pixel value loss calculation, structural similarity loss calculation, and perceptual difference loss calculation are performed; based on the pixel value loss, the three-dimensional Gaussian splash parameter generation model is optimized, and based on the structural similarity loss and perceptual difference loss, the three-dimensional structured feature parameters of the preset three-dimensional virtual object model are optimized.
[0128] In S3, when estimating the parameters of each image in the image sequence, the method of S1 can be adopted, and the target object image in S1 is replaced with the image in the image sequence, and then the steps of S1 are executed to obtain the body posture parameters SMPLX_Pose_params corresponding to each image, and further construct a body posture parameter SMPLX_Pose_params sequence. In addition, it should be noted that since only the body posture parameters need to be obtained in S3, when estimating the parameters of the image, only SMPLX human body prior parameters can be estimated, and FLAME facial prior parameters can be not estimated to obtain facial parameters and body parameters. After that, when training the model based on the obtained facial parameters and body parameters, only joint loss can be calculated, and facial loss can be not calculated, thereby reducing the amount of calculation and improving calculation efficiency.
[0129] Further, as Figure 1 The specific implementation of the method, the embodiment of the present application provides a device for generating a three-dimensional virtual object, such as Figure 8 As shown, the device comprises:
[0130] An image acquisition module 801 is used to acquire a target object image, wherein the target object image includes a target object having a face part and a body part;
[0131] A priori parameter determination module 802 is used to perform parameter estimation on the target object image to obtain target object facial parameters and target object body parameters, and train a virtual object priori model based on the target object facial parameters and the target object body parameters to obtain target object prior parameters corresponding to the trained virtual object prior model;
[0132] A Gaussian splash parameter determination module 803 is used to generate three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image;
[0133] The virtual object generation module 804 is used to perform three-dimensional virtual object rendering according to the target object prior parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object.
[0134] In an optional implementation, the priori parameter determination module 802 is further configured to:
[0135] Performing FLAME prior parameter estimation on the target object image to obtain first prior parameters, wherein the first prior parameters include first facial prior parameters; and performing SMPLX prior parameter estimation on the target object image to obtain second prior parameters, wherein the second prior parameters include second facial prior parameters and joint point parameters;
[0136] The second facial priori parameters in the second priori parameters are replaced with the first facial priori parameters, and the facial parameters and the body parameters of the target object are determined based on the replaced second priori parameters.
[0137] In an optional implementation, the priori parameter determination module 802 is further configured to:
[0138] Initializing the prior parameters of the virtual object prior model, wherein the prior parameters include body prior parameters and face prior parameters;
[0139] Inputting the target object facial parameters and the target object body parameters into the virtual object prior model to obtain the virtual object three-dimensional feature information;
[0140] Performing a two-dimensional projection on the three-dimensional feature information of the virtual object to obtain virtual object feature projection information, performing loss calculation based on the virtual object feature projection information and target object feature point information corresponding to the target object image, and optimizing the prior parameters based on the obtained first loss calculation result until the model training condition is met and the training is completed;
[0141] The prior parameters in the trained virtual object prior model are obtained as the target object prior parameters.
[0142] In an optional implementation, the priori parameter determination module 802 is further configured to:
[0143] Extracting the facial key point coordinates and joint bone point coordinates corresponding to the target object image;
[0144] Performing facial loss calculation based on the facial key point coordinates and the facial key point projection coordinates in the virtual object feature projection information, and performing joint loss calculation based on the joint bone point coordinates and the joint bone point projection coordinates in the virtual object feature projection information;
[0145] The facial prior parameters are optimized based on the facial loss, and the body prior parameters are optimized based on the joint loss.
[0146] In an optional implementation, the Gaussian splash parameter determination module 803 is further used to:
[0147] Initializing three-dimensional structured feature parameters of the preset three-dimensional virtual object model, sampling the three-dimensional structured feature parameters according to a grid based on vertex position coding corresponding to the preset three-dimensional virtual object model, and obtaining vertex features of virtual object vertices corresponding to the preset three-dimensional virtual object model;
[0148] A three-dimensional Gaussian splash parameter generation model is used to generate three-dimensional Gaussian splash parameters corresponding to the vertices of the virtual object according to the vertex features and the target object image.
[0149] In an optional embodiment, the body prior parameters in the prior parameters include body shape parameters, body posture parameters and joint offset parameters, and the face prior parameters in the prior parameters include face shape parameters and face offset parameters;
[0150] The virtual object generation module 804 is further used for:
[0151] Rendering the initialized virtual object prior model using the body shape parameters, the facial shape parameters, the facial offset parameters and the three-dimensional Gaussian splash parameters to obtain an initial three-dimensional virtual object model, and skinning the initial three-dimensional virtual object model based on the body posture parameters and the joint offset parameters to obtain an intermediate three-dimensional virtual object model;
[0152] Perform two-dimensional projection on the intermediate three-dimensional virtual object model to obtain full projection information of the virtual object, perform loss calculation based on the full projection information of the virtual object and the target object image, and optimize the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model based on the obtained second loss calculation result until the three-dimensional virtual object model generated based on the optimized three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model meets the three-dimensional target object rendering conditions.
[0153] In an optional implementation, the virtual object generation module 804 is further configured to:
[0154] extracting a target object mask from the target object image;
[0155] Based on the full projection information of the virtual object and the target object mask, performing pixel value loss calculation, structural similarity loss calculation, and perceptual difference loss calculation;
[0156] The three-dimensional Gaussian splash parameter generation model is optimized based on pixel value loss, and the three-dimensional structured feature parameters of the preset three-dimensional virtual object model are optimized based on structural similarity loss and perceptual difference loss.
[0157] Further, as Figure 6 The specific implementation of the method, the embodiment of the present application provides a redirection device for a three-dimensional virtual object, such as Fig. 9 As shown, the device comprises:
[0158] The video acquisition module 901 is used to acquire a target object video and extract a target object image sequence in the target object video;
[0159] A sequence determination module 902 is used to perform parameter estimation on each target object image sequence to obtain a body parameter sequence corresponding to each target object image sequence;
[0160] The redirection module 903 is used to redirect the action of the target object according to the body parameter sequence, the target object prior parameters corresponding to the target object and the three-dimensional Gaussian splash parameters, wherein the target object prior parameters and the three-dimensional Gaussian splash parameters are obtained by the three-dimensional virtual object generation device as described in claim 10.
[0161] In an optional embodiment, the body parameter sequence includes a body posture parameter sequence, and the target object prior parameters include target object body shape parameters, target object joint offset parameters, target object facial shape parameters and target object facial offset parameters; the redirection module 903 is also used to: render a preset virtual object prior model according to the target object body shape parameters, the target object facial shape parameters, the target object facial offset parameters and the three-dimensional Gaussian splash parameters, and skin the rendered virtual object prior model based on the body posture parameter sequence and the target object joint offset parameters to obtain an action redirection model corresponding to the target object.
[0162] It should be noted that for other corresponding descriptions of the functional units involved in the three-dimensional virtual object generation device and redirection device provided in the embodiment of the present application, reference can be made to Figures 1 to 7 The corresponding description in the method will not be repeated here.
[0163] The embodiment of the present application also provides a computer device, which can be a personal computer, a server, a network device, etc. The computer device includes a bus, a processor, a memory and a communication interface, and can also include an input and output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.
[0164] Those skilled in the art will appreciate that the structure of the above-mentioned computer device is only a partial structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components, or combine certain components, or have a different arrangement of components.
[0165] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0166] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0167] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0168] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0169] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0170] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for generating a three-dimensional virtual object, characterized in that: The method comprises: Acquire a target object image, wherein the target object image includes a target object having a face portion and a body portion; Performing parameter estimation on the target object image to obtain target object facial parameters and target object body parameters, and training a virtual object prior model based on the target object facial parameters and the target object body parameters to obtain target object prior parameters corresponding to the trained virtual object prior model; Generate three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image; A three-dimensional virtual object rendering is performed according to the target object priori parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object.
2. The method according to claim 1, characterized in that The step of performing parameter estimation on the target object image to obtain target object facial parameters and target object body parameters includes: Performing FLAME prior parameter estimation on the target object image to obtain first prior parameters, wherein the first prior parameters include first facial prior parameters; and performing SMPLX prior parameter estimation on the target object image to obtain second prior parameters, wherein the second prior parameters include second facial prior parameters and joint point parameters; The second facial priori parameters in the second priori parameters are replaced with the first facial priori parameters, and the facial parameters and the body parameters of the target object are determined based on the replaced second priori parameters.
3. The method according to claim 1, characterized in that The step of training the virtual object prior model based on the target object facial parameters and the target object body parameters to obtain the target object prior parameters corresponding to the trained virtual object prior model includes: Initializing the prior parameters of the virtual object prior model, wherein the prior parameters include body prior parameters and face prior parameters; Inputting the target object facial parameters and the target object body parameters into the virtual object prior model to obtain the virtual object three-dimensional feature information; Performing a two-dimensional projection on the three-dimensional feature information of the virtual object to obtain virtual object feature projection information, performing loss calculation based on the virtual object feature projection information and target object feature point information corresponding to the target object image, and optimizing the prior parameters based on the obtained first loss calculation result until the model training condition is met and the training is completed; The prior parameters in the trained virtual object prior model are obtained as the target object prior parameters.
4. The method according to claim 3, characterized in that The performing loss calculation based on the virtual object feature projection information and the target object feature point information corresponding to the target object image, and optimizing the priori parameters based on the obtained first loss calculation result, includes: Extracting the facial key point coordinates and joint bone point coordinates corresponding to the target object image; Performing facial loss calculation based on the facial key point coordinates and the facial key point projection coordinates in the virtual object feature projection information, and performing joint loss calculation based on the joint bone point coordinates and the joint bone point projection coordinates in the virtual object feature projection information; The facial prior parameters are optimized based on the facial loss, and the body prior parameters are optimized based on the joint loss.
5. The method according to claim 1, characterized in that The generating of three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image includes: Initializing three-dimensional structured feature parameters of the preset three-dimensional virtual object model, sampling the three-dimensional structured feature parameters according to a grid based on vertex position coding corresponding to the preset three-dimensional virtual object model, and obtaining vertex features of virtual object vertices corresponding to the preset three-dimensional virtual object model; A three-dimensional Gaussian splash parameter generation model is used to generate three-dimensional Gaussian splash parameters corresponding to the vertices of the virtual object according to the vertex features and the target object image.
6. The method according to claim 5, characterized in that The body prior parameters in the prior parameters include body shape parameters, body posture parameters and joint offset parameters, and the face prior parameters in the prior parameters include face shape parameters and face offset parameters; The performing three-dimensional virtual object rendering according to the target object prior parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object includes: Rendering the initialized virtual object prior model using the body shape parameters, the facial shape parameters, the facial offset parameters and the three-dimensional Gaussian splash parameters to obtain an initial three-dimensional virtual object model, and skinning the initial three-dimensional virtual object model based on the body posture parameters and the joint offset parameters to obtain an intermediate three-dimensional virtual object model; Perform two-dimensional projection on the intermediate three-dimensional virtual object model to obtain full projection information of the virtual object, perform loss calculation based on the full projection information of the virtual object and the target object image, and optimize the three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model based on the obtained second loss calculation result until the three-dimensional virtual object model generated based on the optimized three-dimensional structured feature parameters and the three-dimensional Gaussian splash parameter generation model meets the three-dimensional target object rendering conditions.
7. The method according to claim 6, characterized in that The performing loss calculation based on the full projection information of the virtual object and the target object image, and optimizing the three-dimensional structural feature parameters and the three-dimensional Gaussian splash parameter generation model based on the obtained second loss calculation result, includes: extracting a target object mask from the target object image; Based on the full projection information of the virtual object and the target object mask, performing pixel value loss calculation, structural similarity loss calculation, and perceptual difference loss calculation; The three-dimensional Gaussian splash parameter generation model is optimized based on pixel value loss, and the three-dimensional structured feature parameters of the preset three-dimensional virtual object model are optimized based on structural similarity loss and perceptual difference loss.
8. A device for generating a three-dimensional virtual object, characterized in that: The device comprises: An image acquisition module, used for acquiring a target object image, wherein the target object image includes a target object having a face part and a body part; a priori parameter determination module, used to perform parameter estimation on the target object image to obtain target object facial parameters and target object body parameters, and train a virtual object priori model based on the target object facial parameters and the target object body parameters to obtain target object priori parameters corresponding to the trained virtual object prior model; A Gaussian splash parameter determination module, configured to generate three-dimensional Gaussian splash parameters for virtual object vertices of a preset three-dimensional virtual object model based on the target object image; A virtual object generation module is used to perform three-dimensional virtual object rendering according to the target object prior parameters and the three-dimensional Gaussian splash parameters to obtain a three-dimensional virtual object based on the target object.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Virtual digital human generation method and electronic equipment
CN121353562A