Method, apparatus and computer program product for generating three-dimensional object reconstruction model

By determining the characteristics of the two-dimensional face image and generating a rendered image comparison training model, combined with stylized parameter adjustment, the problem of poor three-dimensional face reconstruction in the existing technology is solved, and a three-dimensional face reconstruction with high credibility and diverse styles is achieved.

CN120388128APending Publication Date: 2025-07-29DELL PROD LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410114542.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When existing three-dimensional face reconstruction technology deals with factors such as camera parameters, lighting conditions and background clutter, the reconstruction effect is limited, making it difficult to generate high-confidence and high-quality three-dimensional face models.

Method used

By determining the shape characteristics, texture characteristics and pose characteristics of the two-dimensional face image, a rendered image is generated, and the three-dimensional object reconstruction model is trained by comparing the rendered image with the input image, and fine-tuning the parameters with the stylized parameter adjustment model to generate a realistic and diverse 3-dimensional face model.

Benefits of technology

It realizes high credibility and high-quality three-dimensional face reconstruction that highly matches the input image in terms of shape, texture and posture, while giving the face a specific style, such as sketch, cartoon or oil painting style, improving the variability and expressiveness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388128A_ABST
    Figure CN120388128A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a device and a computer program product for generating a three-dimensional object reconstruction model. The method includes acquiring an input image including a two-dimensional object. The method further includes determining a shape feature, a texture feature, and a pose feature of the two-dimensional object. The method further includes generating a rendered image based on the shape feature, the texture feature, the pose feature, and the input image. The method further includes generating a three-dimensional object reconstruction model based on the input image and the rendered image. The method further includes adjusting the three-dimensional object reconstruction model according to the parameter adjustment model for stylization. The three-dimensional object reconstruction model generated by the method can truly reconstruct a two-dimensional object in the aspects of shape, texture, posture and the like, so that the three-dimensional object output by the model can be matched with an input image in each dimension, and the three-dimensional object with high credibility and high quality is obtained; and a three-dimensional object reconstruction model with diversified styles can be generated in combination with knowledge in a specific style field, so that the variability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of object reconstruction, and more particularly, to methods, devices, and computer program products for generating three-dimensional object reconstructions. Background Art

[0002] With the development of technology, three-dimensional face reconstruction technology has gradually become a popular technology in the field of computer vision, and has a wide range of applications in fields such as computer graphics, gaming, virtual reality, biometrics, and human-computer interaction. At the same time, three-dimensional face reconstruction technology is of great significance for promoting the development of related fields of cognitive science, physiology, and psychology based on face recognition.

[0003] Three-dimensional face reconstruction technology refers to reconstructing a three-dimensional face model of a measured individual based on one or more face images of the individual. During the three-dimensional face reconstruction process, various factors of the face image, such as camera parameters, lighting conditions, background clutter, and compression artifacts, etc., may affect the reconstructed three-dimensional face. Summary of the Invention

[0004] Embodiments of the present disclosure propose a method, device, and computer program product for generating a three-dimensional object reconstruction model. In a first aspect of the embodiments of the present disclosure, a method for generating a three-dimensional object reconstruction model is provided. The method includes obtaining an input image including a two-dimensional object. The method further includes determining the shape feature, texture feature, and pose feature of the two-dimensional object. The method further includes generating a rendered image based on the shape feature, texture feature, pose feature, and the input image. The method further includes generating a three-dimensional object reconstruction model based on the input image and the rendered image. The method further includes adjusting the three-dimensional object reconstruction model according to a tuning parameter model for stylization.

[0005] In a second aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement a method for monitoring a distributed system, the method including obtaining an input image including a two-dimensional object. The method further includes determining the shape feature, texture feature, and pose feature of the two-dimensional object. The method further includes generating a rendered image based on the shape feature, texture feature, pose feature, and the input image. The method further includes generating a three-dimensional object reconstruction model based on the input image and the rendered image. The method further includes adjusting the three-dimensional object reconstruction model according to a tuning parameter model for stylization.

[0006] In a third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, a method for monitoring a distributed system is implemented. The method includes obtaining an input image including two-dimensional objects. The method further includes determining shape features, texture features, and pose features of the two-dimensional objects. The method further includes generating a rendered image based on the shape features, texture features, pose features, and the input image. The method further includes generating a three-dimensional object reconstruction model based on the input image and the rendered image. The method further includes adjusting the three-dimensional object reconstruction model according to a tuning model for stylization.

[0007] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 A schematic diagram of an example environment in which multiple embodiments of the present disclosure can be implemented is shown;

[0010] Figure 2 A flowchart of a method for generating a three-dimensional object reconstruction model according to some embodiments of the present disclosure is shown;

[0011] Figure 3 A schematic diagram of a process for determining shape features, texture features, and pose features according to some embodiments of the present disclosure is shown;

[0012] Figure 4 A schematic diagram of a process for generating a reconstruction loss, an expression loss, and a regularization loss according to some embodiments of the present disclosure is shown;

[0013] Figure 5 A schematic diagram of a process for adjusting a three-dimensional object reconstruction model according to some embodiments of the present disclosure is shown; and

[0014] Figure 6 A block diagram of a device that can implement multiple embodiments of the present disclosure is shown.

[0015] In all the drawings, the same or similar reference numerals represent the same or similar elements. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0017] In the description of the embodiments of the present disclosure, the term "comprising" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless otherwise specified. There may also be other explicit and implicit definitions hereinafter.

[0018] Three-dimensional objects, such as human faces, have highly variable and complex facial shapes and textures, depending on factors such as a person's identity, expression, pose, age, gender, and race, etc. Generally speaking, three-dimensional object reconstruction can be achieved based on deep learning methods. During the reconstruction process, a convolutional neural network can be used to reconstruct a two-dimensional object image into a three-dimensional object model. In this method, the convolutional neural network is trained using a large dataset of three-dimensional real object models and then fine-tuned on real object models with sparse landmarks. This method can process object images with arbitrary poses and expressions, but it requires aligning the object and the landmarks, and the ground truth three-dimensional facial shapes are scarce and expensive, requiring professional equipment such as laser scanners and structured light systems, etc.

[0019] Another method uses an autoencoder to learn the low-dimensional latent space of the three-dimensional object shape, but this method depends on a fixed camera model and lighting conditions. There is also a method that uses a generative adversarial network for three-dimensional object reconstruction. The generative adversarial network is trained on a real object dataset and self-supervised in terms of multi-view Figure 1 consistency and perceptual loss, but this method requires multiple views of the same object for training to generate high-quality three-dimensional objects. Therefore, so far, three-dimensional objects have been limited in terms of expressiveness and photorealism, relying on real image factors, fixed templates, and prior knowledge.

[0020] To this end, embodiments of the present disclosure provide a solution for generating a three-dimensional object reconstruction model. The solution is to determine the shape features, texture features, and pose features of a two-dimensional object in an input image, generate a rendered image based on the shape features, texture features, and pose features of the two-dimensional object, and generate a three-dimensional object reconstruction model by comparing the rendered image with the input image. The three-dimensional object reconstruction model generated by this method can truly reconstruct the two-dimensional object in terms of shape, texture, and pose, so that the three-dimensional object output by the model can match the input image in each dimension, obtaining a three-dimensional object with high credibility and quality. In addition, the solution also fine-tunes the three-dimensional object reconstruction model through a stylized parameter adjustment model. In this way, a three-dimensional object reconstruction model with diverse styles can be generated by combining the knowledge of specific style fields, improving the variability of the generated model.

[0021] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which multiple embodiments of the present disclosure can be implemented. As Figure 1 shown, the exemplary environment 100 may include an input image 101 and a three-dimensional object reconstruction model 104. The input image 101 includes a two-dimensional object. Specifically, when implemented, the input to the three-dimensional object reconstruction model 104 can be multiple images or a video. The three-dimensional object reconstruction model 104 is used to convert the two-dimensional object in the input image 101 into a three-dimensional object 102 that matches it. The object can be a human face. The three-dimensional object reconstruction model 104 can be a machine learning model. A machine learning model is a computational model that has a certain ability after learning from samples. Specifically, it can be a neural network, such as a CNN (Convolutional Neural Networks), an RNN (Recurrent Neural Networks), etc. Of course, the machine learning model can also adopt other types of models. Here, only the architecture and functions in the exemplary environment 100 are described for illustrative purposes.

[0022] In some embodiments, the 3D object reconstruction model 104 may include shape parameters, texture parameters, and pose parameters. After the input image 101 is input into the 3D object reconstruction model 104, the 3D object reconstruction model 104 can determine the shape features, texture features, and pose features in the input image 101. During the training process of the 3D object reconstruction model 104, a rendered image with realistic lighting and shadow effects can be generated based on the shape features, texture features, and pose features. The rendered image can be generated using methods such as rasterization, ray tracing, and neural rendering, and can be specifically selected according to actual needs. By comparing the rendered image with the input image 101, the pixel-level intensity difference between the two images can be obtained, thereby optimizing the shape parameters, texture parameters, and pose parameters in the 3D object reconstruction model 104 to obtain a 3D object reconstruction model 104 with high reconstruction quality, strong expression ability, and high credibility.

[0023] Reference Figure 1 , the example environment 100 may further include a parameter tuning model 103. After the 3D object reconstruction model 104 is trained, the parameter tuning model 103 for stylization is used to fine-tune the shape parameters, texture parameters, and pose parameters in the 3D object reconstruction model 104, so that the 3D object reconstruction model 104 can not only reconstruct a 3D object with high credibility, but also endow the object with a preset style, such as a sketch style, a cartoon character style, and an oil painting style, etc. The purpose of stylization can be achieved by allowing the parameter tuning model 103 to pre-learn images or videos of the required style.

[0024] As can be seen from the above description, the solution of the present disclosure generates a rendered image based on the shape features, texture features, and pose features of a 2D object, and trains the 3D object reconstruction model by comparing the rendered image with the input image. The 3D object reconstruction model generated by this method can truly reconstruct the 2D object in terms of shape, texture, and pose, so that the 3D object output by the model can match the input image in each dimension, obtaining a 3D object with high credibility and high quality. In addition, the solution also fine-tunes the 3D object reconstruction model through a stylized parameter tuning model. In this way, a 3D object reconstruction model with diverse styles can be generated by combining the knowledge in a specific style field, improving the variability of the generated model.

[0025] It should be understood that the architecture and functions in the example environment 100 are described only for exemplary purposes, and do not imply any limitation on the scope of the present disclosure. The embodiments of the present disclosure can also be applied to other environments with different structures and / or functions.

[0026] The following will be combined with Figures 2 to 6Describe the process of the embodiments of the present disclosure in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and do not limit the protection scope of the present disclosure. It can be understood that the embodiments described below may further include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.

[0027] Figure 2 FIG. 200 is a flowchart of a method for generating a three-dimensional object reconstruction model according to some embodiments of the present disclosure. At block 202, an input image including a two-dimensional object is obtained. For example, as Figure 1 shown, the input image 101 can be used as training data for the three-dimensional object reconstruction model 104, and the two-dimensional object can be a human face. The input image 101 can be a plurality of images, which can be multiple images including two-dimensional human faces at different angles and poses, or images generated under different lighting conditions and different camera parameter conditions. In the embodiments of the present disclosure, there is no limitation on the environmental conditions when generating the input image 101, nor on whether the two-dimensional object in the input image 101 is occluded. Shape aggregation can also be performed using complementary information from different images to perform multi-image object reconstruction.

[0028] At block 204, the shape feature, texture feature, and pose feature of the two-dimensional object are determined. For example, as Figure 1 shown, the shape feature, texture feature, and pose feature of the two-dimensional object in the input image 101 are determined, and the determined shape feature, texture feature, and pose feature are used as training data for the three-dimensional object reconstruction model 104, so that the trained three-dimensional object reconstruction model 104 can reconstruct a three-dimensional object that matches the input two-dimensional object in multiple dimensions such as shape, texture, and pose, improving the reconstruction quality, expression ability, and credibility of the three-dimensional object reconstruction model 104.

[0029] At block 206, a rendered image is generated based on the shape feature, texture feature, pose feature, and the input image. For example, as Figure 1 shown, during the training process of the three-dimensional object reconstruction model 104, a rendered image with realistic lighting and shadow effects can be generated according to the shape feature, texture feature, and pose feature. The rendered image can be generated using methods such as rasterization, ray tracing, and neural rendering, and can be specifically selected according to actual needs.

[0030] At block 208, a three-dimensional object reconstruction model is generated based on the input image and the rendered image. For example, as Figure 1As shown, by comparing the rendered image with the input image 101, the pixel-level intensity difference between the two images can be obtained. By minimizing the intensity difference, the shape parameters, texture parameters, and pose parameters in the 3D object reconstruction model 104 are adjusted so that the 3D object 102 output by the 3D object reconstruction model 104 can match the input image 101 in multiple dimensions such as shape, texture, and pose. That is to say, the 3D object reconstruction model 104 trained by the method of the present disclosure can generate a more realistic and expressive 3D object.

[0031] At block 210, the 3D object reconstruction model is adjusted according to the tuning parameter model for stylization. For example, as Figure 1 shown, in order to improve the expressiveness and variability of the 3D object reconstruction model 104, the present disclosure uses the tuning parameter model 103 to adjust the parameters in the 3D object reconstruction model 104 for the purpose of stylization. The present disclosure uses the tuning parameter model 103 for stylization to finely tune the shape parameters, texture parameters, and pose parameters in the 3D object reconstruction model 104 so that the 3D object reconstruction model 104 can not only reconstruct a 3D object with high confidence but also endow the object with a preset style, such as a sketch style, a cartoon character style, and an oil painting style, etc. The purpose of stylization can be achieved by allowing the tuning parameter model 103 to pre-learn images or videos of the required style.

[0032] In this way, the generated 3D object reconstruction model can truly reconstruct the 2D object in terms of shape, texture, and pose, so that the 3D object output by the model can match the input image in each dimension, and a 3D object with high confidence and high quality can be obtained. In addition, the present disclosure also finely tunes the 3D object reconstruction model through the stylized tuning parameter model. In this way, a 3D object reconstruction model with diverse styles can be generated by combining the knowledge in a specific style field, and the variability and expressiveness of the generation model can be improved.

[0033] The following will be combined with Figures 3 to 6 Specifically, the process of generating the 3D object reconstruction model will be described. In the embodiments of the present disclosure, the explanation will be carried out in the order of determining features, determining the model loss, and stylizing the model. Figures 3 to 4 shows a schematic diagram of generating the 3D object reconstruction model. Figure 5 shows a schematic diagram of stylizing the 3D object reconstruction model. The specific data mentioned in the following description are all exemplary and are not used to limit the protection scope of the present disclosure. It can be understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.

[0034] Figure 3FIG. 300 is a schematic diagram showing the process of determining shape features, texture features, and pose features of some embodiments of the present disclosure. After the input image 301 is input into the three-dimensional object reconstruction model 303, the three-dimensional object reconstruction model 303 generates a three-dimensional object 302. In order to make the three-dimensional object 302 match the two-dimensional object in the input image 301 as much as possible, the three-dimensional object reconstruction model 303 needs to be trained. In some embodiments, the shape feature 304, texture feature 305, and pose feature 306 of the two-dimensional object in the input image 301 can be determined, and the determined shape feature 304, texture feature 305, and pose feature 306 are used as the training data of the three-dimensional object reconstruction model 303.

[0035] In the embodiments of the present disclosure, the shape feature 304 can be determined by the average object shape feature, the identity feature of the two-dimensional object, and the expression feature:

[0036]

[0037] where S represents the shape feature, represents the average object shape feature, αi represents the identity parameter, αe represents the expression parameter, Ai represents the identity basis matrix with Ki columns, and Ae represents the expression basis matrix with Ke columns. The shape parameters in the three-dimensional object reconstruction model 303 can include the expression parameter αe and the identity parameter αi, and the average object shape feature specifically refers to the average shape feature learned from a large number of three-dimensional objects.

[0038] The texture feature 305 can be determined according to the average object texture feature and the identity feature of the two-dimensional object:

[0039]

[0040] where T represents the texture feature, represents the average object texture feature, Bi represents the identity basis matrix, and βi represents the texture parameter. The average object texture feature specifically refers to the average texture feature learned from a large number of three-dimensional objects.

[0041] The pose feature 306 can be determined according to the rotation feature, scaling feature, and translation feature of the two-dimensional object:

[0042] R = sR x (θ x )R y (θ y )R z (θ z ) + t (3)

[0043] Among them, R represents the pose feature, Rx represents the rotation matrix around the x-axis, Ry represents the rotation matrix around the y-axis, Rz represents the rotation matrix around the z-axis, θx represents the rotation angle around the x-axis, θy represents the rotation angle around the y-axis, θz represents the rotation angle around the z-axis, s represents the scaling factor, and t represents the translation feature.

[0044] Figure 4 FIG. 400 is a schematic diagram showing a process of generating a reconstruction loss, an expression loss, and a regularization loss according to some embodiments of the present disclosure. In some embodiments, after obtaining the input image 401, the three-dimensional object reconstruction model 403 generates a three-dimensional object 402. In order to make the three-dimensional object 402 match the two-dimensional object in the input image 401 as much as possible, it is necessary to train the three-dimensional object reconstruction model 403. After determining the shape feature 404, the texture feature 405, and the pose feature 406, the three-dimensional object reconstruction model 403 can be trained according to the reconstruction loss 407 for measuring the pixel intensity difference, the expression loss 408 for measuring the emotional expression, and the regularization loss 409 for measuring the object rationality.

[0045] In some embodiments, the rendered image can be determined according to Equation (4):

[0046] I r = Render(I i , S, T, R) (4)

[0047] Among them, Ir represents the rendered image, Render represents the rendering function, Ii represents the input image, S represents the shape feature, T represents the texture feature, and R represents the pose feature. The reconstruction loss 407 can be determined according to the rendered image Ir and the input image Ii:

[0048] L r = λ1|I - I r |1 + λ2|F(I) - F(I r )|1 (5)

[0049] Among them, Lr represents the reconstruction loss, λ1 represents the first weight, λ2 represents the second weight, and F represents feature extraction. In specific implementation, a feature extractor for calculating high-level features in an image or video can be used to extract the input feature F(Ii) in the input image Ii and the rendering feature F(Ir) of the rendered image Ir. The reconstruction loss Lr enables the three-dimensional object reconstruction model 403 to use both low-level and perceptual-level information for supervision during training, so as to improve the clarity and accuracy of three-dimensional object reconstruction.

[0050] In some embodiments, the expression loss 408 can be generated by the true expression label of the two-dimensional object and the expression coefficient in the three-dimensional object reconstruction model 403:

[0051]

[0052] Where Le represents the expression loss, c represents the number of expressions, yc represents the expression coefficient, and pc represents the true expression label. In the embodiments of the present disclosure, six basic expressions are used, for example, anger, disgust, fear, happiness, sadness, and surprise. It should be understood that the maximum value of c is 6. Of course, other types of expressions can also be set, and specific ones can be selected according to actual needs. As Figure 4 shown, a classification expression classifier 410 can be set in the three-dimensional object reconstruction model 403. The classification expression classifier 410 extracts the true expression label of the two-dimensional object, and the classification expression classifier 410 can select a pre-trained CNN model. The three-dimensional object reconstruction model 403 can obtain the expression coefficient according to the extracted features of the two-dimensional object. By calculating the cross-entropy loss of the expression coefficient and the true expression label, that is, the expression loss Le, the expression accuracy of the generated three-dimensional object 402 can be improved.

[0053] In some embodiments, the regularization loss 409 can be generated by the average shape feature, average texture feature, average pose feature, shape feature S, texture feature T, and pose feature R:

[0054]

[0055] Where Lr represents the regularization loss, μi represents the average identity feature, μe represents the expression identity feature, μt represents the average texture feature, and U represents a uniform distribution with a predetermined range of pose parameters. The average identity feature μi refers to the identity feature learned from a large number of three-dimensional objects, the average expression feature μe refers to the expression feature learned from a large number of three-dimensional objects, and the average texture feature μt refers to the texture feature learned from a large number of three-dimensional objects. The regularization loss Lr measures the degree to which the shape parameters, texture parameters, and pose parameters in the three-dimensional object reconstruction model 403 learn the prior distribution from the large dataset. Among them, the shape parameters include the expression parameter αe and the identity parameter αi, the texture parameters include the texture parameter βi, and the pose parameters include the rotation angle θx around the x-axis, the rotation angle θy around the y-axis, the rotation angle θz around the z-axis, the scaling factor s, and the translation feature t.

[0056] In some embodiments, the total loss function of the three-dimensional object reconstruction model 403 is:

[0057] L = L r + λ3L e + λ4L r(8)

[0058] Among them, L represents the total loss, λ3 represents the third weight, and λ4 represents the fourth weight. In this way, the generated three-dimensional object reconstruction model can truly reconstruct the two-dimensional object in terms of shape, texture, and pose, etc., so that the three-dimensional object output by the model can match the input image in each dimension, and a three-dimensional object with high credibility and high quality can be obtained.

[0059] Figure 5 Schematic diagram of the process 500 for adjusting a three-dimensional object reconstruction model according to some embodiments of the present disclosure. After the three-dimensional object reconstruction model 503 obtains the input image 501, it determines the shape feature 504, texture feature 505, and pose feature 506, and optimizes the shape parameter 507, texture parameter 508, and pose parameter 509 based on the determined features to train the three-dimensional object reconstruction model 503. After the three-dimensional object reconstruction model 503 is trained, the tuning parameter model 502 for stylization is used to fine-tune the shape parameter 507, texture parameter 508, and pose parameter 509 in the three-dimensional object reconstruction model 503, so that the three-dimensional object reconstruction model 503 can not only reconstruct a three-dimensional object with high credibility, but also endow the object with a preset style. For example, sketch style, cartoon image style, oil painting style, etc. The purpose of stylization can be achieved by allowing the tuning parameter model 503 to pre-learn images or videos of the required style.

[0060] The tuning parameter model 502 can be a machine learning model for stylizing the three-dimensional object reconstruction model 503. In some embodiments, the tuning parameter model 502 can be a Bayesian model. The tuning process can use the input image Ii, the rendered image Ir, and the style label y of the tuning parameter model as observed variables, and adjust the shape feature, texture feature, and pose feature by maximizing the lower bound of the log marginal likelihood of the observed variables to adjust the three-dimensional object reconstruction model.

[0061] As described above, the 3D object reconstruction model 503 includes an expression parameter αe, an identity parameter αi, a texture parameter βi, a rotation angle θx about the x-axis, a rotation angle θy about the y-axis, a rotation angle θz about the z-axis, a scaling factor s, and a translation feature t. The above parameters can be regarded as latent variables of the observed variables, and the prior distributions of the latent variables can be expressed as p(αi), p(αe), p(βi), p(s), p(θx), p(θy), p(θz), p(t). The likelihood functions of the observed variables can be expressed as p(I|Ir), p(y|αe), p(Ir|αi, αe, βi, s, θx, θy, θz, t). The posterior distributions of the latent variables can be expressed as q(αi|I, y, Ir), q(αe|I, y, Ir), q(βi|I, y, Ir), q(s|I, y, Ir), q(θx|I, y, Ir), q(θy|I, y, Ir), q(θz|I, y, Ir), q(t|I, y, Ir). The lower bound of the log marginal likelihood of the observed variables to be maximized can be expressed as:

[0062]

[0063] where Eq represents the expectation of the posterior distributions of all parameters, and KL(q||p) represents the Kullback-Leibler divergence between q and p. To approximate q with a tractable family of distributions, the present disclosure uses the mean field approximation, which assumes that q factorizes over all latent variables:

[0064] q(α i ,α e ,β i ,s,θ x ,θ y ,θ z ,t|I,y,I r )=q(α i |I,y,I r )q(α e |I,y,I r )q(β i |I,y,I r )q(s|I,y,I r )q(θ x |I,y,I r )q(θ y |I,y,I r )q(θ z |I,y,I r )q(t|I,y,I r ) (10)

[0065] After factorization, it is further assumed that each factorized distribution is a Gaussian distribution with a diagonal covariance matrix:

[0066]

[0067] Where z represents any latent variable, μz represents the mean to which the latent variable is mapped to a Gaussian distribution, and Σz represents the covariance to which the latent variable is mapped to a Gaussian distribution.

[0068] In some embodiments, to adjust the above parameters, the parameter tuning model 502 of the present disclosure uses a Bayesian deep neural network for learning and adjustment, outputs the mean and covariance of each latent variable according to the input image or video, and tunes the parameters by maximizing the lower bound of the log marginal likelihood of the observed variables using the stochastic gradient descent method. The parameter tuning model can combine the knowledge of a specific style domain to generate a style-diverse 3D object reconstruction model, improving the variability of the generation model.

[0069] Figure 6 FIG. shows a schematic block diagram of an exemplary device 600 that can be used to implement embodiments of the present disclosure. As shown, the device 600 includes a computing unit 601, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or computer program instructions loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0070] Multiple components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0071] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method 200 described above may be executed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute method 200 in any other suitable manner (e.g., by means of firmware).

[0072] The functions described above herein can be performed at least in part by one or more hardware logic components. By way of example and not limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), and the like.

[0073] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0074] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. Additionally, although the operations are depicted in a particular order, this should be understood to require that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, the various features that are described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.

[0075] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating a three-dimensional object reconstruction model, comprising: Obtaining an input image including a two-dimensional object; Determining the shape feature, texture feature, and pose feature of the two-dimensional object; Generating a rendered image based on the shape feature, the texture feature, the pose feature, and the input image; Generating a three-dimensional object reconstruction model based on the input image and the rendered image; and Adjusting the three-dimensional object reconstruction model according to a tuning parameter model for stylization.

2. The method according to claim 1, wherein determining the shape feature, texture feature, and pose feature of the two-dimensional object comprises: Determining an average object shape feature, an identity feature of the two-dimensional object, and an expression feature; And Based on the determined average object shape feature, the identity feature, and the expression feature, determining the shape feature.

3. The method according to claim 2, wherein determining the shape feature, texture feature, and pose feature of the two-dimensional object further comprises: Determining an average object texture feature and an identity feature of the two-dimensional object; And Based on the determined average object texture feature and the identity feature of the two-dimensional object, determining the texture feature.

4. The method according to claim 3, wherein determining the shape feature, texture feature, and pose feature of the two-dimensional object further comprises: Determining a rotation feature, a scaling feature, and a translation feature of the two-dimensional object; And Based on the determined rotation feature, the scaling feature, and the translation feature, determining the pose feature.

5. The method according to claim 1, wherein generating a three-dimensional object reconstruction model comprises: Based on the shape feature, the texture feature, and the pose feature, determining a reconstruction loss for measuring the pixel intensity difference, an expression loss for measuring the emotional expression, and a regularization loss for measuring the object rationality; and Based on the reconstruction loss, the expression loss, and the regularization loss, generating the three-dimensional object reconstruction model.

6. The method according to claim 5, wherein determining a reconstruction loss for measuring the pixel intensity difference, an expression loss for measuring the emotional expression, and a regularization loss for measuring the object rationality comprises: Determining an input feature of the input image and a rendering feature of the rendered image; Based on the input image, the rendered image, the input feature, and the rendering feature, generating the reconstruction loss.

7. The method according to claim 5, wherein determining a reconstruction loss for measuring the pixel intensity difference, an expression loss for measuring the emotional expression, and a regularization loss for measuring the object rationality comprises: Based on the shape feature, the texture feature, and the pose feature, generating an emotional expression coefficient; Extracting a true expression label of the two-dimensional object; And Based on the expression coefficient and the true expression label, generating the expression loss.

8. The method according to claim 5, wherein determining a reconstruction loss for measuring the pixel intensity difference, an expression loss for measuring the emotional expression, and a regularization loss for measuring the object rationality comprises: Based on the average features of the object, determine the average shape feature, the average texture feature, and the average pose feature; And Based on the average shape feature, the average texture feature, the average pose feature, the shape feature, the texture feature, and the pose feature, generate the regularization loss.

9. The method according to claim 1, wherein adjusting the three-dimensional object reconstruction model includes: Taking the input image, the rendered image, and the style label of the parameter adjustment model as observed variables; And Based on the lower bound of the log marginal likelihood of the observed variables, adjust the three-dimensional object reconstruction model.

10. The method according to claim 1, further comprising: Based on the adjusted three-dimensional object reconstruction model, reconstruct the two-dimensional object into a three-dimensional object.

11. An electronic device, comprising: At least one processor; And Coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the electronic device to perform actions, the actions including: Obtain an input image including a two-dimensional object; Determine the shape feature, the texture feature, and the pose feature of the two-dimensional object; Based on the shape feature, the texture feature, the pose feature, and the input image, generate a rendered image; Based on the input image and the rendered image, generate a three-dimensional object reconstruction model; and According to a parameter adjustment model for stylization, adjust the three-dimensional object reconstruction model.

12. The device according to claim 11, wherein determining the shape feature, the texture feature, and the pose feature of the two-dimensional object includes: Determine the average object shape feature, the identity feature of the two-dimensional object, and the expression feature; And Based on the determined average object shape feature, the identity feature, and the expression feature, determine the shape feature.

13. The device according to claim 12, wherein determining the shape feature, the texture feature, and the pose feature of the two-dimensional object further includes: Determine the average object texture feature and the identity feature of the two-dimensional object; And Based on the determined average object texture feature and the identity feature of the two-dimensional object, determine the texture feature.

14. The device according to claim 13, wherein determining the shape feature, the texture feature, and the pose feature of the two-dimensional object further includes: Determine the rotation feature, the scaling feature, and the translation feature of the two-dimensional object; And Based on the determined rotation feature, the scaling feature, and the translation feature, determine the pose feature.

15. The device according to claim 11, wherein generating a three-dimensional object reconstruction model includes: Based on the shape feature, the texture feature, and the pose feature, determine a reconstruction loss for measuring the pixel intensity difference, an expression loss for measuring the emotional expression, and a regularization loss for measuring the object rationality; and Based on the reconstruction loss, the expression loss, and the regularization loss, generate the three-dimensional object reconstruction model.

16. The device according to claim 15, wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object rationality includes: Determining input features of the input image and rendering features of the rendered image; Generating the reconstruction loss based on the input image, the rendered image, the input features, and the rendering features.

17. The device according to claim 15, wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object rationality includes: Generating an emotional expression coefficient based on the shape feature, the texture feature, and the pose feature; Extracting a true expression label of the two-dimensional object; And Generating the expression loss based on the expression coefficient and the true expression label.

18. The device according to claim 15, wherein determining a reconstruction loss for measuring a pixel intensity difference, an expression loss for measuring an emotional expression, and a regularization loss for measuring object rationality includes: Determining an average shape feature, an average texture feature, and an average pose feature based on an average feature of the object; And Generating the regularization loss based on the average shape feature, the average texture feature, the average pose feature, the shape feature, the texture feature, and the pose feature.

19. The device according to claim 11, wherein adjusting the three-dimensional object reconstruction model includes: Taking the input image, the rendered image, and a style label of the parameter adjustment model as observed variables; And Adjusting the three-dimensional object reconstruction model based on a lower bound of the log marginal likelihood of the observed variables.

20. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions that, when executed, cause the machine to perform actions, the actions including: Obtaining an input image including a two-dimensional object; Determining a shape feature, a texture feature, and a pose feature of the two-dimensional object; Generating a rendered image based on the shape feature, the texture feature, the pose feature, and the input image; Generating a three-dimensional object reconstruction model based on the input image and the rendered image; and Adjusting the three-dimensional object reconstruction model according to a parameter adjustment model for stylization.