Single-image three-dimensional head reconstruction method and system based on guide diffusion model
The three-dimensional head reconstruction of a single image is solved by guiding diffusion model, which solves the problems of rough texture reconstruction, low fidelity and poor generalization ability in the prior art, and realizes the reconstruction of a high-resolution and high-fidelity three-dimensional head model.
Patent Information
- Application Number
- CN202510194037.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing single-image three-dimensional head reconstruction technology has problems such as rough texture reconstruction, low fidelity, inability to reconstruct invisible areas, poor generalization ability, and poor generation results.
Using a method based on the guided diffusion model, we preprocess and optimize the three-dimensional head model, weak projection mapping and cylinder UV expansion are performed, and the invisible area texture in the UV texture map is repaired using the guided diffusion network model, and the repaired texture is backsticked back to the three-dimensional model.
A high-resolution, high-fidelity, artifact-free three-dimensional head model is realized from a single image, enriching the surface details of the model, improving the accuracy of local detail textures, and adapting to changes in different ages, genders, races and facial expressions.
Smart Images

Figure CN120107481A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and computer graphics, and more specifically, relates to a single image three-dimensional head reconstruction method and system based on a guided diffusion model. Background Art
[0002] Single-image 3D head reconstruction refers to the use of algorithms to automatically generate a 3D head model represented by a mesh from a single portrait image. With the rise of the metaverse, 3D head reconstruction has shown great application prospects. In addition, 3D head reconstruction has broad application prospects in film and television production, game entertainment, smart education and other fields. There are currently two ways to obtain a 3D head model: (1) manual modeling or reconstruction of a 3D head model based on a high-precision 3D scanner or stereo vision system; (2) direct reconstruction of a 3D head model from a portrait image based on deep learning technology. The first method can obtain a high-precision 3D head model, but the equipment used (such as a laser ranging 3D scanner) is expensive, has low operability, and is difficult to popularize; and during the scanning process, individuals are prone to shaking and generating noise, resulting in incomplete scanned 3D head models. The second method based on deep learning directly reconstructs a 3D head model from a single portrait image, which can save costs to the greatest extent and greatly improve the convenience of operation.
[0003] In recent years, researchers in this field have proposed a series of deep learning-based methods to learn prior knowledge from data, but most studies focus on the reconstruction of geometric structures, and rarely on the reconstruction of the entire head texture. Although the prior art proposes methods to estimate the 3D shape and texture details of the head from unrestricted inputs (such as wild images), there are still many problems. The ideal reconstruction should be high-fidelity and artifact-free. More specifically, it should faithfully convey the head posture and texture details of the human in the image, and the generated 3D head model should be a complete head without holes or missing parts, and it should not be a non-human head shape and other artifacts.
[0004] Existing texture reconstruction methods can be mainly divided into two categories: one is a texture reconstruction method based on three-dimensional meshes, and the other is a texture reconstruction method based on new perspective synthesis. In the early days, researchers proposed to reconstruct a three-dimensional head model based on 3DMM (three-dimensional deformable model) combined with UV (two-dimensional texture coordinates corresponding to the vertex information of geometric figures) texture mapping. However, 3DMM itself has limitations. It can only perform a rough geometric estimate and uses a unified template. Even if the complete UV texture can be restored, its fixed template causes the generated result to be far from the real head. The difference in the reconstruction results is quite significant, and the hair part is not fully considered in the reconstruction process. In addition, there is a method that reconstructs the mesh by inputting image features and uses a diffusion model to repair the incomplete texture. Although the current diffusion model has shown certain effectiveness in repairing UV mapping, existing research still does not achieve ideal results, and it does not consider the invisible area on the back when repairing UV. In addition, there is another study that mainly uses image synthesis methods to generate heads from various perspectives, but only a small part of them can generate 360-degree perspectives. There are also many problems such as low resolution, incomplete images, and the generated results are inconsistent with the actual situation.
[0005] In summary, although the single-image 3D head texture reconstruction method has made certain progress in recent years, the following problems still exist: (1) The texture of the reconstructed 3D model is rough and the fidelity is low. Although the texture of the 3D head model reconstructed by some methods looks realistic, it is difficult to match the identity of the person in the input image. (2) The texture of the invisible area of the input image cannot be reconstructed or the details are less and the accuracy is poor. Although some methods can restore the texture of facial expressions, the expression ability of the back texture part other than the facial area is weak. (3) The generalization ability of the algorithm is poor and it is difficult to adapt to changes in age, race, gender, etc. For example, when the model trained on the adult dataset reconstructs a child's portrait, the reconstructed result still looks like an adult. (4) The generated result is flat, distorted, and unlike the person in the input image. Although the new perspective synthesis method can reconstruct a real 360-degree perspective, the resolution is low and the generated result is far from the person in the input image. Summary of the invention
[0006] In view of the above defects or improvement needs of the prior art, the present invention provides a single-image three-dimensional head reconstruction method and system based on a guided diffusion model, which can reconstruct a high-fidelity, artifact-free, and realistic three-dimensional head model with rich texture details from a single portrait image of an individual of different ages, genders, races, and facial expressions.
[0007] To achieve the above object, according to one aspect of the present invention, a single image three-dimensional head reconstruction method based on a guided diffusion model is provided, comprising the steps of:
[0008] Preprocess the single image to be reconstructed to obtain a head portrait;
[0009] reconstructing a three-dimensional head model from the head portrait, and optimizing the reconstructed three-dimensional head model to restore smooth surface details;
[0010] Weakly projecting the three-dimensional head model after restoring the smooth surface details onto the head portrait under the predicted camera matrix, and using the pixel information of the head portrait as the texture information of the corresponding three-dimensional point, thereby obtaining a three-dimensional head model with frontal texture;
[0011] Perform cylindrical UV unfolding on the 3D head model with the front texture to obtain a 2D UV texture map and record the UV mapping relationship;
[0012] Inputting the two-dimensional UV texture map into a trained guided diffusion network model to conditionally predict the texture of the invisible area in the two-dimensional UV texture map, and repairing to obtain a complete two-dimensional UV texture map;
[0013] By using the UV mapping relationship, the complete two-dimensional UV texture map is pasted back to the three-dimensional head model with the front texture.
[0014] Preferably, before pasting the complete two-dimensional UV texture map back onto the three-dimensional head model with the front texture, the method further includes the step of inputting the complete UV texture map into a super-resolution network model to improve the resolution of the complete UV texture map.
[0015] Preferably, the three-dimensional head model after restoring the smooth surface details is weakly projected onto the head portrait under the predicted camera matrix, and the pixel information of the head portrait is used as the texture information of the corresponding three-dimensional point to obtain the three-dimensional head model with frontal texture, which includes the steps of:
[0016] Interpolate the 3D head model after restoring the smooth surface details, mark the coordinates of the 3D points of the 3D head model after the interpolation as V(X, Y, Z), and use the 3D points of the 3D head model with a value greater than 0.5 on the z-axis as the 3D points of the front texture to be determined, and the remaining 3D points as the invisible areas of the texture to be predicted;
[0017] Mapping the three-dimensional point V(X, Y, Z) in the three-dimensional space to the head portrait through weak projection mapping transformation under the predicted camera matrix, and obtaining the two-dimensional pixel coordinate point v(x, y) corresponding to each three-dimensional point V(X, Y, Z);
[0018] The pixel information of the two-dimensional pixel coordinate point v(x, y) of the head portrait is used as the texture information of the corresponding three-dimensional point V(X, Y, Z), so as to obtain a three-dimensional head model with frontal texture.
[0019] Preferably, the calculation formula of the weak projection mapping transformation is:
[0020]
[0021] Where f is the focal length of the camera.
[0022] Preferably, the step of using the pixel information of the two-dimensional pixel coordinate point v(x, y) of the head avatar as the texture information of the corresponding three-dimensional point V(X, Y, Z) comprises the following steps:
[0023] Extract the neighborhood information of the two-dimensional pixel coordinate point v(x, y) of the head portrait, and determine the color of the two-dimensional pixel coordinate point v(x, y) according to the neighborhood information of the two-dimensional pixel coordinate point v(x, y), which is recorded as C x,y ;
[0024] The color C of the two-dimensional pixel coordinate point v(x,y) x,y As the texture information of the corresponding three-dimensional point V(X,Y,Z).
[0025] Preferably, the cylindrical UV unfolding of the three-dimensional head model with the front texture comprises the steps of:
[0026] The 3D head model with the front texture is subjected to cylindrical UV expansion, and the corresponding point of the 3D point V (X, Y, Z) after cylindrical UV expansion in the UV space is recorded as T (u, v);
[0027] The color of each triangular face of the 3D head model with front texture is rendered to the corresponding UV color area in UV space through diffuse reflection. The calculation formula is:
[0028]
[0029] Among them, R(·) represents diffuse reflection rendering, Area(u 1...n ,v 1...n ) represents the color area surrounded by multiple uv coordinates in UV space. Area(V 1...n ) represents the color area of the triangular patch in the 3D head model with front texture. represents the two-dimensional UV texture map, Represents the frontal texture of a 3D head model.
[0030] Preferably, the UV mapping relationship is:
[0031]
[0032] Where ρ is the distance from the three-dimensional point V(X,Y,Z) to the z-axis, θ is the angle of the projection of the point (X,Y) on the xy plane from the positive direction of the x-axis counterclockwise, and z min is the minimum value of the 3D point of the 3D head model in the z-axis direction, max It is the maximum value of the 3D point of the 3D head model in the z-axis direction.
[0033] Preferably, the training of the guided diffusion network model comprises the steps of:
[0034] For the 3D head model samples, a 2D UV texture map is obtained through UV mapping and exposed at different ratios to form a training image sample set. At the same time, the mask of the invisible area corresponding to each training image is obtained.
[0035] The diffusion network model is trained using a training picture sample set.
[0036] Preferably, the loss function of the diffusion network model is:
[0037]
[0038] Among them, x is the training image, P is the step size set during denoising, and γ represents the current noise level. represents the added noise predicted by the diffusion network model based on x and γ, ε refers to the noise added at each time step, Represents the noise image The loss Z obtained by fitting the original training image x after denoising at each time step ε,γ represents the noise of the current time step of the forecast, ||.|| P It means that the noise reduction process continues from step 1 to step P. T Represents matrix transpose.
[0039] According to another aspect of the present invention, a single image three-dimensional head reconstruction system based on a guided diffusion model is provided, comprising:
[0040] A preprocessing module, used for preprocessing the single image to be reconstructed to obtain a head portrait;
[0041] A mesh reconstruction module is used to reconstruct a three-dimensional head model from the head portrait, and optimize the reconstructed three-dimensional head model to restore smooth surface details;
[0042] A weak projection coloring module is used to weakly project the three-dimensional head model after restoring the smooth surface details onto the head portrait under the predicted camera matrix, and use the pixel information of the head portrait as the texture information of the corresponding three-dimensional point, so as to obtain a three-dimensional head model with frontal texture;
[0043] The cylindrical UV mapping module is used to perform cylindrical UV unfolding on the three-dimensional head model with the front texture, obtain a two-dimensional UV texture map and record the UV mapping relationship;
[0044] The UV texture map guided repair module is used to input the two-dimensional UV texture map into the trained guided diffusion network model to conditionally predict the texture of the invisible area in the two-dimensional UV texture map, repair the complete two-dimensional UV texture map, and use the UV mapping relationship to paste the complete two-dimensional UV texture map back to the three-dimensional head model with the front texture.
[0045] In general, the above technical solutions conceived by the present invention have beneficial effects compared with the prior art:
[0046] (1) The present invention uses a guided diffusion model to repair the invisible areas of the incomplete UV texture map, thereby reconstructing the full head texture. The present invention not only retains the facial texture details (such as wrinkles, moles, etc.) of the original input to the greatest extent, but also can maintain the high-resolution output of the three-dimensional head model, thereby enriching the surface details of the reconstructed three-dimensional head model and improving the accuracy of the local detail texture. Compared with other methods for head reconstruction, the present invention can reconstruct a high-resolution, high-fidelity, artifact-free, and texture-rich three-dimensional head from a single image without considering the individual's age, gender, race, and facial expression.
[0047] (2) The guided diffusion model proposed in the present invention is only trained for the invisible area and does not directly embed the mask, but directly adds noise prediction to the invisible area. Therefore, the invisible area of the incomplete UV texture map can be directly repaired to reconstruct the full head texture.
[0048] (3) Compared with manual modeling or reconstruction of a three-dimensional head model based on a three-dimensional scanner or a stereoscopic vision system, the present invention only requires a single portrait image to be input into a computer to complete the reconstruction, which can save costs to the greatest extent and greatly improve the convenience of operation, and has broad application prospects in the future.
[0049] (4) In the method of directly reconstructing a three-dimensional head model from multi-view images based on deep learning technology, the present invention only needs to input one picture to reconstruct a high-fidelity three-dimensional head model, thereby reducing the application conditions to the greatest extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flow chart of a single-image three-dimensional head reconstruction method according to an embodiment of the present invention;
[0051] Figure 2 Schematic diagram of a network framework used in a single-image three-dimensional head reconstruction method according to an embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of the projection relationship between weak projection mapping and UV mapping in an embodiment of the present invention;
[0053] Figure 4 This is an example effect diagram of comparing three-dimensional head reconstruction of a standard public data set with other methods according to an embodiment of the present invention;
[0054] Figure 5 This is an example effect diagram of comparing three-dimensional head reconstruction of field images with other methods according to an embodiment of the present invention;
[0055] Figure 6 It is a 3600 example effect diagram of three-dimensional head reconstruction of field images according to an embodiment of the present invention. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0057] In the description of the embodiments of the present application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product or equipment comprising a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products or equipment.
[0058] The naming or numbering of the steps in the embodiments of the present invention does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0059] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0060] The present invention provides a single image three-dimensional head reconstruction method and system based on a guided diffusion model, which are described below respectively.
[0061] like Figure 1 and Figure 2 As shown, the single image 3D head reconstruction method based on the guided diffusion model according to an embodiment of the present invention comprises the following steps:
[0062] S101, preprocessing the single image to be reconstructed to obtain a head image.
[0063] Specifically, the portrait image to be reconstructed is obtained, a portrait mask is extracted from the portrait image, and background factors are removed to obtain a head image.
[0064] The background of the input image P to be reconstructed is segmented to obtain a black and white image describing the portrait area, denoted as M, and the black and white image is used to restrict the computer to operate the head image I of the image to be reconstructed:
[0065]
[0066] Represents dot product.
[0067] In one embodiment, an API interface provided by an existing platform may be used to remove the background outside the portrait from the input image, and then the image mask may be segmented using code.
[0068] S102, reconstructing a three-dimensional head model from the head image I, and optimizing the reconstructed three-dimensional head model to restore smooth surface details.
[0069] The reconstructed three-dimensional head model in step S102 can adopt any method in the prior art, for example, the three-dimensional head model is reconstructed using the SeIF method. SeIF is a method for generating an animated three-dimensional head model from a single image, and the method includes the following steps: acquiring a head image and generating a parameterized model; transferring semantic information from a normalized semantic template to the parameterized model, generating a specific semantic model for the head image and sampling it, and assigning a semantic code to each sampling point; extracting a normal map and a feature map of the head image; inputting each sampling point, feature vector and semantic code into a deep implicit function, and predicting an occupancy value and an accurate semantic code for each sampling point; extracting equipotential surfaces from the occupancy field, generating a three-dimensional head model, and unifying the structure using accurate semantic codes; and dynamically driving the structurally unified three-dimensional head model using an expression-driven algorithm.
[0070] The reconstructed 3D head model can be optimized by any method in the prior art, such as using a normal geometry optimization algorithm, which is to perform an affine transformation between the normal map rendered by the SeIF method and the head image to be reconstructed to achieve the offset of the 3D points in the reconstructed 3D head model. The position of the 3D points in the 3D head model reconstructed by the SeIF method is moved by calibrating the normal map and the head image to be reconstructed, so as to achieve the optimization of the 3D head model.
[0071] S103, weakly projecting the three-dimensional head model after restoring the smooth surface details in step S102 onto the head portrait under the predicted camera matrix, and then using the pixel information of the head portrait I as the texture information of the corresponding three-dimensional point to obtain a three-dimensional head model with frontal texture.
[0072] In weak projection mapping, the normal direction can be used to determine the positive and negative directions of the three-dimensional model.
[0073] Further, S103 includes sub-steps:
[0074] (1) Interpolation processing is performed on the reconstructed three-dimensional head model to retain the details of the input image I to the greatest extent possible. The coordinates of the three-dimensional points of the three-dimensional head model after the interpolation processing are marked as V(X, Y, Z). At the same time, the three-dimensional points of the three-dimensional model with values greater than 0.5 on the z-axis are selected as the three-dimensional points to be colored, that is, the three-dimensional points of the front texture to be determined, and the remaining points are set to gray, that is, the invisible area of the texture to be predicted.
[0075] Preferably, if Figure 3 As shown, the coordinates of the three-dimensional point in the three-dimensional space are marked as V (X, Y, Z), and each three-dimensional point is mapped to the head portrait I through weak projection mapping transformation under the predicted camera matrix (the camera focal length is f), and the two-dimensional pixel coordinate point v (x, y) corresponding to each three-dimensional point is obtained, and the floor function is used to convert it into integer coordinates to index the corresponding pixel position. The projection relationship is as follows:
[0076]
[0077] When the three-dimensional point V(X,Y,Z) is weakly projected onto the head portrait I, an error may be caused due to the rounding down operation when obtaining the corresponding two-dimensional pixel coordinate point v(x,y). In order to obtain the color information of the accurate position as much as possible, the neighborhood information of the two-dimensional pixel coordinate point v(x,y) of the head portrait is extracted, and the color of the two-dimensional pixel coordinate point v(x,y) is determined according to the neighborhood information of the two-dimensional pixel coordinate point v(x,y) to improve the robustness of the algorithm. The color value of the corresponding pixel position v(x,y) obtained is recorded as C x,y ;
[0078] Furthermore, C x,y The calculation formula is:
[0079] C x,y =(C r (x,y),C g (x,y),C b (x,y)
[0080] C(I,V)=χ(C x,y )
[0081] Among them, C r ,C g ,C b denote the values of the red, green, and blue channels, respectively, and χ(·) denotes the neighborhood function.
[0082] S104: Performing cylindrical UV mapping on the three-dimensional head model with the front texture to obtain a two-dimensional UV texture map and recording the UV mapping relationship.
[0083] UV is a two-dimensional texture coordinate system that uses the letters U and V to indicate axes in two-dimensional space, helping to place two-dimensional image textures on 3D surfaces. The UV mapping relationship defines how the two-dimensional texture coordinates are mapped to the three-dimensional model surface mesh.
[0084] Preferably, if Figure 3 As shown, for the three-dimensional point V (X, Y, Z) of the three-dimensional head model, it is expressed by the surface parameter equation:
[0085] X=χ(s,t),Y=η(s,t),Z=z(s,t)
[0086] Among them, χ(·),η(·),z(·) are the conversion functions of the three-dimensional surface, and (s,t) are the parameters corresponding to a point on the three-dimensional surface.
[0087] For the coordinate point T(u,v) in UV space, it is expressed as:
[0088] u=φ(s,t),
[0089] Among them, φ(·), It is the conversion function of UV space according to the three-dimensional surface parameters (s, t).
[0090] Based on the above representation relationship, the grid three-dimensional point V (X, Y, Z) is projected into the UV space through cylindrical UV mapping. The cylindrical coordinates of the three-dimensional point in the cylindrical space coordinate system are The three-dimensional coordinate point V (X, Y, Z) and its corresponding relationship are:
[0091] X=ρcosθ,Y=ρsinθ,Z=z
[0092] Where ρ is the distance from the point to the z-axis (i.e. ), θ is the angle of the projection of the point (X,Y) on the xy plane from the positive direction of the x-axis to the counterclockwise direction
[0093] For a point T(u,v) in UV space, it is defined as:
[0094]
[0095] Among them, z min is the minimum value of the three-dimensional head model in the z-axis direction, max It is the maximum value of the 3D head model in the z-axis direction.
[0096] After the mesh is UV unfolded, the color of each triangular face of the three-dimensional model is rendered to the UV color area corresponding to the UV space through diffuse reflection. The calculation formula is as follows:
[0097]
[0098] Among them, R(·) represents diffuse reflection rendering, Area(u 1...n ,v 1...n ) represents the color area surrounded by multiple uv coordinates in UV space. Area(V 1...n ) represents the triangular patch color area in the 3D head model.
[0099] The color value of the pixel point T(u,v) on the UV texture image can be recorded as C(u,v), and its calculation formula is as follows:
[0100] C(u,v)=(r(u,v),g(u,v),b(u,v))
[0101] Among them, r, g, b represent the values of the red, green, and blue channels respectively. Through this calculation, the color value of the corresponding pixel point T(u, v) in the UV space can be obtained.
[0102] S105: Inputting the two-dimensional UV texture map into the trained guided diffusion network model to conditionally predict the texture of the invisible area, and repairing to obtain a complete two-dimensional UV texture map.
[0103] The principle of the guided diffusion model is explained in detail below.
[0104] The guided diffusion model adds noise to the texture area that needs to be learned in the real UV texture map through given steps until it becomes a pure noise image, and then denoises it according to a given step size to restore the original image. In the reverse denoising process, it continuously compares it with the noise added in each step of the forward process, calculates the loss difference between the two images before and after denoising, and continuously iterates until the loss is minimized. In this way, the image distribution law and image generation direction in the reverse denoising process are learned, and the purpose of repairing the invisible area of the UV incomplete texture map is achieved.
[0105] Given a real UV texture image x, add noise to the area set to gray after weak projection in step S103. Given a noise image at a certain moment in the noise adding process It can be expressed as:
[0106]
[0107] Among them, γ represents the current noise level, that is, the accumulation of noise added at each time step, and ε refers to the noise added at each time step.
[0108] The noisy image The real UV texture image x and the current noise level γ are used as input parameters to parameterize the neural network model. The network model can be expressed as:
[0109]
[0110] The training process is described in detail below.
[0111] (1) Prepare a training sample set.
[0112] For the existing 3D head model samples, a complete 2D UV texture map is obtained through UV mapping and then exposed at a ratio of 0.7 to 1.3 to form a training image set sample. At the same time, a mask of the invisible area corresponding to the training image is obtained.
[0113] (2) Using the training image sample set to train the diffusion network model.
[0114] In order to predict the noise added at each time step, the loss function designed in the embodiment of the present invention is as follows:
[0115]
[0116] Among them, P is the step size set during denoising, It means that the noise image After denoising at each time step, the loss obtained by fitting the original training image x and the noise added at the current time step are used. During training, the loss obtained at each iteration and the noise at the current time step are re-input into the neural network as weight coefficients, and the iteration is continued until the output loss is minimized. At this time, the predicted noise is the optimal solution. ||.|| P Indicates that the denoising process stops from step 1 to step P, T represents the transposed matrix, and when the result of ||.|| is a vector, the rows and columns must be consistent to be multiplied with the weight coefficient. During training, P is set to 250, 500, and 1000 to verify that the repair efficiency of the model is improved as much as possible while ensuring the training effect, reducing time cost and memory consumption.
[0117] The loss function is used to iteratively compare the noise image and the real UV texture image, and the parameters of the neural network are returned to evaluate whether the noise vector predicted during the training process is accurate. In this way, the model can learn the rules of the UV texture denoising process, that is, when the UV texture image is noisy and becomes a pure noise image, the rules should be followed when denoising and repairing it back to the original UV texture image.
[0118] Furthermore, the trained guided diffusion model is used to repair the UV defective texture map, including the following steps:
[0119] The inverse denoising process learned during the above training is used to implement the UV incomplete texture map repair process. Given a UV incomplete texture map x, the incomplete part of x is denoised with T steps to a pure noise image, and then denoised for P times. The trained guided diffusion model is used to repair the predicted UV incomplete texture map x to obtain its corresponding real image x. 0 , the calculation formula is as follows:
[0120]
[0121] γ p represents the noise at the Pth time step, x p Represents the image after denoising P time steps.
[0122] S106: Input the complete UV texture map into the super-resolution network model to super-resolution to 4096 resolution, so as to fully ensure that the input head portrait can be retained on the three-dimensional model with high fidelity.
[0123] Step S106 is a preferred step but not a necessary step.
[0124] S107: Using the recorded UV mapping relationship, paste the UV texture map back to the three-dimensional head model with the front texture in S103.
[0125] Preferably, if Figure 3As shown, for a three-dimensional point V(X,Y,Z) on the surface of a 3D model, after obtaining T(u,v) through UV mapping, the color value of the UV texture image at point T(u,v) can be assigned to the three-dimensional point V(X,Y,Z) through a predetermined relationship, thereby achieving the attachment of the texture to the surface of the 3D model.
[0126] Thanks to the powerful expressiveness and flexibility of the guided diffusion model in image restoration, the present invention effectively solves the problems of low resolution, unlikeness, low fidelity and poor algorithm generalization ability of the three-dimensional head model reconstructed by the existing methods. Figure 4 Shown is an example effect diagram of three-dimensional head reconstruction on a standard public data set in accordance with an embodiment of the present invention compared with other methods. Figure 5 The figure shows an example effect diagram of three-dimensional head reconstruction of field images by an embodiment of the present invention compared with other methods. Figure 6 The figure shows an example of a 360° viewing angle of a 3D head model of a field image reconstructed from a single portrait image by the present invention. It can be seen that the present invention can reconstruct a 3D head model with high resolution, high fidelity, and rich texture details for portrait images of different ages, races, genders, and facial expressions.
[0127] It should be noted that the above steps S101 and S106 are not necessary steps, and users can flexibly adjust them according to their needs.
[0128] A single-image three-dimensional head reconstruction system based on a guided diffusion model according to an embodiment of the present invention includes:
[0129] A preprocessing module, used for preprocessing the single image to be reconstructed to obtain a head portrait;
[0130] A mesh reconstruction module is used to reconstruct a three-dimensional head model from the head portrait, and optimize the reconstructed three-dimensional head model to restore smooth surface details;
[0131] A weak projection coloring module is used to weakly project the three-dimensional head model after restoring the smooth surface details onto the head portrait under the predicted camera matrix, and use the pixel information of the head portrait as the texture information of the corresponding three-dimensional point, so as to obtain a three-dimensional head model with frontal texture;
[0132] The cylindrical UV mapping module is used to perform cylindrical UV unfolding on the three-dimensional head model with the front texture, obtain a two-dimensional UV texture map and record the UV mapping relationship;
[0133] The UV texture map guided repair module is used to input the two-dimensional UV texture map into the trained guided diffusion network model to conditionally predict the texture of the invisible area in the two-dimensional UV texture map, repair the complete two-dimensional UV texture map, and use the UV mapping relationship to paste the complete two-dimensional UV texture map back to the three-dimensional head model with the front texture.
[0134] The working principle and technical effect of the single-image 3D head reconstruction system are the same as those of the single-image 3D head reconstruction method described above, and will not be described in detail here.
[0135] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A single image three-dimensional head reconstruction method based on a guided diffusion model, characterized in that: Includes steps: Preprocess the single image to be reconstructed to obtain a head portrait; reconstructing a three-dimensional head model from the head portrait, and optimizing the reconstructed three-dimensional head model to restore smooth surface details; Weakly projecting the three-dimensional head model after restoring the smooth surface details onto the head portrait under the predicted camera matrix, and using the pixel information of the head portrait as the texture information of the corresponding three-dimensional point, thereby obtaining a three-dimensional head model with frontal texture; Perform cylindrical UV unfolding on the 3D head model with the front texture to obtain a 2D UV texture map and record the UV mapping relationship; Inputting the two-dimensional UV texture map into a trained guided diffusion network model to conditionally predict the texture of the invisible area in the two-dimensional UV texture map, and repairing to obtain a complete two-dimensional UV texture map; By using the UV mapping relationship, the complete two-dimensional UV texture map is pasted back to the three-dimensional head model with the front texture.
2. The single image three-dimensional head reconstruction method based on the guided diffusion model according to claim 1, characterized in that: Before pasting the complete two-dimensional UV texture map back to the three-dimensional head model with the front texture, the method further includes the step of inputting the complete UV texture map into the super-resolution network model to improve the resolution of the complete UV texture map.
3. The single image 3D head reconstruction method based on guided diffusion model as claimed in claim 1, characterized in that: The three-dimensional head model after restoring the smooth surface details is weakly projected onto the head portrait under the predicted camera matrix, and the pixel information of the head portrait is used as the texture information of the corresponding three-dimensional point to obtain the three-dimensional head model with frontal texture, which includes the following steps: Interpolate the 3D head model after restoring the smooth surface details, mark the coordinates of the 3D points of the 3D head model after the interpolation as V(X, Y, Z), and use the 3D points of the 3D head model with a value greater than 0.5 on the z-axis as the 3D points of the front texture to be determined, and the remaining 3D points as the invisible areas of the texture to be predicted; Mapping the three-dimensional point V(X, Y, Z) in the three-dimensional space to the head portrait through weak projection mapping transformation under the predicted camera matrix, and obtaining the two-dimensional pixel coordinate point v(x, y) corresponding to each three-dimensional point V(X, Y, Z); The pixel information of the two-dimensional pixel coordinate point v(x, y) of the head portrait is used as the texture information of the corresponding three-dimensional point V(X, Y, Z), so as to obtain a three-dimensional head model with frontal texture.
4. The single image three-dimensional head reconstruction method based on the guided diffusion model as claimed in claim 3, characterized in that: The calculation formula of weak projection mapping transformation is: Where f is the focal length of the camera.
5. The single image three-dimensional head reconstruction method based on the guided diffusion model as claimed in claim 3, characterized in that: The step of using the pixel information of the two-dimensional pixel coordinate point v(x, y) of the head avatar as the texture information of the corresponding three-dimensional point V(X, Y, Z) comprises the following steps: Extract the neighborhood information of the two-dimensional pixel coordinate point v(x, y) of the head portrait, and determine the color of the two-dimensional pixel coordinate point v(x, y) according to the neighborhood information of the two-dimensional pixel coordinate point v(x, y), which is recorded as C x,y ; The color C of the two-dimensional pixel coordinate point v(x,y) x,y As the texture information of the corresponding three-dimensional point V(X,Y,Z).
6. The single image 3D head reconstruction method based on guided diffusion model as claimed in claim 1, characterized in that: The cylindrical UV unfolding of the three-dimensional head model with the front texture comprises the following steps: The 3D head model with the front texture is subjected to cylindrical UV expansion, and the corresponding point of the 3D point V (X, Y, Z) after cylindrical UV expansion in the UV space is recorded as T (u, v); The color of each triangular face of the 3D head model with front texture is rendered to the corresponding UV color area in UV space through diffuse reflection. The calculation formula is: Among them, R(·) represents diffuse reflection rendering, Area(u 1...n ,v 1...n ) represents the color area surrounded by multiple uv coordinates in UV space. Area(V 1...n ) represents the color area of the triangular patch in the 3D head model with front texture. represents the two-dimensional UV texture map, Represents the frontal texture of a 3D head model.
7. The single image three-dimensional head reconstruction method based on the guided diffusion model according to claim 6, characterized in that: The UV mapping relationship is: Where ρ is the distance from the three-dimensional point V(X,Y,Z) to the z-axis, θ is the angle of the projection of the point (X,Y) on the xy plane from the positive direction of the x-axis counterclockwise, and z min is the minimum value of the 3D point of the 3D head model in the z-axis direction, max It is the maximum value of the 3D point of the 3D head model in the z-axis direction.
8. The single image three-dimensional head reconstruction method based on the guided diffusion model as claimed in claim 1, characterized in that: The training of the guided diffusion network model includes the steps of: For the 3D head model samples, a 2D UV texture map is obtained through UV mapping and exposed at different ratios to form a training image sample set. At the same time, the mask of the invisible area corresponding to each training image is obtained. The diffusion network model is trained using a training picture sample set.
9. The single image three-dimensional head reconstruction method based on the guided diffusion model as claimed in claim 8, characterized in that: The loss function of the diffusion network model is: Among them, x is the training image, P is the step size set during denoising, and γ represents the current noise level. represents the added noise predicted by the diffusion network model based on x and γ, ε represents the noise added at each time step, Represents the noise image The loss Z obtained by fitting the original training image x after denoising at each time step ε,γ represents the noise of the current time step of the forecast, ||.|| P It means that the noise reduction process continues from step 1 to step P. T Represents matrix transpose.
10. A single image three-dimensional head reconstruction system based on a guided diffusion model, characterized in that: include: A preprocessing module, used for preprocessing the single image to be reconstructed to obtain a head portrait; A mesh reconstruction module is used to reconstruct a three-dimensional head model from the head portrait, and optimize the reconstructed three-dimensional head model to restore smooth surface details; A weak projection coloring module is used to weakly project the three-dimensional head model after restoring the smooth surface details onto the head portrait under the predicted camera matrix, and use the pixel information of the head portrait as the texture information of the corresponding three-dimensional point, so as to obtain a three-dimensional head model with frontal texture; The cylindrical UV mapping module is used to perform cylindrical UV unfolding on the three-dimensional head model with the front texture, obtain a two-dimensional UV texture map and record the UV mapping relationship; The UV texture map guided repair module is used to input the two-dimensional UV texture map into the trained guided diffusion network model to conditionally predict the texture of the invisible area in the two-dimensional UV texture map, repair the complete two-dimensional UV texture map, and use the UV mapping relationship to paste the complete two-dimensional UV texture map back to the three-dimensional head model with the front texture.
Citation Information
Patent Citations
Method and system for generating animatable three-dimensional head model from single image
CN117893673A
Unsupervised three-dimensional face model reconstruction system and method
CN117975525A
System and method for reconstructing three-dimensional human head model from single image
CN118262035A
3D human body model material generation method, system and device based on diffusion model and medium
CN118429537A
Three-dimensional virtual character generation method, electronic equipment, storage medium and program product
CN119251361A
Cited By
Regional color correction method and system based on binocular view
CN120931540A
Binocular view based regional color correction method and system
CN120931540B