Three-dimensional reconstruction method and device based on image set, equipment and medium
By extracting and processing the pose and feature parameters of multiple face images at different angles, the problem of three-dimensional reconstruction under face image blur is solved, and efficient face three-dimensional modeling and scene reconstruction are achieved.
Patent Information
- Application Number
- CN202510622967.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The prior art is difficult to complete the reconstruction of the three-dimensional scene of the face with blurred face images.
By acquiring multiple face image sets at different angles, extracting pose parameters and upsampling processing, the face feature parameters and the initial three-dimensional reconstruction unit are determined. Then, these parameters are input into the three-dimensional reconstruction model, and iterative calculations are performed to generate the target three-dimensional reconstruction unit, and then a three-dimensional face model is established.
It realizes the efficient three-dimensional modeling of the face and scene reconstruction of the character objects when using a relatively blurred face image set, and improves the modeling efficiency and quality.
Smart Images

Figure CN120147559A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of computer vision and 3D modeling, and particularly to a 3D reconstruction method, device, equipment, and medium based on an image set. Background Art
[0002] With the development of computer information technology, especially the development of artificial intelligence technology, more and more traditional problems have been solved by artificial intelligence technology. For example, generating a 3D model based on a set of two-dimensional images has always been a research topic in many fields. Due to the importance and richness of human facial features, how to accurately generate a 3D human face model with a relatively high credibility based on two-dimensional human face images is a problem worthy of research, which can be used in fields such as 3D human face recognition, post-production of movies, and games.
[0003] In the prior art, for the method of establishing a 3D human face model, there is manual 3D modeling, but this method requires the user to have professional skills and will consume a lot of time, with low efficiency and unstable modeling effects; or 3D modeling is performed through the captured human face images, but this method has high requirements for the resolution of the captured human face images. If the captured images are blurred, 3D modeling cannot be performed.
[0004] Currently, for the prior art, if the captured human face image set is blurred, no effective solution has been proposed for how to complete high-quality 3D human face scene reconstruction. Summary of the Invention
[0005] Based on this, it is necessary to provide a 3D reconstruction method, device, equipment, and medium based on an image set for the above technical problems.
[0006] In a first aspect, the present application provides a 3D reconstruction method based on an image set. The method includes: Obtain a to-be-processed human face image set, where the to-be-processed human face image set includes multiple to-be-processed human face images of different angles for the same human object; Extract the pose parameters of the to-be-processed human face image set, perform upsampling processing on the to-be-processed human face image set to obtain a target human face image set, determine human face feature parameters based on the target human face image set; and calculate an initial 3D reconstruction unit based on the geometric structure of the human object determined by the to-be-processed human face images of different angles; Input the pose parameters, human face feature parameters, and the initial 3D reconstruction unit into a 3D reconstruction model, perform iterative calculation on the initial 3D reconstruction unit based on the pose parameters and human face feature parameters to obtain a target 3D reconstruction unit, and establish a 3D human face model corresponding to the human object based on the target 3D reconstruction unit.
[0007] In one embodiment, after obtaining the target three-dimensional reconstruction unit, the method further includes: Project the three-dimensional face model onto a two-dimensional plane through differentiable rasterization to generate a set of deblurred face image sequences.
[0008] In one embodiment, performing upsampling processing on the set of face images to be processed to obtain a set of target face images, including: Successively use each face image to be processed in the set of face images to be processed as the first face image; Map the first face image to a latent space variable; Input the latent space variable and a preset constant tensor into a preset normalization module for normalization processing to obtain a target face image corresponding to the first face image; Based on the target face images corresponding to each image in the set of face images to be processed, obtain the set of target face images.
[0009] In one embodiment, the normalization module includes multiple processing sub-modules, and the data transfer relationship between the multiple processing sub-modules is one-way sequential transfer. Each processing sub-module includes a normalization sub-module, a convolution module, and an upsampling module. Among them, the input data of the first normalization sub-module is the latent space variable and the constant tensor, and the input data of other normalization sub-modules except the first normalization sub-module in all processing sub-modules is the output data of the previous module and the latent space variable. The input data of the convolution module is the output data of the previous module, and the input data of the upsampling module is the output data of the previous module.
[0010] In one embodiment, inputting the latent space variable and the constant tensor into the first normalization sub-module for normalization processing to obtain a first normalization result, including: Map the latent space variable to a scaling factor and a bias factor; Perform weighted summation by combining the scaling factor, the bias factor, and the sub-feature map corresponding to the target face image to obtain the first normalization result. Among them, the scaling factor characterizes the contrast and texture detail intensity of the target face image, and the bias factor characterizes the hue trend of the target face image.
[0011] In one embodiment, performing iterative calculation on the initial three-dimensional reconstruction unit based on the pose parameters and the face feature parameters to obtain the target three-dimensional reconstruction unit, including: Perform calculation on the initial three-dimensional reconstruction unit based on the pose parameters and the face feature parameters to obtain a basic three-dimensional reconstruction unit, and obtain a set of preliminarily deblurred face images according to the basic three-dimensional reconstruction unit; Blur the preliminarily de-blurred face image set to obtain multiple blurred face images, calculate the loss function results between the blurred face images and the corresponding face images to be processed one by one, and backpropagate the gradients of the loss function results into the 3D reconstruction model to perform iterative calculations on the initial 3D reconstruction unit until the 3D reconstruction model including the initial 3D reconstruction unit converges to generate the target 3D reconstruction unit, where the loss function results include the per-pixel loss term and the structural similarity loss term between the blurred face image and the face image to be processed.
[0012] In one embodiment, obtaining the initial 3D reconstruction unit includes: Extract the key feature points in the face image set to be processed, and determine the matching relationship of the key feature points between each face image to be processed; Based on the matching relationship, determine the 3D coordinates of each key feature point to obtain the initial 3D point cloud, where the initial 3D point cloud represents the geometric structure of the face of the person object; Based on the initial 3D point cloud, establish the initial 3D reconstruction unit.
[0013] In a second aspect, the present application also provides a 3D reconstruction device based on an image set. The device includes: An acquisition module for acquiring a face image set to be processed, where the face image set to be processed includes multiple face images to be processed at different angles of the same person object; A calculation module for extracting the pose parameters of the face image set to be processed, upsampling the face image set to be processed to obtain a target face image set, determining face feature parameters based on the target face image set; and calculating the initial 3D reconstruction unit based on the geometric structure of the person object determined from the face images to be processed at different angles; A generation module for inputting the pose parameters, face feature parameters, and the initial 3D reconstruction unit into the 3D reconstruction model, performing iterative calculations on the initial 3D reconstruction unit based on the pose parameters and face feature parameters to obtain the target 3D reconstruction unit, and establishing a face 3D model corresponding to the person object based on the target 3D reconstruction unit.
[0014] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Acquire a face image set to be processed, where the face image set to be processed includes multiple face images to be processed at different angles of the same person object; Extract the pose parameters of the face image set to be processed, upsample the face image set to be processed to obtain the target face image set, and determine the face feature parameters based on the target face image set; and calculate the initial three-dimensional reconstruction unit based on the geometric structure of the person object determined by the face images to be processed at different angles. Input the pose parameters, face feature parameters, and the initial three-dimensional reconstruction unit into the three-dimensional reconstruction model, perform iterative calculations on the initial three-dimensional reconstruction unit based on the pose parameters and face feature parameters to obtain the target three-dimensional reconstruction unit, and establish a three-dimensional face model corresponding to the person object based on the target three-dimensional reconstruction unit.
[0015] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented: Obtain a face image set to be processed, where the face image set to be processed includes multiple face images to be processed at different angles for the same person object; Extract the pose parameters of the face image set to be processed, upsample the face image set to be processed to obtain the target face image set, and determine the face feature parameters based on the target face image set; and calculate the initial three-dimensional reconstruction unit based on the geometric structure of the person object determined by the face images to be processed at different angles. Input the pose parameters, face feature parameters, and the initial three-dimensional reconstruction unit into the three-dimensional reconstruction model, perform iterative calculations on the initial three-dimensional reconstruction unit based on the pose parameters and face feature parameters to obtain the target three-dimensional reconstruction unit, and establish a three-dimensional face model corresponding to the person object based on the target three-dimensional reconstruction unit.
[0016] The above three-dimensional reconstruction method, device, equipment, and medium based on an image set first obtain a face image set to be processed, extract the pose parameters of the face image set to be processed, upsample the face image set to be processed to obtain the target face image set, determine the face feature parameters based on the target face image set, and calculate the initial three-dimensional reconstruction unit based on the geometric structure of the person object determined by the face images to be processed at different angles; finally, input the pose parameters, face feature parameters, and the initial three-dimensional reconstruction unit into the three-dimensional reconstruction model, perform iterative calculations on the initial three-dimensional reconstruction unit based on the pose parameters and face feature parameters to obtain the target three-dimensional reconstruction unit, and establish a three-dimensional face model corresponding to the face object based on the target three-dimensional reconstruction unit. Through the three-dimensional reconstruction method based on face images given in the present application, it is possible to achieve scene reconstruction using a face image set to be processed composed of multiple blurred face images, and it is also possible to efficiently complete the three-dimensional face modeling of a person object. Description of the Drawings
[0017] Figure 1 It is an application environment diagram of a 3D reconstruction method in an embodiment; Figure 2 It is a schematic flowchart of a 3D reconstruction method in an embodiment; Figure 3 It is a schematic diagram of the structure of a normalization module in an embodiment; Figure 4 It is a schematic flowchart of the modeling process of a 3D human face model in a preferred embodiment; Figure 5 It is a schematic flowchart of the process of obtaining a deblurred human face image set based on a 3D model in an embodiment; Figure 6 It is a structural block diagram of a 3D reconstruction device in an embodiment; Figure 7 It is an internal structure diagram of a computer device in an embodiment. Specific embodiments
[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] The 3D reconstruction method for an image set provided by an embodiment of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other network servers. First, obtain a set of to-be-processed human face images, where the set of to-be-processed human face images includes multiple to-be-processed human face images of the same human object from different angles. Then, extract the pose parameters of the set of to-be-processed human face images, perform upsampling processing on the set of to-be-processed human face images to obtain a set of target human face images, determine the human face feature parameters based on the set of target human face images, and determine the geometric structure of the human object based on the to-be-processed human face images, calculate to obtain an initial 3D reconstruction unit. Finally, input the pose parameters, human face feature parameters and the initial 3D reconstruction unit into a trained 3D reconstruction model to obtain a target 3D reconstruction unit, and establish a 3D human face model corresponding to the human object based on the target 3D reconstruction unit. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0020] In one embodiment, as Figure 2 shown, a three-dimensional reconstruction method based on an image set is provided. Taking the server in Figure 1 as an example, the method includes the following steps: Step S210: Obtain a face image set to be processed, where the face image set to be processed includes multiple face images to be processed at different angles for the same person object.
[0021] Specifically, the face image set to be processed includes multiple face image sets to be processed (15 - 50 images) taken at intervals of 15° for the same person object I ori , where the face images to be processed are relatively blurred images. In practical applications, the face images to be processed can be blurred images caused by shooting mistakes or taken in special scenarios. The judgment of blurred images includes, but is not limited to, judging based on the sharpness of the graphic edges in the image. If the graphic edges are sharp and the outline and details of the object can be clearly distinguished, the image is considered relatively clear; otherwise, the image is considered relatively blurred. Or, judging based on the gray-scale change of the image. If the gray-scale change in the image is obvious and there are rich hierarchical details, the image is considered relatively clear; otherwise, the image is considered relatively blurred, and so on.
[0022] Step S220: Extract the pose parameters of the face image set to be processed, perform upsampling processing on the face image set to be processed to obtain a target face image set, determine face feature parameters based on the target face image set; and calculate an initial three-dimensional reconstruction unit based on the geometric structure of the person object determined from the face images to be processed at different angles.
[0023] Specifically, extract the pose parameters of each face image to be processed in the face image set to be processed. Among them, the geometric structure of the three-dimensional scene can be restored from the two-dimensional image set based on the camera localization and mapping algorithm (abbreviated as COLMAP), so as to extract the pose parameters of the person object. And perform upsampling processing on each face image to be processed in the face image set to be processed. In some preferred embodiments, the resolution of the face image to be processed can be increased to 1024 pixels × 1024 pixels, so as to obtain a high-resolution target face image set. Further, determine face feature parameters according to the target face image set. Among them, face feature parameters can be determined from the target face image set based on deep learning methods, or face feature parameters can be calculated by methods such as principal component analysis. Face feature parameters include various parameters that can reflect face information, and these parameters can be geometric features, texture features, and so on.
[0024] Further, determine the geometric structure of the person object based on the face images to be processed at different angles in the face image set. Specifically, the methods for calculating the geometric structure include, but are not limited to, first extracting feature points from each image in the face image set to be processed (or the target face image set), and matching these feature points between different images. Calculate the three-dimensional coordinates of each feature point according to the matching results, and combine the camera parameters (intrinsic and extrinsic parameters) of each image to obtain a preliminary three-dimensional point cloud. This three-dimensional point cloud represents the geometric structure of the above-mentioned person object. Calculate the initial three-dimensional reconstruction unit based on the geometric structure. The initial three-dimensional reconstruction unit is a three-dimensional Gaussian function reconstruction unit. Specifically, the initial three-dimensional reconstruction unit represents each point or a group of points in the three-dimensional point cloud of the face geometric structure modeled as a three-dimensional Gaussian distribution. Among them, it can be realized by calculating the density distribution of the surrounding point cloud for each point or a group of points. For example, a calculation radius can be preset, and then calculate the average position and variance of all points within this radius to obtain the parameters of the three-dimensional Gaussian distribution. Further, many points in the point cloud correspond to a series of three-dimensional Gaussian function reconstruction units. At this time, this series of three-dimensional Gaussian function reconstruction units is the above-mentioned initial three-dimensional reconstruction unit. The initial three-dimensional reconstruction unit is the basic building block for representing the face scene. Its essence is a three-dimensional Gaussian function. For any three-dimensional variable, its multi-dimensional Gaussian distribution probability function can be defined as: ; As can be seen from the above, the initial three-dimensional reconstruction unit is the preliminary result of modeling the point cloud data based on the three-dimensional Gaussian distribution. Each point or a group of points in the point cloud can be regarded as a Gaussian distribution in a three-dimensional space. Each such Gaussian distribution can be regarded as a reconstruction unit, and these reconstruction units together constitute the reconstruction of the face scene.
[0025] Step S230, input the pose parameters, face feature parameters, and the initial three-dimensional reconstruction unit into the three-dimensional reconstruction model, perform iterative calculations on the initial three-dimensional reconstruction unit based on the pose parameters and face feature parameters to obtain the target three-dimensional reconstruction unit, and establish a three-dimensional face model corresponding to the person object based on the target three-dimensional reconstruction unit.
[0026] Specifically, input the obtained pose parameters, face feature parameters, and initial 3D reconstruction unit into a 3D reconstruction model. Based on the pose parameters and face feature parameters, iteratively calculate the trainable parameters in the initial 3D reconstruction unit to obtain the target 3D reconstruction unit. Among them, the 3D reconstruction model can be selected by relevant technicians according to actual needs, and preferably can be a multi-layer perceptron. Input the above three parameters into the multi-layer perceptron, iteratively calculate the trainable parameters in the initial 3D reconstruction unit, and thus output a target 3D reconstruction unit that conforms to a 3D Gaussian distribution. The target 3D reconstruction unit is a set of 3D Gaussian distribution parameters after iterative optimization, which characterizes the geometric structure and feature distribution of the face object in 3D space. Finally, establish a 3D face model of the person object based on the target 3D reconstruction unit. Specifically, the face feature parameters represent a certain feature of the face, such as 2D key point coordinates, extracted high-dimensional feature vectors, texture, or appearance information; the above pose parameters are generally camera pose parameters, and the camera pose parameters characterize the camera projection matrix of each image, including the internal parameter matrix K, rotation matrix R i , translation vector T i . Input the pose parameters, face feature parameters, and initial 3D reconstruction unit into the 3D reconstruction model. By iteratively optimizing the trainable parameters in the initial 3D reconstruction unit, the finally generated target 3D reconstruction unit can conform to the 2D observations of the input images, be consistent with the face features, and maintain the mathematical properties of the 3D Gaussian distribution.
[0027] Furthermore, the trainable parameters in the above initial 3D reconstruction unit include but are not limited to the mean vector μ j , which characterizes the 3D point position of the person object; the covariance matrix ∑j, which characterizes the distribution shape of the person object; and the weight ωj, which is a parameter in the calculation and is used for weight update in the mixture model. And the above pose parameters are used to provide the projection relationship from 3D to 2D to ensure that the trainable parameters in the initial 3D reconstruction unit obtained after iterative calculation are consistent with the observations from all viewpoints; the above face feature parameters are constraint conditions that can guide the adjustment of the trainable parameters in the initial 3D reconstruction unit. For example, if the above face feature parameters are key point coordinates, then the above mean vector μ j should be close to the 3D positions of these points. If the above face feature parameters are texture information, then the above covariance matrix ∑j needs to reflect the local geometric features of the person object.
[0028] Through steps S210 to S230, it is possible to complete the reconstruction of the 3D face model of the person object based on a set of to-be-processed face images (15 - 50 images) that are blurred and of low image quality at different angles for the same person object. Furthermore, this application completes 3D reconstruction based on pose parameters, face feature parameters, and an initial 3D reconstruction unit, with higher efficiency and can save computing resources.
[0029] In some of these embodiments, after obtaining the target three-dimensional reconstruction unit, the method further includes: Projecting the three-dimensional face model onto a two-dimensional plane through differentiable rasterization to generate a set of de-blurred face image sequences.
[0030] Specifically, for observing the three-dimensional model during display, in this embodiment, a set of face image sequences can be generated based on the above three-dimensional face model, that is, the target three-dimensional reconstruction unit is projected onto a two-dimensional plane through differentiable rasterization, so as to generate a set of de-blurred face image sequences. Through this embodiment, it is not only convenient for relevant technical personnel to observe and analyze the model, but also further reduces the requirements for the face image set to be processed. Even if the face image set to be processed is a blurred face image, the de-blurring process of the face image set to be processed can be realized through this application, and finally a clear set of face image sequences can be obtained.
[0031] In some of these embodiments, upsampling the face image set to be processed to obtain a target face image set includes: Sequentially taking each face image to be processed in the face image set to be processed as a first face image; Mapping the first face image to a latent space variable; Inputting the latent space variable and a preset constant tensor into a preset normalization module for normalization processing to obtain a target face image corresponding to the first face image; Based on the target face images corresponding to each image in the face image set to be processed, obtaining the target face image set.
[0032] Specifically, this embodiment provides a method for upsampling the face image set to be processed. Sequentially taking each face image to be processed in the face image set to be processed as a first face image and mapping the first face image to a latent space variable w. Specifically, the first face image can be input into a preset upsampling reconstruction module and passed through multiple (preferably set to 8) fully connected layers, so as to map the image information of the first face image to the latent space variable w.
[0033] Then, the latent space variable and a preset constant tensor are input into a preset normalization module for normalization processing to obtain a target face image corresponding to the first face image. The normalization module includes a normalization sub-module and a convolution sub-module. The normalization sub-module can reduce the bias of data distribution, making it easier for the model to learn and generalize. The convolution sub-module can extract features in the image and generate a high-resolution image through upsampling and further convolution operations. In summary, the first face image is input into the normalization module to obtain the corresponding target face image. Based on the processed target face images corresponding to each image in the face image set to be processed, a target face image set is obtained. In summary, this embodiment provides a method for upsampling a face image to be processed, which can improve the resolution of the face image to be processed and facilitate the processing of the face image in subsequent steps.
[0034] In one embodiment, the normalization module includes a plurality of processing sub-modules, and the data transfer relationship between the plurality of processing sub-modules is unidirectional and sequential. Each processing sub-module includes a normalization sub-module, a convolution module, and an upsampling module. Among them, the input data of the first normalization sub-module is the latent space variable and the constant tensor, and the input data of other normalization sub-modules except the first normalization sub-module in all processing sub-modules is the output data of the previous module and the latent space variable. The input data of the convolution module is the output data of the previous module, and the input data of the upsampling module is the output data of the previous module.
[0035] Specifically, Figure 3 FIG. is a schematic structural diagram of a normalization module in an embodiment. The normalization module includes a plurality of processing sub-modules, and the data transfer relationship between the plurality of processing sub-modules is unidirectional and sequential. Each processing sub-module includes a normalization module, a convolution module, and an upsampling module. In a Figure 3 preferred embodiment taking... as an example, the arrangement order of each module in each processing sub-module is: normalization module - convolution module - normalization module - upsampling module - convolution module. Among them, the convolution module is preferably a 3×3 convolution module. Multiple processing sub-modules can be set in the normalization module, and the number of processing sub-modules can be determined according to actual needs in practical applications. For example, 9 processing sub-modules can increase the resolution of the face image to be processed to 1024×1024 pixels.
[0036] The above constant tensor can be set by relevant technical personnel according to actual needs. The introduction of the constant tensor is used to provide an initial feature basis and control the global consistency of the generation process. In some preferred embodiments, the constant tensor can be set to 4×4×512. In summary, combined with Figure 3It can be known that the input data of the first normalization sub-module are the latent space variable and the constant tensor, and the input data of each convolutional module are the output data of the previous module. Moreover, among all the processing sub-modules, except for the first normalization sub-module, the input data of other normalization sub-modules are the output data of the previous module and the latent space variable. For the 9 processing sub-modules, the resolution of the initially input first face image is improved to 4×4 pixels, 8×8 pixels, 16×16 pixels, 32×32 pixels, 64×64 pixels, 128×128 pixels, 256×256 pixels, 512×512 pixels, and 1024×1024 pixels respectively. The above normalization sub-module can be set as an AdaIN (Adaptive Instance Normalization) module.
[0037] In one embodiment, the latent space variable and the constant tensor are input into the first normalization sub-module for normalization processing to obtain the first normalization result, including: Mapping the latent space variable into a scaling factor and a bias factor; Performing weighted summation by combining the scaling factor, the bias factor and the sub-feature map corresponding to the target face image to obtain the first normalization result, where the scaling factor represents the contrast and texture detail intensity of the target face image, and the bias factor represents the hue trend of the target face image.
[0038] Specifically, each first face image includes i feature maps, and each feature map x i is individually normalized. First, the latent space variable is mapped into a scaling factor y s,i and a bias factor y b,i through a learnable affine transformation. Then, the scaling factor and the bias factor are weighted-summed with the features of the above feature map once. This process is regarded as the process of normalizing the feature map x i once through the above normalization sub-module to obtain the first normalization result. This first normalization result is the output of the normalization sub-module, that is, the input of the convolutional module in the same processing sub-module. The process of the normalization sub-module performing normalization once can be expressed as: ; Among them, represents the calculated mean value of x i , represents the calculated standard deviation of x i .
[0039] In one embodiment, iterative calculation is performed on the initial three-dimensional reconstruction unit based on the pose parameter and the face feature parameter to obtain the target three-dimensional reconstruction unit, including: Calculate the initial 3D reconstruction unit based on the pose parameters and face feature parameters to obtain the basic 3D reconstruction unit, and obtain the preliminary de-blurred face image set according to the basic 3D reconstruction unit; Blur the preliminary de-blurred face image set to obtain multiple blurred face images, calculate the loss function results between the blurred face images and the corresponding face images to be processed one by one, and backpropagate the gradients of the loss function results into the 3D reconstruction model to perform iterative calculations on the initial 3D reconstruction unit until the 3D reconstruction model including the initial 3D reconstruction unit converges to generate the target 3D reconstruction unit, where the loss function results include the per-pixel loss term and the structural similarity loss term between the blurred face image and the face image to be processed.
[0040] Specifically, this embodiment provides a method for iteratively calculating the trainable parameters in the 3D reconstruction unit until the target 3D reconstruction unit is obtained. First, after obtaining the face image set to be processed, obtain the pose parameters, face feature parameters, and initial 3D reconstruction unit based on the face image set to be processed, and then optimize and calculate the parameters of the initial 3D reconstruction unit based on the pose parameters and face feature parameters to obtain the optimized basic 3D reconstruction unit. After obtaining the basic 3D reconstruction unit, obtain the basic face 3D model based on the basic 3D reconstruction unit, and project the basic face 3D model onto the 2D plane through differentiable rasterization to generate the above-mentioned preliminary de-blurred face image set.
[0041] Blur each of the preliminary de-blurred face images in the preliminary de-blurred face image set again based on the motion blur formula, and the motion blur formula is as follows: ; The blurred image here is the average value of n corresponding clear images. In the formula, B(x) represents the above-mentioned blurred face image after re-blurring, and C l (x) represents the above-mentioned preliminary de-blurred face image.
[0042] In summary, multiple blurred face images after re-blurring can be obtained. Calculate the loss function results between the blurred face images and the corresponding face images to be processed, where each blurred face image has a corresponding face image to be processed one by one, the shooting angles of the two are the same, and both are blurred images. Among them, in this embodiment, the loss function is composed of a per-pixel loss term and a structural similarity loss term, that is, the loss function L all has: L all =(1 - λ)L pixel +λL SSIM , where usually λ = 0.2, L pixel is the per-pixel loss term, and L SSIM is the structural similarity loss term.
[0043] Specifically, the above-mentioned per-pixel loss term L pixel can be expressed as: ; where M represents the total number of pixels in each image, B(x) represents the blurred face image, C(x) represents the corresponding face image to be processed, and || || represents the L 1 norm.
[0044] The above-mentioned structural similarity loss term L SSIM can be expressed as: ; where q represents the total number of pixel blocks used to compare the difference values in each image, B i represents the pixel block of the blurred face image, C i represents the pixel block of the face image to be processed, and SSIM (Structural Similarity Index) is an index used to measure the similarity between the blurred face image and the corresponding face image to be processed, and usually completes the similarity comparison through several aspects such as the brightness, contrast, and structure of the two images.
[0045] Calculate the result of the loss function between the blurred face image and the face image to be processed through the above method, and backpropagate the gradient of the loss function result to the three-dimensional reconstruction model to perform iterative calculation on the trainable parameters in the initial three-dimensional reconstruction unit. That is, in the three-dimensional modeling stage, it is necessary to perform iterative calculation on the above-mentioned initial three-dimensional reconstruction unit based on the pose parameters, face feature parameters, and the backpropagated gradient until the three-dimensional reconstruction unit converges. The specific convergence conditions include but are not limited to: the iteration reaches a preset number of iterations (such as 100 times, 300 times, etc.), and the change in the loss function result is less than a preset threshold ε, that is, there is L (t+1) -L (t) < ε, and so on. In some preferred embodiments, the above three-dimensional reconstruction model can be selected to use a Multilayer Perceptron (MLP). In summary, in this embodiment, in the modeling stage, the initial three-dimensional reconstruction unit performs iterative calculation based on the above pose parameters and face feature parameters, and finally outputs a trained target three-dimensional reconstruction unit that conforms to the three-dimensional features of the face.
[0046] In one embodiment, obtaining the initial three-dimensional reconstruction unit includes: Extracting key feature points from the face image set to be processed, and determining the matching relationship of key feature points between each face image to be processed; Based on the matching relationship, determine the three-dimensional coordinates of each key feature point to obtain an initial three-dimensional point cloud, where the initial three-dimensional point cloud represents the geometric structure of the face of the human object; Based on the initial three-dimensional point cloud, an initial three-dimensional reconstruction unit is established.
[0047] Specifically, the Structure from Motion algorithm can be used to extract key feature points from each image in the above-mentioned face image set to be processed. Among them, the image feature points can be extracted based on the SIFT (Scale-Invariant Feature Transform) key point detection algorithm. Match these feature points in different face images to be processed, construct the corresponding relationship of key feature points between each image, and thus obtain the matching relationship of key feature points based on the above corresponding relationship, where the matching relationship represents the matching result between key feature points in each image.
[0048] Then, based on the matching relationship, calculate the three-dimensional coordinates of each key feature point. Among them, the three-dimensional coordinates of each feature point can be calculated by triangulation. Further, combined with the camera parameters (including internal parameters and external parameters) corresponding to each image, an initial three-dimensional point cloud is calculated. The initial three-dimensional point cloud represents the geometric structure of the face of the human object. Finally, each point or a group of points in the point cloud is modeled as a three-dimensional Gaussian distribution. Many points contained in the point cloud correspond to a series of three-dimensional Gaussian function reconstruction units. At this time, this series of three-dimensional Gaussian function reconstruction units is the above-mentioned initial three-dimensional reconstruction unit.
[0049] This application also provides a preferred embodiment of a three-dimensional reconstruction method based on a face image set.
[0050] For the modeling stage of the face three-dimensional model, first obtain the face image set to be processed, and based on the face image set to be processed, obtain the pose parameters, face feature parameters, and initial three-dimensional reconstruction unit. Then, optimize and calculate the parameters of the initial three-dimensional reconstruction unit based on the pose parameters and face feature parameters to obtain a basic three-dimensional reconstruction unit. After obtaining the basic three-dimensional reconstruction unit, obtain a basic face three-dimensional model based on the basic three-dimensional reconstruction unit, project the basic face three-dimensional model onto a two-dimensional plane to obtain a preliminarily de-blurred face image set, and perform re-blurring processing on each preliminarily de-blurred face image in the preliminarily de-blurred face image set to obtain a re-blurred face image set. Then, calculate the loss function between the blurred face image and the corresponding face image to be processed one by one, and perform backpropagation based on the loss function result to iteratively train the trainable parameters in the initial three-dimensional reconstruction unit until the initial three-dimensional reconstruction unit converges, so as to obtain the target three-dimensional reconstruction unit. Figure 4For the modeling stage of the 3D face model in an embodiment, it should be noted that this embodiment mainly trains the trainable parameters in the initial 3D reconstruction unit. Figure 4 Other parts passed by the red arrow in Figure 4 do not contain trainable parameters and only play a role in propagating gradients.
[0051] After the 3D face model is modeled, as Figure 5 shown, Figure 5 As shown in Figure 5 , it is a schematic flowchart of the process for obtaining a 2D deblurred face image sequence set based on the 3D face model in an embodiment. After obtaining the above 3D face model, in order to visually observe the 3D model, the target 3D reconstruction unit can be projected onto a 2D plane through a differentiable rasterization process to generate a clear face image sequence set. In summary, the present application can not only achieve efficient 3D scene reconstruction but also deblur multiple blurred face images to obtain a clear face image set, that is, the deblurred face image set I. Out .
[0052] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0053] Based on the same inventive concept, the embodiments of the present application also provide a 3D reconstruction device for face images for implementing the 3D reconstruction method for face images involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the 3D reconstruction device for face images provided below can refer to the limitations on the 3D reconstruction method for face images in the above text and will not be elaborated here.
[0054] In an embodiment, as Figure 6 shown, a 3D reconstruction device based on an image set is provided, including: an acquisition module 61, a calculation module 62, and a generation module 63, where: The acquisition module 61 is used to acquire a set of face images to be processed, where the set of face images to be processed includes multiple face images to be processed at different angles for the same person object; A calculation module 62 is configured to extract pose parameters of a face image set to be processed, perform upsampling processing on the face image set to be processed to obtain a target face image set, and determine face feature parameters based on the target face image set; and calculate an initial three-dimensional reconstruction unit based on the geometric structure of a person object determined from face images to be processed at different angles. A generation module 63 is configured to input the pose parameters, face feature parameters, and the initial three-dimensional reconstruction unit into a trained three-dimensional reconstruction model, perform iterative calculation on the initial three-dimensional reconstruction unit based on the pose parameters and face feature parameters to obtain a target three-dimensional reconstruction unit, and establish a three-dimensional face model corresponding to the person object based on the target three-dimensional reconstruction unit.
[0055] Each module in the above three-dimensional reconstruction device for image sets can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in the processor of a computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0056] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structural diagram can be as Figure 7 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to the three-dimensional reconstruction of face images. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a three-dimensional reconstruction method based on a face image set.
[0057] Those skilled in the art can understand that Figure 7 the structure shown in
[0058] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Obtain a face image set to be processed, where the face image set to be processed includes multiple face images to be processed at different angles for the same person object; Extract the pose parameters of the face image set to be processed, upsample the face image set to be processed to obtain the target face image set, and determine the face feature parameters based on the target face image set; and calculate the initial 3D reconstruction unit based on the geometric structure of the person object determined by the face images to be processed at different angles. Input the pose parameters, face feature parameters, and the initial 3D reconstruction unit into the 3D reconstruction model, and perform iterative calculations on the initial 3D reconstruction unit based on the pose parameters and face feature parameters to obtain the target 3D reconstruction unit, and establish a 3D face model corresponding to the person object based on the target 3D reconstruction unit.
[0059] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0060] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0061] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0062] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A three-dimensional reconstruction method based on an image set, characterized in that: The method comprises: Acquire a set of face images to be processed, wherein the set of face images to be processed includes a plurality of face images to be processed taken from different angles of the same person object; Extracting the pose parameters of the face image set to be processed, upsampling the face image set to be processed to obtain a target face image set, and determining face feature parameters based on the target face image set; and calculating an initial three-dimensional reconstruction unit based on the geometric structure of the human object determined by the face images to be processed at different angles; The posture parameters, facial feature parameters and the initial three-dimensional reconstruction unit are input into a three-dimensional reconstruction model, the initial three-dimensional reconstruction unit is iteratively calculated based on the posture parameters and the facial feature parameters to obtain a target three-dimensional reconstruction unit, and a three-dimensional facial model corresponding to the character object is established based on the target three-dimensional reconstruction unit.
2. The method according to claim 1, characterized in that After obtaining the target 3D reconstruction unit, the method further includes: The three-dimensional face model is projected onto a two-dimensional plane through differentiable rasterization to generate a deblurred face image sequence set.
3. The method according to claim 1, characterized in that The step of upsampling the face image set to be processed to obtain a target face image set includes: Sequentially taking each of the to-be-processed face images in the to-be-processed face image set as a first face image; Mapping the first face image into a latent space variable; Inputting the latent space variables and the preset constant tensor into a preset normalization module for normalization processing to obtain a target face image corresponding to the first face image; Based on the target face image corresponding to each image in the face image set to be processed, the target face image set is obtained.
4. The method according to claim 3, characterized in that The normalization module includes multiple processing sub-modules, and the data transmission relationship between the multiple processing sub-modules is unidirectional and sequential transmission. Each of the processing sub-modules includes a normalization sub-module, a convolution module and an upsampling module, wherein the input data of the first normalization sub-module is the latent space variable and the constant tensor, and the input data of all normalization sub-modules except the first normalization sub-module in all the processing sub-modules are the output data of the previous module and the latent space variable, the input data of the convolution module is the output data of the previous module, and the input data of the upsampling module is the output data of the previous module.
5. The method according to claim 4, characterized in that The step of inputting the latent space variables and the constant tensor into a first normalization submodule for normalization to obtain a first normalization result includes: Mapping the latent space variables into scaling factors and bias factors; The first normalization result is obtained by combining the scaling factor, the deviation factor and the sub-feature map corresponding to the target face image for weighted summation, wherein the scaling factor represents the contrast and texture detail intensity of the target face image, and the deviation factor represents the tone trend of the target face image.
6. The method according to claim 1, characterized in that The iterative calculation of the initial 3D reconstruction unit based on the posture parameter and the facial feature parameter to obtain a target 3D reconstruction unit includes: The initial 3D reconstruction unit is calculated based on the posture parameters and the facial feature parameters to obtain a basic 3D reconstruction unit, and a preliminary deblurred facial image set is obtained according to the basic 3D reconstruction unit; The face image set after preliminary deblurring is blurred to obtain multiple blurred face images, and the loss function results between the blurred face images and the corresponding face images to be processed are calculated, and the gradient of the loss function result is reversely transmitted to the 3D reconstruction model, and the initial 3D reconstruction unit is iteratively calculated until the 3D reconstruction model including the initial 3D reconstruction unit converges to generate the target 3D reconstruction unit, wherein the loss function result includes a pixel-by-pixel loss term and a structural similarity loss term between the blurred face images and the face images to be processed.
7. The method according to claim 1, characterized in that Acquiring the initial three-dimensional reconstruction unit includes: Extracting key feature points from the set of face images to be processed, and determining matching relationships between the key feature points between the face images to be processed; Based on the matching relationship, determine the three-dimensional coordinates of each key feature point to obtain an initial three-dimensional point cloud, wherein the initial three-dimensional point cloud represents the geometric structure of the face of the person object; The initial three-dimensional reconstruction unit is established based on the initial three-dimensional point cloud.
8. A three-dimensional reconstruction device based on an image set, characterized in that: The device comprises: An acquisition module, used for acquiring a set of face images to be processed, wherein the set of face images to be processed includes a plurality of face images to be processed taken from different angles of the same person object; A calculation module is used to extract the posture parameters of the face image set to be processed, perform upsampling processing on the face image set to be processed to obtain a target face image set, determine the face feature parameters based on the target face image set; and calculate an initial three-dimensional reconstruction unit based on the geometric structure of the human object determined by the face images to be processed at different angles; A generation module is used to input the posture parameters, facial feature parameters and the initial three-dimensional reconstruction unit into a three-dimensional reconstruction model, iteratively calculate the initial three-dimensional reconstruction unit based on the posture parameters and the facial feature parameters to obtain a target three-dimensional reconstruction unit, and establish a three-dimensional facial model corresponding to the character object based on the target three-dimensional reconstruction unit.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Three-dimensional face reconstruction and six-degree-of-freedom pose estimation method, device and equipment
CN116843834A
Face three-dimensional reconstruction method and device, electronic equipment and storage medium
CN116958406A
Human face three-dimensional reconstruction method and device based on artificial intelligence, equipment and medium
CN119228991A
Three-dimensional face reconstruction method, apparatus, and device, medium, and product
WO2024032464A1