Appearance capture
By receiving multiple object images and processing these images using deep learning neural network models, the problem of difficulty in capturing high-quality 3D data using a single camera under ambient lighting conditions in the prior art is solved, and efficient and economical 3D data capture is achieved.
Patent Information
- Application Number
- CN202380078358.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-24
- Filing Date
- 2023-10-20
- Publication Date
- 2025-06-24
AI Technical Summary
Prior art requires expensive highly specialized light field settings when capturing high-quality 3D geometric data and appearance data, and it is difficult to use a single camera for image capture under ambient lighting conditions.
By receiving multiple object images, each corresponding to a different view orientation, these images are processed using a deep learning neural network model to determine diffuse and specular maps of the object surface and generate grid and normal maps.
It enables the capture of high-quality 3D geometric and appearance data using a single camera under ambient lighting conditions, reducing equipment costs and complexity and improving accessibility of image capture.
Smart Images

Figure CN120202497A_ABST
Abstract
Description
Background Art
[0001] In computer graphics, a great deal of attention has been paid to high-quality 3D acquisition of faces or objects / materials of subjects including 3D shapes and appearances for realistic rendering applications ranging from movie visual effects and games to product design / visualization / advertising and AV / VR applications.
[0002] Expensive and highly specialized light field setups for face capture have been described, see for example "Acquiring the reflectance field of a human face", Paul Debevec, Tim Hawkins, Chris Tchou, Haarm-Pieter Duiker, Westley Sarokin and Mark Sagar, Proceedings of the 27th annual conference on Computer graphics and interactive techniques (SIGGRAPH), 2000 (hereinafter referred to as "Debevec 2000"). See also "Rapid acquisition of specular and diffuse normal maps from polarized spherical gradient illumination", Wan-Chun Ma, Tim Hawkins, Pieter Peers, Charles-Felix Chabert, Malte Weiss and Paul Debevec, EGSR'07: Proceedings of the 18th Eurographics conference on Rendering Techniques, pp. 183-194, June 2007 (hereinafter referred to as "Ma 2007"). See also "Multiview face capture using polarized spherical gradient illumination", Abhijeet Ghosh, Graham Fyffe, Borom Tunwattanapong, Jay Busch, Xueming Yu and Paul Debevec, ACM Transactions on Graphics (TOG) 30, 6, (2011) (hereinafter referred to as "Ghosh 2011").See also “Diffuse-specular separation using binary spherical gradient illumination”, Christos Kampouris, Stefanos Zafeiriou, Abhijeet Ghosh, SR'18: Proceedings of the Eurographics Symposium on Rendering: Experimental Ideas & Implementations, July 2018 (hereinafter referred to as “Kampouris 2018”). Such light field setups enable the highest quality and flexibility in rendering / relighting.
[0003] However, while such methods can be used to capture high-quality geometric and appearance data for high-end applications, there is also significant interest in more accessible image capture methods that can use a single camera and operate under ambient lighting conditions. For example, see:
[0004] - “Authentic Volumetric Avatars from a Phone Scan” by Cao et al. ACM Trans. Graph. 41, 4, Article 1 (July 2022), https: / / doi.org / 10.1145 / 3528223.3530143 (hereinafter referred to as “CAO 2022”).
[0005] - “Learning to reconstruct shape and spatially-varying reflectance from a single image” by Li et al., ACM Transactions on Graphics, Vol. 37, No. 6, December 2018, Article 269, pp. 1-11, https: / / doi.org / 10.1145 / 3272127.3275055 (hereinafter referred to as “LI 2018”).
[0006] - “High-Fidelity 3D Digital Human Head Creation from RGB-D Selfies” by Bao et al., arXiv:2010.05562v2, https: / / doi.org / 10.48550 / arXiv.2010.05562 (hereinafter referred to as “BAO 2020”).
[0007] - "AvatarMe++: Facial Shape and BRDF Inference with Photorealistic Rendering-Aware GANs" by Lattas et al., arXiv:2112.05957v1, https: / / doi.org / 10.48550 / arXiv.2112.05957 (hereinafter referred to as "LATTAS2021").
[0008] - "AvatarMe: Realistically Renderable 3D Facial Reconstruction 'in-the-wild'" by Lattas et al., arXiv:2003.13845v1, https: / / doi.org / 10.48550 / arXiv.2003.13845 (hereinafter referred to as "LATTAS2020").
[0009] - "High-Fidelity Facial Reflectance and Geometry Inference From an Unconstrained Image" by Yamaguchi et al., ACM Trans. Graph. 37, 4, Article 162 (July 2018), 14 pages, https: / / doi.org / 10.1145 / 3130800.3130817 (hereinafter referred to as "YAMAGUCHI2018").
[0010] - "Two-Shot Spatially-Varying BRDF and Shape Estimation" by Boss et al., June 2020, Conference: 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), DOI: 10.1109 / CVPR42600.2020.00404 (hereinafter referred to as "BOSS2020"). Summary of the Invention
[0011] According to a first aspect of the present invention, there is provided a method including receiving a plurality of object images of an object. Each object image corresponds to a different view direction. The object images include a first object image and a second object image corresponding to a first direction and a second direction. The method further includes determining a mesh corresponding to a target region of the object surface based on a first subset including two or more of the plurality of object images. The method further includes determining a diffuse map and a specular map corresponding to the target region of the object surface based on processing a second subset of the object images using a deep learning neural network model trained to estimate diffuse and specular albedo components based on an input image. The second subset includes at least the first object image and the second object image. The method further includes determining a normal map corresponding to the target region of the object surface based on the object images of the second subset. The method further includes storing and / or outputting the mesh, the diffuse map, the specular map, and the normal map.
[0012] Each part of the target region of the object surface may be imaged by at least one object image of the second subset. Each part of the target region of the object surface may be imaged by at least one of the first object image and the second object image. The second subset of the object images may include a third object image corresponding to a third direction. The diffuse map and the specular map corresponding to the target region of the object surface may be determined based on processing the second subset including the first object image, the second object image, and the third object image using a deep learning neural network model. The diffuse map and the specular map corresponding to the target region of the object surface may be determined based on processing the second subset including the first object image, the second object image, and the third object image and one or more other object images using a deep learning neural network model. The normal map corresponding to the target region of the object surface may be a tangent normal map determined based on high-pass filtering each of the second subset including the first object image, the second object image, and the third object image. The normal map corresponding to the target region of the object surface in the form of a tangent normal map may be determined based on high-pass filtering each of the second subset including the first object image, the second object image, and the third object image and the one or more other object images. The normal map corresponding to the target region of the object surface may be a photometric normal map.
[0013] In this way, the method is based on at least two object images. However, three, four or more object images can be used to determine the mesh and / or to determine each of the diffuse map and the specular map. The plurality of object images (first subset) for meshing can be independent of the plurality of object images (second subset) for diffuse-specular estimation. The second subset of the object images for determining the diffuse map and the specular map is always the same as the second subset of the object images for determining the normal map (e.g., tangent normal map).
[0014] Storing and / or outputting the mesh, the diffuse map, the specular map and the normal map (e.g., tangent normal map) can include storing and / or outputting a rendering generated based on the diffuse map, the specular map and the normal map.
[0015] Each of the first direction, the second direction and (if used) the third direction can be separated from each of the first direction, the second direction and (if used) the third direction by 30 degrees or more. The first direction can be at an angle of at least 30 degrees with the second direction, and if used, the first direction can be at an angle of at least 30 degrees with the third direction. The second direction can be at an angle of at least 30 degrees with the first direction, and if used, the second direction can be at an angle of at least 30 degrees with the third direction. When used, the third direction can be at an angle of at least 30 degrees with the first direction, and the third direction can be at an angle of at least 30 degrees with the second direction.
[0016] The first direction, the second direction and (if used) the third direction can be substantially coplanar. Substantially coplanar can refer to the possibility of defining a common plane that forms an angle of no more than 10 degrees with each of the first direction, the second direction and the third direction.
[0017] The first subset of the object images on which the mesh is based can include or take the form of the first object image, the second object image and (if used) the third object image. The first subset of the object images on which the mesh is based can exclude the first object image, the second object image and (if used) the third object image. In other words, the first subset and the second subset can intersect, or the first subset and the second subset can be mutually exclusive.
[0018] The target area can correspond to a part of the total surface of the object. The target area can include or take the form of a face.
[0019] The first direction and the second direction may be principal directions. When in use, the third direction may also be a principal direction.
[0020] When the target region is the face, the first direction and the second direction may be substantially coplanar and each at an angle of about ±45° relative to the front view (i.e., the first direction may be at -45° relative to the front view in the plane, while the second direction may be at +45°). The about 45° may correspond to 45° ± 10°. If used, the third direction may correspond to the front view. The front view may correspond to the front principal direction antiparallel to the vector average of the face normals at each surface point of the face. The face normals may be determined based on the mesh.
[0021] Preferably, the mesh is determined based on a first subset including more than three object images among the plurality of object images. Preferably, the mesh is determined based on a first subset including more than ten object images among the plurality of object images.
[0022] The plurality of object images may include video data. The method may further include extracting from the video data two or more object images that form the first subset on which the mesh determination is based.
[0023] The first object image and the second object image may be extracted from the video data. When in use, the third object image may be extracted from the video data. The first object image, the second object image, and (if used) the third object image may not be extracted from the video data.
[0024] The video data may correspond to a perspective that passes through at least 45° of a first arc that is generally centered on the object (centered on the front view for the face) and along at least 45° of a second arc that is generally centered on the object (centered on the front view for the face) and intersects the first arc at an angle between 30° and 90°. The video data corresponding to the first arc and the second arc may belong to a single continuous video clip. The video data corresponding to the first arc and the second arc may belong to separate video clips.
[0025] Determining a diffuse map and a specular map corresponding to a target region of an object surface may include: for each of a second subset of object images (which includes at least a first object image and a second object image), providing the object image as an input to a deep learning neural network model, and obtaining a corresponding camera-space diffuse map and a corresponding camera-space specular map as outputs. The method may further include generating a UV-space diffuse map based on projecting the camera-space diffuse map corresponding to the second subset of object images onto a mesh. The method may further include generating a UV-space specular map based on projecting the camera-space specular map corresponding to the second subset of object images onto a mesh.
[0026] The mesh may include UV-coordinates for each vertex. Projecting a diffuse map or a specular map onto the mesh will associate each pixel with UV-coordinates (which may be interpolated between vertices), resulting in the generation of a UV map.
[0027] When the same UV-coordinates correspond to pixels of more than two diffuse maps corresponding to the second subset, generating the UV-space diffuse map may include blending. When the UV-coordinates correspond to more than two diffuse maps corresponding to the second subset, generating the UV-space specular map may include blending. Blending may include any technique known in the field of multi-view texture capture, e.g., averaging the low-frequency response of more than two maps (or images), and imprinting (overlaying) the high-frequency response (high-pass filtering) from a single map (or image) corresponding to the view direction closest to the mesh normal at that UV coordinate. The low-frequency response of a map (or image) may be obtained by blurring the map (or image) (e.g., Gaussian blur). The high-frequency response may be obtained by subtracting the low-frequency response of the map (or image) from the map (or image).
[0028] Determining a diffuse map and a specular map corresponding to a target region of an object surface may include generating a UV-space input texture based on projecting each of a second subset of object images (which includes at least a first object image and a second object image) onto a mesh. The method may further include providing the UV-space input texture as an input to a deep learning neural network model, and obtaining a corresponding UV-space diffuse map and a corresponding UV-space specular map as outputs. When the UV-coordinates correspond to more than two of the object images of the second subset, generating the UV-space input texture may include blending.
[0029] In either case (camera-space or UV-space diffuse-specular estimation), the second subset may include one or more object images in addition to the first object image and the second object image. For example, a third object image and / or one or more further object images.
[0030] Determining a mesh corresponding to a target region of an object surface may include applying a structure-from-motion technique to a first subset of object images (including more than two object images of the plurality of object images).
[0031] The first subset of object images for the structure-from-motion technique may include a first image and a second image. The first subset of object images for the structure-from-motion technique may include or take the form of a first object image, a second object image, and a third object image. Preferably, the structure-from-motion technique may be applied to a first subset including fifteen to thirty object images (including the endpoints) of the plurality of object images.
[0032] The plurality of object images may include one or more depth maps of the target region and / or one or more structured light images of the target region. The first subset of object images on which the mesh determination is based may include one or more depth maps and / or one or more structured light images.
[0033] Determining a mesh corresponding to a target region of an object surface may include fitting a 3D deformable mesh model (3DMM) to a first subset of object images of the plurality of object images.
[0034] Determining a mesh corresponding to a target region of an object surface may include a neural surface reconstruction technique.
[0035] The neural surface reconstruction technique may take the form of the NeuS method described by Wang et al. in "NeuS: Learning Neural Implicit Surfaces for Multi-View Reconstruction via Volume Rendering", https: / / doi.org / 10.48550 / arXiv.2106.10689.
[0036] The deep learning neural network model for diffuse-specular reflection estimation may include or take the form of a multi-layer perceptron.
[0037] The deep learning neural network model for diffuse-specular reflection estimation may include or take the form of a convolutional neural network.
[0038] The deep learning neural network model may include or take the form of a generative adversarial network (GAN) model. The deep learning neural network model may include or take the form of the model described by Wang et al. in "High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs" at CVPR, 2018 (also known as "Pix2PixHD").
[0039] The deep learning neural network model can include or take the form of a U-net model. An example of a suitable U-net model is described in "Deep Polarimetric Imaging for 3D Shape and SVBRDF Acquisition" by Deschaintre et al., Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021.
[0040] The deep learning neural network model can include or take the form of a diffusion model. The diffusion model can be tile-based. The diffusion model can be patch-based. Examples of suitable diffusion models are described in O., & Legenstein, R. (2022) "Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models", arXiv, https: / / doi.org / 10.48550 / ARXIV.2207.14626.
[0041] The method can also include receiving a number of environmental images. Each environmental image can correspond to a field of view oriented away from the object. The deep learning neural network can also be configured to receive additional input based on the plurality of environmental images.
[0042] Each environmental image can correspond to an object image of the plurality of images, and the environmental image and the corresponding object image can have fields of view oriented in opposite directions. For example, the object image of the plurality of object images may have been captured using the rear camera of a mobile phone (smartphone) or a tablet, while the corresponding environmental image may have been captured simultaneously using the front (or "selfie") camera of the mobile phone or tablet (and vice versa). Each object image of the second subset can correspond to an environmental image.
[0043] The number of environmental images can include a first environmental image and a second environmental image corresponding to a first object image and a second object image, respectively. When a third object image is included, a corresponding third environmental image can be included. The first environmental image can correspond to a field of view opposite to the first object image from substantially the same position. The second environmental image can correspond to a field of view opposite to the second object image from substantially the same position. When used, the third environmental image can correspond to a field of view opposite to the third object image from substantially the same position.
[0044] In this way, by combining the input based on the environmental images, the deep learning neural network model can consider the environmental illumination conditions when estimating the diffuse and specular reflection components corresponding to the target region.
[0045] A deep learning neural network may include a first encoder branch configured to convert an input object image into a first latent representation; and a second encoder branch configured to convert an environmental image corresponding to the input object image into a second latent representation. The first latent representation and the second latent representation may be concatenated and then processed by a common decoder branch to generate an output diffuse map and a specular map.
[0046] The method may further include mapping a plurality of environmental images to an environment map and providing the environment map as an input to the deep learning neural network model.
[0047] The environment map may correspond to a sphere or a portion of a sphere centered approximately on the object. Each pixel of each environmental image may be mapped to a corresponding region on the surface of the sphere. Mapping the plurality of environmental images to the environment map may further include filling in missing regions of the environment map. For example, the method may include filling in regions of the environment map corresponding to the convex hull of the environmental images when projected onto the surface of the sphere. An image filling deep learning model may be applied to generate the filled regions of the environment map. Alternatively, it may not be necessary to use a separate filling deep learning model, and portions of the environment map may be directly encoded along with the camera pose and fed into the deep learning neural network model.
[0048] In other examples, the environment map may correspond to a cylinder or a portion of a cylinder.
[0049] A deep learning neural network may include a first encoder branch and a second encoder branch, the first encoder branch being configured to convert an input object image into a first latent representation, and the second encoder branch being configured to convert the environment map into a second latent representation. The first latent representation and the second latent representation may be concatenated and then processed by a common decoder branch to generate an output diffuse map and a specular map.
[0050] Determining a normal map (whether a tangent normal map or a photometric normal map) may include generating a camera space normal map based on high-pass filtering of a second subset of the object image. Determining the normal map may include generating a UV-space normal map based on projecting the camera space normal map corresponding to the second subset onto a mesh.
[0051] Determining the normal map may further include generating a third camera space normal map based on a third object image (e.g., a tangent normal map obtained based on high-pass filtering), and generating the UV-space normal map may be based on projecting the first camera space normal map, the second camera space normal map, and the third camera space normal map onto the mesh.
[0052] Determining a normal map (whether a tangent normal map or a photometric normal map) can include generating a UV-space input texture based on projecting a second subset of an object image onto a mesh, and generating a UV-space normal map based on high-pass filtering the UV-space input texture. When using a third object image, generating the UV-space input texture can be based on projecting the first image, the second image, and the third image onto the mesh.
[0053] When estimating a diffuse map and a specular map in UV space, the same UV-space input texture can be used as an input to a deep learning neural network model and for calculating the UV-space normal map.
[0054] Each normal map can be in the form of a tangent normal map. For example, when generating a UV-space normal map based on projecting a camera-space normal map, each camera-space normal map can be in the form of a tangent normal map such that the resulting UV-space normal map is also a tangent-space normal map.
[0055] Alternatively, when generating a UV-space normal map based on a UV-space input texture, the generated UV-space normal map can be a tangent-space normal map.
[0056] Each tangent normal map can be determined based on high-pass filtering the corresponding image. For example, when generating a UV-space normal map based on projecting a camera-space normal map, each camera-space tangent normal map can be generated based on high-pass filtering the corresponding object image of the second subset.
[0057] Alternatively, when generating a UV-space normal map based on a UV-space input texture, the UV-space tangent normal map can be generated based on high-pass filtering the UV-space input texture.
[0058] Each tangent normal map can be determined based on processing the corresponding input image using a second deep learning neural network model trained to estimate a rotation map based on the input image, and converting the estimated rotation map to a tangent normal map.
[0059] For example, when generating a UV-space normal map based on projecting a camera-space normal map, each camera-space tangent normal map can be generated based on processing the corresponding object image of the second subset using the second deep learning neural network model and then converting the output rotation map to a tangent normal map.
[0060] Alternatively, when generating a UV-space normal map based on a UV-space input texture, the UV-space tangent normal map can be generated based on processing the UV-space input texture using the second deep learning neural network model and then converting the output rotation map to a tangent normal map.
[0061] Each tangent normal map can be determined based on processing a corresponding input image using a third deep learning neural network model trained to estimate a tangent normal map based on the input image.
[0062] For example, when generating a UV - space normal map based on a projected camera space normal map, each camera space tangent normal map can be generated based on processing a second subset of the corresponding object images using a third deep learning neural network model.
[0063] Alternatively, when generating a UV - space normal map based on a UV - space input texture, the UV - space tangent normal map can be generated based on processing the UV - space input texture using a third deep learning neural network model.
[0064] The method may further include determining a photometric normal map corresponding to a target region of the object surface based on a mesh and a tangent normal map corresponding to the target region of the object surface. The photometric normal map can be determined based on imprinting a mesh normal with the tangent normal map corresponding to the target region of the object surface.
[0065] The photometric normal map can be determined by determining a high - spatial - frequency component of the surface normal as an output of providing the tangent normal map as an input to a fourth deep learning neural network model trained to infer the high - spatial - frequency component of the surface normal. Determining the photometric normal map can include combining the high - spatial - frequency component of the surface normal with the mesh normal.
[0066] Each normal map can be in the form of a photometric normal map. Each photometric normal map can be determined based on processing a corresponding input image using a fifth deep learning neural network model trained to estimate a photometric normal map based on the input image.
[0067] For example, when generating a UV - space normal map based on a projected camera space normal map, each camera space photometric normal map can be generated based on processing a second subset of the corresponding object images using a fifth deep learning neural network model.
[0068] Alternatively, when generating a UV - space normal map based on a UV - space input texture, the UV - space photometric normal map can be generated based on processing the UV - space input texture using a fifth deep learning neural network model.
[0069] The method may further include generating a rendering of the object. The rendering can be based on a mesh, a diffuse map, a specular map, and a normal map (e.g., a tangent normal map or a photometric normal map). When computed, the photometric normal can additionally or alternatively be used Figure 1 in conjunction with the tangent normal map.
[0070] Multiple object images (including any video clips, depth maps, and / or structured light images) can be received from a handheld device used to obtain the multiple object images. When in use, multiple environmental images can also be obtained with and received from the same handheld device.
[0071] The method can also include using a handheld device including a camera to obtain multiple object images (including any video clips, depth maps, and / or structured light images). When in use, the method can also include using the handheld device to obtain multiple environmental images.
[0072] The handheld device can include or take the form of a mobile phone or smartphone. The handheld device can include or take the form of a tablet computer. The handheld device can include or take the form of a digital camera. In other words, the handheld device can be primarily used for taking photos, i.e., a dedicated camera such as a digital single-lens reflex (DSLR) camera or the like.
[0073] The handheld device can be used only to obtain multiple object images and send them to a separate and / or remote (relative to the handheld device) location for processing. For example, steps of the method other than obtaining the multiple object images can be performed by a server or equivalent data processing system communicatively coupled to the handheld device via one or more networks. The network can be wired or wireless. The network can include the Internet.
[0074] A mesh, diffuse map, specular map, and normal map (e.g., tangent normal and / or photometric normal) can be output from the server executing the method to another server and / or to the handheld device used to obtain the multiple object images. A photometric normal map can be output from the server executing the method to another server and / or to the handheld device used to obtain the multiple object images. Rendering can be output from the server executing the method to another server and / or to the handheld device used to obtain the multiple object images.
[0075] Alternatively, the handheld device can be used to perform all steps of the method (i.e., local processing).
[0076] The method can also include processing one or more diffuse maps output by a deep learning neural network model and corresponding to the input image, including: generating a low-frequency image by blurring the input image; generating a high-pass filtered image by subtracting the low-frequency image from the input image; normalizing the high-pass filtered image by pixel-wise dividing by the input image; and generating a refined diffuse map based on a linear function of pixel-wise multiplying the diffuse map with the normalized high-pass filtered image.
[0077] The linear function can be in the form of f(NORM(i,j,k)) = 1 + 0.5·NORM(i,j,k), where NORM(i,j,k) is the pixel value of the normalized high-pass filtered image corresponding to the i-th row, the j-th column, and the k-th color channel.
[0078] A refined diffuse map corresponding to each camera space diffuse map can be generated, and the refined diffuse map can be projected onto a mesh to generate a UV-space diffuse map. Alternatively, the input image can be a UV-space input texture, and the UV-space diffuse map can be refined in the same way to generate a refined UV-space diffuse map.
[0079] According to a second aspect of the present invention, there is provided a non-transitory computer-readable medium storing a computer program. When executed by a digital electronic processor, the computer program causes the digital electronic processor to execute the method of the first aspect.
[0080] The computer program may include features corresponding to any feature of the method. The definitions applicable to the method may also apply equally to the computer program.
[0081] Regarding the feature of obtaining a plurality of object images and / or environmental images using a handheld device, the computer program may include instructions for executing a graphical user interface configured to guide a user through the process of obtaining a plurality of object images and / or environmental images using the handheld device.
[0082] According to a third aspect of the present invention, there is provided a method of imaging an object, including obtaining a plurality of object images of the object using a first camera of a handheld device. Each object image corresponds to a different view direction of the first camera. The method of imaging the object further includes obtaining a plurality of environmental images using a second camera of the handheld device. The second camera is arranged to have a field of view that is generally opposite in orientation to the first camera. Each environmental image corresponds to an object image among the plurality of object images.
[0083] The plurality of object images may include a first object image and a second object image corresponding to a first direction and a second direction respectively, such that each part of a target region of the object surface is imaged by at least one of the first object image and the second object image.
[0084] The plurality of object images may further include a third object image corresponding to a third direction.
[0085] The method of imaging the object may include features corresponding to any feature of the method of the first aspect. The definitions applicable to the method of the first aspect (and / or its features) may also apply equally to the method of imaging the object (and / or its features).
[0086] Object images and environmental images obtained using a method of imaging an object can be processed using the method of the first aspect.
[0087] According to a fourth aspect of the present invention, there is provided an apparatus configured to receive a plurality of object images of an object. Each object image corresponds to a different view direction. The plurality of object images includes a first object image and a second object image corresponding to a first direction and a second direction. The apparatus is further configured to determine a mesh corresponding to a target region of the object surface based on a first subset including two or more of the plurality of object images. The apparatus is further configured to determine a diffuse map and a specular map corresponding to the target region of the object surface based on processing a second subset of the object images using a deep learning neural network model trained to estimate diffuse and specular albedo components based on an input image. The second subset includes at least the first object image and the second object image. The apparatus is further configured to determine a normal map corresponding to the target region of the object surface based on high-pass filtering each object image of the second subset. The apparatus is further configured to store and / or output the mesh, the diffuse map, the specular map, and the normal map.
[0088] The apparatus may include features corresponding to any of the features of the method of the first aspect, the computer program of the second aspect, and / or the method of imaging an object according to the third aspect. The definitions applicable to the method of the first aspect (and / or its features), the computer program of the second aspect (and / or its features), and / or the method of imaging an object of the third aspect (and / or its features) may equally apply to the apparatus.
[0089] The apparatus may include a digital electronic processor, a memory, and a non-volatile storage device storing a computer program which, when executed by the digital electronic processor, causes it to perform the functions the apparatus is configured to perform.
[0090] The system may include the apparatus and a handheld device. The handheld device may be as defined in the method of the first aspect and / or the method of imaging an object according to the third aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Certain embodiments of the present invention will now be described by way of example with reference to the accompanying drawings, in which:
[0092] Figure 1 is a flow chart of a method for processing an image to obtain geometric and reflection characteristics of an imaged object;
[0093] Figure 2 shows Figure 1 an example of the input and output of the method shown;
[0094] Figure 3 Schematically shows an exemplary geometry processed by the method shown for obtaining an image for use Figure 1 ;
[0095] Figure 4 Shows the coordinate system referred to herein;
[0096] Figure 5 Shows for Figure 1 the angular range from which an object is imaged by the method shown;
[0097] Figure 6 Schematically shows a system for performing the method shown; Figure 1 ;
[0098] Figure 7 and Figure 8 schematically shows the UNet architecture of a deep learning neural network model;
[0099] Figure 9 Schematically shows a generative adversarial network model;
[0100] Figure 10 Schematically shows a diffusion model for diffuse - specular reflection estimation of an input image;
[0101] Figure 11 Is a flowchart showing the camera - space implementation of step S3 shown in Figure 1 ;
[0102] Figure 12 Is a flowchart showing the UV - space implementation of step S3 shown in Figure 1 ;
[0103] Figure 13 Schematically shows a first exemplary method for calculating a normal map of an input image in the form of a tangent normal map;
[0104] Figure 14 Is a flowchart showing the camera - space implementation of step S4 shown in Figure 1 ;
[0105] Figure 15 Is a flowchart showing the UV - space implementation of step S4 shown in Figure 1 ;
[0106] Figures 16A to 16C Shows the rendering for training a deep learning network model for the method shown in Figure 1 ;
[0107] Figures 17A to 17C Shows the rendering for training a deep learning network model for Figure 1The ground truth diffuse map of the deep learning network model of the method shown in
[0108] Figures 18A to 18C Shows the ground truth (inverted) specular reflection map for training the deep learning network model for the method shown in Figure 1 The ground truth (inverted) specular reflection map of the deep learning network model of the method shown in
[0109] Figures 19A to 19C Shows the image input to the method shown in Figure 1 The image input to the method shown in
[0110] Figures 20A to 20C Shows the generated diffuse map corresponding to Figures 19A to 19C The generated diffuse map corresponding to
[0111] Figures 21A to 21C Shows the generated (inverted) specular reflection map corresponding to Figures 19A to 19C The generated (inverted) specular reflection map corresponding to
[0112] Figures 22A to 22C Shows the tangent normal map corresponding to Figures 19A to 19C The tangent normal map corresponding to
[0113] Figures 23A to 23C Shows the determined mesh corresponding to the object imaged in Figures 19A to 19C The determined mesh corresponding to the object imaged in
[0114] Figure 24A Shows the UV space (texture space) diffuse map obtained by mapping and blending the diffuse map shown in Figures 20A to 20C onto the mesh shown in Figures 23A to 23C The UV space (texture space) diffuse map obtained by mapping and blending the diffuse map shown in
[0115] Figure 24B Shows the UV space (texture space) specular reflection map obtained by mapping and blending the specular reflection map shown in Figures 21A to 21C onto the mesh shown in Figures 23A to 23C The UV space (texture space) specular reflection map obtained by mapping and blending the specular reflection map shown in
[0116] Figure 24C Shows the UV space (texture space) tangent normal map obtained by mapping and blending the tangent normal map shown in Figures 22A to 22C onto the mesh shown in Figures 23A to 23C The UV space (texture space) tangent normal map obtained by mapping and blending the tangent normal map shown in
[0117] Figure 25A And Figure 25B Shows the photo-realistic rendering generated using the mesh shown in Figures 23A to 23C and the texture map shown in Figures 24A to 24C The photo-realistic rendering generated using the mesh shown in
[0118] Figure 26 Is for obtaining an image for input to Figure 1Schematic block diagram of an exemplary handheld device of the method shown;
[0119] Figure 27 Schematically shows an image capture configuration for obtaining an environmental image for expansion Figure 1 of the method shown;
[0120] Figure 28 Schematically shows a deep learning neural network model adapted to incorporate environmental image data;
[0121] Figures 29A to 29C Presents a comparison between a second exemplary method for calculating a tangent normal map and Figure 13 the first exemplary method shown; and
[0122] FIGS. 31A and 31B show a comparison of renders obtained using tangent normal maps generated using a first ( Figure 13 ) and a second ( Figure 29B and Figure 29C ) method for calculating a tangent normal map. Detailed Description
[0123] Hereinafter, like parts are denoted by like reference numerals.
[0124] Here, a method is described for obtaining a 3D geometry and an optical property map of an object using multi-view capture compatible with acquisition using a single handheld device including a camera. The optical property map includes at least a diffuse reflection map, a specular reflection map, and a normal map, but may also (optionally) include further properties. As described below, in some examples, the normal map may take the form of a tangent normal map, and in further examples, the normal map may be calculated directly in the form of a photometric normal map. The 3D geometry and optical property maps that can be obtained using the methods described herein can be used to generate accurate and realistic renders of imaged objects, including human faces.
[0125] Referring to Figure 1 , a flowchart of the overall method is shown.
[0126] Also referring to Figure 2 , examples of input and output images during the method are shown. In Figure 2 , the specular reflection image SP shown n has been inverted for ease of visualization and reproduction. Similarly, Figure 2 the normal map N in the form of a tangent normal map TN shown in n has been inverted and the contrast rebalanced for ease of visualization. n
[0127] Also referring to Figure 3, schematically shows an example geometry for obtaining an object image of an object 1 in the form of a human head.
[0128] An overview of the method should be presented, followed by further details of each step.
[0129] Receive a plurality of images of the object 1, hereinafter referred to as "object images" (step S1). For the purposes of the following description, the nth of a total of N object images is denoted as IM n . Each object image IM n is composed of pixel values IM n (i, j, k), where 1 ≤ i ≤ I and 1 ≤ j ≤ J represent pixel coordinates in an image plane (also referred to as the "camera plane") with an I×J resolution, and 1 ≤ k ≤ K represents a color channel. For example, for an RGB image IM n , k = 1 can represent red, k = 2 can represent green, and k = K = 3 can represent blue. Some cameras 3 can obtain images IM m with more than three color channels, such as infrared (IR), ultraviolet (UV), additional visible colors, or even depth when the camera 3 is aligned (or adjacent) with a depth sensor 4 ( Figure 26 ). The camera plane of each image IM n will be oriented differently (perpendicular to the corresponding view direction r n ).
[0130] Each object image IM n corresponds to a different view direction r and includes an image of a part of the surface 2 of the object 1 visible from that view direction r. Optionally, the object image IM n can be obtained as a previous step of the method (step S0), but equally, the object image IM n can be obtained in advance and retrieved from a storage device for processing, or transmitted from a remote location. The received object images IM n include at least a first object image IM1 and a second object image IM2 obtained with respect to the object 1. The first object image IM1 and the second object image IM2 can be arranged such that each part of the target region 5 of the object surface 2 is imaged by at least one of the first object image IM1 and the second object image IM2 (in other examples, further object images IM n can be used, and this condition can apply to a subset REFLECT of the object images IM n used for estimating diffuse and specular reflection components). A pair of object images IM1, IM2 is considered to obtain the following-described diffuse map DF n and specular map SP n for a target region 5 such as the face of the subject 1.The minimum quantity required.
[0131] Preferably, in order to obtain the best output quality, the received object image IM n includes at least a first object image IM1, a second object image IM2, and a third object image IM3 obtained with respect to object 1 such that each part of the target region 5 of the object surface 2 is imaged by at least one of the first object image IM1, the second object image IM2, and the third object image IM3. Although the third object image IM3 is not required, the following description will assume that the third object image IM3 is included. Modifications to the minimum set of the first object image IM1 and the second object image IM2 will be obvious.
[0132] The object image IM can be obtained indoors or outdoors under ambient lighting conditions n . In other words, no special control of lighting is required. Although precise control of lighting is not necessary, a lower level of ambient lighting can be supplemented. For example, when a handheld device 12 ( Figure 6 ) such as a mobile phone or a tablet is used to obtain the object image IM n , the flash light-emitting diodes (LEDs) and / or the display screen of the handheld device 12 can be used to illuminate the object 1 with additional light.
[0133] Based on a first subset MESH (discussed further below) that includes two or more object images IM n , a grid 20 corresponding to the target region 5 of the object 1 surface 2 is determined (see Figure 2 ) (step S2). Preferably, a larger number of object images IM n are included in the first subset MESH and are used to determine the grid 20.
[0134] Based on a second subset REFLECT (discussed further below) that uses a deep learning neural network model trained to estimate diffuse and specular reflection components from the input image to process the object image IM n , a diffuse reflection DF UV and a specular reflection SP UV albedo map corresponding to the target region 5 of the object 1 surface 2 are determined (step S3). The second subset REFLECT includes at least the first object image IM1 and the second object image IM2, and preferably also includes the third object image IM3. The term "albedo" herein refers to the property of the diffuse reflection DF UV and the specular reflection SP UV maps as components of the reflectance of the target region 5 of the object / subject 1 surface 2, and for the sake of brevity, the term "albedo" is not distinguished from the diffuse reflection DF UV and the specular reflection SP throughout this specification.UV is associated with the texture map. As explained below, for the camera space object image IM n the diffuse and specular components are estimated to generate the camera space diffuse DF n and specular SP n texture maps, which are then projected (and blended) onto the mesh 20 to obtain the diffuse DF UV and specular SP UV texture maps corresponding to the target region 5 in the UV space (texture space). Alternatively, the camera space object image IM n can be projected (and blended) onto the mesh 20 to obtain the UV texture IM UV which is processed by a deep learning neural network model to directly estimate the diffuse DF UV and specular SP UV texture maps corresponding to the target region 5. UV
[0135] The normal map N corresponding to the target region 5 of the surface 2 of the object 1 is determined based on each of the first object image IM1 and the second object image IM2 and preferably also based on the third object image IM3 UV (step S4). As explained below and similar to the diffuse and specular estimation, the normal map N can be calculated in the camera space or directly in the UV space UV . In most examples herein, the normal map N is calculated in the form of a tangent normal map TN UV e.g., based on high-pass filtering (described below with respect to UV ) or using a trained second or third deep learning neural network model (described below with respect to Figure 13 ) or Figures 29A to 30B ).
[0136] Optionally, the photometric normal PN corresponding to the target region 5 can be calculated based on the mesh 20 geometry and the tangent normal map TN UV (step S5). The photometric normal PN UV is used for rendering and while pre-calculation is not essential, this may speed up subsequent rendering. The determination of the photometric normal PN UV is discussed in further detail below. Herein, the photometric normal PN UV is a normal specified in the global coordinates of the object / subject 1 rather than in the local coordinates of the mesh of the object / subject 1 as is the case for the tangent normal TN UV . UV
[0137] Optionally, one or more renderings of the target region 5 of object 1 can be generated (step S6). For example, the target region 5 can be rendered from several viewpoints and / or under different lighting conditions. The rendering viewpoints and / or lighting conditions can be user-configurable, for example using a graphical user interface (GUI). The rendering can be based on the mesh 20, the diffuse texture map DF UV , the specular texture map SP UV and optionally on the normal N UV texture map (such as the tangent normal texture map TN UV ) to generate the photometric normal texture map PN UV (either precomputed or computed at runtime).
[0138] Store and / or output the mesh 20, the diffuse texture map DF UV , the specular texture map SP UV and the normal texture map N UV (such as TN UV )(step S7). Then when a representation of the target region 5 of object 1 is desired to be rendered, the mesh 20, the diffuse texture map DF UV , the specular texture map SP UV and the normal texture map N UV corresponding to the target region 5 of object 1 can be retrieved as needed. For example, the output can be stored locally on the computer / server 14( Figure 6 ) executing the method and / or sent back to the source, such as the handheld device 12( Figure 6 ), which acquired and / or provided the object image IM n (step S1). Storing and / or outputting the mesh 20, the diffuse texture map DF UV , the specular texture map SP UV and the normal texture map N UV optionally includes storing and / or outputting the photometric normal texture map PN UV (if step S5 is used) and / or the rendering generated based on the mesh 20, the diffuse texture map DF UV , the specular texture map SP UV and the normal texture map N UV (if step S6 is used).
[0139] If there is a further object / subject 1 to be measured (step S8|yes), the method is repeated.
[0140] Object image
[0141] Specifically referring to step S1, a suitable configuration of the object image for subsequent processing will be described.
[0142] N object images IM1,..., IM Nshould include at least a first object image IM1 and a second object image IM2 corresponding respectively to a first direction r1 and a second direction r2 distributed around the object 1, such that each part of the target region 5 of the object surface 2 is imaged by at least one of the first object image IM1 and the second object image IM2. Preferably, in order to obtain a better quality diffuse reflection DF n and specular reflection SP n textures, N object images IM1, …, IM N may further include a third object image IM3 corresponding to a third direction r3. Figure 3 The example shown uses first, second, and third object images IM1, IM2, IM3, and will be described with respect to all three (although it should be remembered that the third object image IM3 is not necessary).
[0143] Specifically referring to Figure 3 , the camera 3 has a corresponding field of view 6 ( Figure 3 the dashed line in). However, when capturing the first object image IM1 along the first view direction r1, the entire surface 2 of the object 1 is usually not visible due to being occluded by other parts of the surface 2. The visible boundary 7 represents the projection of the visible part of the surface 2 onto the camera 3 at the first view direction r1. In Figure 3 the illustration of, the visible boundary 7 is bounded on one side by the ear of the subject 1 (when the object 1 is a person, we will also use the term "subject"), and on the other side (at the maximum angle) by the nose of the subject 1. The visible boundary 7 is not necessarily a single closed curve, and in some cases there may be multiple separate parts to account for the locally occluded parts of the surface 2.
[0144] Therefore, due to occlusion, any single image IM1 usually cannot capture enough information to determine the optical properties of the target region 5 of the surface 2 of the object 1. This is particularly applicable to complex surfaces, such as Figure 3 the case shown, where the target region 5 is the face and the object / subject 1 is the head. Although methods such as image filling networks, assuming symmetry, etc. can be applied, in the method of this specification, multi-view imaging is used to solve occlusion. For example, after obtaining the first image IM1, the camera 3 moves to the second view direction r2 to obtain the corresponding second image IM2, and moves to the third view direction r3 to obtain the corresponding third image IM3. Then, the target region 5 for the method is that part of the surface 2, each part of which has been imaged in at least one of the first, second, and third object images IM1, IM2, IM3.
[0145] The first, second, and third object images IM1, IM2, IM3 can be considered as defining the target region 5. Alternatively, if the target region 5 is known in advance, such as the face of the subject 1, the view directions r1, r2, r3 for capturing the first, second, and third object images IM1, IM2, IM3 can be selected accordingly to meet this condition.
[0146] For Figure 3 visual clarity, the visible boundary 7 is not included in each of the second and third camera 3 positions, but only in the rightmost camera 3 position corresponding to the second direction r2, to show the overall target region 5.
[0147] Figure 3 Illustrated are the first, second, and third object images IM1, IM2, IM3 captured from the corresponding view directions r1, r2, r3 arranged in a single horizontal plane with respect to the target region 5 (face) of the object / subject 1 (head). For the face, this has been found to be generally feasible because the degree of occlusion is more significant during horizontal movement than during vertical movement (thus requiring more horizontal captures from more angles). However, depending on the shape of the target region 5 and the fraction of the total surface 2 represented by the target region (which can be up to 100% in some cases), the object images IM1, …, IM can be captured from a wide range of angles. N .
[0148] Further object images IM can be obtained n>3 , and the first, second, and third object images IM1, IM2, IM3 can be considered those object images with visible boundaries, the union of which corresponds to the target region 5. In other words, the first, second, and third object images IM1, IM2, IM3 are not necessarily continuous, but rather the object images IM corresponding to the outer extent of the target region 5 n , where intermediate object images IM n can optionally be included to provide improved accuracy and reliability (by repeated sampling).
[0149] Although the object images IM1, …, IM can be obtained by shifting a single camera 3 between different viewpoints N , there is no reason why two or more (even N) different cameras 3 cannot be used to obtain the object images IM1, …, IM N , potentially capturing the object images IM1, …, IM simultaneously N .
[0150] Similarly, although the moving camera 3 has been described, the object images from different view directions r can also be obtained by rotating the object / subject 1 and / or by a combination of moving the camera 3 and the object / subject 1 relative to each other nObject image IM n 。
[0151] Although Figure 3 the specific examples shown illustrate obtaining first, second, and third object images IM1, IM2, IM3, in general, only the first and second object images IM1, IM2 are essential.
[0152] Also refer to Figure 4 which shows a useful coordinate system for defining the view direction r for reference.
[0153] The position vector r is shown relative to a conventional right-handed Cartesian coordinate (x, y, z) where its x-axis is considered aligned with the midline 8 of the target region 5 and its origin 9 is located at the centroid of the object / subject 1. For example, in Figure 3 the illustration, the midline 8 corresponds to the middle of the face of the subject 1 and the origin 9 is located in the middle of the head of the subject 1. The image IM n is obtained facing the direction corresponding to the position vector r which terminates at the origin 9 in these coordinates.
[0154] It should be understood that in practice, the object images IM1, …, IM N will not all be obtained such that the view directions r converge at a single point-like origin 9. The following discussion is intended as an approximate guide for the relative positioning of the camera 3 (or cameras 3) for obtaining the object images IM1, …, IM N For a given set of object images IM1, …, IM N the method will include determining a global objective coordinate space based on multiple viewpoints and mapping out the relevant positions and viewpoints corresponding to each object image IM n The global coordinate space can be used in other steps of the method, such as but not limited to the determination of the normal map N UV 。
[0155] The projection of the position vector r onto the x - y (equatorial) plane is parallel to the line 10 in the x - y plane and makes an angle β = φ with the x - axis (and midline 9). In Figure 3 the illustration, this line 10 is parallel to the plane of the figure and roughly corresponds to the horizontal plane. In the latitude - longitude spherical parameterization, this is the longitude angle β, and in spherical polar coordinates it is the azimuth angle φ. The position vector r makes a latitude angle α with the x - y (equatorial) plane and a polar angle θ with the z - axis. The position vector r is represented as (r, α, β) in the latitude - longitude spherical parameterization where r is the length (magnitude) of the position vector r extending back from the common origin 9. The position vector r is represented as (r, α, β) in the latitude - longitude spherical parameterization.
[0156] Also refer to Figure 5, the angular range with respect to the object / subject 1 spanned by the object image IM is shown and discussed. n The angular range with respect to the object / subject 1 spanned by the object image IM.
[0157] To image the surface 2 of the object / subject 1 completely, the object image IM n needs to cover substantially all latitudes -90 ≤ α ≤ 90 and longitudes 180 < β ≤ 180 angles.
[0158] In practice, this degree is usually not required. Specifically regarding the human face, some guidelines can be provided.
[0159] The object images IM1, … IM N with respect to the object / subject 1 span a solid angle region 11 of longitude Δβ and latitude Δα, and they are all concentrated at the midline 8 corresponding to the imaginary front of the face 5 of the subject 1.
[0160] The object images IM1, … IM can be obtained at regular or irregular intervals along the first horizontal arc at α = 0. N , and move between β = -Δβ / 2 and β = +Δβ / 2, then obtain the object images IM1, … IM at regular or irregular intervals along the second vertical arc at β = 0. N , and move between α = -Δα / 2 and α = +Δα / 2. For example, the first video clip can be obtained using a camera 3 moving around the first arc, and the second video clip can be obtained by moving around the second arc (in both cases, pointing the camera 3 at the subject 1). Then the object images IM can be extracted from the video clip data. n . In some examples, only a subset of the object images IM for mesh generation (step S2) is extracted from the video in this way, and a high-resolution still image IM is obtained. n as the input for diffuse - specular reflection estimation (step S3). In practice, the first arc and the second arc will not be such precisely defined paths, and this does not prevent the application of the method described herein (see, for example, the experimental examples starting from n ). Figure 19A .
[0161] Alternatively, the object images IM1, … IM of the periphery of the region 11 can be obtained at regular or irregular intervals, and then one or more object images IM within the region 11 are obtained. N n .
[0162] Depending on the mesh generation method used (step S2), the total number N of the object images IM n can be very small, for example, at least two (more preferably three). When using relatively sparse object images IM nAt this time, region 11 only corresponds to the view direction r n 's bounding rectangle (in longitude-latitude space). When all view directions r n are coplanar, region 11 can take the form of a line (e.g., when using the minimum number of the first object image IM1 and the second object image IM2).
[0163] It should be noted that the first object image IM1, the second object image IM2, and any other image IM n (such as the third object image IM3) for diffuse-specular reflection estimation (step S3) do not have to be located on the periphery of region 11 and usually may not be, because the angular ranges of α and β required for some meshing techniques (step S2) may require an angular range larger than the angular range required for diffuse-specular reflection estimation of the target region 5. In other words, in some examples, the first direction r1, the second direction r2, and preferably the third direction r3 can define region 11, but in other examples, they are only included within region 11.
[0164] In general, each of the first direction r1, the second direction r2, and (when used) the third direction r3 can be separated from each of the first direction r1, the second direction r2, and (when used) the third direction r3 by 30 degrees or more. For the face, the inventors have found that the first direction r1, the second direction r2, and (when used) the third direction r3 can be substantially coplanar, where the first object image IM1 and the second object image IM2 are equidistant from either side in front of the face in the equatorial plane (α = 0) with α = β = 0 (or "front view"). Empirically, the inventors have found that the angles α = 0, β = ±45° work well for the first and second directions r1, r2. This arrangement generally corresponds to the main directions of the face. Preferably, the third object image IM3 corresponding to the front view α = β = 0 is also obtained.
[0165] In some examples, the first object image IM1, the second object image IM2, and (when used) the third object image IM3 can be obtained after meshing (step S2). For example, before returning to obtain (steps S0 and S1) the first, second, and (when used) third object images IM1, IM2, IM3 from the observation directions r1, r2, r3 determined based on the grid, sufficient object images IM nto allow the grid to be determined (step S3). For example, the (third) front view direction r3 may be set to correspond to a frontal principal direction that is antiparallel to the vector average of the facial normals at each surface point of the face determined based on the grid, while the other pair of view directions r1, r2 are spaced to either side of the equatorial plane (α=0). Auditory and / or visual cues may be used to guide the user to correctly position the camera 3. For example, if the camera 3 is a handheld device 12 ( Figure 6 ) such as a mobile phone, the GUI may be used to receive input from a gyroscope / accelerometer (not shown) in the handheld device 12 to guide positioning.
[0166] Although the object images IM1, ..., IM N can be from any source, but one particular application is to use the handheld device 12 to obtain the object images IM1, ..., IM N .
[0167] For example, refer to Figure 6 , system 13 is shown.
[0168] System 13 includes a handheld device 12 that communicates with a server 14 via one or more networks 15 .
[0169] The handheld device 12 comprises a camera 3 and is used to obtain (step S0) object images IM1, ..., IM N Then, the object images IM1, ..., IM N The data is sent to the server 14, and the server 14 performs the more intensive processing steps (steps S1 to S5). Rendering may be performed on either or both of the handheld device 12 and the server 14 (step S6). Storage may be performed on either or both of the handheld device 12 (storage device 16) and the server 14 (storage device 19) (step S7).
[0170] The handheld device 12 may include or take the form of a mobile phone or smartphone, a tablet computer, a digital camera, etc. In some examples, the handheld device 12 may be primarily used for taking pictures, i.e. a dedicated purpose camera 3, such as a digital single lens reflex (DSLR) camera or the like. The handheld device 12 includes a storage device 16, which optionally stores the object images IM1, ..., IM N and / or a local copy of the mesh and diffuse, specular and normal maps transmitted back from the server 14.
[0171] Server 14 can be any suitable data processing device and includes a digital electronic processor 17 and a memory 18 to enable the execution of computer program code (e.g., stored in storage device 19) for performing the method. For simplicity, other common components of server 14 are not shown.
[0172] In other examples, instead of using server 14, a handheld device 12 can perform all steps of the method.
[0173] Mesh generation
[0174] Referring again specifically to Figure 1 and Figure 2 , there are multiple options for the specific method of determining the mesh (step S2).
[0175] Based on at least two object images IM n to determine a mesh corresponding to the target region 5 of the surface 2 of object 1, but depending on the specific method, it can include many more object images. Regardless of the number of input object images IM n , only a single mesh 20 is generated. Each vertex of mesh 20 has associated coordinates, a geometric normal, and a list of other vertices to which it is connected.
[0176] Mesh generation (step S2) is based on a subset of object images IM n . Let ALL = {IM1, IM2,..., IM n ,..., IM N} be the set of all object images, and be the subset of object images IM n used as the input for mesh generation (step S2). The subset MESH includes 2 to N elements, and the exact minimum depends on the method used. Preferably, the subset MESH includes at least three and preferably more object images IM n .
[0177] Depending on the mesh generation method used, the subset MESH can include one or both of the first object image IM1 and the second object image IM2 (and / or the third object image IM3 if used). Alternatively, for other methods, the subset MESH can specifically exclude the first, second, and (when used) third object images IM1, IM2, IM3.
[0178] Structure from motion:
[0179] Structure from motion (stereophotogrammetry) techniques can be applied to a set of two or more object images IM nDetermination of the mesh 20 is performed using a subset MESH (step S2). Preferably, when applying structure-from-motion techniques, the subset MESH includes 15 to 30 object images IM n (including endpoints). Since the quality of structure-from-motion techniques depends to a large extent on the number of input images, the method of obtaining a video clip beforehand and then extracting the object images IM n as frames of video data may be particularly useful.
[0180] As long as an accurate 3D geometry can be generated based on a subset MESH of reasonable size (e.g., 15 to 30 object images IM n ), the choice of a specific structure-from-motion method is not considered critical.
[0181] Depth maps:
[0182] If the set ALL of object images IM n includes one or more depth maps of the target region 5 and / or one or more structured light images of the target region 5 from which depth maps can be calculated, these can be used to generate the mesh 20.
[0183] For example, if the handheld device 12 includes a depth sensor 4( Figure 26 ), then the set ALL of received object images IM n can include processed data in the form of depth maps, or raw data in the form of structured light IR images captured by the depth sensor 4 and from which depth maps can be generated. In both cases, the mesh 20 can be generated based on a single depth map, but better results can be obtained by merging depth maps from multiple viewpoints to generate the mesh 20.
[0184] In some examples, information from depth maps can be mixed with information from structure-from-motion (stereophotogrammetry) to generate the mesh 20.
[0185] 3D deformable mesh model (3DMM):
[0186] The 3D deformable mesh model (3DMM) attempts to model an object as a linear combination of base object geometries. For example, when the object 1 is a person and the target region 5 is their face, the base object geometry will be other faces.
[0187] 3DMM fitting can be performed based on a single image. However, in order to apply this method to the face, the subset MESH should contain at least two object images IM from different viewpoints n in order to more accurately capture the overall shape of the face 5 of the subject 1.
[0188] The quality of 3DMM fitting depends to a large extent on the quality and diversity of the available base object geometries.
[0189] Neural surface reconstruction:
[0190] The determination of the mesh 20 can be performed (step S2) by applying a neural surface reconstruction technique to a subset MESH of the object image IM n . For example, the NeuS method described by Wang et al. in "NeuS: Learning Neural Implicit Surfaces from Volume Rendering for Multi-View Reconstruction", https: / / doi.org / 10.48550 / arXiv.2106.10689 (hereinafter referred to as "WANG2021").
[0191] Diffuse - specular reflection estimation
[0192] Referring specifically again to Figure 1 and Figure 2 , there are multiple options for using a deep learning neural network model to determine the diffuse and specular reflection maps corresponding to the target region (step S3).
[0193] Broadly speaking, the methods can be divided into:
[0194] - Camera space estimation and then merging (see Figure 11 ); or
[0195] - Merging and then UV (texture) space estimation (see Figure 12 ).
[0196] Although context training is required, the same type of deep learning neural network model can be used in either case. Camera space estimation is discussed first.
[0197] The deep learning neural network model is an image - to - image conversion network that is trained to receive an input image (e.g., the object image IM n ) and generate a pair of output images, one corresponding to the diffuse component and the other corresponding to the specular component. Let DF n represent the diffuse component of the object image IM n , and let SP n represent the specular component (see Figure 2 ).
[0198] The deep learning neural network model is not limited to a specific network architecture and can use any image-to-image conversion network if appropriately trained. Preferably, a light stage or equivalent functionality capture arrangement is used to obtain ground truth data (mesh, diffuse, and specular data) for training. For example, the devices and methods described in US17 / 504,070 and / or PCT / GB2022 / 051819, the contents of which are hereby incorporated by reference in their entirety. Alternatively, data captured using a light stage or equivalent functionality can be used to generate photorealistic renders, which can be used to generate training data (see Figures 16A to 18C ).
[0199] For example, the deep learning neural network model can take the form of a multi-layer perceptron, a convolutional neural network, a diffusion network, etc.
[0200] Reference Figure 7 shows a high-level schematic of the UNet model 21.
[0201] An input image 22 (e.g., including red, green, and blue color channels) is convolved by an encoder 24 into a latent vector representation 23. The decoder 25 then takes the latent vector representation 23 and upsamples back to the original dimensions of the input image 22, thereby generating an output image 26. The output image 26 is shown as having the same color channels as the input image 22, but can have more or fewer channels depending on the task for which the model 21 is trained.
[0202] In this context, the UNet model 21 for diffuse specular estimation can have a single encoder 24 and a branched structure with a pair of decoders 25, one for generating a diffuse map DF n based on the same latent representation vector 23, n and another for generating a specular map SP
[0203] based on the same latent representation vector 23. Alternatively, the UNet model 21 can be trained for a single decoder 25 that upsamples the latent representation vector 23 to an output image 26 with four color channels (red diffuse, green diffuse, blue diffuse, and specular). Figure 8 shows a more detailed schematic of a typical UNet model 21.
[0204] As the name implies, the UNet model 21 has a symmetric "U" shape. The encoder side 24 takes the form of a series of cascaded convolutional networks, linked in series by max pooling connections. An m×m max pooling connection works by taking the maximum value of an m×m region. For example, a 2×2 max pooling halves both dimensions of the image representation. The decoder 25 side is similar to the encoder 24, except that the convolutional networks are cascaded in series with up-conversion (or upsampling) connections, which act as the inverse operation of the encoder max pooling connections.
[0205] Newer UNet architectures include copy and crop or "skip" connections (illustrated by the dashed arrows in Figure 8 ) that pass high-frequency information missed during convolution from each encoder 24 layer to the corresponding decoder 25 layer.
[0206] As an example of a suitable UNet model 21, the inventors have confirmed that the method of the present specification works with a U-net model described in Deschaintre et al., "Deep Polarization Imaging for 3D Shape and SVBRDF Acquisition," Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021 (hereinafter referred to as "DESCHAINTRE2021").
[0207] Generative adversarial network training
[0208] A generative adversarial network (GAN) is another deep learning network architecture that has been used for image-to-image conversion.
[0209] Reference Figure 9 , shows a high-level schematic of the GAN model 27.
[0210] The GAN model 27 consists of two parts: a generator network 28 and a discriminator network 29. The generator network 28 takes the input image 22 (e.g., the object image IM n ), and generates an output "fake image" or images. In this context, the output will be the diffuse map DF n and the specular map SP n . The output is passed to the discriminator network, which uses ground truth data 30, 31 from the training set and corresponding to the input object image IM n to output a label 32 indicating whether the image is considered fake or real (by applying the label 31) (e.g., measured directly using active illumination techniques as previously described). The discriminator 29 also receives the object image IM n as an input. The generator 28 generates outputs DF n , SP nand succeed. The original GAN model 27 uses a noise vector as the input to the generator 28, but the illustrated architecture uses the image IM n as the input, which is referred to as a "conditional GAN", and thus the generator 28 is an image-to-image conversion network.
[0211] Any suitable image-to-image conversion network can be used as the generator 28, such as a multi-layer perceptron, or the UNet architecture described above in this document. Similarly, any suitable classifier can be used as the discriminator 29. Once the generator 28 has been trained using a suitable training set, the discriminator 29 can be discarded and the generator 28 applied to unknown images. In this way, the GAN is essentially a specific method for training a deep learning network model (generator) for the task of estimating the diffuse reflection DF n and the specular reflection SP n components. n is a specific method for training a deep learning network model (generator) for the task of estimating the diffuse reflection DF
[0212] As an example of a suitable GAN model 27, the inventors have confirmed that the method of this specification works using the model described by Wang et al. in "High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs", CVPR, 2018 (hereinafter referred to as "Pix2PixHD").
[0213] Diffusion model
[0214] Reference Figure 10 , shows a high-level schematic diagram of the diffusion model 33. Figure 10 The specular reflection image Spec shown in k has been inverted for ease of visualization.
[0215] On the left, we have the ground truth image for training, the input image Im0, and the corresponding diffuse Diff0 and specular Spec0 components. As with all deep learning models discussed in this document, the ground truth data for training is obtained from a light-stage or equivalent capture arrangement, or is a rendering generated using such data.
[0216] For each of the diffuse map and the specular map, there is a noise addition branch and a denoising branch. For example, the ground truth diffuse map Diff0 has noise added by the noise generator 34 to generate the first-generation noise image Diff1. This process is applied sequentially until the Kth-generation noise image Diff K is pure noise. The same process is applied to generate the noise specular reflection image Spec1,..., Spec K。The noise generator 34 can be of any suitable type, such as Gaussian distribution. However, the statistical parameters of the noise generator 34 do not have to be the same for each step (although the actual noise is of course pseudo-random). For example, each noise generator 34 can add a different amount of noise to other noise generators 34. In a preferred implementation, the amount of noise increases linearly with each generation of the noise images Diff k , Spec k .
[0217] The denoising branch of the diffuse texture map is formed by sequential image-to-image denoising networks 35, which are trained separately for each step. For example, the K-th generation denoising network 35 K is trained to recover the (K - 1)-th generation noisy diffuse image Diff K based on inputs in the form of the noise image Diff K-1 and the original image Im0 (serving as the ground truth). Similarly along the denoising branch, until the first-level denoising network 351 is trained to recover the ground truth image Diff0 based on the first-generation noisy diffuse image Diff1 and the original image Im0.
[0218] Once the denoising networks 351,..., 35 K have been trained, they can be applied to images with unknown true diffuse, such as the object image IM n . For example, a pure noise image is generated and input together with the object image IM n into the K-th generation denoising network 35 K , and then backpropagated to generate the estimated diffuse DF n (feeding the object image IM K at each denoising network 35 n ).
[0219] The denoising branch for the specular texture map is similarly formed by sequential image-to-image denoising networks 361,..., 36 K , which are trained and used in the same way as the diffuse denoising networks 351,..., 35 K .
[0220] The denoising networks 35, 36 can be any type of supervised learning image-to-image conversion network, such as, for example, the UNet model.
[0221] As an example of a suitable diffusion model 33, the inventors have confirmed that the method of this specification uses as The model described in O., & Legenstein, R. "Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models", https: / / doi.org / 10.48550 / ARXIV.2207.14626 (hereinafter referred to as "OZDENIZCI2022") works.
[0222] Although the diffusion network 33 may require more effort to train and apply due to the multi-generation denoising networks 35, 36, the inventors have found that the results provide the highest quality among the deep learning neural network models tested so far (see, for example, obtained using the diffusion network 33 with K = 25 generations Figures 19A to 21C . The diffusion network is OZDENIZCI2022, modified to have a four-channel output instead of three channels. The output includes three channels for diffuse albedo and one channel for specular reflection. In fact, the diffuse and specular reflection branches are combined, rather than being Figure 10 shown separately as
[0223] Diffuse-specular estimation in camera space:
[0224] As described above, any of the deep learning neural network models described can be applied to diffuse-specular estimation in camera space or UV space (texture space) estimation.
[0225] Also refer to Figure 11 , which shows further details of step S3 in the case of diffuse-specular estimation in camera space.
[0226] The generation of the diffuse-specular map (step S3) is based on a subset of the object image IM n . Recall that ALL = {IM1, IM2,..., IM n ,..., IM N} represents the set of all object images IM n , and then let represent the subset of the object image IM n used as the input for the generation of the diffuse-specular map (step S3). The subset REFLECT includes Nr elements, where Nr is between 2 and N. The subset REFLECT always includes the first and second object images IM1, IM2. The subset REFLECT can generally intersect with the subset MESH. However, in some implementations, the subset REFLECT and MESH can be mutually exclusive.
[0227] For each of the Nr object images IM n belonging to the subset REFLECT, starting from the first n = 1 (step S9), the object image IM nis provided as input to a deep learning neural network model and an output in the form of a corresponding camera space diffuse map DF n and a corresponding camera space specular map SPn is obtained (step S10).
[0228] If the index n is not yet equal to the subset REFLECT size Nr (step S11 | NO), the next object image IM n+1 (steps S12 and S10) is processed.
[0229] Once all the object images IM n in the subset REFLECT have been processed (step S11 | YES), the camera space diffuse maps {DF1,..., DF Nr} and the camera space specular maps {SP1,..., SP Nr} are projected onto the texture space (step S13). Based on projecting the camera space diffuse maps {DF1,..., DF Nr} corresponding to the subset REFLECT onto the mesh 20, a UV space diffuse map DF UV is generated. Based on projecting the camera space specular maps {SP1,..., SP Nr} corresponding to the subset REFLECT onto the mesh 20, a UV space specular map SP UV is generated.
[0230] The mesh 20 includes UV - coordinates (texture space coordinates) corresponding to each vertex. Projecting the diffuse map DF n or the specular map SP n onto the mesh 20 will associate each pixel in the camera space with a UV - coordinate (which can be interpolated for positions between the vertices of the mesh 20), thereby generating the corresponding UV maps DF UV and SP UV . When the same UV - coordinate corresponds to two or more pixels of the subset REFLECT, generating the UV - space maps DF UV and SP UV involves blending the corresponding pixels of the camera space maps DF n and SP n .
[0231] Any technique known in the field of multi - view texture capture can be used to perform the blending. For example, by averaging the low - frequency responses of two or more camera space diffuse maps DF n , and then stamping (overlaying) the camera space diffuse map DF nHigh-frequency response (high-pass filtering). The same processing can be applied to the specular reflection map SP n . The map DF can be obtained by blurring, for example, by Gaussian blurring the map DF n 、SP n to obtain the map DF n 、SP n of the low-frequency response. The high-frequency response can be obtained by subtracting the low-frequency response of the map DF n ,SP n from the map DF n ,SP n .
[0232] Diffuse-specular estimation in UV-space (texture space):
[0233] Also refer to Figure 12 for further details of step S3 for UV-space (texture space) estimation.
[0234] Based on projecting each object image IM of the subset REFLECT n onto the grid 20, a UV-space input texture IM is generated UV (step S14). This process can include mixing the object images IM corresponding to the same UV coordinates n 、SP n in the same way as mixing the diffuse map and the specular map DF n in the camera space method.
[0235] Then the UV-space input texture IM UV is input into a deep learning neural network model to directly obtain the corresponding UV-space diffuse map DF UV and the UV-space specular map SP UV (step S15).
[0236] Compared with the camera space diffuse-specular estimation, performing the estimation in the UV-space will occupy more memory than the camera space data unless decomposed into smaller patches. Even so, if the UV space data is decomposed into multiple parts, the partial UV space maps must have continuous parameterization because segmented parameterization is not applicable to dividing the UV space data into patches. In practice, continuous parameterization may be difficult to obtain from direct scanning but is usually the result of the mesh template registration process. Therefore, the UV-space input texture IM UV will require this additional registration step or registration to a 3DMM for proper continuous parameterization. Therefore, the camera space diffuse-specular estimation method may be more preferable to reduce the memory and / or computational requirements for performing the diffuse-specular estimation (step S3).
[0237] The first method for calculating tangent normal maps
[0238] Refer specifically again to Figure 1 and Figure 2 for a more detailed explanation of the calculation of the normal map N n 、N UV using the tangent normal maps TN n 、TN UV in the form of an example associated with the normal map N UV (Step S4).
[0239] In the first method for calculating tangent normal maps, a tangent normal map TN n is generated based on a high-pass filtered input image (e.g., the object image IM n 、TN UV ).
[0240] Similar to diffuse - specular estimation, the determination of the tangent normal map TN UV corresponding to the target region 5 of the surface 2 of object 1 can be performed in camera space ( Figure 14 ) or in UV - space ( Figure 15 ).
[0241] Also refer to Figure 13 which shows a first exemplary method for generating the tangent normal map TN n (Step S4). The tangent normal map TN n has been processed to maximize contrast for ease of visualization.
[0242] For the purpose of illustrating the first exemplary method, assume the input image takes the form of the object image IM n although the method does not depend on the nature of the input image.
[0243] The input image IM n is converted to a grayscale image GREY n . A blur filter 37 is applied to the grayscale image GREY n to generate a blurred image BLUR n . For example, a Gaussian blur filter can be applied (for the images shown herein, a kernel of size five, corresponding to a radius of two, was used). Due to the blurring, the blurred image BLUR n only includes low spatial frequencies. Then the grayscale image GREY n and the blurred image BLUR n are input to the difference block 38 to subtract the blurred image BLUR n to generate the tangent normal map TN n which only includes the grayscale image GREYn High-spatial-frequency content.
[0244] This is not the only way to generate high-pass filtered information, and as an alternative, the grayscale image GREY n can be subjected to a two-dimensional Fourier transform, a high-pass filter applied in the frequency domain, and then an inverse Fourier transform to generate the tangent normal map TN n . Alternatively, the tangent normal map TN n can be generated based on a rotation map RN (described below) generated using a second deep learning neural network model n or directly using a third deep learning neural network model (described below).
[0245] Camera space processing:
[0246] As previously described, the input image in camera space or UV-space (texture space) can be used to generate the tangent normal map N n , N UV (e.g., the tangent normal map TN n , TN UV ).
[0247] Reference is also made to Figure 14 which shows further details of step S4 in the case of camera space normal calculation.
[0248] The generation of the normal map N n , N UV (step S4) is based on a subset REFLECT of the same object image IM as the diffuse-specular reflection estimate n .
[0249] For each of the Nr object images IM belonging to the subset REFLECT n , starting from the first n = 1 (step S16), the object image IM n is used as an input to calculate the corresponding camera space normal map N n , e.g., the camera space tangent normal map TN n (step S17). This calculation can use an exemplary method associated with Figure 13 or any other suitable method described herein.
[0250] If the index n is not yet equal to the subset REFLECT size Nr (step S18 | NO), then the next object image IM is processed n+1 (steps S19 and S17).
[0251] Once all the object images IM in the subset REFLECT nhave all been processed (step S11|YES), the camera space normal maps {N1, …, N Nr}(e.g., the camera space tangent normal maps {TN1, …, TN Nr}) are projected onto the texture space (step S20). Based on projecting the camera space normal maps {N1, …, N Nr} corresponding to the subset REFLECT onto the mesh 20, a UV-space normal map N UV (e.g., the UV-space tangent normal map TN UV ) is generated. If necessary, the projection process will include blending as described above regarding the camera space diffuse-specular estimation method ( Figure 11 ).
[0252] UV-Space Processing:
[0253] Also referring to Figure 15 , further details of step S4 for UV-space (texture space) calculation are shown.
[0254] Based on projecting each object image IM of the subset REFLECT n onto the mesh 20, a UV-space input texture IM UV (step S21) is generated. This process may include blending the pixels of the object images IM n 、SP n 、N n corresponding to the same UV coordinates in the same way as for the diffuse, specular, or normal maps DF n .
[0255] Then the UV-space input texture IM UV is processed to directly obtain the corresponding UV-space normal map N UV , e.g., the UV-space tangent normal map TN UV (step S22). This calculation may use the exemplary method explained regarding Figure 13 , or any other suitable method described herein.
[0256] When the diffuse and specular maps DF UV 、SP UV are also estimated in the UV space ( Figure 12 ), the projection need not be repeated (step S21 may be omitted), and the same UV-space input texture IM UV can be used as the input to the deep learning neural network model (step S3) and for calculating the UV-space normal map N UV (step S22).
[0257] Performing calculations in UV - space relies heavily on the coherence of UV coordinates compared to camera - space normal calculations. Using some methods to generate meshes and associated UV coordinates, the resulting UV maps can be very scattered. In such cases, high - pass filtering may cause mixing in non - local regions and degrade the quality of the normal map N UV (e.g., the tangent normal map TN UV ). This may require additional registration steps or registration to a 3DMM for proper continuous parameterization (as discussed above regarding diffuse - specular estimation).
[0258] Photometric normal calculation
[0259] Referring specifically again to Figure 1 and Figure 2 , the calculation of the photometric normal map (step S5) will be explained in more detail.
[0260] Rendering requires the photometric normal PN UV , so the calculation can be omitted unless / until rendering needs to be generated (step S6). However, pre - calculating the photometric normal (step S5) and storing / transmitting it together with the mesh 20 and the parameter maps DF UV , SP UV , N UV may be useful to save computational costs in subsequent rendering steps (optionally, the normal map N UV can be preferentially omitted rather than the photometric normal).
[0261] Direct calculation:
[0262] When the normal map N UV takes the form of the tangent normal map TN UV , it has the characteristics of a high - frequency image, which can be processed as a height map where dark is treated as "deep" and bright is treated as "high". The tangent normal map TN UV can be converted to a high - frequency normal map HFN UV by differentiation.
[0263] Then, the photometric normal map PN UV is generated by superimposing the high - frequency normal map HFN UV and the geometric normal at each UV coordinate. The geometric normal at the UV coordinate is obtained by interpolating the normals corresponding to the surrounding vertices. This process can sometimes be described as "stamping" the high - frequency details of the tangent normal map TN UV onto the geometric normal from the mesh 20.
[0264] Machine - learning calculation:
[0265] Alternatively, instead of directly differentiating the tangent normal map TNUV , the high-frequency normal map HFN can be determined by providing the tangent normal map TN UV as an input to a fourth deep learning neural network model trained to infer the high-frequency spatial components of the surface normal, where the high-frequency spatial components take the form of the high-frequency normal map HFN UV . Similar to the diffuse and specular maps used to train the deep learning neural network model for diffuse-specular estimation, the ground truth photometric normal for training the fourth deep learning neural network model can be obtained from a light stage or an equivalent function measurement and / or a rendering derived therefrom UV .
[0266] Then, the high-frequency normal map HFN UV is superimposed with the geometric normal in the same manner as the direct calculation method to obtain the photometric normal map PN UV .
[0267] In another approach, the fourth deep learning neural network model can alternatively be trained to directly infer the photometric normal map PN UV based on a normal map N UV in the form of the tangent normal map TN UV , the mesh 20, and optionally one or both (or the sum) of the diffuse DF UV and specular SP UV maps
[0268] Experimental examples
[0269] Examples of training data ( Figures 16A to 18C ), object images IM n ( Figures 19A to 19C ), and the output of the method ( Figures 20A to 25B ) will be presented
[0270] These examples relate to the implementation of the method, where the object image IM n includes mutually exclusive subsets MESH and REFLECT. The subset MESH is obtained by capturing frames in a video recorded using a handheld device 12 in the form of an iPhone 13 Pro (RTM), which moves along a first longitudinal β arc and a second transverse α arc, both of which cover 8 ± 45° from the midline corresponding approximately to the front view of the subject 1's face. The total number of object images IM n extracted from the video clip for the subset MESH depends on the object image IM nvaries with the number of acceptable frames that can be extracted from the video clip data. For example, there may be motion blur in the image, which depends on how fast the user moves the handheld device 12 during capture. Blurry images will be discarded (and can be determined using conventional methods applied to autofocus applications). Additionally or alternatively, the subject 1 may blink, and such images will also be discarded (determined through preprocessing analysis), leaving a certain number of object images IM for structure from motion. n , which cannot be determined in advance. Typically, the total number of object images IM n extracted from the video clip for the subset MESH is in the range of 15 to 30.
[0271] The subset REFLECT includes first, second, and third object images IM1, IM2, IM3 (and does not overlap with the subset MESH), which correspond to view directions n1, n2, n3 generally located in the equatorial plane α1 = α2 = α3 = 0° and corresponding to longitudinal angles β1 = -45°, β2 = 0°, and β3 = 45°.
[0272] Mesh 20 generation (step S2) uses structure from motion applied to the subset MESH. The method used is described in Ozyesil, Onur & Voroninski, Vladislav & Basri, Ronen & Singer, Amit. (2017), "A Review of Structure from Motion. Acta Numerica. 26. 10. 1017 / S096249291700006X".
[0273] The deep learning neural network for the presented images is a diffusion model as described in OZDENIZCI2022, which is modified for four-channel output instead of three-channel. The output includes three channels for diffuse albedo and one channel for specular reflection. The diffusion network 33 includes K = 25 generations.
[0274] The method is also applied using the Pix2PixHD model by Wang et al., although the output presented herein is obtained using the diffusion model.
[0275] Diffuse-specular reflection estimation (step S3) and normal calculation (tangent normal calculation in this example) (step S4) are performed in the camera space (see Figure 11 and Figure 14 ).
[0276] Training data
[0277] The diffusion model used was trained on data of 118 faces captured using the apparatus and method described in PCT / GB2022 / 051819 to obtain each corresponding diffuse, specular, and normal map DF UV , SP UV , PN UV . Specifically, the capture setup used 8 iPads (RTM) to provide illumination and 5 iPhones (RTM) to capture images (shown in Figure 1 of PCT / GB2022 / 051819), where the illumination conditions imposed a binary multiplexing pattern (see Figure 9 A to Figure 10 F) of PCT / GB2022 / 051819), and post - processing was performed using linear system analysis (see pages 147, line 18 to page 155, line 20 and Figures 37A to 38B of PCT / GB2022 / 051819).
[0278] For each training face model, three main images were rendered under multiple illumination conditions.
[0279] As synthetic data is generated for training the diffusion model, realistic rendering is important for training to accurately infer the diffuse and specular reflections of the real object image IM n . This was accomplished by embedding a physics - based model in our data rendering system to ensure realistic rendering results. The model used was the Cook - Torrance BRDF as described in Cook, R.L. & Torrance, K.E. (1982), "A Reflectance Model for Computer Graphics", ACM Trans. Graph., 1(1), 7 - 24, https: / / doi.org / 10.1145 / 357290.357293 (hereinafter referred to as "COOK1982").
[0280] To generate realistic illumination scenes for the training images, thirty captured environment maps were used, including three environment maps obtained from the Internet and seven environment maps captured by the inventors. By generating further environment maps using rotation around the x - axis, ten environment maps were augmented to thirty environment maps. The range of the environment illumination maps used covered a variety of real - world illumination scenes, including the most complex and dynamic scenes.
[0281] For example, also refer to Figures 16A to 16C, photorealistic renderings from the training set are shown from view directions roughly corresponding to angles β1 = -45°, β2 = 0°, and β3 = 45°. To counteract the fact that the first, second, and third object images IM1, IM2, IM3 do not actually lie precisely at β1 = -45°, β2 = 0°, and β3 = 45° degrees from the midline 8 (front view), a certain degree of random angular transformation is applied, and then the training data images are generated.
[0282] Also refer to Figures 17A to 17C , which respectively show the diffuse components corresponding to Figures 16A to 16C .
[0283] Also refer to Figures 18A to 18C , which respectively show the specular components corresponding to Figures 16A to 16C , inverted for visualization.
[0284] Figures 17A to 17C The diffuse components shown in UV are obtained directly by texturing the measured diffuse map DF Figures 17A to 17C onto the measured mesh 20 and then reprojecting it onto the corresponding view direction n. n The diffuse components shown in
[0285] Figures 18A to 18C provide the ground truth diffuse values Diff0 for training the diffusion model to infer the diffuse map DF UV The specular components shown in Figures 18A to 18C are obtained directly by texturing the measured specular map SP n onto the measured mesh 20 and then reprojecting it onto the corresponding view direction n.
[0286] Also refer to Figures 19A to 19C , which shows examples of the first, second, and third object images IM1, IM2, IM3 for the human subject 1 and the target region 5 corresponding to the face of the subject 1.
[0287] Before the forward pass, the target region 5 (subject face) is masked in the Figures 19A to 19C first, second, and third object images IM1, IM2, IM3. The masking method used is "pyfacer", https: / / pypi.org / project / pyfacer / , version 0.0.1 released on February 28, 2022.
[0288] Also refer to Figures 20A to 20C , which shows the calculation using the diffusion model and corresponding to Figures 19A to 19CThe camera space diffuse reflection maps DF1, DF2, DF3.
[0289] Also refer to Figures 21A to 21C , which shows the camera space specular reflection maps SP1, SP2, SP3 calculated using a diffusion model and corresponding to Figures 19A to 19C . The specular reflection maps SP1, SP2, SP3 have been inverted for visualization.
[0290] Also refer to Figures 22A to 22C , which shows the camera space tangent normal maps TN1, TN2, TN3 corresponding to Figures 19A to 19C . The contrast of the tangent normal maps TN1, TN2, TN3 has been renormalized for visualization.
[0291] Also refer to Figures 23A to 23C , which shows the mesh 20 reconstructed from the structure-from-motion processing of the subset MESH of the object image IM n .
[0292] The mesh 20 is shown without texture and from viewpoints corresponding to the view directions n1, n2, n3 of the first, second, and third object images IM1, IM2, IM3.
[0293] Also refer to Figure 24A , which shows the UV-space diffuse reflection map DF Figures 20A to 20C obtained by mapping and blending the camera space diffuse reflection maps DF1, DF2, DF3 of Figures 23A to 23C onto the mesh 20 of UV .
[0294] Also refer to Figure 24B , which shows the UV-space specular reflection map SP Figures 21A to 21C obtained by mapping and blending the camera space specular reflection maps SP1, SP2, SP3 of Figures 23A to 23C onto the mesh 20 of UV .
[0295] Also refer to Figure 24C , which shows the UV-space tangent normal map TN Figures 22A to 22C obtained by mapping and blending the camera space tangent normal maps TN1, TN2, TN3 of Figures 23A to 23C onto the mesh 20 of UV .
[0296] Although shown with a continuous UV parameterization, in other UV parameterizations, the UV-space maps DF UV , SP UV , TN UV may be fragmented.
[0297] Also refer toFigure 25A and Figure 25B , based on Figures 23A to 23C the grid 20 shown in Figures 24A to 24C and the UV - space texture DF shown in UV , SP UV , TN UV , shows a photorealistic rendering for a pair of view directions.
[0298] From Figure 25A and Figure 25B it can be observed that the method can capture sufficiently accurate and detailed geometric (mesh) and appearance data (UV - space texture DF n ), SP UV , TN UV , TN UV ) using a relatively small set of input object images IM obtained by the handheld device 12 in the form of a commercial smart phone, for convincingly rendering the model.
[0299] The handheld device
[0300] also refers to Figure 26 , showing a block diagram of an exemplary handheld device 12 in the form of a smart phone, tablet computer, etc.
[0301] The handheld device 12 can be used to obtain the object image IM n , as an initial step (step S0) of the method. Then, the handheld device 12 can perform the subsequent steps of the method locally, or can send the object image IM n to the server 14 or other data processing devices capable of performing the post - processing steps of the method (at least steps S1 to S7).
[0302] The handheld device 12 includes one or more digital electronic processors 39, a memory 40, and a color display 41. The color display 41 is typically a touch screen and can be used to guide the user to capture the object image IM n and / or display the rendering of the completed model. The handheld device 12 also includes at least one camera 3 for obtaining the object image IM nThe camera 3 of the handheld device 12 can be a front camera 3a oriented in a direction substantially the same as that of the color display 41, or a rear camera 3b oriented in a direction substantially opposite to that of the color display 41. The handheld device 12 can include both the front 3a and rear 3b cameras 3, or can include only one of these types. The handheld device 12 can include two or more front cameras 3a and / or two or more rear cameras 3b. For example, recent handheld devices 12 in the form of smart phones can include a front camera 3a and additionally include two or more rear cameras 3b, typically having a higher resolution than the front or “selfie” camera. The front imaging portion 42 can include any front camera 3a forming part of the handheld device 12, and one or more front flash LEDs 43 for providing a camera “flash” when taking pictures in low light conditions. If a flash is required, the ambient lighting may be insufficient for the purposes of the methods described herein. The rear imaging portion 44 can include any rear camera 3b forming part of the handheld device 12, and one or more rear flash LEDs 45. Although described as “portions”, the cameras 3a, 3b and associated flash LEDs 43, 45 need not be integrated into a single device / package, and in many cases can simply be co-located on the respective faces of the handheld device 12.
[0303] The handheld device 12 can also include one or more depth sensors 4, each depth sensor 4 including one or more IR cameras 46 and one or more IR sources 47. When present, the depth map output by the depth sensor 4 can correspond to the object image IM n and is used as an input for grid 20 generation as described above.
[0304] The handheld device 12 further includes a rechargeable battery 48, one or more network interfaces 49, and non-volatile memory / storage 50. The network interface 49 can be of any type and can include a Universal Serial Bus (USB), a wireless network (such as IEEE802.11b or 802.11), Bluetooth (RTM), etc. When processing is not performed locally, at least one network interface 49 is used to provide a wired or wireless link 51 to the network 15.
[0305] The non-volatile storage 50 stores an operating system 52 for the handheld device 12, and program code 53 for implementing the specific functions required to obtain the object image IM n such as a user interface (graphical and / or auditory) for guiding the user through the capture process. If the method is to be performed locally, the non-volatile storage 50 can also store the program code 53 for implementing the method. Optionally, the handheld device 12 can also store a sticker image cache 54. For example, the handheld device 12 can store the object image IM na local copy of, and transmitting the image to the server 14 via the link 51.
[0306] The handheld device 12 also includes a bus 55 interconnecting the other components, and also includes other components well known as part of the handheld device 12.
[0307] When using the front camera 3a on the handheld device 12 including a screen such as the color display 41, the color display 41 can also act as an area light source, preferably illuminating the object 1 with white light. This can be used to supplement low ambient lighting. In some cases, such as if capturing in a dark room or dim environment, a screen such as the color display 41 can be the only light source illuminating the object 1, or at least the main light source. If instead using the rear camera 3b on a handheld device 12 such as a mobile / smart phone or tablet, the flash 45 adjacent to the rear camera 3b can be used to illuminate the object 1. This is again useful in dark or dim environments.
[0308] Ambient lighting capture
[0309] As previously mentioned, the deep learning neural network model is trained using training images corresponding to a series of different environmental conditions to allow inference of accurate diffuse and specular maps DF UV 、SP UV , without considering the specific ambient lighting corresponding to the input object image IM n .
[0310] However, if the ambient lighting can be measured or estimated and taken into account in the diffuse - specular estimation (step S3), the accuracy can be improved.
[0311] Specifically, the method can be extended to include receiving a number M of environmental images EN1,..., EN M , representing the m - th environmental image as EN m . Each environmental image corresponds to a field of view directed away from the object / subject 1. Preferably, each environmental image is directed oppositely to the view direction r of the corresponding object image IM n (although it is not required that each object image IM n has an environmental image).
[0312] The deep learning neural network is also configured (see the description below of Figure 28 ) to receive additional input based on the environmental image EN n . For example, for each object image IM in the subset REFLECT n , the deep learning neural network can also receive the corresponding environmental image EN m . Alternatively, when obtaining an object image IM for extractionn When determining the video segment of the grid 20, a corresponding outward-facing video segment can be obtained and used to determine the environment map corresponding to a part or all of the virtual sphere or cylinder, which virtual sphere or cylinder is focused on the object / subject 1.
[0313] Reference Figure 27 shows the image capture configuration 56.
[0314] The image capture configuration 56 uses a handheld device 12 with a front 3a and a rear 3b camera 3, such as a smartphone or a tablet. Either the front 3a or the rear 3b camera can obtain the object image IM in any of the ways described previously herein. n Preferably, the camera 3 with the best resolution is used for capturing the object image IM n which is usually the rear camera 3b.
[0315] Meanwhile, the reversely oriented camera 3, such as the front camera 3a when the rear camera 3b obtains the object image IM n obtains the environment image EN m Preferably, but not necessarily, each environment image EN m corresponds to and is obtained simultaneously with the corresponding object image IM n In this way, the rear camera 3b captures the object image IM n while the front camera 3a captures the environment image EN m capturing the details of the environment towards the part of the target area 5 being imaged. For example, the handheld device 12 can move on a longitudinal arc 57 centered generally on the object / subject 1 while obtaining the object image IM n and the environment image EN m pair regularly or irregularly. Alternatively, when the handheld device 12 moves around the arc 57, video segments can be obtained from both the front 3a and the rear 3b cameras, and subsequently the object image IM n and the environment image EN m pair are extracted as time-related frames.
[0316] When the front camera 3a is deployed on the handheld device 12, and the handheld device includes a screen such as a color display 41 which is used as an area light source for illuminating the object 1, the ambient lighting can be calibrated as the lighting from the color display 41 size area source to the object 1 in the direction and position from the camera 3 view, which in turn can be determined using structure from motion reconstruction. If instead the rear camera 3b is deployed and the rear flash 45 adjacent to the rear camera 3b is used to illuminate the object 1, then in this case the ambient lighting can be calibrated according to the position and view direction of the camera 3 relative to the object 1.
[0317] In this way, by combining the input based on the environmental image EN m , the deep learning neural network model can consider the environmental lighting conditions when inferring and estimating the diffuse and specular reflection maps DF UV , SP UV . For example, for each object image IM n in the subset REFLECT, the deep learning neural network can also receive the corresponding environmental image EN m that is obtained at the same time and has a field of view with the opposite orientation.
[0318] Reference Figure 28 shows an example of an adapted deep learning neural network model 58.
[0319] The adapted model 58 includes a first (object image) encoder 59 that converts an input image (e.g., object image IM n ) into a first latent vector representation 60. The adapted model 58 also includes a second (environmental image) encoder 61 that converts the environmental image EN n corresponding to the input image IM m into a second latent vector representation 62. The first and second latent vector representations 60, 62 are concatenated and processed by a single common decoder branch 63 to generate an output image having three channels corresponding to the diffuse reflection map DF n and a fourth channel corresponding to the specular reflection map SP n . The encoders 59, 61 and the decoder 63 can be any suitable networks used in image-to-image conversion.
[0320] For example, in the case of a diffusion model, the environmental image EN m can be provided as an additional input to one, some, or all of the denoising networks 35, 36.
[0321] Additionally or alternatively, the method can include mapping the environmental image EN m to an environmental map (not shown) and providing the environmental map as an input to the deep learning neural network model (compared with the example shown in Figure 28 , the only difference being the relative dimensions of the second decoder 61).
[0322] The environmental map can correspond to a sphere or a part of a sphere centered approximately on object 1 in a far-field configuration (not shown). Each pixel of each environmental image EN m can be mapped to a corresponding region on the surface of the sphere. Mapping the environmental image EN mMapping to an environment map may also include filling in missing regions of the environment map, such as using known images to fill a deep learning neural network model. For example, the method may include filling in regions of the environment map that correspond to the convex hull of the environment image EN m when projected onto the surface of a sphere. Alternatively, the environment map may correspond to a far-field cylinder or a portion of a far-field cylinder.
[0323] Enhancing the resolution of the diffuse map
[0324] The diffuse map DF of target region 5 estimated by a deep learning neural network n 、DF UV is often slightly blurred compared to the corresponding original image IM n 、IM UV . This is due to various factors, including (but not limited to) the nature of the deep learning neural network, and also because the original image IM
[0325] 、IM n 、IM UV typically has a higher resolution than the diffuse map DF n 、DF UV predicted by the deep learning neural network. To add very fine details to the output diffuse map DF n 、DF UV , the inventors have developed a final diffuse refinement step.
[0326] The diffuse refinement step will be explained in the context of camera-space diffuse-specular estimation and the camera-space diffuse map DF n , but the process is equally applicable to the UV-space diffuse map DF UV directly output by a deep learning neural network based on UV-space texture IM UV .
[0327] First, a low-frequency image, denoted as LOW n , is generated by blurring the object image IM n . This is slightly different from the blurred image generated during the first method of calculating the tangent normal map TN n because it should not be converted to grayscale, but should retain the same number of color channels as the object image IM n .
[0328] A high-pass filtered image, denoted as HIGH n , is obtained as IMn - LOW n . A normalized high-pass filtered image, denoted as NORM n , is generated by dividing the high-pass filtered image HIGH n by the input object image IM pixel by pixel.n Finally, a refined diffuse map, denoted as REF, is generated by linearly multiplying the diffuse map DF n by the normalized high-pass filtered image NORM n on a per-pixel basis. According to experience, the inventors have found that good results can be obtained using the following function: n
[0329] REF n (i, j, k) = DF n (i, j, k) * [1 + 0.5 * NORM n (i, j, k)]
[0330] where 1 ≤ i ≤ I and 1 ≤ j ≤ J represent the pixel coordinates in the image plane with I×J resolution, and 1 ≤ k ≤ K represents the color channels. For example, for an RGB image IM n k = 1 can represent red, k = 2 can represent green, and k = K = 3 can represent blue. If the resolution of the diffuse map DF n is lower than I×J, it should be upsampled to the same resolution before calculating the refined diffuse map REF n .
[0331] The use of the normalized high-pass filtered image NORM n allows the reintroduction of high-frequency details while minimizing or avoiding the reintroduction of specular reflection components due to the previous normalization of the original object image IMn. Then, the refined diffuse map REF n can be used to replace the diffuse map DF n in any of the subsequent processing steps described herein.
[0332] As previously described, the same processing can be applied to the UV-space texture IM UV and the corresponding diffuse map DF UV generated by a deep learning neural network model to obtain a refined UV-space diffuse map REF UV .
[0333] Although an example of a linear function of the normalized high-pass filtered image NORM n has been shown, the relative weighting of the normalized high-pass filtered image NORM n need not be 0.5 and can vary / adjust according to the specific application.
[0334] Modify
[0335] It will be understood that the above-described embodiments can be variously modified. Such modifications can involve equivalent features and other features and / or methods already known in the design, manufacture, and use of lighting and in the design, manufacture, and use of devices and / or their component parts for lighting and / or performing image / video processing techniques, and these features can be used in place of or in addition to the features already described herein. The features of one embodiment can be replaced or supplemented with the features of another embodiment.
[0336] Second method for tangent normal map calculation
[0337] In the previous example, the first method described in association with Figure 13 has been used to describe the calculation of the normal map N n in the form of tangent normal maps TN UV and TN n and N UV and N n However, this is not the only method for calculating the normal map N UV in the form of tangent normal maps TN n and TN UV and N
[0338] Instead, in the second method for tangent normal map calculation, each tangent normal map TN n and TN UV is based on processing the corresponding input image IM n and IM UV using a second deep learning neural network model (not shown) trained to estimate the rotation map RN n ( Figure 29B ) and RN UV , and the estimated rotation map RN n and IM UV is converted to the tangent normal map TN n ( Figure 29B ) and RN UV and TN n and TN UV to determine.
[0339] Directly predicting the photometric normal PN n and PN UV of the face of the subject 1 with the head in an arbitrary direction requires knowledge of the direction and a good understanding of facial anatomy (global shape). To simplify the problem, for the second method of tangent normal map calculation, the output prediction is the rotation map RN n , where W and H are the width and height of the input image IM n and IM UV , and the two angles correspond to the longitudinal (Figure 4 rotation in the angle β) and in the latitudinal direction ( Figure 4 rotation in the angle α) to capture the high-frequency details found in the specular reflection normal map. This representation only encodes local surface information, where the global shape and camera direction are irrelevant. Thus, the second method of tangent normal map calculation can equally be applied to camera space ( Figure 14 ) or UV-space ( Figure 15 ), although preferably the second deep learning neural network model is specifically trained for the intended usage scenario.
[0340] to rotate the map RN n 、RN UV The output prediction in the form of is bi-directionally convertible with the tangent space normal representation TN n 、TN UV (using the rotation of the unit vector pointing in the z-direction (as Figure 4 shown) using two angles α, β results in the tangent normal TN n 、TN UV ).
[0341] The second deep learning neural network model can be of any type described herein regarding the (first) deep learning neural network model for estimating the diffuse DF n 、DF UV and specular SP n 、SP UV maps. It has been found that the diffusion model 33 provides particularly good qualitative output for the tangent normal maps TN n 、TN UV .
[0342] Regardless of the type of deep learning neural network model used, the same training database described hereinabove regarding Figures 16A to 18C the mesh 20 and the corresponding diffuse, specular, and normal maps DF UV 、SP UV 、PN UV can be used to generate synthetic training data. Equivalently, comparable facial geometries and BRDF datasets can be used.
[0343] For example, also referring to Figures 29A to 29C , an example of using the second deep learning neural network model in the form of a diffusion model is shown.
[0344] Figure 29A Shows, for comparison, the components of the rotation map RN n obtained using a high-frequency filtering method similar to the first method. Figure 29B Shows the rotation map RN of the second deep learning neural network model obtained in the form of a diffusion model ncomponents Figure 29C shows a photometric normal map obtained based on Figure 29B the rotation map RN shown in n The diffusion model that provides the second deep learning neural network model is passed an image IM of the occluded face as a condition
[0345] The diffusion model that provides the second deep learning neural network model is passed an image IM of the occluded face as a condition n Using the mesh 20 and the corresponding diffuse, specular, and normal maps DF Figures 16A to 18C described above in this article UV , SP UV , PN UV The same training database is used to generate conditional images for training. Specifically, our 3D face model (mesh 20) database and the corresponding BRDF maps (DFUV, SPUV, PNUV) are used to generate realistic images using OpenGL. The specular and diffuse normals mapped to the face geometry are rendered. Subsequently, the rotation map RN of the training image is calculated as a post-processing step n For the conditional images used to train the diffusion model, realistic faces are rendered in multiple different lighting environments (as described above in this article), and the head is randomly rotated around three main views at approximately β = 45°, 0°, and -45° longitude and α = 0° latitude
[0346] Also refer to Figure 30A and 30B , a close comparison between the rendered facial (skin) patches obtained using the first method (high-pass filtering) ( Figure 30A ) and the second method (the second deep learning neural network model) ( Figure 30B ) calculated using the tangent normal map is presented
[0347] It can be observed that the high-frequency details obtained using the second method (the second deep learning neural network model) ( Figure 30B ) calculated using the tangent normal map exhibit qualitatively enhanced local details, resulting in a more realistic skin appearance and clearer skin features
[0348] When generating the UV space tangent normal map TN n (see Figure 14 ) based on the projection camera space normal map PN UV , each camera space tangent normal map TN n is generated based on processing the corresponding object image IM of the second subset using the second deep learning neural network model (step S17), and then the output rotation map RN n is converted to the camera space tangent normal map TN n . Then the camera space tangent normal map TN n . Then the camera space tangent normal map TN nProjected onto the UV space tangent normal map TN UV (Step S20), as described above in this article.
[0349] Alternatively, when based on the UV space input texture IM UV generating the UV space tangent normal map TN UV at this time (see Figure 15 ), based on processing the UV space input texture IM using a second deep learning neural network model UV , and then rotating the output map RN UV converting it to the tangent normal map TN UV to generate the UV space tangent normal map TN UV .
[0350] The third method of tangent normal map calculation
[0351] Given the two-way convertibility of the rotation map RN n , RN UV and the tangent normal map TN n , TN UV An alternative third method of tangent normal map calculation is to train a third deep learning neural network model (not shown) to process the input image IM n , IM UV to directly estimate the corresponding tangent normal map TN n , TN UV .
[0352] The third method of tangent normal map calculation is the same as the second method of tangent normal map calculation in other aspects.
[0353] The method of directly calculating the photometric normal
[0354] As described above in this article, given the head in an arbitrary direction, directly predicting the photometric normal PN of the face of the subject 1 n , PN UV in the figure is a more difficult problem. Nevertheless, this is possible, and the calculation of the tangent normal map TN n , TN UV can be completely omitted in favor of directly estimating the photometric normal map PN n , PN UV .
[0355] In the method of directly calculating the photometric normal, each normal map N n , N UV adopts the form of the photometric normal map PN n , PN UV . By using a training based on the input image IM n , IM UVEstimated photometric normal map PN n 、PN UV is processed by a fifth deep learning neural network model (not shown) for the corresponding input image IM n 、IM UV to determine each photometric normal map PN n 、PN UV 。
[0356] The fifth deep learning neural network model can be any type described herein for the first deep learning neural network model for estimating diffuse DF n 、DF UV and specular SP n 、SP UV maps, or either the second or third deep learning neural network model. Similarly, synthetic data (such as the mesh 20 and the corresponding diffuse, specular, and normal maps DF Figures 16A to 18C described hereinabove with respect to UV 、SP UV 、PN UV database) can be used to train the fifth deep learning neural network model. Equivalently, comparable facial geometry and BRDF datasets can be used.
[0357] When generating a UV space normal map TN UV (see Figure 14 ) from a projected camera space normal map, each camera space photometric normal map PN n is generated based on processing a second subset of the corresponding object images IM n using the fifth deep learning neural network model (step S17). Then the camera space photometric normal map PN n is projected onto the UV space photometric normal map PN UV (step S20) as described above.
[0358] Alternatively, when generating a UV space photometric normal map PN UV from a UV space input texture IM UV (see Figure 15 ), the UV space photometric normal map PN UV is generated based on processing the UV space input texture IM UV using the fifth deep learning neural network model.
[0359] Although the claims in this application have been formulated as specific combinations of features, it should be understood that the scope of disclosure of the present invention also includes any novel feature or any novel combination of features that are explicitly or implicitly disclosed herein, as well as any generalization thereof, whether or not related to the same invention claimed in any current claim, and whether or not alleviating any or all of the technical problems of the present invention. The applicant hereby notifies that during the examination of this application or in any further application derived therefrom, new claims may be formulated for such features and / or such combinations of features.
Claims
1. A method, comprising: Receiving a plurality of object images of an object, each object image corresponding to a different view direction, wherein the plurality of object images includes a first object image and a second object image corresponding to a first direction and a second direction; Determining a mesh corresponding to a target region of the object surface based on a first subset of the plurality of object images, the first subset including two or more object images of the plurality of object images; Determining a diffuse map and a specular map corresponding to a target region of the object surface based on processing a second subset of the object images using a deep learning neural network model trained to estimate diffuse and specular albedo components based on an input image, wherein the second subset includes at least the first object image and the second object image; Determining a normal map corresponding to a target region of the object surface based on the object images of the second subset; Storing and / or outputting the mesh, the diffuse map, the specular map, and the normal map.
2. The method according to claim 1, wherein the plurality of object images includes video data, and wherein the method includes extracting the first subset of object images on which the mesh determination is based from the video data.
3. The method according to claim 1 or 2, wherein determining the diffuse map and the specular map corresponding to a target region of the object surface includes: For each of the object images in the second subset: Providing the object image as an input to the deep learning neural network model and obtaining a corresponding camera space diffuse map and a corresponding camera space specular map as outputs; Generating a UV space diffuse map based on projecting the camera space diffuse map corresponding to the second subset of the object images onto the mesh; Generating a UV space specular map based on projecting the camera space specular map corresponding to the second subset of the object images onto the mesh.
4. The method according to claim 1 or 2, wherein determining the diffuse map and the specular map corresponding to a target region of the object surface includes: Generating a UV space input texture based on projecting each of the object images in the second subset onto the mesh; Providing the UV space input texture as an input to the deep learning neural network model and obtaining a corresponding UV space diffuse map and a corresponding UV space specular map as outputs.
5. The method according to any one of claims 1 to 4, wherein determining the mesh corresponding to a target region of the object surface includes applying a structure from motion technique to the first subset of the object images.
6. The method according to any one of claims 1 to 5, wherein the plurality of object images includes one or more depth maps of the target region and / or one or more structured light images of the target region; wherein the first subset of object images on which the mesh determination is based includes one or more depth maps and / or one or more structured light images.
7. The method according to any one of claims 1 to 4, wherein determining the mesh corresponding to the target region of the object surface includes fitting a 3D deformable mesh model 3DMM to a first subset of the object image.
8. The method according to any one of claims 1 to 4, wherein determining the mesh corresponding to the target region of the object surface includes a neural surface reconstruction technique.
9. The method according to any one of claims 1 to 8, wherein the deep learning neural network model includes a multi-layer perceptron.
10. The method according to any one of claims 1 to 8, wherein the deep learning neural network model includes a convolutional neural network.
11. The method according to any one of claims 1 to 10, further comprising receiving a plurality of environmental images, each environmental image corresponding to a field of view oriented away from the object; wherein the deep learning neural network is further configured to receive additional input based on the plurality of environmental images.
12. The method according to claim 11, further comprising mapping the plurality of environmental images to an environmental texture map and providing the environmental texture map as an input to the deep learning neural network model.
13. The method according to any one of claims 1 to 12, wherein determining the normal map includes: generating a camera space normal map based on a second subset of the object image; generating a UV space normal map based on projecting the camera space normal map corresponding to the second subset onto the mesh.
14. The method according to any one of claims 1 to 12, wherein determining the normal map includes: generating a UV space input texture based on projecting a second subset of the object image onto the mesh; generating a UV space normal map based on the UV space input texture.
15. The method according to any one of claims 1 to 14, wherein each normal map includes a tangent normal map.
16. The method according to claim 15, wherein each tangent normal map is determined based on high-pass filtering the corresponding image.
17. The method according to claim 15, wherein each tangent normal map is determined based on: processing the corresponding input image using a second deep learning neural network model trained to estimate a rotation map based on the input image; converting the estimated rotation map to a tangent normal map.
18. The method according to claim 15, wherein each tangent normal map is determined based on processing the corresponding input image using a third deep learning neural network model trained to estimate a tangent normal map based on the input image.
19. The method according to any one of claims 15 to 14, further comprising determining a photometric normal map corresponding to the target region of the object surface based on the mesh and the tangent normal map corresponding to the target region of the object surface.
20. The method according to claim 19, wherein the photometric normal map is determined by: determining a high spatial frequency component of the surface normal as an output of providing the tangent normal map as an input to a fourth deep learning neural network model trained to infer the high spatial frequency component of the surface normal; The photometric normal map is determined by combining the high spatial frequency components of the surface normals with the mesh normals.
21. The method according to any one of claims 1 to 14, wherein each normal map includes a photometric normal map; wherein each photometric normal map is determined based on processing a corresponding input image using a fifth deep learning neural network model trained to estimate a photometric normal map based on the input image.
22. The method according to any one of claims 1 to 21, further comprising using a handheld device including a camera to acquire the plurality of object images.
23. The method according to any one of claims 1 to 24, further comprising processing one or more diffuse maps output by a deep learning neural network model and corresponding to the input image, including: generating a low-frequency image by blurring the input image; generating a high-pass filtered image by subtracting the low-frequency image from the input image; normalizing the high-pass filtered image by pixel-wise dividing by the input image; and generating a refined diffuse map based on a linear function of pixel-wise multiplying the diffuse map and the normalized high-pass filtered image.
24. A method for imaging an object, comprising: acquiring a plurality of object images of the object using a first camera of a handheld device, each object image corresponding to a different view direction of the first camera, acquiring a plurality of environment images using a second camera of the handheld device, wherein the second camera is arranged with a field of view oriented generally opposite to the first camera, and wherein each environment image corresponds to an object image of the plurality of object images; wherein the plurality of object images includes a first object image and a second object image corresponding to a first direction and a second direction.
25. An apparatus configured to: receive a plurality of object images of an object, each object image corresponding to a different view direction, wherein the plurality of object images includes a first object image and a second object image corresponding to a first direction and a second direction; determine a mesh corresponding to a target region of the object surface based on a first subset of the plurality of object images, the first subset including two or more object images of the plurality of object images; determine a diffuse map and a specular map corresponding to a target region of the object surface based on processing a second subset of the object images using a deep learning neural network model trained to estimate diffuse and specular albedo components based on the input image, wherein the second subset includes at least the first object image and the second object image; determine a normal map corresponding to a target region of the object surface based on high-pass filtering each object image of the second subset; and store and / or output the mesh, the diffuse map, the specular map, and the normal map.
Citation Information
Patent Citations
Acquisition of optical characteristics
US12288375B2