Generating Reflectance Maps for Relightable 3D Models

JP2025535119APending Publication Date: 2025-10-22SONY GROUP CORP +1
View PDF -1 Cites -1 Cited by

Patent Information

Application Number
JP2025521144
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-09-21
Publication Date
2025-10-22

Smart Images

  • Figure 2025535119000001_ABST
    Figure 2025535119000001_ABST
Patent Text Reader

Abstract

An electronic device and method for generating a reflectance map for a relightable 3D model are disclosed. The electronic device acquires multi-view image data including a set of images of an object and generates a 3D mesh of the object based on the multi-view image data. The electronic device acquires a set of motion-compensated images based on minimizing rigid body motion associated with the object between images of the set of images and generates a texture map in UV space based on the motion-compensated set of images and the 3D mesh. The electronic device acquires a specular reflectance map and a diffuse reflectance map based on separating specular and diffuse reflectance components from the texture map and acquires a relightable 3D model of the object based on the specular reflectance map and the diffuse reflectance map.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE / INCORPORATION BY REFERENCE TO RELATED APPLICATIONS]

[0001] This application claims the benefit of priority to U.S. Patent Application No. 17 / 965,425, filed with the United States Patent Office on October 13, 2022. Each of the above applications is incorporated herein by reference in its entirety.

[0002] Various embodiments of the present disclosure relate to 3D modeling, and more particularly, to electronic devices and methods for the generation of reflectance maps for reilluminatable 3D models. [Background technology]

[0003] Advances in the field of three-dimensional (3D) computer graphics have led to the ability to create 3D models to visualize real-world objects in a 3D computer graphics environment. A 3D model is a static 3D mesh that resembles the shape of a particular object. Typically, such 3D models are manually designed by computer graphics artists, commonly known as modelers, using modeling software applications. Such 3D models may not be equally usable in animation or various virtual reality systems or applications. Texture mapping is an important method for texturing 3D models by defining texture details to be applied to the 3D model. Creating realistic 3D models and high-fidelity texture / reflectance maps has been a challenging problem in computer graphics and computer vision. With increasing applications in the fields of virtual reality, 3D human avatars, games, and virtual simulations, it is becoming increasingly important to generate accurate, high-fidelity texture or reflectance maps to impart photorealism to 3D models.

[0004]

[0004] The limitations and disadvantages of conventional methods will become apparent to those skilled in the art by comparing the described system with certain aspects of the present disclosure illustrated in the remainder of this application with reference to the drawings. Summary of the Invention [Problem to be solved by the invention]

[0005]

[0005] An electronic device and method for generating reflectance maps for relightable 3D models is provided, as substantially shown in and / or described in connection with at least one figure and more fully set forth in the claims.

[0006]

[0006] These and other features and advantages of the present disclosure can be understood by considering the following detailed description of the disclosure in conjunction with the accompanying drawings in which like elements are designated by like reference numerals throughout. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary network environment for generating reflectance maps for reilluminatable 3D models, according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram illustrating the example electronic device of FIG. 1, according to an embodiment of the present disclosure. [Figure 3A] 3B, together with FIG. 3C, illustrate an exemplary processing pipeline for generating a reflectance map according to an embodiment of the present disclosure. [Figure 3B] 3A and 3B illustrate an exemplary processing pipeline for generating a reflectance map according to an embodiment of the present disclosure. [Figure 4] FIG. 2 illustrates an exemplary processing pipeline for obtaining a set of color-corrected images, according to an embodiment of the present disclosure. [Figure 5] FIG. 1 illustrates an exemplary processing pipeline for generation of an improved normal map, according to an embodiment of the present disclosure. [Figure 6]FIG. 1 illustrates an exemplary processing pipeline for the generation of a reflectance map for a reilluminatable 3D model, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0008]

[0013] The implementations described below can be found in electronic devices and methods for generating a reflectance map for a relightable 3D model. An exemplary aspect of the present disclosure can provide an electronic device (e.g., a server, desktop, laptop, or personal computer) capable of performing operations for generating a reflectance map for a relightable 3D model. The electronic device can acquire multi-view image data including a set of images of an object. The object can be exposed to a set of lighting patterns (e.g., an omnidirectional lighting pattern, a polarized lighting pattern, and / or a gradient lighting pattern) within a capture period of the multi-view image data. The electronic device can generate a 3D mesh of the object based on the multi-view image data and acquire a set of motion-compensated images based on minimizing rigid motion associated with the object between images of the set of images. The electronic device can generate a texture map in UV space based on the set of motion-compensated images and the 3D mesh. The electronic device can acquire the specular reflectance map and the diffuse reflectance map based on separating the specular reflectance component and the diffuse reflectance component from the texture map. The electronic device can obtain a relightable 3D model of the object based on the specular reflectance map and the diffuse reflectance map.

[0009]

[0014] Typically, 3D models can be manually designed by computer graphics artists, commonly known as modelers, using modeling software applications. Such 3D models may not be equally usable in animation or various virtual reality systems or applications. Texture mapping can be used to define texture details to be applied to the 3D model and texture the 3D model. Creating realistic models and texture maps has been a challenging problem in the fields of computer graphics and computer vision. Also, estimating full-head skin reflectance is important for generating relightable 3D head models for photorealistic game and film production.

[0010]

[0015] To address these requirements, this disclosure describes methods for generating high-quality, high-resolution skin reflectance maps, including diffuse, specular, color, normal, and height maps, for objects such as 3D head scans using polarized spherical gradient illumination patterns. This disclosure further describes robust diffuse and specular separation methods that allow cameras to be placed away from the equator of the light cage. This disclosure further describes operations for generating color maps to match unpolarized scan results and a pipeline for generating and refining normal maps. This disclosure further describes fast color correction methods and operations for performing rapid frame-to-frame motion compensation.

[0011]

[0016] 1 is a block diagram illustrating an exemplary network environment for generating a reflectance map for a reilluminatable 3D model, according to an embodiment of the present disclosure. Referring to FIG. 1, a network environment 100 is shown. The network environment 100 can include an electronic device 102, a server 104, a database 106, an imaging setup 108, and a communication network 110. The database 106 can include multi-view image data 112. The imaging setup 108 can include a first structure 114A, a second structure 114B, and an Nth structure 114N.

[0012]

[0017] 1 further illustrates a plurality of image capture devices 116 that may be installed on a 3D cage structure including a first structure 114A, a second structure 114B, and an Nth structure 114N. The plurality of image capture devices 116 may include, for example, a first image capture device 116A, a second image capture device 116B, and an Nth image capture device 116N. The electronic device 102 and the server 104 may be communicatively coupled to each other via a communications network 110. Also illustrated in FIG. 1 are objects (e.g., actors 118) and users 120 (e.g., 3D artists or developers) that may be associated with the electronic device 102.

[0013]

[0018] The electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to acquire multi-view image data 112 including a set of images of an object (e.g., actor 118). Based on the image data, the electronic device 102 may perform a set of operations to generate a set of reflectance maps that may be necessary to acquire a re-illuminatable 3D model of the object. The object may be exposed to a set of lighting patterns within the capture period of the multi-view image data. Examples of the electronic device 102 may include, but are not limited to, a computing device, a smartphone, a cellular phone, a mobile phone, a gaming device, a mainframe machine, a server, a computer workstation, and / or a consumer electronics (CE) device.

[0014]

[0019] Server 104 may include suitable logic, circuitry, interfaces, and / or code that may be configured to perform operations such as data / file storage, 3D rendering, or 3D reconstruction operations (e.g., photogrammetric reconstruction operations). In one or more embodiments, server 104 may store multi-view image data and perform at least one operation associated with electronic device 102. Server 104 may be implemented as a cloud server and may perform operations through web applications, cloud applications, HTTP requests, repository operations, file transfers, etc. Other implementations of server 104 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, or a cloud computing server.

[0015]

[0020] In at least one embodiment, the server 104 can be implemented as multiple distributed cloud-based resources using several techniques known to those skilled in the art. Those skilled in the art will understand that the scope of the present disclosure is not limited to implementing the server 104 and the electronic device 102 as two separate entities. In certain embodiments, the functionality of the server 104 can be incorporated in whole or at least in part into the electronic device 102 without departing from the scope of the present disclosure. In certain embodiments, the server 104 can host the database 106. Alternatively, the server 104 can be separate from the database 106 and can be communicatively coupled to the database 106.

[0016]

[0021] Database 106 may include suitable logic, interfaces, and / or code that can be configured to store multi-view image data 112 or metadata associated with the multi-view image data 112. For example, the metadata may include an identifier for the image capture device that captured the image, the lighting pattern used during capture, or an identifier for the viewpoint from which the image was captured, or an index value to indicate the position of the image within a set of images (included in the multi-view image data 112). Database 106 may be stored or cached on a device such as a server (e.g., server 104) or electronic device 102. The device that stores database 106 may be configured to receive queries for the multi-view image data 112 or the metadata. In response, the device that stores database 106 may retrieve and provide the multi-view image data 112 or the metadata to electronic device 102.

[0017]

[0022] In some embodiments, database 106 may be hosted on multiple servers stored in the same or different locations. Operations of database 106 may be performed using hardware, including a processor, a microprocessor (e.g., performing or controlling the execution of one or more operations), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). In other cases, database 106 may be implemented using software.

[0018]

[0023] The imaging setup 108 may correspond to a 3D cage structure on which multiple image capture devices 116 may be disposed and oriented to scan an object within the 3D cage structure from multiple viewpoints. The imaging setup 108 may include multiple structures 114, each connected at specific locations to form a cage-like structure (e.g., a 3D dome structure as shown in the figures). The present disclosure is not limited to any specific shape of the 3D cage structure. In some embodiments, the shape of the cage-like structure may be cylindrical, rectangular, or any shape depending on the requirements of the volumetric studio / capture. In some embodiments, each of the multiple structures 114 may have the same or different dimensions depending on the requirements of the volumetric studio / capture. In addition to the multiple image capture devices 116, multiple sound capture devices (not shown) and / or multiple light sources (not shown) may be disposed at specific locations on the multiple structures 114 to form the imaging setup 108.

[0019]

[0024] By way of example and not limitation, each structure may include a mount for holding at least one image capture device (represented by a circle in FIG. 1 ) and at least one processing device. As shown in FIG. 1 , each structure (e.g., a truss) may include a frame of a particular material (e.g., metal, plastic, or fabric) for holding at least one of an image capture device, a processing device, an audio capture device, and a light source (e.g., a flash). Different 3D structures of the same or different shapes may be connected to form an imaging setup 108. In an embodiment, the processing device may be the electronic device 102.

[0020]

[0025] In some embodiments, a mobile imaging setup can be created. In such implementations, each of the multiple structures 114 of the mobile imaging setup can correspond to an unmanned aerial vehicle (UAV), and the multiple image capture devices 116, multiple light sources, and / or other devices can be mounted on the multiple unmanned aerial vehicles (UAVs).

[0021]

[0026] The communication network 110 may include a communication medium that enables the electronic device 102 and the server 104 to communicate with each other. The communication network 110 may be either a wired or wireless connection. Examples of the communication network 110 may include, but are not limited to, the Internet, a cloud network, a cellular or wireless mobile network (such as Long Term Evolution and Fifth Generation (5G) New Radio (NR)), a Wireless Fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), or a metropolitan area network (MAN). The various devices in the network environment 100 may be configured to connect to the communication network 110 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols include, but are not limited to, at least one of Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11, Light Fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.

[0022]

[0027] In operation, the electronic device 102 can be configured to acquire multi-view image data 112 that includes a set of images of an object (e.g., an actor 118). The object can be exposed to a set of illumination patterns during the capture of the multi-view image data 112. By way of example and not limitation, the set of illumination patterns can include one or more of a cross-polarized omnidirectional illumination pattern, a gradient illumination pattern, and a polarized illumination pattern (including a cross-polarized illumination pattern and a parallel polarized illumination pattern).

[0023]

[0028] In an exemplary embodiment, the object can be a human head (including a face), and the multi-view image data 112 can be acquired from an imaging setup 108 that can operate as a polarization-based light cage. The object can be scanned from multiple viewpoints via one or more cameras in the imaging setup 108 to acquire the multi-view image data 112. To acquire high-fidelity reflectance and normal / height maps for the object, the object must be exposed to different illumination patterns while capturing images of the object from different viewpoints. Details regarding the multi-view image data 112 are further shown, for example, in FIG. 3A .

[0024]

[0029] The electronic device 102 can be configured to generate a 3D mesh of an object based on the multi-view image data 112. By way of example and not limitation, the 3D mesh can be generated from a set of images using a photogrammetry-based method (such as structure from motion (SfM)), a method requiring stereoscopic images, or a method requiring monocular cues (such as shape from shading (SfS), photometric stereo, or shape from texture (SfT)). Details of such methods are omitted from this disclosure for brevity. The 3D mesh can be an untextured mesh that resembles the 3D shape of the object. The 3D mesh can use polygons to define the shape or geometry of the object. An example 3D mesh for a human head is shown, for example, in FIG. 3A.

[0025]

[0030] When scanning an object to capture a set of images, the object may need to remain stationary throughout the scanning phase. However, the object may have unavoidable motion (e.g., head motion). The actual frame-to-frame motion may be assumed to be small. Rigid body motion may be estimated and removed by performing patch matching between images or frames to obtain a motion-compensated set of images. Specifically, the electronic device 102 may obtain a motion-compensated set of images based on minimizing rigid body motion associated with the object between images in the set of images. Details regarding the motion-compensated set of images are further shown, for example, in 306 of FIG. 3A .

[0026]

[0031] In some cases, it may be necessary to perform color correction of the motion compensated images to obtain accurate albedo values ​​for both the diffuse and specular components. Therefore, color correction can be applied to a subset of images in the set of motion compensated images to obtain a set of color compensated images.

[0027]

[0032] To obtain a relightable 3D model, it may be necessary to separate the specular and diffuse components from the images to obtain an albedo for the 3D mesh. The accuracy of the separation typically decreases with an increase in camera view. Therefore, a texture map can be generated in UV space, and then the diffuse and specular separation can be performed in UV space. The electronic device 102 can generate the texture map in UV space based on a set of motion-compensated images (or color-compensated images) and a 3D mesh. The texture map can help include high-frequency details of an object in the 3D model of the object, and such a map can be obtained based on mapping a set of motion-compensated images to UV space. Examples of texture maps include a cross-polarized UV texture map and a parallel-polarized UV texture map. Details regarding texture maps are further shown, for example, in FIG. 3B at 308.

[0028]

[0033] The electronic device 102 can obtain the specular reflectance map and the diffuse reflectance map based on separating the specular reflectance component and the diffuse reflectance component from the texture map. The specular reflectance map can represent the glossiness of the surface of the object, and the diffuse reflectance map can represent the reflection from the object without any atmospheric reflection. The specular reflectance component and the diffuse reflectance component can be separated from each texture map, i.e., the cross-polarized UV texture map and the parallel-polarized UV texture map. Details regarding the specular reflectance map and the diffuse reflectance map are shown, for example, in FIG. 3B.

[0029]

[0034] The electronic device 102 can further obtain a relightable 3D model of the object based on the specular and diffuse reflectance maps. The relightable 3D model can be a static 3D mesh that resembles the shape of the object (e.g., the head of the actor 118). Details regarding the relightable 3D model are shown, for example, in 314 of FIG. 3B.

[0030]

[0035] Figure 2 is a block diagram illustrating the example electronic device of Figure 1, in accordance with an embodiment of the present disclosure. The description of Figure 2 is provided with reference to elements of Figure 1. With reference to Figure 2, an electronic device 102 is shown. The electronic device 102 may include circuitry 202, memory 204, input / output (I / O) devices 206, and a network interface 208. The input / output (I / O) devices 206 may include a display device 210.

[0031]

[0036] Circuitry 202 may include suitable logic, circuits, and / or interfaces that can be configured to execute program instructions associated with different operations to be performed by electronic device 102. Circuitry 202 may include one or more processing units that may be implemented as separate processors. In some embodiments, one or more processing units may be implemented as an integrated processor or a collection of processors that collectively perform the functions of one or more dedicated processing units. Circuitry 202 may be implemented based on several processor technologies known in the art. Example implementations of circuitry 202 may include an x86-based processor, a graphics processing unit (GPU), a reduced instruction set computing (RISC) processor, an application-specific integrated circuit (ASIC) processor, a complex instruction set computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other control circuitry.

[0032]

[0037] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions to be executed by the circuitry 202. The memory 204 may be configured to store multi-view image data. Example implementations of the memory 204 may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), hard disk drive (HDD), solid state drive (SSD), CPU cache, and / or a secure digital (SD) card.

[0033]

[0038] The I / O device 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive input and provide output based on the received input. For example, the I / O device 206 may receive a first user input indicating a selection of multi-view image data 112. The I / O device 206 may be further configured to display a set of images included in the multi-view image data 112. The I / O device 206 may include a display device 210. Examples of the I / O device 206 may include, but are not limited to, a touchscreen, a keyboard, a mouse, a joystick, a microphone, or a speaker.

[0034]

[0039] The network interface 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communications between the electronic device 102 and the server 104 over the communications network 110. The network interface 208 may be implemented using various known technologies to support wired or wireless communications between the electronic device 102 and the communications network. The network interface 208 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuit.

[0035]

[0040] The network interface 208 may be configured to communicate via wireless communication with a network such as the Internet, an intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a number of communication standards, protocols, and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), Fifth Generation (5G) New Radio (NR), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, or IEEE 802.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), protocols for email, instant messaging, and short message service (SMS).

[0036]

[0041] The display device 210 may include suitable logic, circuitry, and interfaces that may be configured to display the multi-view image data 112 and / or the set of images included in the 3D mesh. The display device 210 may be a touchscreen that may enable a user (e.g., user 120) to provide user input via the display device 210. The touchscreen may be at least one of a resistive touchscreen, a capacitive touchscreen, or a thermal touchscreen. The display device 210 may be implemented using several known technologies, such as, but not limited to, at least one of a liquid crystal display (LCD) display, a light-emitting diode (LED) display, a plasma display, or an organic LED (OLED) display technology, or other display devices. According to an embodiment, the display device 210 may refer to a display screen of a head-mounted device (HMD), a smart glasses device, a see-through display, a projection display, an electrochromic display, or a transparent display. Various operations of the circuit 202 for generating a reflectance map for a reilluminatable 3D model are further described, for example, in FIGS. 3A and 3B .

[0037]

[0042] Figures 3A and 3B jointly illustrate an exemplary processing pipeline for generating a reflectance map, according to an embodiment of the present disclosure. The descriptions of Figures 3A and 3B are provided with reference to elements in Figures 1 and 2. Referring to Figures 3A and 3B, an exemplary processing pipeline 300 is shown illustrating exemplary operations 302 through 314. The exemplary operations 302 through 314 can be performed by any computing system, such as the electronic device 102 of Figure 1 or the circuit 202 of Figure 2. The exemplary processing pipeline 300 further illustrates multi-view image data 302A, a 3D mesh 304A, a set of motion-compensated images 306A, a specular reflectance map 312A, a diffuse reflectance map 312B, and a relightable 3D model 314A. The multi-view image data 302A can include N images, such as image 316A, image 316B, and image 316N. Image 316N includes a ColorChecker image 318 along with the face or head of an object. The number of images shown in FIG. 3A are provided by way of example only, and such example should not be construed as limiting the disclosure.

[0038]

[0043] Acquisition of multi-view image data can be performed at 302. The circuit 202 can acquire multi-view image data 302A including a set of images of an object. The object can be exposed to a set of lighting patterns within a capture period of the multi-view image data 302A. The object can be any animate or inanimate object. An example of an object as a human head is shown in FIG. 3A. One or more cameras can scan the object and capture one or more images from each viewpoint while the object is exposed to the set of lighting patterns within the capture period. Multiple images of the object can be captured from each camera view. The electronic device 102 can receive a set of images (i.e., a multi-view high-resolution image) from one or more cameras. The set of images can include image(s) of different lighting patterns and viewpoints.

[0039]

[0044] In one embodiment, the object can be a human head, and the multi-view image data 302A can be acquired from an imaging setup (e.g., imaging setup 108) operating as a polarization-based light cage. The light cage can be, for example, a dome-shaped cage structure, which can include several movable or static lighting devices and one or more image capture devices positioned at different positions on the cage structure. The lighting devices can emit different lighting patterns based on one or more control signals from the electronic device 102 or a standalone controller device. In the case of a polarization-based light cage, the lighting devices can emit polarized light (cross-polarized or parallel-polarized). The object, i.e., an actor, can sit in the center of the light cage, and each image capture device can capture an image of the human head while the human head is exposed to a set of lighting patterns. For example, the object can be an actor, and one or more image capture devices can capture a set of images of the actor's head under different lighting patterns.

[0040]

[0045] In some embodiments, the set of illumination patterns can include one or more of a cross-polarized omnidirectional illumination pattern, a gradient illumination pattern, and a polarized illumination pattern (including a cross-polarized illumination pattern and a parallel polarized illumination pattern). In some embodiments, the set of illumination patterns can include at least six polarized gradient illumination patterns, including cross-polarized illumination patterns and parallel polarized illumination patterns under three axes, namely, the "X" axis, the "Y" axis, and the "Z" axis. In some cases, using nine gradient illumination patterns may be preferable to provide higher quality normal generation rather than using a minimum of six or a maximum of twelve gradient illumination patterns without performing motion compensation.

[0041]

[0046] At 304, a 3D mesh can be generated. The circuit 202 can be configured to generate a 3D mesh 304 of the object based on the multi-view image data 302A. The 3D mesh 304 can be an untextured base mesh that can be used in operations associated with generating texture, reflectance, or normal / height maps for the object. As described, the 3D mesh can be generated from a set of images using a photogrammetry-based method (such as structure from motion (SfM)), a method requiring stereo images, or a method requiring monocular cues (such as shape from shading (SfS), photometric stereo, or shape from texture (SfT)). Details of such methods are omitted from this disclosure for brevity.

[0042]

[0047] At 306, a motion-compensated set of images can be generated. In many cases, the object may not be stationary throughout the entire capture period of the set of images. For example, the object may be the head of an actor whose images may be captured. During the capture period, head movement may be unavoidable. Typically, rigid body motion can be estimated and removed based on estimating and aligning the 3D position of markers or coded targets placed on a hat (worn by the actor). However, many studios prefer to capture images of actors with hair (i.e., without a hat). In such situations, coded targets or markers may not be suitable. If the object is assumed to be stationary for at least one second, the actual frame-to-frame motion can be assumed to be negligible. Rigid body motion can be removed by performing patch matching between images to obtain the motion-compensated set of images 306A. The circuit 202 can be configured to obtain the motion-compensated set of images 306A based on minimizing rigid body motion between images in the set of images.

[0043]

[0048] At 308, color correction can be performed. The circuit 202 can be configured to acquire a ColorChecker image that can be exposed to a cross-polarized omnidirectional illumination pattern. The circuit 202 can estimate a color matrix based on a comparison of the color values ​​of the ColorChecker image 308A with reference color values. The circuit 202 can apply the estimated color matrix to a subset of images in the set of motion-corrected images 306A to obtain a set of color-corrected images. The subset of images can be associated with one of an omnidirectional cross-polarized illumination pattern or a parallel-polarized illumination pattern.

[0044]

[0049] Typically, the generation of a reflectance map involves dividing the specular and diffuse reflectance components into the input image (e.g., cross-polarized light). JPEG2025535119000002.jpg6150 image and parallel polarization The specular and diffuse reflectance components can be directly separated from the multi-view cross-polarized and parallel-polarized images (JPEG2025535119000003.jpg6150 image) to generate view-dependent specular and diffuse reflectance components for all camera views. Each of the specular and diffuse reflectance components can be mapped to UV space to generate a texture map. The above processing order may heavily rely on the accuracy of the polarization separation. Unfortunately, the accuracy of the polarization separation may decrease with increasing camera views, which may be far from the equator of the polarization-based light cage, thereby leading to view-dependent specular maps. As a result, the generated specular maps are generally noisy and may sometimes contain significant artifacts. Instead of the above approach, the present disclosure can first perform UV mapping on the multi-view cross-polarized and parallel-polarized images, and then perform diffuse and specular separation in UV space, for example, as described in 310 and 312. This can generate higher-quality reflectance maps.

[0045]

[0050] At 310, a texture map can be generated. The circuit 202 can generate the texture map in UV space based on the set of motion-compensated images and the 3D mesh. Alternatively, the texture map can be generated based on the set of color-compensated images and the 3D mesh. By way of example and not limitation, the 3D mesh can be unwrapped into two-dimensional (2D) UV space to obtain the UV map. The UV mapping operation can apply color values ​​to the UV map based on an affine transformation between multiple triangles of the 3D mesh in the UV map and corresponding multiple triangles in the set of motion-compensated (or color-compensated) images. Because the set of images includes one or more images captured under a cross-polarized lighting pattern and one or more images captured under a parallel-polarized lighting pattern, UV mapping can be performed on multi-view cross-polarized and parallel-polarized images. The view-dependent specular component is considered to be relatively small compared to the view-dependent color information. One or more images captured under the cross-polarized lighting pattern can be merged in UV space to obtain the cross-polarized UV texture map. Additionally, one or more images captured under a parallel polarized lighting pattern can be merged in UV space to obtain a parallel polarized UV texture map. Diffuse and specular separation can be performed in UV space using texture maps, as described at 312.

[0046]

[0051] At 312, a reflectance map can be obtained. The circuit 202 can be configured to obtain a specular reflectance map 312A and a diffuse reflectance map 312B based on separating specular and diffuse reflectance components from a texture map. More specifically, the specular and diffuse reflectance components can be separated from a cross-polarized UV texture map and a parallel-polarized UV texture map. The specular reflectance map 312A can represent the surface gloss of an object, such as human skin, while the diffuse reflectance map 312B can represent reflections from an object without any atmospheric reflections. Separating the specular and diffuse reflectance components from the texture map can enable the generation of higher-quality reflectance maps. Furthermore, the above technique allows the camera to be positioned 45 to 60 degrees away from the equator in a polarization-based light cage without producing polarization-related artifacts in the albedo. Also, separating the specular and diffuse reflectance components from the texture map does not require the camera to be centered around the equator.

[0047]

[0052] The separated diffuse and specular components may include ambient occlusion (AO), which may need to be removed to generate albedo. Ambient occlusion may be shadows present in a set of images due to occlusion of points on the surface of an object by light sources. Certain points on an object may be occluded from light sources, and the shadows of such points may darken certain portions of the relightable 3D model. Therefore, the specular reflectance map 312A and the diffuse reflectance map 312B may not display accurately. According to one embodiment, ambient occlusion can be roughly estimated from the 3D shape by a path tracing operation. Specifically, hemispherical rays are considered to be emanating from a given point, and the path of each ray can be traced to check for intersections of the ray with occlusion. Once the ambient occlusion is estimated, it can be removed from the specular reflectance map 312A and the diffuse reflectance map 312B to obtain improved specular and diffuse reflectance maps. In one embodiment, the circuit 202 can be configured to estimate AO based on the 3D shape of the 3D mesh 304A and refine the specular reflectance map 312A and the diffuse reflectance map 312B based on removing ambient occlusion from the specular reflectance map 312A and the diffuse reflectance map 312B.

[0048]

[0053] At 314, a relightable 3D model 314A can be obtained. The circuit 202 can be configured to obtain a relightable 3D model 314A of the object based on the specular reflectance map 312A and the diffuse reflectance map 312B. The relightable 3D model 314A can be a model of a 3D object that can be illuminated with a lighting pattern that differs from the set of lighting patterns in which the set of images are captured.

[0049]

[0054] In some embodiments, circuit 202 can be further configured to generate a color map based on a linear combination of specular reflectance map 312A and diffuse reflectance map 312B. In certain situations, game and movie studios may prefer to maintain consistency in terms of color or contrast between texture maps generated from multi-view image data 302A (obtained from imaging setup 108 (i.e., polarization-based light cage)) and traditional 3D / 4D scanning systems using non-polarized illumination. However, diffuse reflectance map 312B generated from multi-view image data 302A may have lower contrast and more saturated colors due to polarized illumination. Color matching using a ColorChecker image may not work in this situation due to differences between the ColorChecker image (shown in image 316N) and the actor's skin. Therefore, Color Map I i Color can be generated based on a linear combination of the specular reflectance map 312A and the diffuse reflectance map 312B, as shown in equation (1). JPEG2025535119000004.jpg7170In the above formula, I i Color is the colormap, I i specular is the specular reflectance component, I i diffuse can be the diffuse reflectance component of the 3D surface point. λ can be a constant. If λ is 1, a linear combination of the specular reflectance map 312A and the diffuse reflectance map 312B can match a texture map generated using a conventional 3D / 4D scanning system that uses unpolarized illumination.

[0050]

[0055] Figure 4 illustrates an exemplary processing pipeline for obtaining a set of color corrected images, according to an embodiment of the present disclosure. Figure 4 is described with reference to elements of Figures 1, 2, 3A, and 3B. Referring to Figure 4, an exemplary processing pipeline 400 is shown illustrating exemplary operations 402 through 406 for obtaining a set of color corrected images. The exemplary operations 402 through 406 can be performed by any computing system, such as the electronic device 102 of Figure 1 or the circuit 202 of Figure 2. The exemplary processing pipeline 400 further illustrates a color checker image 402A and a set of color corrected images 406A.

[0051]

[0056] Acquisition of a ColorChecker image can be performed at 402. The circuit 202 can be configured to acquire a ColorChecker image 402A. The ColorChecker image 402A can be an x-rite ColorChecker image that can be captured under omnidirectional cross-polarized illumination. The ColorChecker image 402A can include various color gradients and can be used as a reference for color correction.

[0052]

[0057] At 404, a color matrix estimation can be performed. Conventionally, color correction of each image in the set of images can be performed using the ColorChecker image 402A. One image of the object can be captured along with the ColorChecker image 402A. The color of each image can be matched to the ColorChecker image 402A to balance the color of each image in the set of images. However, this process can generally be slow for large image data sets because the color matrix estimation process is repeated, file loading operations are repeated, and unnecessary color correction is performed on certain frames. In contrast, the circuit 202 can estimate the color matrix based on a comparison of the color values ​​of the ColorChecker image 402A to reference color values. The values ​​of each color in the ColorChecker image 402A can be compared to the reference color values, and a least-squares method can be used to estimate the color matrix. As an example, the color matrix can include three rows and three columns.

[0053]

[0058] At 406, the estimated color matrix may be applied to a subset of the images. The circuit 202 may apply the estimated color matrix to only a subset of the images in the set of motion-compensated images to obtain a set of color-compensated images. The subset of images may be associated with one of an omnidirectional cross-polarized illumination pattern or a parallel-polarized illumination pattern. Because the normal map and height map generation process has little correlation with the color accuracy of the images, color correction may not be performed on other motion-compensated images associated with the polarized illumination pattern. By correcting a subset of images, the time it takes to perform color correction may be shorter compared to traditional color correction techniques.

[0054]

[0059] In one embodiment, the set of motion compensated images includes 14-bit raw files that can be converted to 16-bit demosaiced images to preserve dynamic range. A linear color space can be chosen because generating normal and height maps may require linear intensity measurements.

[0055]

[0060] Figure 5 illustrates an exemplary processing pipeline for generating an improved normal map, according to an embodiment of the present disclosure. Figure 5 is described with reference to elements of Figures 1, 2, 3A, 3B, and 4. Referring to Figure 5, an exemplary processing pipeline 500 is shown illustrating exemplary operations 502 through 518 for generating a reflectance map for a relightable 3D model. The exemplary operations 502 through 514 may be performed by any computing system, such as the electronic device 102 of Figure 1 or the circuit 202 of Figure 2. The exemplary processing pipeline 500 further illustrates a set of specularly separated gradient images 502A, surface normals 504A, geometry normals 506A, a rotation matrix 508A, a normal map 514A, a height map 516A, and an improved normal map 518A.

[0056]

[0061] Specularly separated gradient images may be acquired at 502. The circuit 202 may be configured to acquire a set of specularly separated gradient images 502A based on removing a diffuse component from each image that may be associated with a gradient illumination pattern and included in a set of motion-compensated images (acquired in FIGS. 1 and 3A-3B). The diffuse component may be removed from each image that is associated with the gradient illumination pattern and included in the set of motion-compensated images, while the specular component may be retained in each image that is associated with the gradient illumination pattern.

[0057]

[0062] At 504, a surface normal calculation can be performed. The circuit 202 can be further configured to calculate a surface normal 504A corresponding to each pixel of an image in the set of specularly separated gradient images 502A. The surface normal 504A can be a vector perpendicular to a given pixel of an image in the set of specularly separated gradient images 502A. The calculated surface normal 504A can include much finer micro-geometry details. However, the calculated surface normal 504A may exist in world space, which may not correspond to the mesh space of the 3D mesh. Furthermore, the calculated surface normal 504A may have a non-uniform low-frequency bias due to the nature of photometry-based normal estimation. Therefore, the calculated surface normal 504A may need to be transformed from world space to object space. The transformed surface normal can then be corrected based on removing the low-frequency bias from the transformed surface normal.

[0058]

[0063] At 506, a calculation of a geometric normal can be performed. The circuit 202 can be configured to calculate the geometric normal 506A based on a linear combination of a set of vertex normals corresponding to the 3D mesh. The combination can be performed based on barycentric coordinates of each pixel of the generated texture map in UV space. The barycentric coordinates can define the position of each pixel of the generated texture map in UV space relative to a reference triangle or tetrahedron. The geometric normal 506A can have the same level of geometric detail as the 3D mesh.

[0059]

[0064] At 508, a transformation of the surface normal may be performed. The circuit 202 may be configured to transform the surface normal 504A from world space to object space based on applying a rotation matrix 508A to the surface normal 504A. The surface normal 504A may exist in world space (i.e., the "X", "Y", and "Z" axes of the set of lighting patterns) and may contain finer microgeometry details of the object. Because the surface normal 504A exists in world space (which does not correspond to mesh space (i.e., 3D mesh space)), it may be necessary to transform the surface normal 504A from world space to object space. Thus, a rotation matrix 508A may be required to transform the surface normal 504A from world space to object space.

[0060]

[0065] In one embodiment, the circuit 202 can be configured to estimate the rotation matrix 508A based on applying a least-squares difference operation to the Gaussian-filtered surface normal and the geometric normal 506A. To obtain the Gaussian-filtered surface normal, the surface normal 504A can first be passed through a Gaussian filter. The Gaussian filter can remove finer micro-geometry details of the surface normal 504A. The Gaussian-filtered surface normal may be smooth and may not contain low-frequency bias. Because the geometric normal 506A may contain less geometric details compared to the surface normal 504A, the geometric normal 506A may not contain low-frequency bias. Therefore, the geometric normal 506A may be smooth and may not need to be passed through a Gaussian filter. A least-squares difference operation can be used to compare the Gaussian-filtered surface normal with the geometric normal 506A to obtain the rotation matrix 508A. Note that surface normal 504A cannot be directly compared to geometric normal 506A because there are some outliers in surface normal 504A that may prevent surface normal 504A from resembling geometric normal 506A. Rotation matrix 508A can transform surface normal 504A from world space to object space based on the rotation of surface normal 504A. Geometric normal 506A can then be discarded, as it may only be needed to transform surface normal 504A from world space to object space.

[0061]

[0066] At 510, a correction of the surface normals can be performed. The circuit 202 can be further configured to correct the transformed surface normals based on removing low-frequency biases from the transformed surface normals. As described, the transformed surface normals may include finer microgeometry details and low-frequency biases that may need to be removed. To remove the low-frequency biases, the transformed surface normals can be passed through a low-frequency Gaussian filter to remove the low-frequency biases from the transformed surface normals to obtain corrected surface normals. The corrected surface normals may be smooth because they may not include low-frequency biases.

[0062]

[0067] At 512, the corrected surface normals can be transformed from object space to tangent space. The circuit 202 can be configured to transform the corrected surface normals from object space to tangent space based on a 3D mesh. The corrected surface normals can exist in object space, and therefore, the orientation of the corrected surface normals can be relative to the orientation of the object. Similarly, the orientation of the transformed surface normal vectors can be relative to the surface of the object and can be independent of the geometry of the object.

[0063]

[0068] Normal map generation can be performed at 514. The circuit 202 can be configured to generate a normal map 514A based on the transformed surface normal vectors. The generated normal map 514A can be associated with the surface and can be independent of the geometry of the object.

[0064]

[0069] An optimization operation can be performed at 516. The circuit 202 can be configured to perform the optimization operation to generate a height map 516A in UV space. The optimization operation can be performed such that the tangent vectors of the height map 516A in UV space remain perpendicular to the normal vector corresponding to the generated normal map 514A. Note that the height map 516A can be a raster image used to represent heights such as bumps in the 3D mesh. The generated height map 516A can be displayed as a grayscale image, can directly reflect high-frequency bumps and cavities on the geometry, and can be easily used for geometry cleanup and refinement. The generated height map 516A and the generated normal map 514A may contain microgeometry details that may be missing in the 3D mesh (i.e., the base mesh generated at 304 in FIG. 3). The generated height map 516A and the generated normal map 514A can be used to generate realistic renderings without subdividing the 3D mesh. However, the generated normal map 514A may generally be preferable.

[0065]

[0070] It should be noted that the 3D meshes prepared by game and / or movie studios may not exactly match the scanned shape of the object, making it difficult to directly use the normal map 514A generated from the polarization-based light cage. Instead, a generated height map 516A may be preferred.

[0066]

[0071] Typically, a height map can be generated based on refining a high-resolution 3D mesh using a separately measured normal map. After refinement, a height map can be calculated based on the high-resolution 3D mesh. However, a height map generated using the above operations may be too smooth and lack detail compared to the generated normal map 514A. In contrast, the circuit 202 can generate a height map 516A in UV space based on an optimization operation. The optimization operation can assume that the 3D mesh has a well-organized topology, and a neighborhood of UV pixels can be approximately equivalent to the corresponding 3D neighborhood. The height map 516A can also be optimized so that the tangent vectors of the height map 516A in UV space remain perpendicular to the normal vector. The number and direction of the tangent vectors can be selectable, with a minimum of two providing a maximum feasibility space for optimizing the height map 516A. Finally, a regularization term can be used to reduce variations in height from the surface of the 3D mesh.

[0067]

[0072] A gradient operation may be performed at 518. The circuit 202 may be configured to generate an improved normal map 518A based on applying the gradient operation to the height map 516A. After the height map 516A is generated, the improved normal map 518A may be generated. The improved normal map 518A may include the details and characteristics present in the normal map 514A. Additionally, the improved normal map 518A may have finer details and less overall noise compared to the normal map 514A.

[0068]

[0073] Figure 6 is a flowchart illustrating operations of an exemplary method for generating a reflectance map for a reilluminatable 3D model, according to an embodiment of the present disclosure. The description of Figure 6 is provided with reference to elements of Figures 1, 2, 3A, 3B, 4, and 5. Referring to Figure 6, a flowchart 600 is shown. The flowchart 600 may include operations 602 through 614 and may be implemented by the electronic device 102 of Figure 1 or the circuit 202 of Figure 2. The flowchart 600 may start at 602 and proceed to 604.

[0069]

[0074] At 604, multi-view image data 112 including a set of images of an object can be acquired. The circuit 202 can be configured to acquire the multi-view image data 112 including the set of images of the object. The object can be exposed to a set of illumination patterns within a capture period of the multi-view image data 112. Details regarding the acquisition of the multi-view image data 112 are shown, for example, in FIG. 3A.

[0070]

[0075] At 606, a 3D mesh of the object can be generated based on the multi-view image data 112. The circuit 202 can be configured to generate a 3D mesh of the object based on the multi-view image data 112. Details regarding the 3D mesh are shown, for example, in Figures 1 and 3A at 304.

[0071]

[0076] At 608, a set of motion-compensated images may be obtained based on minimizing rigid body motion associated with objects between images in the set of images. The circuit 202 may be configured to obtain a set of motion-compensated images (such as the set of motion-compensated images 306A of FIG. 3A) based on minimizing rigid body motion between images. Details regarding the compensation of the set of images are shown, for example, in 306 of FIG. 3A.

[0072]

[0077] At 610, a texture map can be generated in UV space based on the set of motion-compensated images and the 3D mesh. Circuit 202 can be configured to generate a texture map in UV space based on the set of motion-compensated images and the 3D mesh. Details regarding the generation of the texture map are further shown, for example, in FIG. 3B at 310.

[0073]

[0078] At 612, a specular reflectance map and a diffuse reflectance map may be obtained based on separating the specular reflectance component and the diffuse reflectance component from the texture map. The circuit 202 may be configured to obtain the specular reflectance map and the diffuse reflectance map based on separating the specular reflectance component and the diffuse reflectance component from the texture map. The specular reflectance component and the diffuse reflectance component may be separated from the cross-polarized UV texture map and the parallel-polarized UV texture map. Details regarding the generation of the specular reflectance map and the diffuse reflectance map are shown, for example, in FIG. 3B at 312.

[0074]

[0079] At 614, a relightable 3D model of the object can be obtained based on the specular reflectance map and the diffuse reflectance map. The circuit 202 can be configured to obtain a relightable 3D model of the object based on the specular reflectance map and the diffuse reflectance map. Control can proceed to an end.

[0075]

[0080] Although flowchart 600 is depicted as separate operations such as 604, 606, 608, 610, 612, and 614, the disclosure is not limited in this respect. Thus, in particular embodiments, such separate operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation, without departing from the essence of the disclosed embodiments.

[0076]

[0081] Various embodiments of the present disclosure may provide a non-transitory computer-readable medium and / or storage medium having stored thereon computer-executable instructions executable by a machine and / or computer to operate an electronic device (e.g., electronic device 102 of FIG. 1 ). Such instructions may cause electronic device 102 to perform operations that may include acquiring multi-view image data (e.g., multi-view image data 302A of FIG. 3A ) including a set of images of an object, where the object may be exposed to a set of lighting patterns within a capture period of the multi-view image data. The operations may further include generating a 3D mesh of the object (e.g., 3D mesh 304A of FIG. 3A ) based on the multi-view image data. The operations may further include acquiring a set of motion-compensated images (e.g., motion-compensated image set 306A of FIG. 3A ) based on minimizing rigid body motion associated with the object between images of the set of images. The operations may further include generating a texture map in UV space based on the set of motion-compensated images and the 3D mesh. The operations may further include obtaining a specular reflectance map and a diffuse reflectance map (e.g., specular reflectance map 312A and diffuse reflectance map 312B of FIG. 3B) based on separating the specular reflectance component and the diffuse reflectance component from the texture map. The operations may further include obtaining a relightable 3D model of the object (e.g., relightable 3D model 314A of FIG. 3B) based on the specular reflectance map and the diffuse reflectance map.

[0077]

[0082] An exemplary aspect of the present disclosure may provide an electronic device (such as electronic device 102 of FIG. 1) including a circuit (such as circuit 202). The circuit 202 may be configured to acquire multi-view image data (e.g., multi-view image data 302A of FIG. 3A) including a set of images of an object, where the object may be exposed to a set of illumination patterns within a capture period of the multi-view image data. The circuit 202 may be configured to generate a 3D mesh of the object (e.g., 3D mesh 304A of FIG. 3A) based on the multi-view image data. The circuit 202 may be configured to acquire a set of motion-compensated images (e.g., motion-compensated image set 306A of FIG. 3A) based on minimizing rigid body motion associated with the object between images of the set of images. The circuit 202 may be configured to generate a texture map in UV space based on the set of motion-compensated images and the 3D mesh. The circuit 202 can be configured to obtain a specular reflectance map and a diffuse reflectance map (e.g., the specular reflectance map 312A and the diffuse reflectance map 312B of FIG. 3B) based on separating the specular reflectance component and the diffuse reflectance component from the texture map. The circuit 202 can be configured to obtain a relightable 3D model of the object (e.g., the relightable 3D model 314A of FIG. 3B) based on the specular reflectance map and the diffuse reflectance map.

[0078]

[0083] In some embodiments, the set of illumination patterns may include one or more of a cross-polarized omnidirectional illumination pattern, a gradient illumination pattern, and a polarized illumination pattern (including a cross-polarized illumination pattern and a parallel polarized illumination pattern).

[0079]

[0084] In one embodiment, the circuit 202 can be configured to acquire a ColorChecker image exposed to a cross-polarized omnidirectional illumination pattern. The circuit 202 can be configured to estimate a color matrix based on a comparison of color values ​​of the ColorChecker image with reference color values. The circuit 202 can be configured to apply the estimated color matrix to a subset of images in the set of motion-compensated images to acquire a set of color-compensated images, where the subset of images can be associated with one of an omnidirectional cross-polarized illumination pattern or a parallel-polarized illumination pattern.

[0080]

[0085] In some embodiments, the set of color-corrected images can be obtained further based on transforming each image in the subset into a set of demosaiced images.

[0081]

[0086] In one embodiment, the circuit 202 can be configured to estimate ambient occlusion based on the 3D shape of the 3D mesh. The circuit 202 can be configured to refine the specular and diffuse reflectance maps based on removing ambient occlusion from the specular and diffuse reflectance maps. Here, a relightable 3D model of the object can be obtained based on the refined specular and diffuse reflectance maps.

[0082]

[0087] In some embodiments, the circuit 202 can be configured to generate a color map based on a linear combination of the specular reflectance map and the diffuse reflectance map.

[0083]

[0088] In an embodiment, the circuit 202 can be configured to obtain a set of specularly separated gradient images based on removing the diffuse component from each image associated with a gradient illumination pattern and included in the set of motion-compensated images.

[0084]

[0089] In one embodiment, the circuit 202 can be configured to calculate a surface normal corresponding to each pixel of an image in the set of specularly separated gradient images. The circuit 202 can be configured to calculate a geometric normal by linearly combining a set of vertex normals corresponding to the 3D mesh, where the combination is performed based on the barycentric coordinates of each pixel of the generated texture map in UV space.

[0085]

[0090] In one embodiment, the circuit 202 can be configured to transform the surface normals from world space to object space by applying a rotation matrix to the surface normals. The circuit 202 can be configured to correct the transformed surface normals based on removing low-frequency bias from the transformed surface normals. The circuit 202 can be configured to transform the corrected surface normal vectors from object space to tangent space based on a 3D mesh. The circuit 202 can be configured to generate a normal map based on the transformed surface normal vectors.

[0086]

[0091] In one embodiment, the circuit 202 may be configured to estimate the rotation matrix based on applying a least squares difference operation to the Gaussian filtered surface normal and the geometric normal.

[0087]

[0092] In an embodiment, the circuitry 202 can be configured to perform an optimization operation to generate a height map in UV space, where the optimization operation can be performed such that the tangent vectors of the height map in UV space are perpendicular to the normal vectors corresponding to the generated normal map.

[0088]

[0093] In one embodiment, the circuit 202 may be configured to generate an improved normal map based on applying a gradient operation to the height map.

[0089]

[0094] The present disclosure may also be situated in a computer program product that includes all features that enable the implementation of the methods described herein and that is capable of executing these methods when loaded into a computer system. A computer program in this context means any expression in any language, code or notation of a set of instructions that is intended to cause a system having information processing capabilities to perform a particular function, either directly or after a) conversion into another language, code or notation, b) reproduction in a different content form, or both.

[0090]

[0001] Although the present disclosure has been described with reference to specific embodiments, those skilled in the art will recognize that various modifications can be made and equivalents can be substituted without departing from the scope of the present disclosure. Many modifications can also be made to adapt a particular situation or content to the teachings of the present disclosure without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the disclosed embodiments, but is intended to include all embodiments falling within the scope of the claims. [Explanation of symbols]

[0091] 100 Network Environment 102 Electronic Devices 104 Server 106 databases 108 Imaging Setup 110 Communication Network 112 Multi-view image data 114 Multiple Structures 114A First Structure 114B Second Structure 114N Nth Structure 116 Multiple Image Capture Devices 116A first image capturing device 116B Second Image Capture Device 116N Nth image capture device 118 actors 120 users 202 circuits 204 memory 206 Input / Output (I / O) Devices 208 Network Interface 210 Display Devices 300 Exemplary Processing Pipeline 302 Acquisition of Multi-view Image Data 302A Multi-view image data 304 3D Mesh Generation 304A 3D Mesh 306 Generating a Set of Motion-Compensated Images 306A Motion-compensated image set 308 Color Correction 308A Color Checker Image 310 Texture Map Generation 312 Obtaining a Reflectance Map 312A Specular Reflectance Map 312B Diffuse Reflectance Map 314 Obtaining Relightable 3D Models 314A Relightable 3D Model 316A Image 316B Image 316N Images 318 Color Checker Images 400 Exemplary Processing Pipeline 402 Get Color Checker Image 402A Color Checker Image 404 Estimating Color Matrices 406 Applying Color Matrix 406A Set of color-corrected images 500 Example Processing Pipeline 502 Obtaining specularly separated gradient images 502A Set of specularly separated gradient images 504 Calculating surface normals 504A Surface Normal 506 Geometry Normal Calculation 506A Geometry Normal 508 Surface normal transformation 508A Rotation Matrix 510 Surface normal correction 512 Corrected Surface Normal Transform 514 Normal Map Generation 514A Normal Map 516 Optimization Calculation Execution 516A Height Map 518 Applying Gradient Operations 518A Improved normal maps 600 Flowchart 602 start 604 acquires multi-view image data including a set of images of an object exposed to a set of illumination patterns within a multi-view image data capture period. 606 Generate 3D mesh of object based on multi-view image data 608. Obtaining a motion-compensated set of images based on minimizing rigid body motion associated with an object between images of the set of images. 610 Generate a texture map in UV space based on a set of motion-compensated images and a 3D mesh 612 Obtaining specular and diffuse reflectance maps based on separating the specular and diffuse reflectance components from a texture map 614 Obtain a relightable 3D model of an object based on specular and diffuse reflectance maps

Claims

1. 1. An electronic device comprising: acquiring multi-view image data comprising a set of images of an object; the object is exposed to a set of illumination patterns during the capture of the multi-view image data; generating a 3D mesh of the object based on the multi-view image data; obtaining a set of motion-compensated images based on minimizing rigid motion associated with the object between images in the set of images; generating a texture map in UV space based on the set of motion-compensated images and the 3D mesh; obtaining a specular reflectance map and a diffuse reflectance map based on separating specular and diffuse reflectance components from the texture map; obtaining a relightable 3D model of the object based on the specular reflectance map and the diffuse reflectance map; a circuit configured to: An electronic device comprising:

2. 2. The electronic device of claim 1, wherein the object is a human head and the multi-view image data is acquired from an imaging setup operating as a polarization-based light cage.

3. 10. The electronic device of claim 1, wherein the set of illumination patterns includes one or more of a cross-polarized omnidirectional illumination pattern, a gradient illumination pattern, and a polarized illumination pattern (including a cross-polarized illumination pattern and a parallel polarized illumination pattern).

4. The circuit comprises: acquiring a ColorChecker image exposed to a cross-polarized omnidirectional illumination pattern; estimating a color matrix based on a comparison of the color values ​​of the ColorChecker image with reference color values; applying the estimated color matrix to a subset of images in the set of motion compensated images to obtain a set of color compensated images; the subset of images is associated with one of an omnidirectional cross-polarized illumination pattern or a parallel polarized illumination pattern; further configured to:

2. The electronic device according to claim 1 .

5. 5. The electronic device of claim 4, wherein the set of color-corrected images is obtained further based on transforming each image of the subset into a set of demosaiced images.

6. The circuit comprises: estimating ambient occlusion based on a 3D shape of the 3D mesh; and refining the specular reflectance map and the diffuse reflectance map based on removing the ambient occlusion from the specular reflectance map and the diffuse reflectance map, the relightable 3D model of the object is obtained based on the refined specular and diffuse reflectance maps; and further configured to:

2. The electronic device according to claim 1 .

7. 10. The electronic device of claim 1, wherein the circuitry is further configured to generate a color map based on a linear combination of the specular reflectance map and the diffuse reflectance map.

8. 10. The electronic device of claim 1, wherein the circuitry is further configured to obtain a set of specularly separated gradient images based on removing a diffuse component from each image in the set of motion-compensated images associated with a gradient illumination pattern.

9. The circuit comprises: calculating a surface normal corresponding to each pixel of an image in the set of specularly separated gradient images; calculating a geometric normal by linearly combining a set of vertex normals corresponding to the 3D mesh, the combination being performed based on barycentric coordinates of each pixel of the texture map generated in UV space; further configured to:

9. The electronic device according to claim 8.

10. The circuit comprises: Transforming the surface normal from world space to object space by applying a rotation matrix to the surface normal; correcting the transformed surface normals based on removing low frequency bias from the transformed surface normals; transforming the corrected surface normal vector from the object space to a tangent space based on the 3D mesh; generating a normal map based on the transformed surface normal vectors; further configured to:

10. The electronic device according to claim 9.

11. 11. The electronic device of claim 10, wherein the circuitry is further configured to estimate the rotation matrix based on applying a least squares difference operation to a Gaussian filtered surface normal and the geometric normal.

12. the circuitry is further configured to perform an optimization operation to generate a height map in the UV space; the optimization operation is performed such that a tangent vector of the height map in UV space is perpendicular to a normal vector corresponding to the generated normal map; 11. The electronic device according to claim 10.

13. 13. The electronic device of claim 12, wherein the circuitry is further configured to generate an improved normal map based on applying a gradient operation to the height map.

14. 1. A method comprising: In electronic devices, acquiring multi-view image data comprising a set of images of an object, exposing the object to a set of illumination patterns during a period of capture of the multi-view image data; generating a 3D mesh of the object based on the multi-view image data; obtaining a set of motion-compensated images based on minimizing rigid body motion associated with the object between images in the set of images; generating a texture map in UV space based on the set of motion-compensated images and the 3D mesh; obtaining a specular reflectance map and a diffuse reflectance map based on separating specular and diffuse reflectance components from the texture map; obtaining a relightable 3D model of the object based on the specular reflectance map and the diffuse reflectance map; A method comprising:

15. 15. The method of claim 14, wherein the object is a human head and the multi-view image data is acquired from an imaging setup operating as a polarization-based light cage.

16. 15. The method of claim 14, wherein the set of illumination patterns includes one or more of a cross-polarized omnidirectional illumination pattern, a gradient illumination pattern, and a polarized illumination pattern (including a cross-polarized illumination pattern and a parallel polarized illumination pattern).

17. acquiring a ColorChecker image exposed to a cross-polarized omnidirectional illumination pattern; estimating a color matrix based on a comparison of the color values ​​of the ColorChecker image with reference color values; applying the estimated color matrix to a subset of images in the set of motion compensated images to obtain a set of color compensated images, the subset of images is associated with one of an omnidirectional cross-polarized illumination pattern or a parallel polarized illumination pattern; 15. The method of claim 14, further comprising:

18. 18. The method of claim 17, wherein the set of color-corrected images is obtained further based on transforming each image in the subset into a set of demosaiced images.

19. estimating ambient occlusion based on a 3D shape of the 3D mesh; refining the specular reflectance map and the diffuse reflectance map based on removing the ambient occlusion from the specular reflectance map and the diffuse reflectance map, the relightable 3D model of the object is obtained based on the refined specular and diffuse reflectance maps; 15. The method of claim 14, further comprising:

20. A non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed by an electronic device, cause the electronic device to perform operations, the operations including: acquiring multi-view image data comprising a set of images of an object; the object is exposed to a set of illumination patterns during the capture of the multi-view image data; generating a 3D mesh of the object based on the multi-view image data; obtaining a set of motion-compensated images based on minimizing rigid body motion associated with the object between images in the set of images; generating a texture map in UV space based on the set of motion-compensated images and the 3D mesh; obtaining a specular reflectance map and a diffuse reflectance map based on separating specular and diffuse reflectance components from the texture map; obtaining a relightable 3D model of the object based on the specular reflectance map and the diffuse reflectance map; Including, 1. A non-transitory computer-readable medium comprising: