Light adding based relighting framework
The light adding based relighting framework addresses computational bottlenecks and flickering artifacts in current techniques by computing normal and light maps and generating relit images, resulting in high-quality and consistent relighting with reduced computational complexity.
Patent Information
- Application Number
- PCT/US2024/058753
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
Current relighting techniques face computational bottlenecks due to multiple sequential neural networks and suffer from flickering artifacts, making them difficult to implement on devices and resulting in unstable relighting results.
A light adding based relighting framework that computes a normal map and a light map using a machine learning model, generates images based on these maps, and outputs a relit image, while discarding complex networks to reduce computational complexity.
The framework achieves high-quality, realistic, and consistent relighting results with limited computing resources, eliminating flickering artifacts and improving temporal consistency.
Smart Images

Figure US2024058753_12062025_PF_FP_ABST
Abstract
Description
LIGHT ADDING BASED RELIGHTING FRAMEWORKCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to Israel Patent Application Serial No. 309188, entitled “LIGHT ADDING BASED RELIGHTING FRAMEWORK” and filed on December 7, 2023, which is expressly incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to processing systems, and more particularly, to one or more techniques for graphics processing.INTRODUCTION
[0003] Computing devices often perform graphics and / or display processing (e.g., utilizing a graphics processing unit (GPU), a central processing unit (CPU), a display processor, etc.) to render and display visual content. Such computing devices may include, for example, computer workstations, mobile phones such as smartphones, embedded systems, personal computers, tablet computers, and video game consoles. GPUs are configured to execute a graphics processing pipeline that includes one or more processing stages, which operate together to execute graphics processing commands and output a frame. A central processing unit (CPU) may control the operation of the GPU by issuing one or more graphics processing commands to the GPU. Modern day CPUs are typically capable of executing multiple applications concurrently, each of which may need to utilize the GPU during execution. A display processor may be configured to convert digital information received from a CPU to analog values and may issue commands to a display panel for displaying the visual content. A device that provides content for visual presentation on a display may utilize a CPU, a GPU, and / or a display processor.
[0004] Current techniques for relighting may be based on multiple sequential neural networks which may cause computational bottlenecks. There is a need for improved relighting techniques.BRIEF SUMMARY
[0005] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements of all aspects nor delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.
[0006] In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus for graphics processing are provided. The apparatus includes a memory; and a processor coupled to the memory and, based on information stored in the memory, the processor is configured to: compute, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image; compute a light map based on the normal map and a high dynamic range (HDR) map; generate an image based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image, where the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level; and output an indication of the generated image.
[0007] To the accomplishment of the foregoing and related ends, the one or more aspects include the features hereinafter fully described and particularly pointed out in the claims. The following description and the annexed drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of but a few of the various ways in which the principles of various aspects may be employed, and this description is intended to include all such aspects and their equivalents.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is a block diagram that illustrates an example content generation system, in accordance with one or more techniques of this disclosure.
[0009] FIG. 2 illustrates an example GPU, in accordance with one or more techniques of this disclosure.
[0010] FIG. 3 illustrates an example image or surface, in accordance with one or more techniques of this disclosure.
[0011] FIG. 4A-4B are diagrams illustrating example aspects pertaining to relighting, in accordance with one or more techniques of this disclosure.
[0012] FIG. 5A is a diagram illustrating an example of a light adding based relighting framework, in accordance with one or more techniques of this disclosure.
[0013] FIG. 5B is a diagram illustrating an example of a total relighting framework, in accordance with one or more techniques of this disclosure.
[0014] FIG. 6 is a diagram illustrating an example rendering engine, in accordance with one or more techniques of this disclosure.
[0015] FIG. 7 is a diagram illustrating an example relighting engine, in accordance with one or more techniques of this disclosure.
[0016] FIG. 8 is a diagram illustrating an example pre-processing engine, in accordance with one or more techniques of this disclosure.
[0017] FIG. 9 is a diagram illustrating an example background composition engine, in accordance with one or more techniques of this disclosure.
[0018] FIG. 10 is a diagram illustrating an example relighting framework, in accordance with one or more techniques of this disclosure.
[0019] FIG. 11 is a diagram illustrating an example pre-processing engine, in accordance with one or more techniques of this disclosure.
[0020] FIG. 12 is a diagram illustrating an example background engine, in accordance with one or more techniques of this disclosure.
[0021] FIG. 13 is a diagram illustrating an example background composition engine, in accordance with one or more techniques of this disclosure.
[0022] FIG. 14 is a diagram illustrating an example relighting framework, in accordance with one or more techniques of this disclosure.
[0023] FIG. 15 is a diagram illustrating example aspects pertaining to a pre-processing module, in accordance with one or more techniques of this disclosure.
[0024] FIG. 16A is a diagram illustrating an example aspect pertaining to a relighting module, in accordance with one or more techniques of this disclosure.
[0025] FIG. 16B is a diagram illustrating an example aspect pertaining to a relighting module, in accordance with one or more techniques of this disclosure.
[0026] FIG. 17 is a diagram illustrating example aspects pertaining to high dynamic range (HDR) relighting, in accordance with one or more techniques of this disclosure.
[0027] FIG. 18A-18C are diagrams illustrating further example aspects pertaining to HDR relighting, in accordance with one or more techniques of this disclosure.
[0028] FIG. 19 is a call flow diagram illustrating example communications between a CPU and a GPU in accordance with one or more techniques of this disclosure.
[0029] FIG. 20 is a flowchart of an example method of graphics processing, in accordance with one or more techniques of this disclosure.DETAILED DESCRIPTION
[0030] Various aspects of systems, apparatuses, computer program products, and methods are described more fully hereinafter with reference to the accompanying drawings. This disclosure may, however, be embodied in many different forms and should not be construed as limited to any specific structure or function presented throughout this disclosure. Rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art. Based on the teachings herein one skilled in the art should appreciate that the scope of this disclosure is intended to cover any aspect of the systems, apparatuses, computer program products, and methods disclosed herein, whether implemented independently of, or combined with, other aspects of the disclosure. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method which is practiced using other structure, functionality, or structure and functionality in addition to or other than the various aspects of the disclosure set forth herein. Any aspect disclosed herein may be embodied by one or more elements of a claim.
[0031] Although various aspects are described herein, many variations and permutations of these aspects fall within the scope of this disclosure. Although some potential benefits and advantages of aspects of this disclosure are mentioned, the scope of this disclosure is not intended to be limited to particular benefits, uses, or objectives. Rather, aspects of this disclosure are intended to be broadly applicable to different wireless technologies, system configurations, processing systems, networks, and transmission protocols, some of which are illustrated by way of example in the figures and in the following description. The detailed description and drawings are merely illustrativeof this disclosure rather than limiting, the scope of this disclosure being defined by the appended claims and equivalents thereof.
[0032] Several aspects are presented with reference to various apparatus and methods. These apparatus and methods are described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, and the like (collectively referred to as “elements”). These elements may be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.
[0033] By way of example, an element, or any portion of an element, or any combination of elements may be implemented as a “processing system” that includes one or more processors (which may also be referred to as processing units). Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems-on-chip (SOCs), baseband processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionality described throughout this disclosure. One or more processors in the processing system may execute software stored on one or more memory components. Software can be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software components, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0034] The term application may refer to software. As described herein, one or more techniques may refer to an application (e.g., software) being configured to perform one or more functions. In such examples, the application may be stored in a memory (e.g., on-chip memory of a processor, system memory, or any other memory). Hardware described herein, such as a processor may be configured to execute the application. For example, the application may be described as including code that,when executed by the hardware, causes the hardware to perform one or more techniques described herein. As an example, the hardware may access the code from a memory and execute the code accessed from the memory to perform one or more techniques described herein. In some examples, components are identified in this disclosure. In such examples, the components may be hardware, software, or a combination thereof. The components may be separate components or subcomponents of a single component.
[0035] In one or more examples described herein, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored on or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can include a random access memory (RAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the aforementioned types of computer-readable media, or any other medium that can be used to store computer executable code in the form of instructions or data structures that can be accessed by a computer.
[0036] As used herein, instances of the term “content” may refer to “graphical content,” an “image,” etc., regardless of whether the terms are used as an adjective, noun, or other parts of speech. In some examples, the term “graphical content,” as used herein, may refer to a content produced by one or more processes of a graphics processing pipeline. In further examples, the term “graphical content,” as used herein, may refer to a content produced by a processing unit configured to perform graphics processing. In still further examples, as used herein, the term “graphical content” may refer to a content produced by a graphics processing unit.
[0037] Some relighting techniques may include multiple sequential neural networks which may cause computational bottlenecks that may cause the relighting techniques to be difficult to implement on a device. Other relighting techniques may operate on a perframe basis and may suffer from flickering artifacts originating from a de-lighting network. Removing the flickering artifacts may entail additional processing, which may be computationally burdensome on a device.
[0038] Various technologies pertaining to a light adding based relighting framework are described herein. In an example, an apparatus may compute, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image. A masked image (e.g., a masked gamma-corrected image) may refer to an image where some pixels of the image are set to zero by setting the background or hiding some image parts with some text, image, cropping, and editing. A masked gammacorrected foreground image may be a foreground image that a system performs gamma correction upon, for example gamma compression of a raw captured image. The system may apply a mask to an image to capture the foreground image, for example by analyzing an image to identify a foreground object in the image (e.g., capturing a face within the image), and masking the background to obtain an image of the face without the background of the image. A normal map may include a red green blue (RGB) image where RGB components correspond to X, Y, and Z coordinates, respectively, of a surface normal. Such a normal map may include a texture map of an image based on normal mapping, or bump mapping, where the normal angles, or bumps, may represent how light reflects off of an object in the image. The apparatus may compute a light map based on the normal map and a high dynamic range (HDR) map. The HDR map may include brightness values for each pixel of an image having a high dynamic range value (e.g., represented by 32 bits) instead of a low dynamic range value (e.g., represented by 8 bits). As used herein, a high dynamic range value may be a value that is comparatively greater than a low dynamic range value. The light map may include a texture map that simulates complex lighting. For example, a light map may include a set of light sources illuminating the object in the normal map, where each light source may be represented by a brightness value correlated with the HDR map. The apparatus may generate an image (e.g., an HDR image) based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image. The the first cropped image may be associated with a first saturation level and the second cropped image may associated with a second saturation level that is less than the first saturation level. The saturation level may represent the intensity of a color in a picture between a minimum value and a maximum value, for example between 0 and 99, or between 0 and 255. The apparatus may output an indication of the generated image. Vis-a-vis computing a light map based on the normal map and a high dynamicrange (HDR) map and generating an image based on the light map, a first cropped image associated with the foreground image, and a second cropped image associated with the foreground image, the apparatus may be able to produce high-quality, realistic, and consistent relighting results with limited computing resources. Furthermore, unlike some other relighting approaches, the relighting results may be free from flickering artifacts.
[0039] In some aspects, to reduce the computation complexity, a lightweight relighting engine may contain a single normal estimation network and diffuse light map calculation. A relighting framework may use the relighting engine, along with an optional a rendering equation, which may be configured to add a diffuse light map to a captured frame (or set of frames) to enhance temporal consistency of the lighting.
[0040] The examples describe herein may refer to a use and functionality of a graphics processing unit (GPU). As used herein, a GPU can be any type of graphics processor, and a graphics processor can be any type of processor that is designed or configured to process graphics content. For example, a graphics processor or GPU can be a specialized electronic circuit that is designed for processing graphics content. As an additional example, a graphics processor or GPU can be a general purpose processor that is configured to process graphics content.
[0041] FIG. 1 is a block diagram that illustrates an example content generation system 100 configured to implement one or more techniques of this disclosure. The content generation system 100 includes a device 104. The device 104 may include one or more components or circuits for performing various functions described herein. In some examples, one or more components of the device 104 may be components of a SOC. The device 104 may include one or more components configured to perform one or more techniques of this disclosure. In the example shown, the device 104 may include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, the device 104 may include a number of components (e.g., a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131). Display(s) 131 may refer to one or more displays 131. For example, the display 131 may include a single display or multiple displays, which may include a first display and a second display. The first display may be a left-eye display and the second display may be a right-eye display. In some examples, the first display and the second display may receive differentframes for presentment thereon. In other examples, the first and second display may receive the same frames for presentment thereon. In further examples, the results of the graphics processing may not be displayed on the device, e.g., the first display and the second display may not receive any frames for presentment thereon. Instead, the frames or graphics processing results may be transferred to another device. In some aspects, this may be referred to as split-rendering.
[0042] The processing unit 120 may include an internal memory 121. The processing unit 120 may be configured to perform graphics processing using a graphics processing pipeline 107. The content encoder / decoder 122 may include an internal memory 123. In some examples, the device 104 may include a processor, which may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120 before the frames are displayed by the one or more displays 131. While the processor in the example content generation system 100 is configured as a display processor 127, it should be understood that the display processor 127 is one example of the processor and that other types of processors, controllers, etc., may be used as substitute for the display processor 127. The display processor 127 may be configured to perform display processing. For example, the display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120. The one or more displays 131 may be configured to display or otherwise present frames processed by the display processor 127. In some examples, the one or more displays 131 may include one or more of a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, a projection display device, an augmented reality display device, a virtual reality display device, a head-mounted display, or any other type of display device.
[0043] Memory external to the processing unit 120 and the content encoder / decoder 122, such as system memory 124, may be accessible to the processing unit 120 and the content encoder / decoder 122. For example, the processing unit 120 and the content encoder / decoder 122 may be configured to read from and / or write to external memory, such as the system memory 124. The processing unit 120 may be communicatively coupled to the system memory 124 over a bus. In some examples, the processing unit 120 and the content encoder / decoder 122 may be communicatively coupled to the internal memory 121 over the bus or via a different connection.
[0044] The content encoder / decoder 122 may be configured to receive graphical content from any source, such as the system memory 124 and / or the communication interface 126. The system memory 124 may be configured to store received encoded or decoded graphical content. The content encoder / decoder 122 may be configured to receive encoded or decoded graphical content, e.g., from the system memory 124 and / or the communication interface 126, in the form of encoded pixel data. The content encoder / decoder 122 may be configured to encode or decode any graphical content.
[0045] The internal memory 121 or the system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, internal memory 121 or the system memory 124 may include RAM, static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, a magnetic data media or an optical storage media, or any other type of memory. The internal memory 121 or the system memory 124 may be a non-transitory storage medium according to some examples. The term “non- transitory” may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term “non-transitory” should not be interpreted to mean that internal memory 121 or the system memory 124 is non-movable or that its contents are static. As one example, the system memory 124 may be removed from the device 104 and moved to another device. As another example, the system memory 124 may not be removable from the device 104.
[0046] The processing unit 120 may be a CPU, a GPU, a GPGPU, or any other processing unit that may be configured to perform graphics processing. In some examples, the processing unit 120 may be integrated into a motherboard of the device 104. In further examples, the processing unit 120 may be present on a graphics card that is installed in a port of the motherboard of the device 104, or may be otherwise incorporated within a peripheral device configured to interoperate with the device 104. The processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuitry, or any combinations thereof. If the techniques are implemented partially in software, the processing unit 120 may store instructions for the software in a suitable, non-transitory computer-readable storage medium, e.g., internal memory 121, and may execute the instructions in hardware using one or more processors toperform the techniques of this disclosure. Any of the foregoing, including hardware, software, a combination of hardware and software, etc., may be considered to be one or more processors.
[0047] The content encoder / decoder 122 may be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 may be integrated into a motherboard of the device 104. The content encoder / decoder 122 may include one or more processors, such as one or more microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuitry, or any combinations thereof. If the techniques are implemented partially in software, the content encoder / decoder 122 may store instructions for the software in a suitable, non-transitory computer-readable storage medium, e.g., internal memory 123, and may execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Any of the foregoing, including hardware, software, a combination of hardware and software, etc., may be considered to be one or more processors.
[0048] In some aspects, the content generation system 100 may include a communication interface 126. The communication interface 126 may include a receiver 128 and a transmitter 130. The receiver 128 may be configured to perform any receiving function described herein with respect to the device 104. Additionally, the receiver 128 may be configured to receive information, e.g., eye or head position information, rendering commands, and / or location information, from another device. The transmitter 130 may be configured to perform any transmitting function described herein with respect to the device 104. For example, the transmitter 130 may be configured to transmit information to another device, which may include a request for content. The receiver 128 and the transmitter 130 may be combined into a transceiver 132. In such examples, the transceiver 132 may be configured to perform any receiving function and / or transmitting function described herein with respect to the device 104.
[0049] Referring again to FIG. 1, in certain aspects, the processing unit 120 may include a light adder 198 configured to compute, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image; compute a light mapbased on the normal map and a high dynamic range (HDR) map; generate an image based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image, where the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level; and output an indication of the generated image. Although the following description may be focused on graphics processing, the concepts described herein may be applicable to other similar processing techniques.
[0050] A device, such as the device 104, may refer to any device, apparatus, or system configured to perform one or more techniques described herein. For example, a device may be a server, a base station, a user equipment, a client device, a station, an access point, a computer such as a personal computer, a desktop computer, a laptop computer, a tablet computer, a computer workstation, or a mainframe computer, an end product, an apparatus, a phone, a smart phone, a server, a video game platform or console, a handheld device such as a portable video game device or a personal digital assistant (PDA), a wearable computing device such as a smart watch, an augmented reality device, or a virtual reality device, a non-wearable device, a display or display device, a television, a television set-top box, an intermediate network device, a digital media player, a video streaming device, a content streaming device, an in-vehicle computer, any mobile device, any device configured to generate graphical content, or any device configured to perform one or more techniques described herein. Processes herein may be described as performed by a particular component (e.g., a GPU) but in other embodiments, may be performed using other components (e.g., a CPU) consistent with the disclosed embodiments.
[0051] GPUs can process multiple types of data or data packets in a GPU pipeline. For instance, in some aspects, a GPU can process two types of data or data packets, e.g., context register packets and draw call data. A context register packet can be a set of global state information, e.g., information regarding a global register, shading program, or constant data, which can regulate how a graphics context will be processed. For example, context register packets can include information regarding a color format. In some aspects of context register packets, there can be a bit or bits that indicate which workload belongs to a context register. Also, there can be multiple functions or programming running at the same time and / or in parallel. For example,functions or programming can describe a certain operation, e.g., the color mode or color format. Accordingly, a context register can define multiple states of a GPU.
[0052] Context states can be utilized to determine how an individual processing unit functions, e.g., a vertex fetcher (VFD), a vertex shader (VS), a shader processor, or a geometry processor, and / or in what mode the processing unit functions. In order to do so, GPUs can use context registers and programming data. In some aspects, a GPU can generate a workload, e.g., a vertex or pixel workload, in the pipeline based on the context register definition of a mode or state. Certain processing units, e.g., a VFD, can use these states to determine certain functions, e.g., how a vertex is assembled. As these modes or states can change, GPUs may need to change the corresponding context. Additionally, the workload that corresponds to the mode or state may follow the changing mode or state.
[0053] FIG. 2 illustrates an example GPU 200 in accordance with one or more techniques of this disclosure. As shown in FIG. 2, GPU 200 includes command processor (CP) 210, draw call packets 212, VFD 220, VS 222, vertex cache (VPC) 224, triangle setup engine (TSE) 226, rasterizer (RAS) 228, Z process engine (ZPE) 230, pixel interpolator (PI) 232, fragment shader (FS) 234, render backend (RB) 236, L2 cache (UCHE) 238, and system memory 240. Although FIG. 2 displays that GPU 200 includes processing units 220-238, GPU 200 can include a number of additional processing units. Additionally, processing units 220-238 are merely an example and any combination or order of processing units can be used by GPUs according to the present disclosure. GPU 200 also includes command buffer 250, context register packets 260, and context states 261.
[0054] As shown in FIG. 2, a GPU can utilize a CP, e.g., CP 210, or hardware accelerator to parse a command buffer into context register packets, e.g., context register packets 260, and / or draw call data packets, e.g., draw call packets 212. The CP 210 can then send the context register packets 260 or draw call packets 212 through separate paths to the processing units or blocks in the GPU. Further, the command buffer 250 can alternate different states of context registers and draw calls. For example, a command buffer can simultaneously store the following information: context register of context N, draw call(s) of context N, context register of context N+l, and draw call(s) of context N+l.
[0055] GPUs can render images in a variety of different ways. In some instances, GPUs can render an image using direct rendering and / or tiled rendering. In tiled rendering GPUs, an image can be divided or separated into different sections or tiles. After the division of the image, each section or tile can be rendered separately. Tiled rendering GPUs can divide computer graphics images into a grid format, such that each portion of the grid, i.e., a tile, is separately rendered. In some aspects of tiled rendering, during a binning pass, an image can be divided into different bins or tiles. In some aspects, during the binning pass, a visibility stream can be constructed where visible primitives or draw calls can be identified. A rendering pass may be performed after the binning pass. In contrast to tiled rendering, direct rendering does not divide the frame into smaller bins or tiles. Rather, in direct rendering, the entire frame is rendered at a single time (i.e., without a binning pass). Additionally, some types of GPUs can allow for both tiled rendering and direct rendering (e.g., flex rendering).
[0056] In some aspects, GPUs can apply the drawing or rendering process to different bins or tiles. For instance, a GPU can render to one bin, and perform all the draws for the primitives or pixels in the bin. During the process of rendering to a bin, the render targets can be located in GPU internal memory (GMEM). In some instances, after rendering to one bin, the content of the render targets can be moved to a system memory and the GMEM can be freed for rendering the next bin. Additionally, a GPU can render to another bin, and perform the draws for the primitives or pixels in that bin. Therefore, in some aspects, there might be a small number of bins, e.g., four bins, that cover all of the draws in one surface. Further, GPUs can cycle through all of the draws in one bin, but perform the draws for the draw calls that are visible, i.e., draw calls that include visible geometry. In some aspects, a visibility stream can be generated, e.g., in a binning pass, to determine the visibility information of each primitive in an image or scene. For instance, this visibility stream can identify whether a certain primitive is visible or not. In some aspects, this information can be used to remove primitives that are not visible so that the non-visible primitives are not rendered, e.g., in the rendering pass. Also, at least some of the primitives that are identified as visible can be rendered in the rendering pass.
[0057] In some aspects of tiled rendering, there can be multiple processing phases or passes. For instance, the rendering can be performed in two passes, e.g., a binning, a visibility or bin-visibility pass and a rendering or bin-rendering pass. During a visibility pass, aGPU can input a rendering workload, record the positions of the primitives or triangles, and then determine which primitives or triangles fall into which bin or area. In some aspects of a visibility pass, GPUs can also identify or mark the visibility of each primitive or triangle in a visibility stream. During a rendering pass, a GPU can input the visibility stream and process one bin or area at a time. In some aspects, the visibility stream can be analyzed to determine which primitives, or vertices of primitives, are visible or not visible. As such, the primitives, or vertices of primitives, that are visible may be processed. By doing so, GPUs can reduce the unnecessary workload of processing or rendering primitives or triangles that are not visible.
[0058] In some aspects, during a visibility pass, certain types of primitive geometry, e.g., position-only geometry, may be processed. Additionally, depending on the position or location of the primitives or triangles, the primitives may be sorted into different bins or areas. In some instances, sorting primitives or triangles into different bins may be performed by determining visibility information for these primitives or triangles. For example, GPUs may determine or write visibility information of each primitive in each bin or area, e.g., in a system memory. This visibility information can be used to determine or generate a visibility stream. In a rendering pass, the primitives in each bin can be rendered separately. In these instances, the visibility stream can be fetched from memory and used to remove primitives which are not visible for that bin.
[0059] Some aspects of GPUs or GPU architectures can provide a number of different options for rendering, e.g., software rendering and hardware rendering. In software rendering, a driver or CPU can replicate an entire frame geometry by processing each view one time. Additionally, some different states may be changed depending on the view. As such, in software rendering, the software can replicate the entire workload by changing some states that may be utilized to render for each viewpoint in an image. In certain aspects, as GPUs may be submitting the same workload multiple times for each viewpoint in an image, there may be an increased amount of overhead. In hardware rendering, the hardware or GPU may be responsible for replicating or processing the geometry for each viewpoint in an image. Accordingly, the hardware can manage the replication or processing of the primitives or triangles for each viewpoint in an image.
[0060] FIG. 3 illustrates image or surface 300, including multiple primitives divided into multiple bins in accordance with one or more techniques of this disclosure. As shownin FIG. 3, image or surface 300 includes area 302, which includes primitives 321, 322, 323, and 324. The primitives 321, 322, 323, and 324 are divided or placed into different bins, e.g., bins 310, 311, 312, 313, 314, and 315. FIG. 3 illustrates an example of tiled rendering using multiple viewpoints for the primitives 321-324. For instance, primitives 321-324 are in first viewpoint 350 and second viewpoint 351. As such, the GPU processing or rendering the image or surface 300 including area 302 can utilize multiple viewpoints or multi-view rendering.
[0061] As indicated herein, GPUs or graphics processors can use a tiled rendering architecture to reduce power consumption or save memory bandwidth. As further stated above, this rendering method can divide the scene into multiple bins, as well as include a visibility pass that identifies the triangles that are visible in each bin. Thus, in tiled rendering, a full screen can be divided into multiple bins or tiles. The scene can then be rendered multiple times, e.g., one or more times for each bin.
[0062] In aspects of graphics rendering, some graphics applications may render to a single target, i.e., a render target, one or more times. For instance, in graphics rendering, a frame buffer on a system memory may be updated multiple times. The frame buffer can be a portion of memory or random access memory (RAM), e.g., containing a bitmap or storage, to help store display data for a GPU. The frame buffer can also be a memory buffer containing a complete frame of data. Additionally, the frame buffer can be a logic buffer. In some aspects, updating the frame buffer can be performed in bin or tile rendering, where, as discussed above, a surface is divided into multiple bins or tiles and then each bin or tile can be separately rendered. Further, in tiled rendering, the frame buffer can be partitioned into multiple bins or tiles.
[0063] As indicated herein, in some aspects, such as in bin or tiled rendering architecture, frame buffers can have data stored or written to them repeatedly, e.g., when rendering from different types of memory. This can be referred to as resolving and unresolving the frame buffer or system memory. For example, when storing or writing to one frame buffer and then switching to another frame buffer, the data or information on the frame buffer can be resolved from the GMEM at the GPU to the system memory, i.e., memory in the double data rate (DDR) RAM or dynamic RAM (DRAM).
[0064] In some aspects, the system memory can also be system-on-chip (SoC) memory or another chip-based memory to store data or information, e.g., on a device or smart phone. The system memory can also be physical data storage that is shared by theCPU and / or the GPU. In some aspects, the system memory can be a DRAM chip, e.g., on a device or smart phone. Accordingly, SoC memory can be a chip-based manner in which to store data.
[0065] In some aspects, the GMEM can be on-chip memory at the GPU, which can be implemented by static RAM (SRAM). Additionally, GMEM can be stored on a device, e.g., a smart phone. As indicated herein, data or information can be transferred between the system memory or DRAM and the GMEM, e.g., at a device. In some aspects, the system memory or DRAM can be at the CPU or GPU. Additionally, data can be stored at the DDR or DRAM. In some aspects, such as in bin or tiled rendering, a small portion of the memory can be stored at the GPU, e.g., at the GMEM. In some instances, storing data at the GMEM may utilize a larger processing workload and / or consume more power compared to storing data at the frame buffer or system memory.
[0066] FIG. 4A-4B include diagrams that illustrate example aspects pertaining to relighting in accordance with one or more techniques of this disclosure. A goal of relighting may be to re-illuminate subjects with novel lighting given a single image and a target high dynamic range image (HDRI) map. Relighting may be useful in many use cases. Such use cases may include improving low-light / unbalanced lighting conditions in portrait images / videos, providing consistent lighting when applying background replacement during video conferencing, and / or realistically embedding reconstructed avatars into virtual reality (VR) environments.
[0067] In some aspects, an input sequence may include an input frame and a processed frame produced by applying relighting techniques applied to the input frame to improve lighting conditions of the input frame. The relighting techniques applied to the input frame may improve lighting conditions, for example low light applied to the facial features of the person shown in the input frame. A relighting framework may increase light in areas of the face which were not illuminated in the input frame.
[0068] In another aspect, an input sequence may include an input frame, and a processed frame produced by applying relighting techniques applied to the input frame to improve lighting conditions of the input frame. The relighting techniques applied to the input frame may improve lighting conditions, for example unbalanced light applied to the facial features of the person shown in the input frame. A relighting framework may balance the light applied to areas of the face that are unbalanced in the input frame.
[0069] In another aspect, an input sequence that may include an input video and a processed video produced by applying relighting techniques applied to the input video to improve lighting conditions of the input video. An input video and an output video may each include a plurality of frames that may be displayed in a sequence. The relighting techniques applied to the input frame may improve lighting conditions, for example the input video may have a virtual background that is unnaturally dark, or unnaturally bright, as compared with the lighting in the foreground that illuminates the face of the person in the input video. A relighting framework may accurately relight subjects in either the foreground or the background of the input video based on target lighting (e.g., rebalance the subjects in the foreground based on the lighting in the background, or rebalance the subjects in the background based on the lighting in the foreground) to enhance immersive experiences having a virtual background.
[0070] In FIG. 4 A, diagram 430 illustrates an input sequence that may include an input frame 432 and a processed frame 434 produced by applying relighting techniques applied to the input frame 432 to improve lighting conditions of the input frame 432. In FIG. 4B, diagram 440 illustrates an input sequence of an input frame 442 and a processed frame 444 produced by applying relighting techniques applied to the input frame 442 to improve lighting conditions of the input frame 442. The input frame 432 and the input frame 442 may be produced by embedding an avatar of a face (e.g., a face captured by a camera) into a virtual reality (VR) environment. The relit sequence in the processed frames may match the lighting of the avatar with the lighting in the VR environment.
[0071] Some relighting techniques may incorporate reflectance information into a network design (i.e., a neural network design). Such relighting techniques may model illumination through a per-pixel lighting representation compared to other techniques that use latent codes. However, relighting techniques that incorporate reflectance information into a network design may include multiple sequential networks. The multiple sequential networks may cause computational bottlenecks that may make the multiple sequential networks difficult to implement on devices. In addition, per-frame relighting techniques may suffer from flickering artifacts that originate from a delighting network. The flickering artifacts may be removed by further processing, which may be computationally burdensome. Aspects presented herein pertain to alight adding based relighting framework to address challenges of computational bottlenecks and / or temporal consistency.
[0072] In one aspect described herein, the light adding based relighting framework may include a lightweight relighting module, or engine, which may include a single normal estimation network and a diffuse light map calculation. In another aspect described herein, a novel rendering equation is proposed which may add a diffuse light map on top of camera captures in order to enhance temporal consistency. The light adding based relighting framework may be applied to difference scenarios, such as light source adding, room light adding, lighting condition improvement, high dynamic range (HDR) relighting, etc.
[0073] FIG. 5A is a diagram 500 illustrating an example of a light adding based relighting framework, in accordance with one or more techniques of this disclosure. FIG. 5B is a diagram 550 illustrating an example of a total relighting framework, in accordance with one or more techniques of this disclosure. The light adding-based relighting framework shown in FIG. 5A may be configured to maintain a balance between relighting performance and computational efficiency. The light adding based relighting framework may achieve photorealistic video relighting using a combination of a single light-weight normal estimation network, a diffuse light map calculation, and a novel rendering equation. For example, a system may input a gamma-corrected foreground image 502 into the light-adding-based relighting framework shown in FIG. 5A. A geometry net 504 may process the gamma-corrected foreground image to generate a normal map 506. A pre-convoluted HDR map 508 may apply a diffuse light map calculation to the normal map 506 to generate a diffuse light map 510. The light adding-based relighting framework may apply a rendering equation 512 to the diffuse light map 510 to generate a relit foreground 514 based on the diffuse light map 510. The relighting framework may then composite a virtual background 516 with the relit foreground 514 based on a mask 518 to generate a relit image 520 with a virtual foreground.
[0074] FIGs. 5A-5B illustrate differences between an exemplary light adding-based relighting framework and a total relighting framework. A system may input a gammacorrected foreground image 552 into the total relighting framework shown in FIG. 5B. A geometry net 504 may process the gamma-corrected foreground image to generate a normal map 506. A pre-convoluted HDR map 558 may apply a diffuselight map calculation to the normal map 506 to generate a set of diffuse, or specular, light maps 560. The system may apply a de-lighting net 554 to the normal map 506, and optionally the foreground image 552, to generate an albedo, low-light image 556. The system may apply a set of diffuse light map calculations based on the set of diffuse, or specular, light maps 560 to the albedo, low-light image 556, and optionally the foreground image 552, and filter the result through a shading net 558 to generate a relit foreground 564. The relighting framework may then composite a virtual background 516 with the relit foreground 564 based on a mask 518 to generate a relit image 570 with a virtual foreground.
[0075] FIG. 6 is a diagram of a rendering engine 600, in accordance with one or more techniques of this disclosure. The light adding based relighting framework may render an image frame according to equation (I) below.(I) relit image = scale X face cropped image + scale2X low saturation face image X masked diffuse light map
[0076] Each scalei and scale? may include default scale values, which may be statically defined or dynamically defined by an admit user. The low- saturation face image and the masked diffuse light map may be based on the face-cropped image.
[0077] In one aspect, the rendering engine 600 may crop a face-cropped image 602 from an input interface, for example from camera captures or a pre-processing engine. The rendering engine 600 may use a masked diffuse light map 606 to create physically convincing light that is to be added to the face-cropped image 602. A physical based rendering (PBR) pipeline, for example one used in three-dimensional (3D) graphics software, may calculate the masked diffuse light map 606. In addition, the rendering engine 600 may obtain a low- saturation face image 604, for example from a preprocessing engine. The rendering engine 600 may multiply the low- saturation face image 604 in a rendering equation 610 to adjust diffuse lights by adding more light to lighter areas and less light to darker areas. A set of input scales 608 may be used to scale the face-cropped image 602, the low-saturation face image 604, and / or the masked diffuse light map 606 relative to one another. The rendering engine 600 may use the rendering equation 610 to generate a relit image 612 in standard red, green blue (sRGB) space.
[0078] FIG. 7 is a diagram of a relighting engine 700, in accordance with one or more techniques of this disclosure. The relighting engine 700 may accept a masked gammacorrected foreground image 702 as an input, for example from a pre-processing engine. The relighting engine 700 may process the masked gamma-corrected foreground image 702 via a geometry net 704 to generate a normal map 706 of the masked gamma-corrected foreground image 702. The relighting engine 700 may process the normal map 706 in a variety of ways, for example by resizing the normal map 706, by cropping a region of interest (ROI) of the normal map 706, and / or by applying a temporal filter to the normal map 706. In some aspects, the ROI may be, for example, a fixed-sized ROI centered on a camera (e.g., a 2 in. x 2 in. area centered on a camera-captured image, a 200 pixel x 200 pixel area centered on a cameraOcaptured image). In some aspetes, the ROI may be, for example, an adaptive ROI that is cropped about a detected face in a captured image. For example, the relighting engine 700 may be configured to detect facial features in an image and center the ROI about the most prominent centered detected facial feature, (e.g., a nose of a person). The temporal filter may be, for example, an average temporal filter. In some aspects, the temporal filter may calculate an average normal prediction based on a number of consecutive frames, for example three consecutive frames.
[0079] In some aspects, the relighting engine 700 may accept as an input a pre-integrated diffuse HDR map as an input 708. A pre-integrated diffuse HDR map may be a diffuse HDR map that is integrated into an HDR map before the relighting engine 700 executes. In other words, the pre-integrated diffuse HDR map may be fed as an input to the relighting engine 700 upon startup or may be stored in an accessible memory and retrieved by the relighting engine 700 upon startup. The relighting engine 700 may be configured to load the HDR map once for an entire session, once for a number of frames (e.g., 10 frames, 20 frames), or once for every frame processed by the relighting engine 700. In some aspects, the relighting engine 700 may accept as an input a light color (e.g., in RGB format) as an input 710, which may be used to calculate a light source (e.g., a green light source for a green color input into the relighting engine 700, a red light source for a red color input into the relighting engine 700). In some aspects, the input 710 may accept a light color as an input, which may be used as a basis to determine a light source to add, or may use an entire HDR map for relighting if no light source is output to the input 710 (e.g., input 710 is a nullvalue). In some aspects, the relighting engine 700 may accept an HDR rotation matrix 712, for example as an output from a pre-processing engine. The relighting engine 700 may perform a diffuse light map calculation to generate a diffuse light map 714 based on the normal map 706, the input 708 (e.g., the HDR map), and the HDR rotation matrix 712.
[0080] FIG. 8 is a diagram of a pre-processing engine 800, in accordance with one or more techniques of this disclosure. The pre-processing engine 800 may accept a set of inputs 802, for example an input image, a segmentation map (seg) and a face bounding box (BB). A segmentation map may be a visual representation that divides an image into distinct regions based on a feature or a characteristic, for example a brightness or a luminosity threshold. A BB may include coordinates used to designate where a face of a person is located. A face BB, or facial BB, may bound the edges of the face, or may be a simple BB defined by four (x,y) coordinates that estimate the bounds of the face, within which a different type of processing may be applied than outside the face BB. The face BB may be applied to an image of a face, or a facial image, to indicate what portions of the image may be processed using a facial lightning algorithm, as opposed to other portions of the image that may be processed using other lighting algorithms (e.g., a background lighting algorithm, a wall algorithm). A facial lighting algorithm may be more complex, or may use more processing power, than other lighting algorithms, which may be less complex, or may use less processing power, than the facial lighting algorithm.
[0081] The pre-processing engine 800 may crop the input image based on the face BB and the segmentation map to generate a cropped image 804. In some aspects, the preprocessing engine 800 may further process the cropped image 804, for example by generating a low- saturation face image by converting an RGB value of the cropped image 804 to an hue saturation and value (HSV), multiply a saturation channel by a scale (e.g., 0.6) in HSV space, and / or convert the multiplied value back to an RGB value. In some aspects, the pre-processing engine 800 may apply gamma correction 2.6 (e.g., via a lookup table), generate a set of masks (e.g., a foreground object mask, a hair mask), and / or process segmentation of the cropped image 804 (e.g., identify segments of the image). For example, the pre-processing engine 800 may generate a foreground segmentation map and a hair segmentation map based on the cropped image 804. The segmentation maps may be applied to the cropped image 804 togenerate a masked gamma-corrected foreground image 808, which may be processed to generate a processed foreground segmentation map 810. A segmentation map may include a mask that designates coordinates having a first property (e.g., a foreground object) and coordinates having a second property (e.g., a background landscape). Here, the map has white pixels that designate the first property and black pixels that designate the second property, although other pixel colors may be used. In some aspects, a segmentation map may designate more than two properties, although binary segmentation maps may be efficient for simple comparisons and masks. A foreground segmentation map may be a segmentation map that indicates coordinates for a set of foreground objects from a background landscape. A hair segmentation map may be a segmentation map that indicates coordinates for a set of hair pixels from other pixels of a frame or image.
[0082] The pre-processing engine 800 may generate an HDR rotation matrix (3 x 3) based on an input HDR rotation degrees (yaw, pitch, and roll). An HDR rotation may include a rotation of an HDR map based on a change of view. For example, a person’s face may rotate to one side between a first frame and a second frame of a video stream, and the HDR rotation may apply a previous HDR map based on the first frame rotated in the same manner to the second frame. An HDR rotation matrix may be a matrix that may be used to rotate an HDR map accordingly. The relighting framework may generate a low saturation face image. For instance, the pre-processing engine 800 may convert from red green blue (RGB) to hue saturation and value (HSV). The preprocessing engine 800 may multiply a saturation channel by 0.6 in HSV space. The pre-processing engine 800 may then convert from HSV to RGB. The pre-processing engine 800 may generate a masked gamma-corrected foreground image. For instance, when a pixel value is 0, the pre-processing engine 800 may assign the value to be 1. The device may apply a gamma correction of 2.6 using a lookup table. The preprocessing engine 800 may set background pixels to a value of 0. The pre-processing engine 800 may generate an HDR rotation matrix (3 x 3) given input HDR rotation degrees (yaw, pitch, and roll).
[0083] FIG. 9 is a diagram of a background composition engine 900, in accordance with one or more techniques of this disclosure. The background composition engine 900 may accept a pre-rendered HDR background 902 as an input. An HDR background may be a background stored using brightness values having a high dynamic range value(e.g., represented by 32 bits) instead of a low dynamic range value (e.g., represented by 8 bits). In some aspects, the pre-rendered HDR background 902 may be generated based on an entire HDR image. A pre-rendered HDR background may be stored in a storage before a background composition engine executes. In other words, the prerendered HDR background may be stored in a library and retrieved for use as a background for superimposing behind an object of a video s. The background composition engine 900 may composite a cropped image, for example the cropped image 804 in FIG. 8, with the pre-rendered HDR background 902 based on a processed foreground segmentation map 904 to generate a relit image 906 having a virtual background. In some aspects, where the background composition engine 900 does not accept a pre-rendered HDR background 902 as an input (e.g., the input is null, a flag for turning off the background composition engine 900 is set), the background composition engine may use a single light source (e.g., an HDR panorama generated based on a small white point light source, or a small green point light source, and a black background).
[0084] In one aspect, a lightweight and stable relighting framework may achieve high-quality and time-consistent video relighting that may run in real-time on devices. Other relighting techniques may not be able to produce high-quality and realistic relighting results with limited computational resources. For instance, image relighting techniques may be prone to flickering artifacts since relighting is applied to each frame without considering consecutive frames. In an example, a total relighting framework may not be able to produce stable relit videos, particularly for videos that include clothing regions. Some video relighting techniques may include a light-weight encoder-decoder architecture that may enable mobile computing (e.g., 15 frames per second (FPS) on a smartphone); however, such architecture may lose facial details and / or may reduce photorealism. Other video relighting techniques may extend a total relighting framework with one or more temporal residual networks to obtain a temporally smoother relit video; however, high network complexity may make such techniques difficult to implement on mobile devices (e.g., such high complexity may use many power resources, which are limited on mobile devices).
[0085] The light adding based relighting framework described herein may realistically relight subjects without de-lighting and without synthesis of high-frequency details. The light adding based relighting framework may add diffuse lighting effects to make subjectsblend into a target HDR environment. Compared to a total relighting framework, the light adding based relighting framework may discard an albedo net engine and / or a shading net engine while using a geometry net to reduce computational complexity while preserving relighting quality. The light adding based relighting framework may utilize a novel rendering equation that may add high-quality and consistent diffuse lighting effects. Some other relighting techniques may use networks to predict intermediate albedo and a direct relit output and hence may be prone to flickering artifacts. In the light adding based relighting framework described herein, by adding light on face cropped images, temporal consistency may be improved due to the stable nature of camera captures.
[0086] The light adding based relighting framework described herein may be used in various image and video relighting scenarios. Such scenarios may include relighting images / videos using single or multiple light sources, relighting images / videos using a given HDR panorama map, improving lighting conditions on images / videos (including low light images / videos, unbalanced light images / videos, and / or shadow images / videos), synthesizing physically convincing lighting effects on a foreground when applying background replacement during video conferencing, and / or embedding avatars into VR environments with consistent lighting.
[0087] The light adding based relighting framework may produce relit images that are physically convincing while preserving facial details in an input image. Physically convincing properties may be observed by checking if a lighting effect is consistent with a direction of a dominant light source and / or if an overall relit color tone is close to a target HDR environment.
[0088] Some relighting techniques may include multiple sequential networks and computational bottlenecks that cause difficulty in implementation on a device. Portrait image relighting techniques may suffer from flickering artifacts that primarily originate from a de-lighting network and that entail further processing in order to be removed. Other relighting techniques may utilize a light-weight encoder-decoder architecture which may enable mobile computing (e.g., 15 FPS on a smartphone); however, facial details may be lost and photorealism may be reduced. Additional relighting techniques may extend total relighting with temporal residual networks to obtain temporally smoother relit video; however, high network complexity may make this technique difficult to implement on mobile devices.
[0089] The light adding based relighting framework described herein may add realistic diffuse lighting effects to camera captures. The light adding based relighting framework may include four main modules: a pre-processing module, a relighting module, a rendering module, and a background composition module. Relighting scenarios may be divided into two types based on algorithm differences: single light source relighting and whole HDR relighting. A single source relighting may relight an image using a single, or one, light source. A whole relighting may relight an image using a plurality of light sources. A whole HDR relighting may relight an image using a plurality of light sources, where each light source is represented by a high dynamic range value (e.g., represented by 32 bits) instead of a low dynamic range value (e.g, represented by 8 bits or 16 bits).
[0090] FIG. 10 is a diagram 1000 illustrating an example of a relighting framework, having the relighting engine 700 of FIG. 7, the pre-processing engine 800 of FIG. 8, the rendering engine 600 of FIG. 6, and the background composition engine 900 of FIG. 9, in accordance with one or more techniques of this disclosure. The pre-processing engine may obtain a set of inputs 1002, for example an input image, foreground and hair segmentation maps, and / or a face BB. The pre-processing engine 800 may process the set of inputs 1002 to generate a masked gamma-corrected foreground image 1004 that is processed by the relighting engine 700, and a foreground segmentation map 1008 that is processed by the background composition engine 900. The rendering engine 600 may process the diffuse light map 1006 from the relighting engine 700 and the foreground segmentation map 1008 generated from the preprocessing engine 800 to generate a masked diffuse light map 1010, which may then be used to generate a relit image 1012 (e.g., in sRGB space). The background composition engine 900 may then compose the relit image 1014 on a background based on the foreground segmentation map 1008 and the relit image 1012.
[0091] FIG. 11 is a diagram of a pre-processing engine 1100, in accordance with one or more techniques of this disclosure. Similar to the pre-processing engine 800 in FIG. 8, the pre-processing engine 1100 may accept a set of inputs 1102. The set of inputs 1102 may include an input image (e.g., a frame from a camera live feed), a set of segmentation maps (e.g., segmentation map(s) for a foreground segmentation and a hair segmentation) and a face BB. The pre-processing engine 1100 may process the input image in a variety of ways. In some aspects, the pre-processing engine 1100may pad a ROI of the input image. In some aspects, the pre-processing engine 1100 may may apply gamma correction 2.6 (e.g., via a lookup table). In some aspects, the pre-processing engine 1100 may apply a set of masks, for example based on a set of segmentation maps received as the set of inputs 1102, to the input image (e.g., to filter out a background, to separate a face from a hair). In some aspects, the pre-processing engine 1100 may resize the input image (e.g., increase the size of the image before further processing, or decrease the size of the image before further processing). In some aspects, the pre-processing engine 1100 may process segmentation of the input image (e.g., identify segments of the image). In some aspects, the pre-processing engine 1100 may generate a low-saturation face image, for example by multiplying a saturation channel by a scale (e.g., 0.6), by adding a color offset (e.g., 0.05), and / or by multiplying a greyscale image by a brightness factor (e.g., 0.4). The pre-processing engine 1100 may generate a masked gamma-corrected foreground image 1108 based on the processing. The pre-processing engine 1100 may generate a processed foreground segmentation map 1110 based on the processing. The pre-processing engine 1100 may generate an HDR rotation matrix (3 x 3) given input HDR rotation degrees (yaw, pitch, and roll).
[0092] FIG. 12 is a diagram of a background engine 1200, in accordance with one or more techniques of this disclosure. The background engine 1200 may obtain an input of a generated HDR image 1202, for example a 360° image for use in a background. In some aspects, the background engine 1200 may generate the HDR image based on a selection from a user, for example, in response to detecting that a user has selected the text “stadium,” the background engine 1200 may generate a stadium model, or in response to detecting that a user has selected the text “Times Square,” the background engine 1200 may generate a model of Times Square. The background engine 1200 may be configured to generate a set of background images once based on a stimulus, for example a selection of text by a user or upon startup of an application. The background engine 1200 may perform diffuse light integration to generate an integrated diffuse HDR map 1204, which may be used for relighting (e.g., by a relighting engine). The background engine 1200 may perform tone mapping to generate a low dynamic range (LDR) map 1206, which may be used for background rendering (e.g., by a background composition engine).
[0093] FIG. 13 is a diagram of a background composition engine 1300, in accordance with one or more techniques of this disclosure. The background composition engine 1300 may accept a rendered background 1302 as an input. For example, the background composition engine 1300 may accept a rendered background from a background engine, such as the background engine 1200 of FIG. 12. In some aspects, the rendered background 1302 may be generated based on an HDR rotation matrix and a generated LDR map. The background composition engine 1300 may composite a relighted image with the rendered background 1302 based on a processed foreground segmentation map 1304 to generate a relit image 1306 having a virtual background.
[0094] FIG. 14 is a diagram 1400 illustrating an example of a relighting framework, having the background engine 1200 of FIG. 12, the relighting engine 700 of FIG. 7, the preprocessing engine 1100 of FIG. 11, the rendering engine 600 of FIG. 6, and the background composition engine 1300 of FIG. 13, in accordance with one or more techniques of this disclosure. The pre-processing engine may obtain a set of inputs 1402, for example an input image, foreground and hair segmentation maps, and / or a face BB. The pre-processing engine 800 may process the set of inputs 1402 to generate a masked gamma-corrected foreground image 1404 that is processed by the relighting engine 700, and a foreground segmentation map 1408 that is processed by the background composition engine 1300. The background engine 1200 may generate an integrated diffuse HDR map for use by the relighting engine 700 to perform relighting on the masked gamma-corrected foreground image 1404 based on the integrated diffuse HDR map of the generated background. The relighting engine 700 may generate a diffuse light map 1406 based on the integrated diffuse HDR map from the background engine 1200 and the masked gamma-corrected foreground image 1404 from the pre-processing engine 1100. The rendering engine 600 may process the diffuse light map 1406 from the relighting engine 700 and the foreground segmentation map 1408 generated from the pre-processing engine 800 to generate a masked diffuse light map 1410, which may then be used to generate a relit image 1412 (e.g., in sRGB space). The background composition engine 1300 may compose the relight image 1414 on a background based on the foreground segmentation map 1408 and the LDR map from the background engine 1200. The background composition engine 1300 may generate the rendered background 1413 based on the foregroundsegmentation map 1408 (e.g., based on an HDR rotation matrix) from the preprocessing engine 1100 and the LDR map from the background engine 1200.
[0095] A light adding based relighting framework may include a single network (e.g., one geometry net). A geometry net may predict a normal map with details given a masked gamma-corrected foreground image. A geometry net may include a lightweight UNet architecture.
[0096] The rendering module may be associated with a rendering equation. The rendering equation may define the addition of two terms: (1) the original captured image and (2) the light that is to be added. The scales scalei and scale2 may be specified to control a ratio between the original captured image and the light that is to be added.
[0097] The background composition module may composite backgrounds. The backgrounds may be pre-rendered or rendered at run-time. The backgrounds may be composited using a processed foreground segmentation map.
[0098] FIG. 15 is a diagram 1500 illustrating example aspects pertaining to a pre-processing engine in accordance with one or more techniques of this disclosure. A relighting framework may use segmentation processing on a foreground segmentation 1502 to hide potential segmentation artifacts around hair boundaries. The relighting framework may perform the segmentation processing using a foreground segmentation map and a hair segmentation map. For example, the relighting framework may generate a hair segmentation map 1506 based on the foreground segmentation 1502 and an input image of a person’s head, and a face segmentation map 1508 based on the foreground segmentation 1502 and an input image of a person’s head. In some aspects, the system may generate the hair segmentation map 1506 before generating the face segmentation map 1508. The relighting framework may perform the segmentation processing using erosion and Gaussian blur parameters. The erosion and Gaussian blur parameters may be adaptive parameters that consider a size of a hair region. When the relighting framework is to perform single light source relighting, the relighting framework may set a larger kernel size and a larger upper bound for the erosion and Gaussian blur parameters. When the relighting framework is to perform whole HDR relighting, the relighting framework may set a smaller kernel size and a smaller upper bound for the erosion and Gaussian blur parameters The relighting framework may perform gaussian blur on the foreground segmentation 1502 to generate a blurred segmentation 1504, and applythe blurred segmentation 1504 to the hair of an image, and not to the face of an image, based on the hair segmentation map 1506 and the face segmentation map 1508 to generate a composite segmentation map 1510 including both gaussian blurred and non-gaussian blurred mask ROIs.
[0099] FIG. 16A is a diagram 1600 illustrating an example aspect pertaining to a first relighting module, and FIG. 16B is a diagram 1650 illustrating an example aspect pertaining to a second relighting module, in accordance with one or more techniques of this disclosure. HDR panoramas may be different for different relighting scenarios. For example, with respect to FIG. 16 A, for single light source relighting, a relighting framework may generate an HDR panorama 1604 based on a small white point light source 1602 and a black background. In another example, with respect to FIG. 16B, for whole HDR relighting, a relighting framework may generate an HDR panorama 1654 based on an 360° environment illumination 1652. The HDR panorama may be pre-defined and pre-convoluted. A pre-convolved diffuse HDR map may support onetime loading when running relighting in a sequence. A relighting framework may not re-convolve a diffuse HDR map in the following scenarios: a change of color of a single light source, a change of rotation degree of an HDRI map, a change of intensity of an HDRI map.
[0100] FIG. 17 is a diagram 1700 illustrating example aspects pertaining to high dynamic range (HDR) relighting in accordance with one or more techniques of this disclosure. A face-cropped sequence 1702 (e.g., a set of consecutive frames of a cropped face) may be processed using a relighting framework. As shown, the relighting framework may be configured to perform no relighting on the face-cropped sequence 1702 to generate a sequence without any relighting 1704, and just a background added to an object (e.g., a user’s face). In another aspect, the relighting framework may be configured to perform basic relighting on the face-cropped sequence 1702 to generate a sequence with relighting 1706, based on the background added to an object. For our whole HDR relighting, the light adding based relighting framework may be temporally smoother and may embed subjects in a desired HDR scene better compared to other relighting techniques.
[0101] FIGs. 18A-18C illustrate various exemplary aspects pertaining to HDR relighting in accordance with one or more techniques of this disclosure. The light adding based relighting framework may be applied in different scenarios, such as light sourceadding, room light adding, lighting condition improvement, HDR relighting, etc. For example, with respect to FIG. 18 A, a relighting framework may process an input image 1802 to perform relighting on the room light illuminating a face of a user to generate a relit image 1804. In another example, with respect to FIG. 18B, a relighting framework may process an input image 1812 to perform relighting on a colored light (e.g., a red light, a light of a certain hue) illuminating a face of a user to generate a relit image 1814 based on the colored light. A user may select a colored light to virtualize illumination on the face, to provide more of a certain hue to the face of the user. In another example, with respect to FIG. 18C, a relighting framework may process an input image 1842 to add a background to generate a composite image without relighting 1844, and perform relighting based on an HDR image (HDRI) based on a synthetic dataset (e.g., based on a series of background images, such as the previous 3 consecutive background images) to generate a relit image 1846 based on total relighting effects of the synthetic dataset.
[0102] In other aspects, a relighting framework may process an input image to improve lighting conditions (e.g., increase the amount of light illuminating a face of a foreground object) to generate a relit image based on the existing detected lighting conditions, or , a relighting framework may process an input image to improve lighting conditions (e.g., rebalance the light currently illuminating a face of a foreground object) to generate a relit image based on an average or mean lighting condition. However, such relighting frameworks that do not perform relighting based on a synthetic dataset may not produce a realistic lighting effects, as a simple increase in luminosity or balance may not achieve realistic effects.
[0103] FIG. 19 is a call flow diagram 1900 illustrating example communications between a CPU 1902 and a GPU 1904 in accordance with one or more techniques of this disclosure. The CPU 1902 may output an indicator 1906 of a set of images to the GPU 1904. The GPU 1904 may receive the indicator 1906 of the set of images from the CPU 1902. The set of images may include captured images of a face of a user having a foreground of the face of the user and a background. At 1908, the GPU 1904 may compute a normal map based on the set of images received from the CPU 1902. The GPU 1904 may use a machine learning (ML) model to generate the normal map based on a masked gamma-corrected foreground image of an object in an image. At 1910, the GPU 1904 may compute a light map based on the normal map computed at 1908and an HDR map (e.g., based on the background image, or a set of background images, for the object). At 1912, the GPU 1904 may generate a set of images based on the light map, a first cropped image of the foreground image (e.g., the user’s face), and a second cropped image of the foreground image (e.g., the user’s hair), where each cropped image may be associated with a different saturation level to correct for relighting effects. At 1914, the GPU 1904 may generate a consolidated image based on the set of foreground images and a set of background images and may output an indicator 1916 of the generated image to the CPU 1902 (e.g., for transmission to a distal device), or a DPU for output to a display.
[0104] FIG. 20 is a flowchart 2000 of an example method of graphics processing in accordance with one or more techniques of this disclosure. The method may be performed by an apparatus, such as an apparatus for graphics processing, a GPU, a CPU, the device 104, a wireless communication device, and the like, as used in connection with the aspects of FIGs. 1-6 and 16-18. The method may be associated with various advantages, such as producing high-quality relighting results with limited computational resources. In an example, the method (including the various aspects detailed below) may be performed by the light adder 198.
[0105] At 2002, the apparatus may compute, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image. In an example, 2002 may be performed by the light adder 198.
[0106] At 2004, the apparatus may compute a light map based on the normal map and a high dynamic range (HDR) map. In an example, 2004 may be performed by the light adder 198.
[0107] At 2006, the apparatus may generate an image based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image, where the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level. In an example, 2006 may be performed by the light adder 198.
[0108] At 2008, the apparatus may output an indication of the generated image. In an example, 2008 may be performed by the light adder 198.
[0109] In one aspect, the apparatus may obtain an input image, a set of segmentation maps, and a bounding box.
[0110] In one aspect, the apparatus may crop the input image based on the bounding box.[OHl] In one aspect, the apparatus may obtain the first cropped image based on the cropped input image.
[0112] In one aspect, the apparatus may adjust a saturation level of the first cropped image.
[0113] In one aspect, the apparatus may obtain the second cropped image based on the adjustment to the saturation level of the first cropped image.
[0114] In one aspect, the apparatus may generate the masked gamma-corrected foreground image based on the input image and the set of segmentation maps.
[0115] In one aspect, the apparatus may obtain an indication of an HDR rotation.
[0116] In one aspect, the apparatus may generate an HDR rotation matrix based on the indication of the HDR rotation, where computing the light map may include computing the light map further based on the HDR rotation matrix.
[0117] In one aspect, the input image may include a first facial image of a user, where the set of segmentation maps may include a foreground segmentation map and a hair segmentation map, where the bounding box may include a facial bounding box, and where the image may include a second facial image of the user.
[0118] In one aspect, the apparatus may obtain an HDR background.
[0119] In one aspect, the apparatus may combine (e.g., composite) the HDR background and the image based on the set of segmentation maps.
[0120] In one aspect, obtaining the HDR background may include obtaining a pre-rendered HDR background, and combining the HDR background and the image may include combining the pre-rendered HDR background and the image.
[0121] In one aspect, obtaining the HDR background may include rendering the HDR background, and where combining the HDR background and the image may include combining the rendered HDR background and the image.
[0122] In one aspect, combining the HDR background and the image may include combining the HDR background and the image based on whole HDR relighting.
[0123] In one aspect, the apparatus may generate a processed segmentation map based on the set of segmentation maps, where combining the HDR background and the image based on the set of segmentation maps may include combining the HDR background and the image based on the processed segmentation map.
[0124] In one aspect, the HDR map may include a pre-integrated diffuse HDR map, where computing the light map may include computing the light map further based on thepre-integrated diffuse HDR map, and where the pre-integrated diffuse HDR map may correspond to a single added light source.
[0125] In one aspect, the apparatus may obtain a first scaling factor corresponding to an input image associated with the foreground image and a second scaling factor corresponding to light that is to be added to the image, where generating the image may include generating the image further based on the first scaling factor and the second scaling factor.
[0126] In one aspect, the ML model may include a single neural network.
[0127] In one aspect, outputting the indication of the generated image may include: storing the indication of the generated image in at least one of a memory, a buffer or a cache; or transmitting the indication of the generated image.
[0128] As used herein, the term “normal map” may refer to a red green blue (RGB) image where RGB components correspond to X, Y, and Z coordinates, respectively, of a surface normal. As used herein, the term “a masked image” (e.g., a masked gammacorrected image) may refer to an image where some pixels of the image are set to zero by setting the background or hiding some image parts with some text, image, cropping, and editing. As used herein, the term “light map” may refer to a texture map that simulates complex lighting. As used herein, the term “HDR rotation” may refer to a roll, a pitch, and a yaw used for HDR lighting purposes. As used herein, the term “whole HDR relighting” may refer to HDR relighting that is not a single light source relighting. As used herein, a “pre-integrated diffuse HDR map” may refer to a diffuse HDR map that has been integrated prior to run-time in order to determine radiance associated information. In contrast, a diffuse HDR map that is not pre-integrated may be integrated at run-time. As used herein, a diffuse HDR map may refer to an HDR associated texture map that defines a color and a pattern of an object.
[0129] In configurations, a method or an apparatus for graphics processing is provided. The apparatus may be a GPU, a CPU, or some other processor that may perform graphics processing. In aspects, the apparatus may be the processing unit 120 within the device 104, or may be some other hardware within the device 104 or another device. The apparatus may include means for computing, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image. The apparatus may further include means for computing a light map based on the normal map and a high dynamic range (HDR) map. The apparatus may further include means forgenerating an image based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image, where the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level. The apparatus may further include means for outputting an indication of the generated image. The apparatus may further include means for obtaining an input image, a set of segmentation maps, and a bounding box. The apparatus may further include means for cropping the input image based on the bounding box. The apparatus may further include means for obtaining the first cropped image based on the cropped input image. The apparatus may further include means for adjusting a saturation level of the first cropped image. The apparatus may further include means for obtaining the second cropped image based on the adjustment to the saturation level of the first cropped image. The apparatus may further include means for generating the masked gamma-corrected foreground image based on the input image and the set of segmentation maps. The apparatus may further include means for obtaining an indication of an HDR rotation. The apparatus may further include means for generating an HDR rotation matrix based on the indication of the HDR rotation, where computing the light map includes computing the light map further based on the HDR rotation matrix. The apparatus may further include means for obtaining an HDR background. The apparatus may further include means for combining the HDR background and the image based on the set of segmentation maps. The apparatus may further include means for generating a processed segmentation map based on the set of segmentation maps, where combining the HDR background and the image based on the set of segmentation maps includes combining the HDR background and the image based on the processed segmentation map. The apparatus may further include means for obtaining a first scaling factor corresponding to an input image associated with the foreground image and a second scaling factor corresponding to light that is to be added to the image, where generating the image includes generating the image further based on the first scaling factor and the second scaling factor.
[0130] It is understood that the specific order or hierarchy of blocks / steps in the processes, flowcharts, and / or call flow diagrams disclosed herein is an illustration of example approaches. Based upon design preferences, it is understood that the specific order or hierarchy of the blocks / steps in the processes, flowcharts, and / or call flow diagramsmay be rearranged. Further, some blocks / steps may be combined and / or omitted. Other blocks / steps may also be added. The accompanying method claims present elements of the various blocks / steps in a sample order, and are not meant to be limited to the specific order or hierarchy presented.
[0131] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language of the claims, where reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.
[0132] Unless specifically stated otherwise, the term “some” refers to one or more and the term “or” may be interpreted as “and / or” where context does not dictate otherwise. Combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ include any combination of A, B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combinations may contain one or more member or members of A, B, or C. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. The words “module,” “mechanism,” “element,” “device,” and the like may not be a substitute for the word “means.” As such, no claim element is to be construed as a means plus function unless the element is expressly recited using the phrase “means for.” Unless stated otherwise, the phrase “aprocessor” may refer to “any of one or more processors” (e.g., one processor of one or more processors, a number (greater than one) of processors in the one or more processors, or all of the one or more processors) and the phrase “a memory” may refer to “any of one or more memories” (e.g., one memory of one or more memories, a number (greater than one) of memories in the one or more memories, or all of the one or more memories).
[0133] In one or more examples, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term “processing unit” has been used throughout this disclosure, such processing units may be implemented in hardware, software, firmware, or any combination thereof. If any function, processing unit, technique described herein, or other module is implemented in software, the function, processing unit, technique described herein, or other module may be stored on or transmitted over as one or more instructions or code on a computer-readable medium.
[0134] Computer-readable media may include computer data storage media or communication media including any medium that facilitates transfer of a computer program from one place to another. In this manner, computer-readable media generally may correspond to: (1) tangible computer-readable storage media, which is non-transitory; or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, compact disc-read only memory (CD-ROM), or other optical disk storage, magnetic disk storage, or other magnetic storage devices. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while discs usually reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. A computer program product may include a computer-readable medium.
[0135] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs, e.g., a chip set. Various components, modules or units are described in this disclosureto emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily need realization by different hardware units. Rather, as described above, various units may be combined in any hardware unit or provided by a collection of inter-operative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. Also, the techniques may be fully implemented in one or more circuits or logic elements.
[0136] Aspect 1 is a method of graphics processing, including: computing, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image; computing a light map based on the normal map and a high dynamic range (HDR) map; generating an image based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image, wherein the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level; and outputting an indication of the generated image.
[0137] Aspect 2 may be combined with aspect 1, further including: obtaining an input image, a set of segmentation maps, and a bounding box; cropping the input image based on the bounding box; and obtaining the first cropped image based on the cropped input image.
[0138] Aspect 3 may be combined with aspect 2, further including: adjusting a saturation level of the first cropped image; and obtaining the second cropped image based on the adjustment to the saturation level of the first cropped image.
[0139] Aspect 4 may be combined with any of aspects 2-3, further including: generating the masked gamma-corrected foreground image based on the input image and the set of segmentation maps.
[0140] Aspect 5 may be combined with any of aspects 2-4, further including: obtaining an indication of an HDR rotation; and generating an HDR rotation matrix based on the indication of the HDR rotation, wherein computing the light map includes computing the light map further based on the HDR rotation matrix.
[0141] Aspect 6 may be combined with any of aspects 2-5, wherein the input image includes a first facial image of a user, wherein the set of segmentation maps includes a foreground segmentation map and a hair segmentation map, wherein the bounding box includes a facial bounding box, and wherein the image includes a second facial image of the user.
[0142] Aspect 7 may be combined with any of aspects 2-6, further including: obtaining an HDR background; and combining the HDR background and the image based on the set of segmentation maps.
[0143] Aspect 8 may be combined with aspect 7, wherein obtaining the HDR background includes obtaining a pre-rendered HDR background, and wherein combining the HDR background and the image includes combining the pre-rendered HDR background and the image.
[0144] Aspect 9 may be combined with aspect 7, wherein obtaining the HDR background includes rendering the HDR background, and wherein combining the HDR background and the image includes combining the rendered HDR background and the image.
[0145] Aspect 10 may be combined with any of aspects 7-9, wherein combining the HDR background and the image includes combining the HDR background and the image based on whole HDR relighting.
[0146] Aspect 11 may be combined with any of aspects 7-10, further including: generating a processed segmentation map based on the set of segmentation maps, wherein combining the HDR background and the image based on the set of segmentation maps includes combining the HDR background and the image based on the processed segmentation map.
[0147] Aspect 12 may be combined with any of aspects 1-9 and 11, wherein the HDR map includes a pre-integrated diffuse HDR map, wherein computing the light map includes computing the light map further based on the pre-integrated diffuse HDR map, and wherein the pre-integrated diffuse HDR map corresponds to a single added light source.
[0148] Aspect 13 may be combined with any of aspects 1-12, further including: obtaining a first scaling factor corresponding to an input image associated with the foreground image and a second scaling factor corresponding to light that is to be added to theimage, wherein generating the image includes generating the image further based on the first scaling factor and the second scaling factor.
[0149] Aspect 14 may be combined with any of aspects 1-13, wherein the ML model includes a single neural network.
[0150] Aspect 15 may be combined with any of aspects 1-14, wherein outputting the indication of the generated image includes: storing the indication of the generated image in at least one of a memory, a buffer or a cache; or transmitting the indication of the generated image.
[0151] Aspect 16 is an apparatus for graphics processing comprising a memory and a processor coupled to the memory and, based on information stored in the memory, the processor is configured to implement a method as in any of aspects 1-15.
[0152] Aspect 17 is the apparatus of aspect 16, wherein the apparatus is a wireless communication device comprising at least one of a transceiver or an antenna coupled to the processor, wherein to obtain the input image, the processor is configured to obtain the input image via at least one of the transceiver or the antenna.
[0153] Aspect 18 is an apparatus for graphics processing including means for implementing a method as in any of aspects 1-15.
[0154] Aspect 19 is a computer-readable medium (e.g., a non-transitory computer-readable medium) storing computer executable code, the computer executable code, when executed by a processor, causes the processor to implement a method as in any of aspects 1-15.
[0155] Various aspects have been described herein. These and other aspects are within the scope of the following claims.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. An apparatus for graphics processing, comprising: a memory; and a processor coupled to the memory and, based on information stored in the memory, the processor is configured to: compute, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image; compute a light map based on the normal map and a high dynamic range (HDR) map; generate an image based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image, wherein the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level; and output an indication of the generated image.
2. The apparatus of claim 1, wherein the processor is further configured to: obtain an input image, a set of segmentation maps, and a bounding box; crop the input image based on the bounding box; and obtain the first cropped image based on the cropped input image.
3. The apparatus of claim 2, wherein the apparatus comprises a wireless communication device comprising at least one of a transceiver or an antenna coupled to the processor, wherein, to obtain the input image, the processor is configured to: obtain the input image via at least one of the transceiver or the antenna.
4. The apparatus of claim 2, wherein the processor is further configured to: adjust a saturation level of the first cropped image; and obtain the second cropped image based on the adjustment to the saturation level of the first cropped image.
5. The apparatus of claim 2, wherein the processor is further configured to: generate the masked gamma-corrected foreground image based on the input image and the set of segmentation maps.
6. The apparatus of claim 2, wherein the processor is further configured to: obtain an indication of an HDR rotation; and generate an HDR rotation matrix based on the indication of the HDR rotation, wherein, to compute the light map, the processor is configured to: compute the light map further based on the HDR rotation matrix.
7. The apparatus of claim 2, wherein the input image comprises a first facial image of a user, wherein the set of segmentation maps comprises a foreground segmentation map and a hair segmentation map, wherein the bounding box comprises a facial bounding box, and wherein the image comprises a second facial image of the user.
8. The apparatus of claim 2, wherein the processor is further configured to: obtain an HDR background; and combine the HDR background and the image based on the set of segmentation maps.
9. The apparatus of claim 8, wherein, to obtain the HDR background, the processor is configured to: obtain a pre-rendered HDR background, and wherein, to combine the HDR background and the image, the processor is configured to: combine the pre-rendered HDR background and the image.
10. The apparatus of claim 8, wherein, to obtain the HDR background, the processor is configured to: render the HDR background, and wherein, to combine the HDR background and the image, the processor is configured to: combine the rendered HDR background and the image.
11. The apparatus of claim 8, wherein, to combine the HDR background and the image, the processor is configured to: combine the HDR background and the image based on whole HDR relighting.
12. The apparatus of claim 8, wherein the processor is further configured to: generate a processed segmentation map based on the set of segmentation maps, wherein, to combine the HDR background and the image based on the set of segmentation maps, the processor is configured to: combine the HDR background and the image based on the processed segmentation map.
13. The apparatus of claim 1, wherein the HDR map comprises a pre-integrated diffuse HDR map, wherein, to compute the light map, the processor is configured to: compute the light map further based on the pre-integrated diffuse HDR map, and wherein the pre-integrated diffuse HDR map corresponds to a single added light source.
14. The apparatus of claim 1, wherein the processor is further configured to: obtain a first scaling factor corresponding to an input image associated with the foreground image and a second scaling factor corresponding to light that is to be added to the image, wherein, to generate the image, the processor is configured to: generate the image further based on the first scaling factor and the second scaling factor.
15. The apparatus of claim 1, wherein the ML model comprises a single neural network.
16. The apparatus of claim 1 , wherein, to output the indication of the generated image, the processor is configured to: store the indication of the generated image in at least one of the memory, a buffer or a cache; or transmit the indication of the generated image.
17. A method of graphics processing, comprising: computing, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image; computing a light map based on the normal map and a high dynamic range (HDR) map; generating an image based on the light map, a first cropped image associated with a foreground image, and a second cropped image associated with the foreground image, wherein the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level; and outputting an indication of the generated image.
18. The method of claim 17, further comprising: obtaining an input image, a set of segmentation maps, and a bounding box; cropping the input image based on the bounding box; and obtaining the first cropped image based on the cropped input image.
19. The method of claim 18, further comprising: adjusting a saturation level of the first cropped image; and obtaining the second cropped image based on the adjustment to the saturation level of the first cropped image.
20. The method of claim 18, further comprising: generating the masked gamma-corrected foreground image based on the input image and the set of segmentation maps.
21. The method of claim 18, further comprising: obtaining an indication of an HDR rotation; and generating an HDR rotation matrix based on the indication of the HDR rotation, wherein computing the light map comprises computing the light map further based on the HDR rotation matrix.
22. The method of claim 18, wherein the input image comprises a first facial image of a user, wherein the set of segmentation maps comprises a foreground segmentation map and a hair segmentation map, wherein the bounding box comprises a facial bounding box, and wherein the image comprises a second facial image of the user.
23. The method of claim 18, further comprising: obtaining an HDR background; and combining the HDR background and the image based on the set of segmentation maps.
24. The method of claim 23, wherein obtaining the HDR background comprises obtaining a pre-rendered HDR background, and wherein combining the HDR background and the image comprises combining the pre-rendered HDR background and the image.
25. The method of claim 23, wherein obtaining the HDR background comprises rendering the HDR background, and wherein combining the HDR background and the image comprises combining the rendered HDR background and the image.
26. The method of claim 23, wherein combining the HDR background and the image comprises combining the HDR background and the image based on whole HDR relighting.
27. The method of claim 23, further comprising: generating a processed segmentation map based on the set of segmentation maps, wherein combining the HDR background and the image based on the set of segmentation maps comprises combining the HDR background and the image based on the processed segmentation map.
28. The method of claim 17, wherein the HDR map comprises a pre-integrated diffuse HDR map, wherein computing the light map comprises computing the light map further based on the pre-integrated diffuse HDR map, and wherein the pre-integrated diffuse HDR map corresponds to a single added light source.
29. The method of claim 17, further comprising: obtaining a first scaling factor corresponding to an input image associated with the foreground image and a second scaling factor corresponding to light that is to be added to the image, wherein generating the image comprises generating the image further based on the first scaling factor and the second scaling factor.
30. A computer-readable medium storing computer executable code, the computer executable code, when executed by a processor, causes the processor to: compute, via a machine learning (ML) model, a normal map based on a masked gamma-corrected foreground image; compute a light map based on the normal map and a high dynamic range (HDR) map; generate an image based on the light map, a first cropped image associated with the foreground image, and a second cropped image associated with a foreground image, wherein the first cropped image is associated with a first saturation level and the second cropped image is associated with a second saturation level that is less than the first saturation level; and output an indication of the generated image.