Systems and methods for spatial white-balance correction for digital cameras

A machine learning model predicts a camera-independent CCT map to correct spatially varying lighting conditions, addressing the challenge of undesirable color casts in image capture devices by decoupling from camera-specific characteristics and ensuring consistent color balance.

WO2026039027A1PCT designated stage Publication Date: 2026-02-19GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/041884
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing image capture devices struggle with white-balance correction in scenes with spatially varying lighting conditions, leading to undesirable color casts due to reliance on global illuminant assumptions and camera-dependent characteristics, which are challenging to generalize across different cameras.

Method used

A machine learning model is trained to predict a camera-independent calibrated color representation format, using a CCT map and delta map to correct spatially varying lighting, decoupling from camera-specific characteristics and leveraging existing calibration efforts.

Benefits of technology

The solution provides effective spatial white-balance correction that is independent of camera characteristics, reducing sensitivity to factors like sensor sensitivity and lens transmission, and achieves consistent color balance across different imaging devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024041884_19022026_PF_FP_ABST
    Figure US2024041884_19022026_PF_FP_ABST
Patent Text Reader

Abstract

An example method includes receiving, from an image sensor of a camera device, an input image of a scene. The method also includes determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene. The method additionally includes providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.
Need to check novelty before this filing date? Find Prior Art

Description

Atty. Docket: 24-0609 -WOSYSTEMS AND METHODS FOR SPATIAL WHITE-BALANCE CORRECTION FOR DIGITAL CAMERASBACKGROUND

[0001] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices, such as still and / or video cameras. The image capture devices can capture images, such as images that include people, animals, landscapes, and / or objects. Such objects may appear at different depths in the image.SUMMARY

[0002] This application generally relates to a learning-based spatial white-balance correction method for digital images. For example, white balancing may be applied to scenes with spatially varying lighting colors, free from any dependency on segmentation requirements. In some aspects, a pixel-wise illuminant map may be generated that is independent of captured scene semantics or object structures. The illuminant map may be used to spatially whitebalance the colors in the captured input image under mixed and / or varying lighting conditions.

[0003] In one aspect, a computer-implemented method is provided. The method includes receiving, from an image sensor of a camera device, an input image of a scene. The method also includes determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene. The method additionally includes providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

[0004] In another aspect, a system is provided. The system may include one or more processors. The system may also include data storage, where the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the system to carry out operations. The operations may include receiving, from an image sensor of a camera device, an input image of a scene. The operations may also include determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene. The operations may additionally include providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

[0005] In another aspect, a computing device is provided. The device includes a primary camera and a secondary camera that share a common field of view. The device also includesAtty. Docket: 24-0609 WO one or more processors and data storage that has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out operations. The operations may include receiving, from an image sensor of a camera device, an input image of a scene. The operations may also include determining a cameraindependent calibrated color representation format for the input image based on a target global chromaticity for the scene. The operations may additionally include providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

[0006] In another aspect, an article of manufacture is provided. The article of manufacture may include a non-transitory computer-readable medium having stored thereon program instructions that, upon execution by one or more processors of a computing device, cause the computing device to carry out operations. The operations may include receiving, from an image sensor of a camera device, an input image of a scene. The operations may also include determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene. The operations may additionally include providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

[0007] In another aspect, a program is provided. The program upon execution by one or more processors of a computing device, causes the computing device to carry out operations. The operations may include receiving, from an image sensor of a camera device, an input image of a scene. The operations may also include determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene. The operations may additionally include providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

[0008] In another aspect, a computer-implemented method is provided. The method includes receiving training data comprising a plurality of pairs, each pair comprising of an image of a scene and an associated camera-independent calibrated color representation format based on a target global chromaticity for the scene. The method also includes training, based on the training data, a machine learning (ML) model to predict a particular camera-independent calibrated color representation format for a given input image. The method additionally includes providing the trained ML model.

[0009] In another aspect, a system is provided. The system may include one or more processors. The system may also include data storage, where the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause theAtty. Docket: 24-0609 -WO system to carry out operations. The operations may include receiving training data comprising a plurality of pairs, each pair comprising of an image of a scene and an associated cameraindependent calibrated color representation format based on a target global chromaticity for the scene. The operations may also include training, based on the training data, a machine learning (ML) model to predict a particular camera-independent calibrated color representation format for a given input image. The operations may additionally include providing the trained ML model.

[0010] In another aspect, a computing device is provided. The device includes a primary camera and a secondary camera that share a common field of view. The device also includes one or more processors and data storage that has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out operations. The operations may include receiving training data comprising a plurality of pairs, each pair comprising of an image of a scene and an associated camera-independent calibrated color representation format based on a target global chromaticity for the scene. The operations may also include training, based on the training data, a machine learning (ML) model to predict a particular camera-independent calibrated color representation format for a given input image. The operations may additionally include providing the trained ML model.

[0011] In another aspect, an article of manufacture is provided. The article of manufacture may include a non-transitory computer-readable medium having stored thereon program instructions that, upon execution by one or more processors of a computing device, cause the computing device to carry out operations. The operations may include receiving training data comprising a plurality of pairs, each pair comprising of an image of a scene and an associated camera-independent calibrated color representation format based on a target global chromaticity for the scene. The operations may also include training, based on the training data, a machine learning (ML) model to predict a particular camera-independent calibrated color representation format for a given input image. The operations may additionally include providing the trained ML model.

[0012] In another aspect, a program is provided. The program upon execution by one or more processors of a computing device, causes the computing device to carry out operations. The operations may include receiving training data comprising a plurality of pairs, each pair comprising of an image of a scene and an associated camera-independent calibrated color representation format based on a target global chromaticity for the scene. The operations may also include training, based on the training data, a machine learning (ML) model to predict aAtty. Docket: 24-0609 -WO particular camera-independent calibrated color representation format for a given input image. The operations may additionally include providing the trained ML model.

[0013] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description and the accompanying drawings.BRIEF DESCRIPTION OF THE FIGURES

[0014] Figure 1 is an illustration of front, right-side, and rear views of a digital camera device, in accordance with example embodiments.

[0015] Figure 2 is an example overview of a white-balance correction system, in accordance with example embodiments.

[0016] Figure 3 is an example illustration of an adjust operation in a white-balance correction system, in accordance with example embodiments.

[0017] Figure 4 is an example illustration of an output image of a white-balance correction system, in accordance with example embodiments.

[0018] Figure 5 is an example illustration of a residual convolutional network for a whitebalance correction system, in accordance with example embodiments.

[0019] Figure 6 is an example illustration of a U-net architecture for a white-balance correction system, in accordance with example embodiments.

[0020] Figure 7 is an example illustration of processing additional input scalars in a whitebalance correction system, in accordance with example embodiments.

[0021] Figure 8 is an example overview of a photofinishing pipeline, in accordance with example embodiments.

[0022] Figure 9 is an example illustration of a user interface for white-balance correction, in accordance with example embodiments.

[0023] Figure 10 is an example illustration of utilizing a user interface for white-balance correction, in accordance with example embodiments.

[0024] Figure 11 illustrates an example of a ground truth image, in accordance with example embodiments.

[0025] Figure 12 illustrates another example of a ground truth image, in accordance with example embodiments.

[0026] Figure 13 illustrates another example of a ground truth image, in accordance with example embodiments.Atty. Docket: 24-0609 -WO

[0027] Figure 14 illustrates an example editing process for the white-balance gains, in accordance with example embodiments.

[0028] Figure 15 illustrates examples of spatially white-balanced “ground-truth” image, in accordance with example embodiments.

[0029] Figure 16 illustrates diversity in scene conditions and the variability in brightness values across different scenes, in accordance with example embodiments.

[0030] Figure 17 illustrates spatial region class statistics in training, validation, and testing sets, in accordance with example embodiments.

[0031] Figure 18 illustrates lighting condition statistics in training, validation, and testing sets, in accordance with example embodiments.

[0032] Figure 19 illustrates light CCT / tint statistics in training, validation, and testing sets, in accordance with example embodiments.

[0033] Figure 20 illustrates CCT-tint distribution in training, validation, and testing sets, in accordance with example embodiments.

[0034] Figure 21 illustrates example images with white-balance correction, in accordance with example embodiments.

[0035] Figure 22 is a diagram illustrating training and inference phases of a machine learning model, in accordance with example embodiments.

[0036] Figure 23 depicts a distributed computing architecture, in accordance with example embodiments.

[0037] Figure 24 is a block diagram of an example computing device, in accordance with example embodiments.

[0038] Figure 25 is a flowchart of a method, in accordance with example embodiments.

[0039] Figure 26 is another flowchart of a method, in accordance with example embodiments.DETAILED DESCRIPTION

[0040] Example methods, devices, and systems are described herein. It should be understood that the words “example” and “exemplary” are used herein to mean “serving as an example, instance, or illustration.” Any embodiment or feature described herein as being an “example” or “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or features. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein.

[0041] Thus, the example embodiments described herein are not meant to be limiting. Aspects of the present disclosure, as generally described herein, and illustrated in the figures, can beAtty. Docket: 24-0609 WO arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are contemplated herein.

[0042] Further, unless context suggests otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should be generally viewed as component aspects of one or more overall embodiments, with the understanding that not all illustrated features are necessary for each embodiment.Overview

[0043] Automatic white balancing (AWB) is generally applied to a raw image to produce a final image in standard RGB color space after removing undesirable color casts. Traditionally, camera white balancing assumes a single global light color in the scene for simplicity. However, a majority of captured images are illuminated by mixed lighting conditions. Consequently, certain areas in the final image may exhibit incorrect white-balance correction, often resulting in undesirable bluish or reddish color casts that were not perceived by the human eye in the real scene. While the problem can be mitigated by segmenting the image into a few segments, such as foreground object and background segments, and applying global whitebalance correction for each segment separately, such a segmentation-based approach is constrained by the accuracy of the segmentation method and limited to the target number (and category, e.g., person) of segments.

[0044] Traditional low-level computer vision techniques and heuristic techniques may not perform well in different scenes. Camera characteristics, such as sensor sensitivity, infrared (IR) cut-off filter, and lens transmission, may vary from one camera to another. Due to the impact on raw images from different cameras (e.g., varying RGB values for the same object captured by different cameras), it is challenging to generate a collection of labeled training data with ground-truth labels to train a deep learning model for AWB as each camera may require a separate training dataset.

[0045] A machine learning (ML) model is described that can be implemented onboard an ISP of the camera. The ML model can be trained independent of camera characteristics. The ML model may be trained to receive an input raw image of a scene and predict a correlated color temperature (CCT) for each input color in the scene. For example, the input raw image may be a processed version of the raw RGB statistics (stats), typically represented as a downsized version of the captured raw image (usually 64 X 48 pixels). In some embodiments, the RGB stats associated with the raw image space of a camera may be mapped to a camera-independent model using a RGB stats processing block. Some embodiments involve receiving additional metadata that describes the captured scene. The ML model receives as input the CCT of theAtty. Docket: 24-0609 WO estimated global illuminant chromaticity. In some embodiments, the estimated global illuminant chromaticity may be provided by a global AWB model. The CCT is an intermediate standardized calibrated color representation format and is indicative of a target global illuminant color to make the ML model’ s prediction aligned with global illuminant corrections.

[0046] Traditionally, the conversion to CCT is typically applied to the chromaticity of raw colors, which can be represented in different forms, such as (R / G, B / G) or rg-chromaticity, where R, G, and B refer to the red, green, and blue color value, respectively. To obtain the corresponding CCT for chromaticity values, a calibration object (e.g., a color chart) is typically used by capturing it under a set of known lighting conditions with known CCT values. By measuring the chromaticity of gray patches, it is possible to build correspondences between the camera’s chromaticity and CCT values. The mapping from raw chromaticity to / from the CCT can then be applied using a lookup table (LuT).

[0047] The RGB stats processing block aims to reduce reliance on camera characteristics, reducing the AWB’s sensitivity to factors like sensor sensitivity and other device-dependent characteristics of the camera response function. Such device independence may be achieved in several ways. For example, RGB stats values from the camera raw space (camera dependent) may be converted to a linear standard RGB (LsRGB) space (camera independent). The LsRGB stats can be then post-processed by normalizing the R, G, B of the LsRGB version of the stats. This normalization may be determined by dividing the R, G, B by (R+G+B). This results in a reduction of the intensity factor of the R, G, B colors (even in the standard LsRGB ), where the intensity of colors of the same captured objects may vary from camera to camera. As another example, RGB stats values may be represented by correlated color temperature (CCT) values corresponding to the chromaticity of the raw RGB colors. To achieve this, the R / G and B / G values (or any other chromaticity representation of raw RGB colors) of the RGB stats may be converted into the CCT value using a calibration lookup table (LuT). The calibration LuT is computed as a component of the white-balance module for tuning. Consequently, the AWB pipeline described herein leverages the pre-existing calibration effort to minimize reliance on the camera-dependent raw values.

[0048] Another way to reduce camera dependence is to represent raw RGB stats as a weighted average (luma stats) of the red, green, blue color channel referring to the scene “luminance.” While there may be some camera dependency in the luma stats, the luma stats can guide the ML model in obtaining insights into spatially linear brightness levels.

[0049] The ML model may be a multimodal network that accepts input in different forms (e.g. , image-like, metadata scalars) to predict a W x H CCT map, where W and H represent the widthAtty. Docket: 24-0609 -WO and height of the image-like reshaped stats. The output of the ML model is in the CCT calibrated space, which is camera-independent. In some aspects, the predicted CCT map may be aligned with the global white balance correction, which is designed and controlled by the existing camera white-balance module. Although the output CCT may be aligned with the global white balance correction, the output CCT can also include potentially spatially varying CCT values that can factor in regions lit with different illuminants than the global illumination source.

[0050] Generally, the CCT map alone may not be sufficient to reconstruct the RGB values of the light for downstream spatial white-balancing of each pixel. Accordingly, a delta produced by the global illuminant estimation method may be used to reconstruct a CCT-delta map (W x H x 2). Delta, being camera-dependent, can vary from one camera raw space to another. Unlike the use of the CCT-delta map (which is camera dependent) in existing approaches, the techniques described herein decouple the delta value from the ML model’s prediction to maintain camera independence. The CCT-delta map is then converted to the corresponding R / G, B / G chroma map (or any other chromaticity values used in the original camera ISP design) in the camera raw space. This conversion can be implemented using a lookup table (LuT) constructed offline during the camera development stage. The chroma map of spatial illuminant colors is then used to compute the white-balance gain map, where each spatial value in the map contains the white-balance correction vector.

[0051] The ML model is designed to perform an image-to-image task. This may be achieved in two ways: an image stream and a scalar stream. The image stream comprises a 3 x 3 convolutional (Conv) layer followed by a series of residual blocks. Each residual block consists of a 1 x 1 conv layer, a 5 x 5 depth conv (DConv), and another 1 X 1 conv layer. The output from each residual block is added to the input (padding may be applied to maintain dimensionality). The scalar stream includes fully connected layers (FC), each followed by a nonlinear function such as ReLU except for the final layer that is followed by a sigmoid function. The final layer in the image stream is a 1 X 1 conv layer producing the CCT map. In some embodiments, the ML model may be an encoder-decoder network (z.e., U-Net). The scalar stream consists of three FC layers and the image stream includes 5 residual convolutional blocks. This consumes approximately 300 KB of memory and approximately 100 KB of memory for optimized version using TF Lite. Such a lightweight network makes the solution highly practical for real-time processing on a camera ISP.Atty. Docket: 24-0609 WO

[0052] The white-balance gain map has a smaller size than the input raw image. Accordingly, the raw image is downsampled to match the dimensions of the white-balance gain map. Additional pixel-wise color correction may be applied to generate a spatially corrected linear sRGB (LsRGB) version of the downsampled raw image. A photofinishing pipeline may be applied to the spatially corrected high-resolution LsRGB image to generate a final sRGB image with spatially corrected colors.Example Camera Systems

[0053] As image capture devices, such as cameras, become more popular, they may be employed as standalone hardware devices or integrated into various other types of devices. For instance, still and video cameras are now regularly included in wireless computing devices (e.g., mobile devices, such as mobile phones), tablet computers, laptop computers, video game interfaces, home automation devices, and even automobiles and other types of vehicles.

[0001] The physical components of a camera may include one or more apertures through which light enters, one or more recording surfaces for capturing the images represented by the light, and lenses positioned in front of each aperture to focus at least part of the image on the recording surface(s). The apertures may be of a fixed size or may be adjustable. In an analog camera, the recording surface may be a photographic film. In a digital camera, the recording surface may include an electronic image sensor (e.g., a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) sensor) to transfer and / or store captured images in a data storage unit (e.g., memory).

[0002] One or more shutters may be coupled to, or positioned near, the lenses or the recording surfaces. Each shutter may either be in a closed position, in which it blocks light from reaching the recording surface, or an open position, in which light is allowed to reach the recording surface. The position of each shutter may be controlled by a shutter button. For instance, a shutter may be in the closed position by default. When the shutter button is triggered (e.g., pressed), the shutter may change from the closed position to the open position for a period of time, known as the shutter cycle. During the shutter cycle, an image may be captured on the recording surface. At the end of the shutter cycle, the shutter may change back to the closed position.

[0003] Alternatively, the shuttering process may be electronic. For example, before an electronic shutter of a CCD image sensor is “opened,” the sensor may be reset to remove any residual signal in its photodiodes. While the electronic shutter remains open, the photodiodes may accumulate charge. When or after the shutter closes, these charges may be transferred toAtty. Docket: 24-0609 WO longer-term data storage. Combinations of mechanical and electronic shuttering may also be possible.

[0004] Regardless of type, a shutter may be activated and / or controlled by something other than a shutter button. For instance, the shutter may be activated by a softkey, a timer, or some other trigger. Herein, the term “capture” may refer to any mechanical and / or electronic shuttering process that results in one or more images being recorded, regardless of how the shuttering process is triggered or controlled.

[0005] The exposure of a captured image may be determined by a combination of the size of the aperture, the brightness of the light entering the aperture, and the length of the shutter cycle (also referred to as the shutter length, the exposure length, or the exposure time). Additionally, a digital and / or analog gain (e.g., based on an ISO setting) may be applied to the image, thereby influencing the exposure. In some embodiments, the term “exposure length,” “exposure time,” or “exposure time interval” may refer to the shutter length multiplied by the gain for a particular aperture size. Thus, these terms may be used somewhat interchangeably, and should be interpreted as possibly being a shutter length, an exposure time, and / or any other metric that controls the amount of signal response that results from light reaching the recording surface.

[0006] In some implementations or modes of operation, a camera may capture one or more still images each time image capture is triggered. In other implementations or modes of operation, a camera may capture a video image by continuously capturing images at a particular rate (e.g., 24 frames per second) as long as image capture remains triggered (e.g., while the shutter button is held down). Some cameras, when operating in a mode to capture a still image, may open the shutter when the camera device or application is activated, and the shutter may remain in this position until the camera device or application is deactivated. While the shutter is open, the camera device or application may capture and display a representation of a scene on a viewfinder (sometimes referred to as displaying a “preview frame”). When image capture is triggered, one or more distinct payload images of the current scene may be captured.

[0054] Cameras, including digital and analog cameras, may include software to control one or more camera functions and / or settings, such as aperture size, exposure time, gain, and so on. Additionally, some cameras may include software that digitally processes images during or after image capture. While the description above refers to cameras in general, it may be particularly relevant to digital cameras. Digital cameras may be standalone devices (e.g., a DSLR camera) or may be integrated with other devices.

[0055] Either or both of a front-facing camera and a rear-facing camera may include or be associated with an ALS that may continuously or from time to time determine the ambientAtty. Docket: 24-0609 WO brightness of a scene that the camera can capture. In some devices, the ALS can be used to adjust the display brightness of a screen associated with the camera (e.g., a viewfinder). When the determined ambient brightness is high, the brightness level of the screen may be increased to make the screen easier to view. When the determined ambient brightness is low, the brightness level of the screen may be decreased, also to make the screen easier to view as well as to potentially save power. Additionally, the ambient light sensor’s input may be used to determine an exposure time of an associated camera, or to help in this determination.

[0056] Figure 1 is an illustration of front, right-side, and rear views of a digital camera device 100, in accordance with example embodiments. Digital camera device 100 may be, for example, a mobile device (e.g., a mobile phone), a tablet computer, or a wearable computing device. However, other embodiments are possible. Digital camera device 100 may include various elements, such as a body 102, a front-facing camera 104, a multi-element display 106, a shutter button 108, and other buttons 110. Digital camera device 100 could further include one or more rear-facing cameras 112, 114. Front-facing camera 104 may be positioned on a side of body 102 typically facing a user while in operation, or on the same side as multi-element display 106. Rear-facing cameras 112, 114 may be positioned on a side of body 102 opposite front-facing camera 104. Referring to the cameras as front-facing and rear-facing is arbitrary, and digital camera device 100 may include multiple cameras positioned on various sides of body 102.

[0057] Multi-element display 106 could represent a cathode ray tube (CRT) display, a lightemitting diode (LED) display, a liquid crystal display (LCD), a plasma display, or any other type of display known in the art. In some embodiments, multi-element display 106 may display a digital representation of the current image being captured by front-facing camera 104 and / or rear-facing cameras 112, 114, or an image that could be captured or was recently captured by either or both of these cameras. Thus, multi-element display 106 may serve as a viewfinder for either camera. Multi-element display 106 may also support touchscreen and / or presencesensitive functions that may be able to adjust the settings and / or configuration of any aspect of digital camera device 100.

[0058] Multi-element display 106 may include additional features related to a camera application. For example, multiple modes may be available for a user, including, a motion mode, portrait mode, video mode, video bokeh mode, and so forth. The camera application may be in camera mode and provide additional features, such as a reverse icon to activate reverse camera view, a trigger button to capture a previewed image, and a photo stream icon to access a database of captured images. Also for example, a magnification ratio slider may beAtty. Docket: 24-0609 WO displayed and a user can move a virtual object along the magnification ratio slider to select a magnification ratio. In some embodiments, a user may use the multi-element display 106, also referred to herein as the display screen, to adjust the magnification ratio (e.g., by moving two fingers on display screen in an outward motion away from each other), and magnification ratio slider may automatically display the magnification ratio.

[0059] Front-facing camera 104 may include an image sensor and associated optical elements such as lenses. Front-facing camera 104 may offer zoom capabilities or could have a fixed focal length. In other embodiments, interchangeable lenses could be used with front-facing camera 104. Front-facing camera 104 may have a variable mechanical aperture and a mechanical and / or electronic shutter. Front-facing camera 104 also could be configured to capture still images, video images, or both. Further, front-facing camera 104 could represent a monoscopic, stereoscopic, or multiscopic camera. Rear-facing cameras 112, 114 may be similarly or differently arranged. Additionally, front-facing camera 104, rear-facing cameras 112, 114, or both, may be an array of one or more cameras.

[0060] Either or both of front-facing camera 104 and rear-facing cameras 112, 114 may include or be associated with an illumination component that provides a light field to illuminate a target object. For instance, an illumination component could provide flash or constant illumination of the target object (e.g., using one or more LEDs). An illumination component could also be configured to provide a light field that includes one or more of structured light, polarized light, and light with specific spectral content. Other types of light fields known and used to recover three-dimensional (3D) models from an object are possible within the context of the embodiments herein.

[0061] In some digital camera devices 100, either or both of front-facing camera 104 and rearfacing cameras 112, 114 may include or be associated with an ambient light sensor that may continuously or from time to time determine the ambient brightness of a scene that the camera can capture. In some devices, the ambient light sensor can be used to adjust the display brightness of a screen associated with the camera (e.g., a viewfinder). When the determined ambient brightness is high, the brightness level of the screen may be increased to make the screen easier to view. When the determined ambient brightness is low, the brightness level of the screen may be decreased, also to make the screen easier to view as well as to potentially save power. Additionally, the ambient light sensor’s input may be used to determine an exposure time of an associated camera, or to help in this determination.

[0062] Digital camera device 100 could be configured to use multi-element display 106 and either front-facing camera 104 or rear-facing cameras 112, 114 to capture images of a targetAtty. Docket: 24-0609 WO object (e.g., a subject within a scene). The captured images could be a plurality of still images or a video image e.g., a series of still images captured in rapid succession with or without accompanying audio captured by a microphone). The image capture could be triggered by activating shutter button 108, pressing a softkey on multi-element display 106, or by some other mechanism. Depending upon the implementation, the images could be captured automatically at a specific time interval, for example, upon pressing shutter button 108, upon appropriate lighting conditions of the target object, upon moving digital camera device 100 a predetermined distance, or according to a predetermined capture schedule.

[0063] As noted above, the functions of digital camera device 100 (or another type of digital camera) may be integrated into a computing device, such as a wireless computing device, cell phone, tablet computer, laptop computer, and so on. For example, a camera controller may be integrated with the digital camera device 100 to control one or more functions of the digital camera device 100.Example White-Balance Correction Systems

[0064] White balance (WB) correction is a significant part of a camera's image signal processor (ISP). Generally speaking, WB correction refers to a reduction and / or elimination of undesirable color casts that appear in captured images. This may arise from an interplay of scene lighting and the sensitivity of an image sensor. Image white balancing is typically applied to raw images earlier in the camera ISP, and may be described by the following relation:Iwb (c) — I (c) X lc / lG'(Eqn. 1)

[0065] where IwbE Z?wx / ix3is a white-balanced version of the raw image, / 6 / jwxhxs > [lR, IG> IBT ER3isanilluminant vector that is typically produced by an illuminant estimator method running as a part of the white-balance module, T refers to a vector transpose, and c E {R, G, B} is a color channel referring to the red, green, and blue channels, respectively. Several approaches are directed to an estimation of the global illuminant color. This includes statisticalbased approaches and learning-based approaches.

[0066] As indicated by Eqn. 1, existing approaches to white-balancing generally assumes that the entire captured scene is illuminated by a single global light source, as Z is a single color vector, representing the global illuminant color in the scene. While such a global illuminant assumption may be satisfied in many scenarios, such an assumption may not always apply, especially when captured scenes exhibit mixed illuminant colors that are spatially varying across the scene. While such a problem may be mitigated by segmenting the image into a fewAtty. Docket: 24-0609 -WO segments, such as foreground person and background segments, and applying global whitebalance correction for each segment separately, such a segmentation-based approach is constrained by the accuracy of the segmentation method and limited to the target number (and category, e.g., person) of segments.

[0067] Described herein is a framework for white-balancing scenes with spatially varying lighting colors, free from dependencies on segmentation requirements. The framework described herein is designed to estimate a pixel-wise illuminant map, irrespective of the captured scene semantics or object structures, enabling it to spatially white-balance the colors in the captured input image I 6unc[erdifferent mixed illumination conditions. In some embodiments, the approach may be based on a machine learning (ML) algorithm. In some embodiments, the ML approach may be executed onboard a camera ISP. As described herein, dependency on camera characteristics, such as sensor sensitivity, may be reduced. The approach makes it a practical solution for deployment on camera ISPs, providing a plug-and- play solution to the problem.

[0068] As part of a camera's auto white-balance (AWB) module, cameras have an illuminant estimation process that is based on an illuminant estimator algorithm. This illuminant estimator algorithm may depend on various inputs (e.g., light sensor measurements, estimated scene brightness level, and so forth), in addition to the statistics of input image colors, which can be represented in different forms, such as, for example, color histograms or small sampled pixel colors from a full-resolution image.

[0069] An objective of the illuminant estimator algorithm is to predict the illuminant color of the captured scene, assuming it is illuminated by a single global light source. Developing a robust illuminant estimator algorithm, applicable to different cameras with diverse characteristics, typically demands tuning efforts during the camera manufacturing stage. This is because each camera may exhibit a distinct response function, usually represented by factors like sensor sensitivity, infrared (IR) cut-off filter, and lens transmission, among others. These differences among cameras can result in varying RGB values for the same object captured by different cameras. Consequently, a tuning effort is typically undertaken to ensure the illuminant estimator algorithm functions appropriately with any new camera featuring a distinct response function.

[0070] Part of such a tuning effort involves converting raw RGB colors into an intermediate standard calibrated color representation format, such as correlated color temperature (CCT) that is less dependent on (or independent of) the camera response function. Such a conversion from raw RGB to CCT may be typically applied to the chromaticity of raw colors, which canAtty. Docket: 24-0609 -WO be represented in different forms, such as (R / G, B / G) or rg-chromaticity, where R, G, and B refer to the red, green, and blue color value, respectively. To obtain the corresponding CCT for chromaticity values, a calibration object (e.g., a color chart) is typically used by capturing it under a set of known lighting conditions with known CCT values. By measuring the chromaticity of gray patches, it is possible to build correspondences between the camera’s chromaticity and CCT values. The mapping from raw chromaticity to / from the CCT can then be applied using a lookup table (LuT).

[0071] Unlike global image white balancing, the techniques described herein targets white balancing images captured under spatially varying light colors. The techniques leverage the tuning process described above that is already invested in developing the global illuminant estimator, and that is configured to be run on existing digital cameras.

[0072] The techniques described herein are designed to be camera-independent, requiring minimal effort when deployed on new cameras. For example, for an ML based approach, training data may be collected once to train the ML model operating in a standard calibrated space. This allows for seamless reuse on new, unseen cameras with different camera response functions.

[0073] Figure 2 is an example overview of a white-balance correction system 200, in accordance with example embodiments. As illustrated in Figure 2, raw RGB statistics (stats) 205 may be processed by a processing block 220, and a processed version of the raw RGB stats 205 may be input to a neural network 230 (e.g., a lightweight deep neural network (DNN)). RGB stats 205 may be represented as a downsized version of a captured raw image and may comprise 64 X 48 pixels. Additional information 225 that describes the captured scene (e.g., metadata related to CCT, brightness level, etc.) may be provided to neural network 230.

[0074] A CCT of the estimated global illuminant chromaticity (typically represented by [R / G, B / G]TE R2) and a delta map may also be input to processing step 220, and provided to other downstream components. The input CCT provides a significant clue to neural network 230 about the target global illuminant color to make the prediction of neural network 230 aligned with global illuminant correction. The global illuminant chromaticity may be represented by CCT and Delta 215. Unlike a calibrated CCT representation, Delta may not be entirely camera-independent. This means that if two different cameras were to capture the same object with some chromaticity, they could yield different Delta values when converting the object's chromaticity to the standard and calibrated CCT and Delta space; however the CCT would be the same. Optionally, neural network 230 may be configured to incorporate additional metadata related to the captured scene, such as scene brightness value, ‘is-flash’, ‘is-Atty. Docket: 24-0609 -WO artificial -light’, ‘has-face’ flags, and more. Such additional information 225 may be typically generated by other modules or hardware components in the camera ISP to determine the current scene’s conditions (e.g., whether there is a face in the captured scene, if the scene is under artificial lighting, or if the flashlight is turned on, etc.).

[0075] The RGB stats processing block 220 can be configured to diminish reliance on camera characteristics, reducing the solution's sensitivity to factors like sensor sensitivity and other device-dependent characteristics of the camera response function. This includes various options, including but not limited to:

[0076] LsRGB stats'. This may involve converting the RGB stats 205 from the camera raw space (camera dependent) to the linear standard RGB (LsRGB) space (camera independent). This can be described by the following relation:(Eqn. 2)

[0077] where rey(. ) and reb(. ) refer to reshaping the image into n x 3 matrix and reshape the n x 3 back to the original shape of the RGB stats, respectively. The term Istatsrefers to raw statistics, which can be computed by downsampling the captured high-resolution raw image. Generally speaking, there may be no standard method for computing raw statistics. However, the raw statistics generally represent the RGB statistics of a high-resolution raw image in a significantly lower resolution (e.g., 64 X 48 pixels) to facilitate fast computations. The symbol n refers to a total number of pixels in the RGB stats 205. The diag .) constructs a 3 x 3 diagonal global white-balance correction matrix, and the CCM is a 3 x 3 color correction matrix that maps white-balanced raw colors to the corresponding LsRGB colors.

[0078] In some embodiments, the value of I may be determined by an illuminant estimator, which is configured to predict the global illuminant color of the light present in the scene. This is employed by camera ISPs for estimating the global illuminant color in a given scene for image white balancing. The values of CCM may be determined by calibration (typically using a calibration color chart captured under different lighting conditions), which may be performed as a part of the global white-balance module in camera ISP, and may differ based on the estimated light color of the scene, brightness level, and other factors. That is, the term hsrgb represents a version of the raw RGB stats, but in the LsRGB space, subsequent to an application of the global correction based on the global illuminant estimator module in the camera ISP.Atty. Docket: 24-0609 -WO

[0079] Unlike the term Istats, whose colors inherently depend on the camera characteristics, the term Iisrgbrepresents the colors in a standard color space (z.e., linear sRGB) after global white-balancing. In order to mitigate potential differences in intensities between different cameras, the LsRGB stats can be then represented by chromaticity: R, G, B -> R / G, B / G or R, G, B -> R / R + G + B), G / R + G + B), or other chromaticity representations.

[0080] CCT stats'. This may involve representing raw RGB stats 205, Istats, by associating correlated color temperature (CCT) values to the chromaticity of the raw RGB colors. To achieve this, the R / G and B / G values (or other chromaticity representations of raw RGB colors) of the input statistics may be converted into the CCT value using a calibration lookup table (LuT). The calibration LuT may be computed as a component of the white-balance module (as described previously) for tuning and may be specific to each camera, meaning that different cameras may have distinct LuT parameters. Consequently, neural network 230 may be configured to leverage the pre-existing calibration effort to minimize reliance on the cameradependent raw values.

[0081] Luma stats'. This may involve representing raw RGB stats 205, Istats, by a weighted average of the red, green, blue color channel referring to the scene “luminance”. While there may be some camera dependency in the luma stats, such a processed version of the RGB stats 205 may assist neural network 230 in obtaining insights into the linear brightness level associated spatially. This weighted average may be implemented using different weights. In case of cameras that produce raw RGB images (and thus raw RGB stats 205) with major differences in intensity, the use of Luma stats may be omitted.

[0082] All (or some of) the processed versions of the raw RGB stats 205 may be concatenated in a third dimension of the image-like input to neural network 230, creating a tensor of W x H x N, where W and H represent the width and height of the image-like reshaped stats (e.g., 64 X 48 pixels), and N is the total number of channels for the processed versions of the RGB statistics (e.g., LsRGB stats, CCT stats, luma stats, and so forth). For instance, when neural network 230 is configured to receive LsRGB stats (3 channels), CCT stats (1 channel), and luma stats (1 channel), then N = 5.

[0083] In some embodiments, neural network 230 may be a multimodal network that can receive inputs in different forms (e.g., image-like, metadata scalars, etc.) to predict a W X H CCT map 235. This CCT map 235 represents the correlated color temperature (CCT) of the corresponding light color at each spatial position in the input RGB stats 205. Since the outputAtty. Docket: 24-0609 -WO is in the CCT calibrated space, this makes the WB correction camera-independent as it learns to produce output in a “standard” space of light (ie., CCT).

[0084] In some embodiments, the predicted CCT map 235 may be adjusted by an adjust block 240 during the inference phase (not during training). The adjust block 240 may be configured to reduce outlier cases and may be implemented in various ways. One approach is to ensure that the predicted CCT map 235 aligns with the global white balance correction, which may be well-designed and controlled by the existing camera white-balance module. This alignment may be computed as follows:(Eqn. 4)

[0085] where CCT' and CCT'adjustedrefer to the predicted W x H CCT map 235 and the adjusted W X H CCT map after the optional processing. CCT' refers to the average CCT value in the predicted CCT map 235 and epsilon or eps is a small number added for numerical stability. The cctgiobaisymbol refers to the CCT scalar representation of the estimated global illuminant color (as part of CCT and Delta 215). The term |.| computes the absolute value and Tcctglobalisathreshold value that is specific for each of the global CCT values. The intuition here is as follows: the relative CCT map may be scaled by dividing it by its average value and then multiplying by the global CCT value (Eqn. 3). Next, outliers (e.g., that fall below a certain threshold value) may be clipped in the adjusted CCT map, CCT map’ 245, setting them equal to the global CCT value. The threshold value may be dynamically adjusted based on the global CCT value estimated by the white-balance module of the current scene. This approach can provide greater control over determining suitable values for trimming small outliers based on the scene's lighting conditions. For instance, for low CCT values (e.g., below 4000 Kelvin), small changes are more noticeable than in higher CCT values. Accordingly, different threshold values may be used for clipping based on the global CCT value. In practice, the thresholds may be fine-tuned to achieve specific behaviors. In some corner cases, as illustrated in Figure 3, this optional operation can be instrumental in mitigating the impact of small outliers on the quality of the final image. The adjusted CCT map, CCT map’ 245, can also be optionally smoothed using a lightweight smoothing operator, such as convolving a median filter.Atty. Docket: 24-0609 -WO

[0086] Figure 3 is an example illustration of an adjust operation in a white-balance correction system, in accordance with example embodiments. The impact of the optional adjust block 240 on the final resultant image is illustrated. The example illustrated herein may be considered to be a corner case and the network output (without adjustment) generally functions well. However, the adjust block 240 effectively manages this corner case. Image 305 is a result of an application of a global white-balance correction. Image 310 is a result of an application of a spatial white-balance correction without an adjust operation. Image 315 is a result of an application of a spatial white-balance correction with an adjust operation. These images demonstrate a special case when the CCT adjustment is not applied as a post-processing step. Although the results of the global correction in image 305 and spatial correction in image 315 appear similar, this is expected as the dominant lighting condition is the same.

[0087] It is important to notice that in most cases, the network produces an output that is aligned with the global CCT value (e.g., estimated by the existing camera white-balance module - specifically the illuminant estimator), with potentially spatially varying CCT values to handle regions lit with different illuminants than the global illumination source. That means, Eqn. 3 can be omitted in most cases and the input CCT value that represents the global illuminant color can be used in determining the mean value of the predicted CCT map 235 and the optional step aims to reduce any outlier cases.

[0088] Figure 4 is an example illustration of an output image of a white-balance correction system, in accordance with example embodiments. Image 405 is a result of an application of a global white-balance correction (3800K). Image 410 is a result of an application of a spatial white-balance correction with real input CCT scalar value of global illuminant color (3800K). Image 415 is a result of an application of a spatial white-balance correction with altered warmer CCT scalar value (7000K). Image 410 corresponds to a real CCT value inputted by the global white-balance module, and image 415 corresponds to an intentionally altered, higher CCT value (warmer). Generally, an increase in the CCT value can significantly impact the final CCT map produced by neural network 230, thereby altering the entire appearance of the final output image.

[0089] For images 405 and 410, there is a single lighting condition, resulting in identical images. In images 405 and 410, the global lighting is consistent, and the network is provided with the CCT used for the global correction in image 405. In image 415, a different color rendering is observed because the network was provided with a different CCT than what was estimated by the global white-balance module, showing how input CCT influences the network's prediction.Atty. Docket: 24-0609 -WO

[0090] Referring back to Figure 2, the CCT map (with or without adjustment) may not be sufficient to reconstruct the RGB values of the light for later spatial white-balancing of each pixel. Therefore, the delta produced by the global illuminant estimation method may be used to reconstruct the CCT-delta map (VF x H x 2).

[0091] Delta is the second dimension in the 2D CCT-delta space in CCT Delta 215 that represents the chroma (or chromaticity) values (e.g., R / G, B / G) of the illumination by the corresponding CCT and the length of the perpendicular to the nearest CCT. Because delta can vary from one camera raw space to another, being camera-dependent, the delta value may be decoupled from the neural network’s prediction to maintain camera independence.

[0092] Broadcast block 250 may broadcast the scalar delta value produced by the global illuminant estimator. A CCT-delta map, CCT-del Map’ 255, that may be optionally padded to preserve a ratio between the original raw image and the CCT-delta map (CCT map 235 or CCT map’ 245). Padding may be applied optionally at padding block 260 when the received raw RGB stats 205, Istats, has been cropped before reaching the white-balance stage in the camera ISP.

[0093] Existing approaches use a CCT-delta map 265 (padded or otherwise) in white balance correction; however, as described herein, a decoupled version of the CCT-delta values provides camera independence. Generally, neural network 230 processes the CCT values (cameraindependent). Subsequently, the delta value produced by the camera's global white balance module is broadcast by broadcast block 250 to construct the CCT-delta map 265. This enables the WB-correction system 200 to be camera-independent, unlike existing approaches that either directly rely on CCT-delta or incorporate CCT to perform global white balance correction.

[0094] In some embodiments, a corresponding 3 x 3 CCM matrix of each value in the CCT- delta map 265 may be optionally computed. Otherwise, the computed global CCM map may be used to correct the image for simplicity. Note that computing the spatial CCM map 270 (which is an optional step) is an operation similar to those applied onboard any camera ISP for global white balance. The only difference is that, instead of computing a single global 3 x 3 CCM matrix, a 3 X 3 CCM matrix may be computed for each value in CCT-delta map 265. Also, for example, depending on the camera ISP design, computing the CCM values may require projecting the CCT-delta map 265 back to the chromaticity values of illuminant colors. Alternatively, the CCM values may be directly computed from the CCT-delta values.

[0095] In some embodiments, a CCT-del to WB gains operation may be performed so that the CCT-delta map 265 may be converted to a corresponding R / G, B / G map (or any otherAtty. Docket: 24-0609 -WO chromaticity values used in the original camera ISP design, such as rg-chromaticity) in the camera raw space. This conversion can be implemented using a lookup table (LuT) constructed offline during the camera development stage. The chroma map of spatial illuminant colors is then used to compute the white-balance (WB) gain map 285, where each spatial value in the map contains the [R / G, 1, / ? / G]Twhite-balance correction vector. This map, along with the CCM map 275, are then provided for use by the white-balance and color correction module.

[0096] The neural network 230 may be designed to perform an image-to-image task (the input: stacked input images (e.g., linear sRGB, CCT, etc.), and scalar(s) and the output: CCT map 235). There are different designs that can perform the task.

[0097] Figure 5 is an example illustration of a residual convolutional network 500 for a whitebalance correction system, in accordance with example embodiments. Residual convolutional network 500 may include two streams. A two-stream design is commonly used in computer vision tasks. Residual convolutional network 500 may include: 1) an image(s) stream 505, and 2) a scalar(s) stream 510. The image stream 505 may include a 3 x 3 convolutional (Conv) layer 515 followed by a series of residual blocks 520(1), 520(2), ..., 520(N). Each residual block may include a 1 x 1 conv layer, a 5 x 5 depth conv (DConv), and another 1 X 1 conv layer. The output from each residual block may be added to the input. In some embodiments, padding may be applied to maintain dimensionality.

[0098] The scalar stream 510 may include fully connected layers (FC) 530(1), 530(2), ..., 530(M), each followed by a nonlinear function such as ReLU except for the final layer that is followed by a sigmoid function 535. The image stream 505 may process image-like versions of the RGB stats, while the scalar stream 510 may process additional metadata (e.g., global illuminant CCT). The output of the scalar stream 510 may be an attention weighting vector used for depth-wise weighting of the latent feature before the last residual block 520(N) in the image stream 505. Alternatively, the attention weighting vector produced by the scalar stream 510 may be multiplied by the spatial dimensions of the latent feature before the last residual block 520(N) in the image stream 505.

[0099] In some embodiments, the final layer in the image stream 505 may be a 1 x 1 conv layer 525 that outputs the CCT map. Note that, except for the last Conv layer 525 in the image stream 505 and theFC layers 530(1), 530(2), ..., 530(M) in the scalar stream 510, each Conv / FC layer is followed by a ReLU (or other nonlinear functions, such as leaky ReLU, or additional functions).Atty. Docket: 24-0609 -WO

[0100] Figure 6 is an example illustration of a U-net architecture 600 for a whitebalance correction system, in accordance with example embodiments. An encoder-decoder network (z.e., U-Net shape may be used) that includes an encoder 605 and a decoder 610. In U-net architecture 600, scalar metadata may be processed by some fully convolutional layers 625(1), 625(2), ..., 625(M) followed by a sigmoid function 630, and then multiplied (either element-wise or channel -wise) by the latent features in the bottleneck 635 of the encoderdecoder network.

[0101] Figure 7 is an example illustration of processing additional input scalars in a white-balance correction system, in accordance with example embodiments. Figure 7 demonstrates that the scalar input(s) 705 may be alternatively broadcast at block 710 to match the width and height of the input image(s) 715 and 720, and stacked with the input image(s) 715 and 720. This can reduce and / or eliminate a need for a fully connected network stream (z.e., a scalar stream) to process the scalar(s) as illustrated by scalar(s) stream 510 of Figures 5 and / or scalar(s) stream 620 of Figure 6. Instead, the scalar(s) 705 may be processed by the convolutional layers that handle the image-like input(s), including the broadcasted scalar(s).

[0102] In some embodiments, the ML model may be a lightweight network with, for example, 36 channels in the latent representation of the image-like data (z.e., the processed image tensor from the first convolutional layer outputting 36 channels and the subsequent processed image-like data within the network until reaching the last convolutional layer). The scalar stream may include of three fully connected (FC) layers and the image stream may include 5 residual convolutional blocks, utilizing approximately 300 KB of memory and around 100 KB of memory for optimized version using Tensor Flow (TF)-Lite (or other optimization representations). Such a lightweight network enables the WB correction system to be highly practical for real-time processing on a camera ISP.

[0103] Figure 8 is an example overview of a photofinishing pipeline 800, in accordance with example embodiments. After following the flow illustrated in Figure 2, a white-balance gain map 865 (e.g., WB gain map 285 of Figure 2) may be generated in the stats dimensionality (z.e., VF X H), typically representing a low-resolution version of the high-resolution raw image that is to be corrected. Some example steps in the photofinishing pipeline 800 to produce the final high-resolution image with spatial color correction are illustrated below.

[0104] In a typical photofinishing pipeline (e.g., HDR+ pipeline or other camera pipelines), a high-resolution raw image 805 may undergo lens-shading correction (LSC) 810 to obtain a raw’ image 815, followed by an application of a color correction module 820. The color correction module 820 involves applying global white-balance gains and color correctionAtty. Docket: 24-0609 WO matrices (CCM), similar to applying Eqn. 2, but applied to a high-resolution raw image. Such a process results in a linear sRGB version of the raw image with global white-balance correction, denoted as LsRGB (global) 825. Subsequent operations such as gamma correction, local tone mapping, etc., may be performed to obtain a final sRGB image with global whitebalance correction.

[0105] To integrate the WB gain map 865 into the photofinishing pipeline 800, some modifications may be applied. The raw’ image 815 may be sampled (or downsized) to generate raw” 850 that matches the dimensions of the small-sized WB gain map 865 (and the spatial CCM map 860 if computed).

[0106] Subsequently, additional pixel-wise color correction 855 may be applied, generating a spatially corrected linear sRGB version of the sampled raw, referred to as LsRGB (S) 870. This spatial color correction 855 involves applying the spatial WB gain map 865 and either global or spatial CCM values and may be described by the following relations:(Eqn. 6)

[0107] where V andrefer to the sampled raw image, raw’ 815, (after lensshading correction 810) and the spatially white-balanced version raw” 850, x and y refer to the spatial index in the x-axis and y-axis in each image, c is the color channel, M is the W x H x 3 WB gain map 865. After computing the spatially white-balanced image, the linear sRGB image LsRGB(S) 870 denoted LsRGBsmay be computed by multiplying the spatially white-balanced image by the CCM matrix (either the global one or the spatially computed map). In Eqn. 6, <z(x,y, i, c) refers to the element in the i — th row and c — th column in the CCM matrix. If CCM is computed spatially, then a(x,y, i, c) is V / x H X 3 x 3, while if the CCM is a single global matrix, then a(x, y, i, c) is a 3 x 3 matrix (i.e., the same as the CCM in Eqn. 2). In this situation, Eqn. 6 may be expressed as:(Eqn. 7)

[0108] Another sampling (or downsizing) operator 830 may be applied to LsRGB (global) 825 to produce a small-sized linear sRGB image with global correction, denoted asAtty. Docket: 24-0609 WOLsRGB (G) 835. Accordingly, a full-resolution LsRGB image with global correction may be obtained (e.g., a sampled version), and a sampled version of the LsRGB image with spatial color correction applied.

[0109] To transfer the spatial color correction from the downsampled image, LsRGB (S) 870, to the high-resolution image, LsRGB (global) 825, a guided upsampling operator 840 (e.g., the Bilateral Guided Upsampling (BGU) method) may be applied. The upsampling operator 840 may effectively transfer colors from the low-resolution image to the high- resolution image, providing spatially corrected imagery in high resolution within a low computational cost. Note that the BGU is one example of the upsampling operator 840, and other methods that can effectively model color changes in the low-resolution spatial image and transfer them to the high-resolution globally white-balanced image may be used. Alternatively, the white-balance gain map 865, which is in low resolution, may serve as an input to the upsampling operator 840 along with the low-resolution and high-resolution images corrected with global white balance. This process can produce a high-resolution white-balance gain map, which can then be applied to the high-resolution raw image. Subsequently, the color correction matrix (e.g., CCM map 860) may be applied to generate the high-resolution spatially white- balanced image in linear sRGB (LsRGB) space.

[0110] The photofinishing pipeline 800 is applied to the spatially corrected high- resolution LsRGB image to generate the final sRGB image with spatially corrected colors, denoted as LsRGB (Spatial) 875.

[0111] Figure 9 is an example illustration of a user interface 900 for white-balance correction, in accordance with example embodiments. The user interface 900 may include a side panel 910 that displays several user-adjustable parameters that can be applied to an input image displayed in display 905. Side panel 910 is a labeling / editing tool that may be used to generate a training dataset to train a ML model. For example, WB correction features 915 may allow a WB correction to be applied globally or spatially. WB correction features 915 may include a display of user-selectable values for CCT, Tint, R / G, B / G, and so forth. A spatial mask tool 920 may enable a user to automatically generate segmentation masks for various regions within an image, including adjustments for threshold, blur, etc. A visualization tool 925 may allow a user to select whether to display an image with a global EB correction, a spatial WB correction, an RGB mask, etc. including an ability to adjust a transparency, enable rotations, and enable a view of ISP rendering.

[0112] Figure 10 is an example illustration of utilizing a user interface 900 for whitebalance correction, in accordance with example embodiments. Image 1005 is an example of aAtty. Docket: 24-0609 -WO globally corrected image. Image 1010 is an example of a spatially corrected image after labeling. Image 1015 is an example of a segmentation mask used to generate a “ground-truth” image displayed in display portion 1020.

[0113] Training of the ML model (e.g., a DNN) may be performed in a supervised manner using a labeled dataset comprising input raw RGB statistics (stats) or equivalent raw image representations paired, along with additional required input data (global white-balance gains and global CCM to compute the input linear sRGB, global illuminant CCT, and other scalar inputs) with corresponding ground-truth CCT maps. These maps represent the spatially correlated color temperature (CCT) of the light in the scene.

[0114] Generally, spatially white-balanced images may be generated by compositing different versions of an image, each version with a different global white-balance correction. Such an approach may be used to design a labeling tool (as illustrated in Figure 9) to enable spatial annotation of images for white-balance correction. The annotation process can involve several steps. For example, the image may be segmented into different sections based on the color of lights present in each scene (see Figure 11). The labeling tool can be configured to offer various segmentation options, including freehand drawing, object-based auto-selection using an off-the-shelf instant object-based segmentation method (e.g., the Simpleclick method, skin and sky auto-segmentation using models trained to detect these regions, and autocompletion using the GrabCut techniques, as illustrated in Figures 11-13.

[0115] Figure 11 illustrates an example of a ground truth image, in accordance with example embodiments. An object selection feature of the labeling tool may be used. Image 1105 corresponds to before an object selection is performed, and image 1110 displays an image with an auto-selected object (e.g., selected with a single click). For example, object 1115 in image 1105 is auto-selected and displayed as selected object 1120 in image 1110.

[0116] Figure 12 illustrates another example of a ground truth image, in accordance with example embodiments. A skin segmentation feature of the labeling tool may be used. Image 1205 corresponds to before a skin segmentation selection is performed, and image 1210 displays an image with an auto-segmented skin (e.g., selected with a single click). For example, skin portion 1215 in image 1205 is auto-segmented and displayed as skin segmentation 1220 in image 1210.

[0117] Figure 13 illustrates another example of a ground truth image, in accordance with example embodiments. A sky segmentation feature of the labeling tool may be used. Image 1305 corresponds to before a sky segmentation selection is performed, and image 1310 displays an image with an auto-segmented sky (e.g., selected with a single click). For example,Atty. Docket: 24-0609 WO sky portion 1315 in image 1305 is auto-segmented and displayed as sky segmentation 1320 in image 1310.

[0118] In some embodiments, the labeling / editing tool may provide a functionality to optionally apply blur to each segment based on the light's fall-off to create smoothly overlapped regions that accurately match the light fall-off in the scene segment. The blur may be applied using any image blurring technique, such as Gaussian blur, box blur, bilateral filter, and so forth.

[0119] In some embodiments, the labeling tool may provide a functionality to manually adjust the white-balance gains of each region to eliminate any undesirable color casts resulting from incorrect white-balance correction using global white-balance gains (produced by the camera illuminant estimator). The adjustment can be applied by tuning the illuminant color (or chromaticity or the white-balance gains) or by adjusting the corresponding CCT and delta values of the global white-balance settings. This process may be iterated over each region until undesirable color casts are spatially removed (e.g., applying color grading to achieve a subjective target for the final image colors).

[0120] Figure 14 illustrates an example editing process 1400 for the white-balance gains, in accordance with example embodiments. Editing process 1400 may be applied either globally or for one of the selected regions in the segmented image (as illustrated in Figures I lls). The editing process 1400 describes one or more functionalities of the user interface 900 of Figure 9. As illustrated in Figure 14, in some embodiments, for spatial white-balance editing, the same image may be rendered multiple times using a camera pipeline (e.g., HDR+ pipeline, each with a different global white-balance gain determined by the annotator for one of the regions in the segmented image).

[0121] For example, the process may begin at step 1405. At step 1410, a new WB setting may be received. The color space may be checked at step 1415. At step 1420, it may be determined whether the color space is RGB. Upon a determination that the color space is RGB, the process moves to step 1425. At step 1425, the RGB is converted to CCT-delta.

[0122] Upon a determination that the color space is not RGB, the process moves to step 1430. At step 1430, the CCT-delta is converted to RGB.

[0123] The output from step 1425 / 1430 is provided to step 1435. At step 1435, additional data may be retrieved. The additional data would have required a determination of a CCM based on the new WB setting.

[0124] At step 1440, a new CCM is computed. At step 1445, labels may be updated with new WB and CCM. At step 1450, the visualization mode may be checked.Atty. Docket: 24-0609 WO

[0125] At step 1455, it may be determined whether the visualization mode is spatial. Upon a determination that the visualization mode is not spatial (is global), the process moves to step 1460. At step 1460, the camera ISP may be run to render the raw image with a global WB and CCM.

[0126] Upon a determination that the visualization mode is spatial, the process moves to step 1465. At step 1465, a list of WB gains and CCM values may be generated.

[0127] At step 1470, the camera ISP may be run multiple times, each time with a global WB gain and CCM value from the list of WB gains and CCM values generated at step 1465. At step 1475, the final image may be composited by linearly combining the rendered images from step 1470 based on the segmentation mask created by the annotator. This segmentation mask may include blurred regions as determined by the annotator if needed.

[0128] Subsequent to steps 1460, and / or 1475, the process terminates at step 1480.

[0129] A “ground-truth” spatially varying CCT map of the corresponding RGB illuminant map may be obtained, which was generated by the annotator to correct each image. The conversion from RGB illuminant colors to the corresponding CCT value may be performed using a calibrated lookup table (LuT) that accepts R / G, B / G chroma values (or an alternative chromaticity representation of illuminant RGB colors), and outputs the corresponding CCT value. This calibrated LuT is typically computed offline as a part of the calibration / tuning process of each camera.

[0130] The generated “ground-truth” CCT maps and the associated segmentation maps that are used in the loss function may be provided by the editing tool. Furthermore, the editing tool can enable a user to label each scene based on lighting condition along with a label associated with each segmented region based on the semantic information. Such information may be useful in applying batch editing for data augmentation.

[0131] The editing / labeling tool provides functionality to label each segmented region based on semantic information and apply batch editing for data augmentation, which helps expand the training dataset through automatic augmentation.

[0132] Figure 15 illustrates examples of spatially white-balanced “ground-truth” image, in accordance with example embodiments. Image 1505 corresponds to a global correction using a single global illuminant vector. Image 1520 corresponds to a spatially corrected version using the labeling / editing tool described herein. Image 1515 illustrates a segmentation mask, dividing the image into distinct spatial regions, such as, for example, indicated by labels “A,” “B,” and “C ” Each region may be assigned a different illuminant color, which is then utilized for separate white balancing of each region. A "ground-truth" RGBAtty. Docket: 24-0609 -WO illuminant map may be generated, as indicated in image 1520, which may be converted to CCT values in order to generate the “ground-truth” CCT map.

[0133] In some embodiments, the labeling / editing tool may be used to annotate about 2,383 images, which may be divided into 1,791 images for training a ML model, 137 images for validation (e.g., to prevent overfitting and monitor training progress against accuracy on unseen images), and 455 images for testing the ML model. The images in the dataset may be captured by different phone cameras including, for example, Pixel 8 Pro’s main camera (for testing set), and different cameras of Pixel 8, Pixel 7, Pixel 7a, Pixel Fold, and Pixel 7 Pro (for training and validation sets). The testing set may be captured by a camera not previously used (e.g., Pixel 8 Pro with the main camera’s image sensor) during training. This allows an assessment of how well the trained model performs on cameras that have not been used during training.

[0134] As the color labeling process is inherently subjective, a singular color expert annotator may execute the entire color correction process to ensure consistency across the entire dataset. Various statistics of the collected dataset are presented in Figures 16-20.

[0135] Figure 16 illustrates diversity in scene conditions and the variability in brightness values across different scenes, in accordance with example embodiments. Graph 1605 illustrates results for scene classes (e.g., food, indoor, outdoor, under-the-sea). Graph 1610 illustrates results for scene brightness values.

[0136] Figure 17 illustrates spatial region class statistics in training, validation, and testing sets, in accordance with example embodiments. Graph 1705 illustrates spatial region class statistics for hard shadow, sky, snow, skin, light source, monitor, lit by indoor light, lit by flashlight, flare, and others.

[0137] Figure 18 illustrates lighting condition statistics in training, validation, and testing sets, in accordance with example embodiments. Graph 1805 illustrates statistics for lighting conditions such as daylight, sunset / sunrise, night, indoor (single), indoor (multiple), and multiple (indoor / outdoor). Graph 1810 illustrates statistics for mixed lighting conditions such as daylight / indoor (single), sunset / sunrise-indoor (single), night-indoor (single), daylightindoor (multiple), sunset / sunrise-indoor (multiple), and night-indoor (multiple).

[0138] These spatial region categories are labeled during the annotation process. In addition to facilitating analysis, they can serve as valuable resources for augmenting the training data. For example, for the semantic category of ‘sky,’ the raw colors of all images containing ‘sky’ may be augmented by adjusting the sky colors in the raw image, using chromatic adaptation with random scalars (that can be adjusted based on each semanticAtty. Docket: 24-0609 -WO category) in the diagonal values of a chromatic adaptation matrix. Subsequently, the spatial “ground-truth” CCT map may be updated accordingly to introduce new training samples. This augmentation process can be applied in a batch processing manner for each semantic category, allowing several variations of images in the training data with realistic variations based on each semantic category. This batch processing can also be applied using the lighting condition category for data augmentation.

[0139] Figure 19 illustrates light CCT / tint statistics in training, validation, and testing sets, in accordance with example embodiments. Here, "tint" refers to the normalized delta, which is the perpendicular distance to the nearest CCT, of each R / G and B / G chroma value. The normalization may be determined per device to obtain the min / max values of the delta axis, and then a min-max normalization may be applied to compute the tint from each delta value. Graph 1905 illustrates statistics for light CCT values. Graph 1910 illustrates statistics for light tint values.

[0140] Figure 20 illustrates CCT-tint distribution in training, validation, and testing sets, in accordance with example embodiments. Graph 2005 illustrates distribution for light CCT-tint values for training data. Graph 2310 illustrates distribution for light CCT-tint values for testing data. Graph 2015 illustrates distribution for light CCT-tint values for validation data.Model Training

[0141] In some embodiments, a DNN (e.g., neural network 230 of Figure 2) may be trained using an Adam optimizer on randomly extracted patches from high-resolution training raw RGB images for 250 epochs using learning rate = 10A-4, followed by fine-tuning on training raw RGB stats samples for additional 250 epochs. This approach aims to achieve reasonable parameter weights on raw RGB images before tuning the model on actual raw RGB stats with a limited number of samples to improve generalization.

[0142] To further enhance model generalization, random geometric augmentations (e.g., flipping, translation) may be optionally applied to input image-like data (e.g., CCT stats, LsRGB stats, etc.). Note that when geometric augmentation is applied, the same operator is applied to the ground-truth CCT map and its associated data (e.g., the labeling mask that is used in computing the loss function).

[0143] In some embodiments, batch augmentation may be applied to images by adding synthetic light to raw images and adjusting the ground truth accordingly. This synthetic light can be applied to each segmented region using the labeling / editing tool described herein, which allows for batch editing of specific semantic regions with specific CCT-tint shifts.Atty. Docket: 24-0609 -WO

[0144] Such augmentation strategies, along with a two-stage training that involves training on patches extracted from the high-resolution raw images, followed by training on raw stats, enables an effective use of a relatively small dataset. For example, although some existing spatial white-balance datasets include approximately 7.5K images, the training dataset described herein comprises only about 1.8K images. Despite this smaller size, the ML model may be trained effectively.

[0145] In some embodiments, the loss function may include a weighted LI loss and a total variation loss that can be described as follows:L(CCT,CCT, G) = wLl(CCT, CCT, G) + TV(CCT')(Eqn. 8)(Eqn. 10)

[0146] where CCT' and CCT refer to the predicted and ground-truth light CCT map, respectively, G is the associated spatial mask obtained during data labeling, that includes k spatial regions, card . ) computes to cardinality of the iAth region in the mask G.

[0147] The weighted LI, denoted wLl(.) in Eqn. 9, calculates the weighted LI loss to ensure that the ML model assigns substantially equal attention to all regions in the ground-truth CCT map, irrespective of their size. Additionally, the total variation loss term 7V(CC'T'(x,y)) in Eqn. 10 encourages the ML model to generate smooth CCT maps.

[0148] As described previously, the images in the testing set may be captured by a previously unseen (by the ML model) device, which includes scenes that have not previously been provided to the ML model during training. Quantitative comparisons may be made with other alternative approaches designed to perform post-capture color correction (either globally or spatially) to correct improperly white-balanced images. Additionally, the camera's global white-balance correction may be compared with the ground-truth images. Different metrics may be used for evaluation, such as, for example, peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), Spatial CIELAB, CD-Flow, and so forth. The metrics compare the final result images of each method (including the camera output using global white-balance correction) against the ground-truth final images produced by the labeling / editing tool described as part of the annotation phase of the training dataset. To ensureAtty. Docket: 24-0609 -WO fair comparisons, models for approaches that require training may be trained using the training dataset generated herein. Such evaluation can confirm that the approach described herein achieves the best values across all metrics when compared with other methods, including the global camera white-balance correction, while using only 300 KB of parameters. This is an affordable number of parameters for running on-board cameras in real-time.Table 1. For Delta E 2000, Spatial CIELAB, and CD-Flow, lower values indicate better performance, whereas for PSNR and SSIM, higher values indicate better results.

[0149] Figure 21 illustrates example images with white-balance correction, in accordance with example embodiments. Images in column 21C1 correspond to a global correction. Images in column 21C2 correspond to a spatial correction as described herein. Images in column 21C3 illustrate the respective predicted CCT maps. The results were produced by using the following input versions of the RGB stats: luma stats, CCT stats, and LsRGB; and the following scalars: brightness value (BV), global light CCT value (predicted by the camera global illuminant estimator), is flash flag, is artifi cial light flag.Atty. Docket: 24-0609 -WO

[0150] The predicted CCT maps displayed in column 21 C3 are visualized after resizing to aid visualization. Global and spatial corrected images can be produced by the HDR+ pipeline in the sRGB color space using the global camera white-balance correction and the proposed spatial AWB correction framework.

[0151] As can be seen from the results, the spatial WB correction described herein successfully achieves spatial white balance correction under various lighting conditions. For instance, in scenes illuminated by a mix of indoor and outdoor lights, the approach described herein effectively removes the bluish cast resulting from incorrect white balance in regions lit by outdoor light. Even when outdoor light appears in reflected regions, spatial white balance correction can correct this.

[0152] Another scenario involving multiple illuminant scenes is when parts are under shaded regions. Spatial white balance correction can enhance the color of shaded regions by accurately depicting the color. Also, for example, indoor scenes (or night scenes) with different light colors (e.g., LED / fluorescent, incandescent lights, mobile or LCD screens) often suffer from unrealistic color casts when it is globally corrected due to incorrect white balance in regions lit by one of the light sources. Spatial white balance correction can spatially correct such scenes.Training Machine Learning Models for Generating Inferences / Predictions

[0153] Figure 22 shows diagram 2200 illustrating a training phase 2202 and an inference phase 2204 of trained machine learning model(s) 2232, in accordance with example embodiments. Some machine learning techniques involve training one or more machine learning algorithms on an input set of training data to recognize patterns in the training data and provide output inferences and / or predictions about (patterns in the) training data. The resulting trained machine learning algorithm can be termed as a trained machine learning model. For example, Figure 22 shows training phase 2202 where machine learning algorithm(s) 2220 are being trained on training data 2210 to become trained machine learning model(s) 2232. Then, during inference phase 2204, trained machine learning model(s) 2232 can receive input data 2230 and one or more inference / prediction requests 1940 (perhaps as part of input data 2230) and responsively provide as an output one or more inferences and / or prediction(s) 2250.

[0154] As such, trained machine learning model(s) 2232 can include one or more models of machine learning algorithm(s) 2220. Machine learning algorithm(s) 2220 may include, but are not limited to: an artificial neural network (e.g., a herein-described convolutional neural networks, a recurrent neural network, a Bayesian network, a hiddenAtty. Docket: 24-0609 -WOMarkov model, a Markov decision process, a logistic regression function, a support vector machine, a suitable statistical machine learning algorithm, and / or a heuristic machine learning system). Machine learning algorithm(s) 2220 may be supervised or unsupervised, and may implement any suitable combination of online and offline learning.

[0155] In some examples, machine learning algorithm(s) 2220 and / or trained machine learning model(s) 2232 can be accelerated using on-device coprocessors, such as graphic processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), and / or application specific integrated circuits (ASICs). Such on-device coprocessors can be used to speed up machine learning algorithm(s) 2220 and / or trained machine learning model(s) 2232. In some examples, trained machine learning model(s) 2232 can be trained, reside and execute to provide inferences on a particular computing device, and / or otherwise can make inferences for the particular computing device.

[0156] During training phase 2202, machine learning algorithm(s) 2220 can be trained by providing at least training data 2210 as training input using unsupervised, supervised, semisupervised, and / or reinforcement learning techniques. Unsupervised learning involves providing a portion (or all) of training data 2210 to machine learning algorithm(s) 2220 and machine learning algorithm(s) 2220 determining one or more output inferences based on the provided portion (or all) of training data 2210. Supervised learning involves providing a portion of training data 2210 to machine learning algorithm(s) 2220, with machine learning algorithm(s) 2220 determining one or more output inferences based on the provided portion of training data 2210, and the output inference(s) are either accepted or corrected based on correct results associated with training data 2210. In some examples, supervised learning of machine learning algorithm(s) 2220 can be governed by a set of rules and / or a set of labels for the training input, and the set of rules and / or set of labels may be used to correct inferences of machine learning algorithm(s) 2220.

[0157] Semi-supervised learning involves having correct results for part, but not all, of training data 2210. During semi-supervised learning, supervised learning is used for a portion of training data 2210 having correct results, and unsupervised learning is used for a portion of training data 2210 not having correct results. Reinforcement learning involves machine learning algorithm(s) 2220 receiving a reward signal regarding a prior inference, where the reward signal can be a numerical value. During reinforcement learning, machine learning algorithm(s) 2220 can output an inference and receive a reward signal in response, where machine learning algorithm(s) 2220 are configured to try to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a valueAtty. Docket: 24-0609 -WO function that provides a numerical value representing an expected total of the numerical values provided by the reward signal over time. In some examples, machine learning algorithm(s) 2220 and / or trained machine learning model(s) 2232 can be trained using other machine learning techniques, including but not limited to, incremental learning and curriculum learning.

[0158] In some examples, machine learning algorithm(s) 2220 and / or trained machine learning model(s) 2232 can use transfer learning techniques. For example, transfer learning techniques can involve trained machine learning model(s) 2232 being pre-trained on one set of data and additionally trained using training data 2210. More particularly, machine learning algorithm(s) 2220 can be pre-trained on data from one or more computing devices and a resulting trained machine learning model provided to a particular computing device, where the particular computing device is intended to execute the trained machine learning model during inference phase 2204. Then, during training phase 2202, the pre-trained machine learning model can be additionally trained using training data 2210, where training data 2210 can be derived from kernel and non-kernel data of the particular computing device. This further training of the machine learning algorithm(s) 2220 and / or the pre-trained machine learning model using training data 2210 of the particular computing device’s data can be performed using either supervised or unsupervised learning. Once machine learning algorithm(s) 2220 and / or the pre-trained machine learning model has been trained on at least training data 2210, training phase 2202 can be completed. The trained resulting machine learning model can be utilized as at least one of trained machine learning model(s) 2232.

[0159] In particular, once training phase 2202 has been completed, trained machine learning model(s) 2232 can be provided to a computing device, if not already on the computing device. Inference phase 2204 can begin after trained machine learning model(s) 2232 are provided to the particular computing device.

[0160] During inference phase 2204, trained machine learning model(s) 2232 can receive input data 2230 and generate and output one or more corresponding inferences and / or prediction(s) 2250 about input data 2230. As such, input data 2230 can be used as an input to trained machine learning model(s) 2232 for providing corresponding inference(s) and / or prediction(s) 2250 to kernel components and non-kernel components. For example, trained machine learning model(s) 2232 can generate inference(s) and / or prediction(s) 2250 in response to one or more inference / prediction requests 1940. In some examples, trained machine learning model(s) 2232 can be executed by a portion of other software. For example, trained machine learning model(s) 2232 can be executed by an inference or prediction daemon to be readily available to provide inferences and / or predictions upon request. Input data 2230 canAtty. Docket: 24-0609 WO include data from the particular computing device executing trained machine learning model(s) 2232 and / or input data from one or more computing devices other than the particular computing device.

[0161] Input data 2230 can include an input image of a scene. Other types of input data are possible as well. Inference(s) and / or prediction(s) 2250 can include a a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene. Inference(s) and / or prediction(s) 2250 can include other output data produced by trained machine learning model(s) 2232 operating on input data 2230 (and training data 2210). In some examples, trained machine learning model(s) 2232 can use output inference(s) and / or predict! on(s) 2250 as input feedback 1960. Trained machine learning model(s) 2232 can also rely on past inferences as inputs for generating new inferences.

[0162] Convolutional neural networks and / or deep neural networks used herein can be an example of machine learning algorithm(s) 2220. For example, machine learning algorithm(s) 2220 may include ML model 600. After training, the trained version of a convolutional neural network can be an example of trained machine learning model(s) 2232. In this approach, an example of the one or more inference / prediction requests 1940 can be a request to predict a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene, and a corresponding example of inferences and / or prediction(s) 2250 can be the camera-independent calibrated color representation format.Example Data Network

[0163] Figure 23 depicts a distributed computing architecture 2300, in accordance with example embodiments. Distributed computing architecture 2300 includes server devices 2308, 2310 that are configured to communicate, via network 2306, with programmable devices 2304a, 2304b, 2304c, 2304d, 2304e. Network 2306 may correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WWAN, a corporate intranet, the public Internet, or any other type of network configured to provide a communications path between networked computing devices. Network 2306 may also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.

[0164] Although Figure 23 only shows five programmable devices, distributed application architectures may serve tens, hundreds, or thousands of programmable devices. Moreover, programmable devices 2304a, 2304b, 2304c, 2304d, 2304e (or any additional programmable devices) may be any sort of computing device, such as a mobile computing device, desktop computer, wearable computing device, head-mountable device (HMD),Atty. Docket: 24-0609 WO network terminal, a mobile computing device, and so on. In some examples, such as illustrated by programmable devices 2304a, 2304b, 2304c, 2304e, programmable devices can be directly connected to network 2306. In other examples, such as illustrated by programmable device 2304d, programmable devices can be indirectly connected to network 2306 via an associated computing device, such as programmable device 2304c. In this example, programmable device 2304c can act as an associated computing device to pass electronic communications between programmable device 2304d and network 2306. In other examples, such as illustrated by programmable device 2304e, a computing device can be part of and / or inside a vehicle, such as a car, a truck, a bus, a boat or ship, an airplane, etc. In other examples not shown in Figure 23, a programmable device can be both directly and indirectly connected to network 2306.

[0165] Server devices 2308, 2310 can be configured to perform one or more services, as requested by programmable devices 2304a-2304e. For example, server device 2308 and / or 2310 can provide content to programmable devices 2304a-2304e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content can include compressed and / or uncompressed content. The content can be encrypted and / or unencrypted. Other types of content are possible as well.

[0166] As another example, server devices 2308 and / or 2310 can provide programmable devices 2304a-2304e with access to software for database, search, computation, graphical, audio, video, World Wide Web / Internet utilization, and / or other functions. Many other examples of server devices are possible as well.Computing Device Architecture

[0167] Figure 24 is a block diagram of an example computing device 2400, in accordance with example embodiments. In particular, computing device 2400 shown in Figure 24 can be configured to perform at least one function described herein, including methods 2500, and / or 2600.

[0168] Computing device 2400 may include a user interface module 2401, a network communications module 2402, one or more processors 2403, data storage 2404, one or more cameras 2418, one or more sensors 2420, and power system 2422, all of which may be linked together via a system bus, network, or other connection mechanism 2405.

[0169] User interface module 2401 can be operable to send data to and / or receive data from external user input / output devices. For example, user interface module 2401 can be configured to send and / or receive data to and / or from user input devices such as a touch screen, a computer mouse, a keyboard, a keypad, a touch pad, a trackball, a joystick, a voice recognition module, and / or other similar devices. User interface module 2401 can also beAtty. Docket: 24-0609 WO configured to provide output to user display devices, such as one or more cathode ray tubes (CRT), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices, either now known or later developed. User interface module 2401 can also be configured to generate audible outputs, with devices such as a speaker, speaker jack, audio output port, audio output device, earphones, and / or other similar devices. User interface module 2401 can further be configured with one or more haptic devices that can generate haptic outputs, such as vibrations and / or other outputs detectable by touch and / or physical contact with computing device 2400. In some examples, user interface module 2401 can be used to provide a graphical user interface (GUI) for utilizing computing device 2400.

[0170] Network communications module 2402 can include one or more devices that provide one or more wireless interfaces 2407 and / or one or more wireline interfaces 2408 that are configurable to communicate via a network. Wireless interface(s) 2407 can include one or more wireless transmitters, receivers, and / or transceivers, such as a Bluetooth™ transceiver, a Zigbee® transceiver, a Wi-Fi™ transceiver, a WiMAX™ transceiver, an LTE™ transceiver, and / or other type of wireless transceiver configurable to communicate via a wireless network. Wireline interface(s) 2408 can include one or more wireline transmitters, receivers, and / or transceivers, such as an Ethernet transceiver, a Universal Serial Bus (USB) transceiver, or similar transceiver configurable to communicate via a twisted pair wire, a coaxial cable, a fiberoptic link, or a similar physical connection to a wireline network.

[0171] In some examples, network communications module 2402 can be configured to provide reliable, secured, and / or authenticated communications. For each communication described herein, information for facilitating reliable communications (e.g., guaranteed message delivery) can be provided, perhaps as part of a message header and / or footer (e.g., packet / message sequencing information, encapsulation headers and / or footers, size / time information, and transmission verification information such as cyclic redundancy check (CRC) and / or parity check values). Communications can be made secure (e.g., be encoded or encrypted) and / or decry pted / decoded using one or more cryptographic protocols and / or algorithms, such as, but not limited to, Data Encryption Standard (DES), Advanced Encryption Standard (AES), a Rivest-Shamir-Adelman (RSA) algorithm, a Diffie-Hellman algorithm, a secure sockets protocol such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms can be used as well or in addition to those listed herein to secure (and then decry pt / decode) communications.Atty. Docket: 24-0609 -WO

[0172] One or more processors 2403 can include one or more general purpose processors (e.g., central processing unit (CPU), etc.), and / or one or more special purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application specific integrated circuits, etc.). One or more processors 2403 can be configured to execute computer-readable instructions 2406 that are contained in data storage 2404 and / or other instructions as described herein.

[0173] Data storage 2404 can include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of one or more processors 2403. The one or more computer-readable storage media can include volatile and / or non-volatile storage components, such as optical, magnetic, organic or other memory or disc storage, which can be integrated in whole or in part with at least one of one or more processors 2403. In some examples, data storage 2404 can be implemented using a single physical device (e.g., one optical, magnetic, organic or other memory or disc storage unit), while in other examples, data storage 2404 can be implemented using two or more physical devices.

[0174] Data storage 2404 can include computer-readable instructions 2406 and perhaps additional data. In some examples, data storage 2404 can include storage required to perform at least part of the herein-described methods, scenarios, and techniques and / or at least part of the functionality of the herein-described devices and networks. In particular, computer- readable instructions 2406 can include instructions that, when executed by processor(s) 2403, enable computing device 2400 to provide for some or all of the functionality described herein.

[0175] In some embodiments, computer-readable instructions 2406 can include instructions that, when executed by processor(s) 2403, enable computing device 2400 to carry out operations. The operations may include receiving, from an image sensor of a camera device, an input image of a scene. The operations may also include determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene. The operations may additionally include providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

[0176] In some embodiments, computer-readable instructions 2406 can include instructions that, when executed by processor(s) 2403, enable computing device 2400 to carry out operations. The operations may include receiving training data comprising a plurality of pairs, each pair comprising an image of a scene and an associated camera-independent calibrated color representation format based on a target global chromaticity for the scene. The operations may also include training, based on the training data, a machine learning (ML)Atty. Docket: 24-0609 WO model to predict a particular camera-independent calibrated color representation format for a given input image. The operations may additionally include providing the trained ML model.

[0177] In some examples, computing device 2400 can include white-balance (WB) correction module 2412. WB correction module 2412 can be configured to determine a cameraindependent calibrated color representation format for the input image based on a target global chromaticity for the scene. Also, for example, WB correction module 2412 can be configured to train a machine learning (ML) model to predict a particular camera-independent calibrated color representation format for a given input image.

[0178] In some examples, computing device 2400 can include one or more cameras 2418. Camera(s) 2418 can include one or more image capture devices, such as still and / or video cameras, equipped to capture light and record the captured light in one or more images; that is, camera(s) 2418 can generate image(s) of captured light. The one or more images can be one or more still images and / or one or more images utilized in video imagery. Camera(s) 2418 can capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as one or more other frequencies of light. Camera(s) 2418 can include a wide camera, a tele camera, an ultrawide camera, and so forth. Also, for example, camera(s) 2418 can be front-facing or rear-facing cameras with reference to computing device 2400. Camera(s) 2418 can include camera components such as, but are not limited to, an aperture, shutter, recording surface (e.g., photographic film and / or an image sensor), lens, and / or shutter button. The camera components may be controlled at least in part by software executed by one or more processors 2403.

[0179] In some examples, computing device 2400 can include one or more sensors 2420. Sensors 2420 can be configured to measure conditions within computing device 2400 and / or conditions in an environment of computing device 2400 and provide data about these conditions. For example, sensors 2420 can include one or more of: (i) sensors for obtaining data about computing device 2400, such as, but not limited to, a thermometer for measuring a temperature of computing device 2400, a battery sensor for measuring power of one or more batteries of power system 2422, and / or other sensors measuring conditions of computing device 2400; (ii) an identification sensor to identify other objects and / or devices, such as, but not limited to, a Radio Frequency Identification (RFID) reader, proximity sensor, one-dimensional barcode reader, two-dimensional barcode (c.g, Quick Response (QR) code) reader, and a laser tracker, where the identification sensors can be configured to read identifiers, such as RFID tags, barcodes, QR codes, and / or other devices and / or object configured to be read and provide at least identifying information; (iii) sensors to measure locations and / or movements ofAtty. Docket: 24-0609 -WO computing device 2400, such as, but not limited to, a tilt sensor, a gyroscope, an accelerometer, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser-displacement sensor, and a compass; (iv) an environmental sensor to obtain data indicative of an environment of computing device 2400, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor (e.g., an ambient light sensor), a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a movement sensor, a microphone, a sound sensor, an ultrasound sensor and / or a smoke sensor; and / or (v) a force sensor to measure one or more forces (e.g., inertial forces and / or G-forces) acting about computing device 2400, such as, but not limited to one or more sensors that measure: forces in one or more dimensions, torque, ground force, friction, and / or a zero moment point (ZMP) sensor that identifies ZMPs and / or locations of the ZMPs. Many other examples of sensors 2420 are possible as well.

[0180] Power system 2422 can include one or more batteries 2424 and / or one or more external power interfaces 2426 for providing electrical power to computing device 2400. Each battery of the one or more batteries 2424 can, when electrically coupled to the computing device 2400, act as a source of stored electrical power for computing device 2400. One or more batteries 2424 of power system 2422 can be configured to be portable. Some or all of one or more batteries 2424 can be readily removable from computing device 2400. In other examples, some or all of one or more batteries 2424 can be internal to computing device 2400, and so may not be readily removable from computing device 2400. Some or all of one or more batteries 2424 can be rechargeable. For example, a rechargeable battery can be recharged via a wired connection between the battery and another power supply, such as by one or more power supplies that are external to computing device 2400 and connected to computing device 2400 via the one or more external power interfaces. In other examples, some or all of one or more batteries 2424 can be non-rechargeable batteries.

[0181] One or more external power interfaces 2426 of power system 2422 can include one or more wired-power interfaces, such as a USB cable and / or a power cord, that enable wired electrical power connections to one or more power supplies that are external to computing device 2400. One or more external power interfaces 2426 can include one or more wireless power interfaces, such as a Qi wireless charger, that enable wireless electrical power connections, such as via a Qi wireless charger, to one or more external power supplies. Once an electrical power connection is established to an external power source using one or more external power interfaces 2426, computing device 2400 can draw electrical power from the external power source the established electrical power connection. In some examples, powerAtty. Docket: 24-0609 -WO system 2422 can include related sensors, such as battery sensors associated with the one or more batteries or other types of electrical power sensors.

[0182] One or more external power interfaces 2426 of power system 2422 can include one or more wired-power interfaces, such as a USB cable and / or a power cord, that enable wired electrical power connections to one or more power supplies that are external to computing device 2400. One or more external power interfaces 2426 can include one or more wireless power interfaces, such as a Qi wireless charger, that enable wireless electrical power connections, such as via a Qi wireless charger, to one or more external power supplies. Once an electrical power connection is established to an external power source using one or more external power interfaces 2426, computing device 2400 can draw electrical power from the external power source the established electrical power connection. In some examples, power system 2422 can include related sensors, such as battery sensors associated with the one or more batteries or other types of electrical power sensors.Example Methods of Operation

[0183] Figure 25 is a flowchart of a method, in accordance with example embodiments. Method 2500 may include various blocks or steps. The blocks or steps may be carried out individually or in combination. The blocks or steps may be carried out in any order and / or in series or in parallel. Further, blocks or steps may be omitted or added to method 2500.

[0184] The blocks of method 2500 may be carried out by various elements of computing device 2400 as illustrated and described in reference to Figure 24.

[0185] Block 2510 involves receiving, from an image sensor of a camera device, an input image of a scene.

[0186] Block 2520 involves determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene.

[0187] Block 2530 involves providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

[0188] In some embodiments, the color representation format may include a correlated color temperature (CCT) map.

[0189] In some embodiments, the input image may be a downsampled version of a raw image captured by the image sensor.

[0190] In some embodiments, the camera-independent calibrated color representation format may be based on a calibration chart that maps a chromaticity of the camera device in known lighting conditions to a camera-specific color representation format.Atty. Docket: 24-0609 -WO

[0191] In some embodiments, the calibration chart may be stored as a look-up-table on the camera device.

[0192] In some embodiments, the determining of the camera-independent calibrated color representation format involves applying a machine learning (ML) model to the input image, the ML model having been trained to predict the camera-independent calibrated color representation format for the input image based on the target global chromaticity for the scene.

[0193] Some embodiments involve providing, to the ML model, luma statistics for the scene, wherein the luma statistics are weighted averages of red, green, and blue (RGB) color channel values indicative of spatially linear brightness levels for the scene.

[0194] Some embodiments involve converting red, green, and blue (RGB) color channel values for the scene to camera-independent linear sRGB (LsRGB) statistics. Such embodiments involve providing, to the ML model, the camera-independent LsRGB statistics.

[0195] Some embodiments involve providing, to the ML model, a concatenation of linear sRGB (LsRGB) statistics, luma statistics, and correlated color temperature (CCT) statistics.

[0196] Some embodiments involve determining, based on the predicted color representation format, a delta map for the scene. Such embodiments involve mapping the delta map to a standardized camera-independent calibrated chroma map based on a look-up-table. Such embodiments additionally involve determining, based on the chroma map, a whitebalance gain map. Such embodiments also involve generating, based on pixel-wise color correction, a spatially corrected high resolution linear sRGB (LsRGB) version of the input image. The AWB correction may be based on the LsRGB version of the input image.

[0197] In some embodiments, the input image may be captured under spatially varying light colors.

[0198] In some embodiments, the ML model may be a convolutional neural network (CNN).

[0199] In some embodiments, the CNN may be based on a U-Net architecture.

[0200] In some embodiments, the CNN may be a residual CNN.

[0201] In some embodiments, the ML model may reside on the camera device.

[0202] Figure 26 is another flowchart of a method, in accordance with example embodiments. Method 2600 may include various blocks or steps. The blocks or steps may be carried out individually or in combination. The blocks or steps may be carried out in any order and / or in series or in parallel. Further, blocks or steps may be omitted or added to method 2600.Atty. Docket: 24-0609 -WO

[0203] The blocks of method 2600 may be carried out by various elements of computing device 2400 as illustrated and described in reference to Figure 24.

[0204] Block 2610 involves receiving training data comprising a plurality of pairs, each pair comprising an image of a scene and an associated camera-independent calibrated color representation format based on a target global chromaticity for the scene.

[0205] Block 2620 involves training, based on the training data, a machine learning (ML) model to predict a particular camera-independent calibrated color representation format for a given input image.

[0206] Block 2630 involves providing the trained ML model.

[0207] In some embodiments, the ML model may be a convolutional neural network(CNN).

[0208] In some embodiments, the CNN may be based on a U-Net architecture.

[0209] In some embodiments, the CNN may be a residual CNN.

[0210] In some embodiments, the ML model may reside on the camera device.

[0211] The particular arrangements shown in the Figures should not be viewed as limiting. It should be understood that other embodiments may include more or less of each element shown in a given Figure. Further, some of the illustrated elements may be combined or omitted. Yet further, an illustrative embodiment may include elements that are not illustrated in the Figures.

[0212] A step or block that represents a processing of information can correspond to circuitry that can be configured to perform the specific logical functions of a herein-described method or technique. Alternatively or additionally, a step or block that represents a processing of information can correspond to a module, a segment, or a portion of program code (including related data). The program code can include one or more instructions executable by a processor for implementing specific logical functions or actions in the method or technique. The program code and / or related data can be stored on any type of computer readable medium such as a storage device including a disk, hard drive, or other storage medium.

[0213] The computer readable medium can also include non-transitory computer readable media such as computer-readable media that store data for short periods of time like register memory, processor cache, and random access memory (RAM). The computer readable media can also include non-transitory computer readable media that store program code and / or data for longer periods. Thus, the computer readable media may include secondary or persistent long-term storage, like read only memory (ROM), optical or magnetic disks, compact disc read only memory (CD-ROM), for example. The computer readable media can also be any otherAtty. Docket: 24-0609 -WO volatile or non-volatile storage systems. A computer readable medium can be considered a computer readable storage medium, for example, or a tangible storage device.

[0214] While various examples and embodiments have been disclosed, other examples and embodiments will be apparent to those skilled in the art. The various disclosed examples and embodiments are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

Atty. Docket: 24-0609 -WOCLAIMSWhat is claimed is:

1. A computer-implemented method, comprising: receiving, from an image sensor of a camera device, an input image of a scene; determining a camera-independent calibrated color representation format for the input image based on a target global chromaticity for the scene; and providing the determined color representation format to an image processing pipeline of the camera device for auto white balance (AWB) correction.

2. The computer-implemented method of claim 1, wherein the color representation format comprises a correlated color temperature (CCT) map.

3. The computer-implemented method of any of claims 1 or 2, wherein the input image is a downsampled version of a raw image captured by the image sensor.

4. The computer-implemented method of any of claims 1-3, wherein the cameraindependent calibrated color representation format is based on a calibration chart that maps a chromaticity of the camera device in known lighting conditions to a camera-specific color representation format.

5. The computer-implemented method of claim 4, wherein the calibration chart is stored as a look-up-table on the camera device.

6. The computer-implemented method of any of claims 1-5, wherein the determining of the camera-independent calibrated color representation format further comprises: applying a machine learning (ML) model to the input image, the ML model having been trained to predict the camera-independent calibrated color representation format for the input image based on the target global chromaticity for the scene.Atty. Docket: 24-0609 -WO7. The computer-implemented method of claim 6, further comprising: providing, to the ML model, luma statistics for the scene, wherein the luma statistics are weighted averages of red, green, and blue (RGB) color channel values indicative of spatially linear brightness levels for the scene.

8. The computer-implemented method of any of claims 6 or 7, further comprising: converting red, green, and blue (RGB) color channel values for the scene to cameraindependent linear sRGB (LsRGB) statistics; and providing, to the ML model, the camera-independent LsRGB statistics.

9. The computer-implemented method of any of claims 6-8, further comprising: providing, to the ML model, a concatenation of linear sRGB (LsRGB) statistics, luma statistics, and correlated color temperature (CCT) statistics.

10. The computer-implemented method of any of claims 6-9, further comprising: determining, based on the predicted color representation format, a delta map for the scene; mapping the delta map to a standardized camera-independent calibrated chroma map based on a look-up-table; determining, based on the chroma map, a white-balance gain map; and generating, based on pixel-wise color correction, a spatially corrected high resolution linear sRGB (LsRGB) version of the input image, and wherein the AWB correction is based on the LsRGB version of the input image.

11. The computer-implemented method of claim 10, wherein the input image is captured under spatially varying light colors.

12. The computer-implemented method of any of claims 6-11, wherein the ML model is a convolutional neural network (CNN).

13. The computer-implemented method of claim 12, wherein the CNN is based on a U-Net architecture.

14. The computer-implemented method of claim 12, wherein the CNN is a residualAtty. Docket: 24-0609 -WOCNN.

15. The computer-implemented method of any of claims 6-14, wherein the ML model resides on the camera device.

16. A computer-implemented method, comprising: receiving training data comprising a plurality of pairs, each pair comprising of an image of a scene and an associated camera-independent calibrated color representation format based on a target global chromaticity for the scene; training, based on the training data, a machine learning (ML) model to predict a particular camera-independent calibrated color representation format for a given input image; and providing the trained ML model.

17. The computer-implemented method of claim 16, wherein the ML model is a convolutional neural network (CNN).

18. The computer-implemented method of claim 17, wherein the CNN is based on a U-Net architecture.

19. The computer-implemented method of claim 17, wherein the CNN is a residual CNN.

20. The computer-implemented method of any of claims 16-19, wherein the ML model resides on the camera device.

21. A computing device, comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out functions that comprise the computer-implemented method of any one of claims 1-Atty. Docket: 24-0609 -WO22. An article of manufacture comprising one or more non-transitory computer readable media having computer-readable instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to carry out functions that comprise the computer-implemented method of any one of claims 1-20.

23. A program that, when executed by one or more processors of a computing device, causes the computing device to carry out functions that comprise the computer- implemented method of any one of claims 1-20.

Citation Information

Patent Citations

  • Method and system of deep learning-based automatic white balancing

    US20190045163A1

  • Techniques for acquiring and processing statistics data in an image signal processor

    US8531542B2