Image or video content auto-conform techniques

EP4751229A2Pending Publication Date: 2026-06-03DOLBY LABORATORIES LICENSING CORP

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-07-16
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively manage and combine image/video content from different cameras, particularly when dealing with mixed formats such as standard dynamic range (SDR) and high dynamic range (HDR), leading to visually perceptible rendering mismatches.

Method used

An auto content conform system that automatically modifies input image/video contents from different cameras to match a reference visual appearance, using techniques such as tone-matching, noise characterization, and color space transformations, to achieve a uniform visual appearance similar to that of a selected reference camera.

Benefits of technology

The system achieves a relatively uniform and satisfactory visual appearance in the output image/video contents, matching the reference visual appearance with minimal manual input, and reduces noise visibility by adjusting codeword gains based on noise characterization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024038225_30012025_PF_FP_ABST
    Figure US2024038225_30012025_PF_FP_ABST
Patent Text Reader

Abstract

Reference images and source images in a working color space are determined. The reference images are derived from a reference camera. The source images are derived from a source camera. Initial gain values between the reference images and the source images are derived based on reference codeword values of the reference images and reference images. The initial gain values are adjusted into modified gain values based on results of noise characterization performed with the source images. The modified gain values are applied to the source images to generate leveled source images. Based on source-to-reference tone mappings, the leveled source images are converted to tone-matched source images.
Need to check novelty before this filing date? Find Prior Art

Description

IMAGE OR VIDEO CONTENT AUTO-CONFORM TECHNIQUESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 515,642, filed on 26 July 2023, European Patent Application No. 23190847.6, filed on 10 August 2023, and U.S. Provisional Application No. 63 / 568,340, filed on 21 March 2024, each of which is incorporated by reference herein in its entirety.TECHNOLOGY

[0002] The present disclosure relates generally to image / video processing, and in particular, to image / video content auto-conform techniques.BACKGROUND

[0003] In recent years, “prosumer” content creation has proliferated, particularly among social media influencers. In those cases, multiple consumer-grade cameras may be used to capture content. As opposed to professional workflows that work with raw image sequences, this modality instead involves ingesting, combining, editing, and grading processed bitstreams, each of which may have already been passed through an image processing pipeline. For content creation involving multiple clips, each camera’s image signal processing pipeline will render a scene perceptually differently - for example one camera may choose to impose a greater amount of saturation and / or contrast than another camera. Editing / grading tools such as commercially available Resolve, Final Cut Pro, or Adobe Premiere allow for importing some of the common standard dynamic range (SDR) formats into their platforms, but these tools are ill-equipped to effectively manage high dynamic range formats that have different tone scales and color space representations peculiar to the underlying cameras. Furthermore, this problem becomes compounded and even worsened visually when a mixture of SDR and high dynamic range (HDR) clips from different cameras are edited into a combined video clip.

[0004] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this sectionqualify as prior art merely by virtue of their inclusion in this section. Similarly, issues identified with respect to one or more approaches should not assume to have been recognized in any prior art on the basis of this section, unless otherwise indicated.BRIEF DESCRIPTION OF DRAWINGS

[0005] An embodiment of the present invention is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which:

[0006] FIG. 1 illustrates an example auto content conform system;

[0007] FIG. 2A and FIG. 2B illustrate example reference and non-reference source images; FIG. 2C illustrates example reference image generated by a reference camera and auto-leveled non-reference source image;

[0008] FIG. 3A illustrates example spatial correlations or correspondence relationships between common image features of two images; FIG. 3B illustrates two example binary maps constructed from performing spatial transformation between images; FIG. 3C illustrates an example codeword mapping curve; FIG. 3D illustrates example pre-saturation-leveled and postsaturation-leveled tone-mapped non-reference images;

[0009] FIG. 4A through FIG. 4F illustrate example process flows;

[0010] FIG. 5 illustrates an example hardware platform on which a computer or a computing device as described herein may be implemented;

[0011] FIG. 6A through FIG. 6C illustrate example process flows;

[0012] FIG. 7A through FIG. 7C illustrate example reference and source images as well as tone-matched source images; and

[0013] FIG. 8 illustrates an example spline relationship to control white point gains based on noise.DESCRIPTION OF EXAMPLE EMBODIMENTS

[0014] In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the present disclosure may be practiced without these specific details.In other instances, well-known structures and devices are not described in exhaustive detail, in order to avoid unnecessarily occluding, obscuring, or obfuscating the present disclosure.Summary

[0015] Input image / video contents previously acquired or created from a mix of different cameras or image acquisition / creation devices may have to be processed or converted into corresponding output image / video contents of a relatively uniform or conforming visual appearance, for example in a combined video clip or image sequence of a common timeline. The relatively uniform or conforming visual appearance of the output image / video contents may, but is not necessarily limited to only, a Dolby Vision visual appearance with a relatively high dynamic range and wide color gamut.

[0016] In such mixed content scenarios, the creative - which may consist of or one or more content creators or authors - can select a reference or “hero” camera from among a plurality of different cameras some or all of which may be used to acquire or create the input image / video contents. The selected reference camera generates image / video contents of a reference visual appearance (also referred to as a reference camera look) specific to the reference camera with its particular image acquisition / processing settings.

[0017] After the input image / video contents of different visual appearances (or different camera looks) are created with different image pipelines implemented by the different cameras, further image processing / conversion operations - which are referred to as auto-conforming or auto-conform operations - can be performed to modify or convert the input image / video contents into the corresponding output image / video content of the same reference visual appearance of the selected reference camera.

[0018] The input image / video contents may have been generated with a wide variety of cameras implementing different image creation pipelines and / or different image signal processors. Example cameras may include, but are not necessarily limited to only, any, some or all of: Canon C300, Apple iPhone 14, GoPro Hero 8, etc. The input image / video contents to be performed with the auto-conform operations may be originally acquired or captured from the same and / or different physical environments or visual scenes in which the cameras are deployed.

[0019] The auto-conformed input image / video contents may be further processed orincorporated into a combined edited or playback timeline with the same relative uniform reference camera visual appearance. It should be noted that some or all techniques as described herein may be implemented or applied regardless of whether the selected reference camera may be deployed in the physical environments or visual scenes to capture or generate any of the input image / video contents.

[0020] Hence, an auto content conform system implementing some or all of the techniques as described herein can automatically conform source image / video contents acquired with different cameras to a reference visual appearance associated with image / video contents acquired with a reference camera, with no or little manual input or manipulation needed for performing the autoconform operations. The system can achieve a relatively satisfactory match of the reference visual appearance, as if all the automatically conformed source image / video contents were derived from the same reference camera or a corresponding image creation pipeline thereof. The auto content conforming operations as described herein may be performed with no or little follow-up manual or user modifications to the resultant output image / video contents, which as automatically conformed have already achieved a relatively uniform visual appearance corresponding to that of the selected reference camera.

[0021] In some operational scenarios, tone-matching or auto-conforming source content of a source camera to an appearance of a reference camera may involve dark-to-bright scene (portion) changes. To reduce noise visibility from amplifying noise grains present or inherent in the source content, additional noise characterization and noise-based adjustment may be made to better control the amount of codeword or white point gain that are to be applied in white (point) leveling operations. More specifically, the noise characterization can be used to generate an estimate of (total) noise in the source content before the tone-matching operation. The noise may be computed through a noise metric generated for image blocks in the source content. The noise adjustment can be used to adjust or limit the amount of the codeword or white point gains. As a result, noise visibility in the tone-matched or auto-conformed content can be much reduced or prevented.

[0022] Example embodiments described herein relate to automatically conforming image content to a reference visual appearance of a reference camera. One or more reference images in a working color space and one or more source images in a selected working color space are determined. The one or more reference images are originally captured by a reference camera.The one or more source images are originally captured by a source camera different from the reference camera. One or more common image features are identified between the one or more reference images and the one or more source images. The one or more common image features in the one or more reference images include a plurality of reference pixels. The one or more common image features in the one or more source images include a plurality of source pixels. A distribution of reference codeword values of the plurality of reference pixels and a distribution of source codeword values of the plurality of source pixels, as represented in the selected color space, are determined to generate one or more source-to-reference codeword mappings. Based at least in part on the one or more source-to-reference codeword mappings, a specific source image originally captured by the source camera converted to a corresponding auto-conformed source image conforming to a reference visual appearance associated with the one or more reference images originally captured by the reference camera.

[0023] Example embodiments described herein relate to automatically conforming image content to a reference visual appearance of a reference camera. One or more reference images and one or more source images in a working color space are determined. The one or more reference images are derived from a reference camera. The one or more source images are derived from a source camera different from the reference camera. Initial gain values between the one or more reference images and the one or more source images are derived. The initial gain values are determined based on reference codeword values of the one or more reference images and source codeword values of the one or more reference images. The initial gain values are adjusted into modified gain values based at least in part on one or more results of noise characterization performed with the one or more source images. The modified gain values are applied to the one or more source images to generate one or more leveled source images. Based at least in part on one or more source-to-reference tone mappings, the one or more leveled source images are converted to one or more tone-matched source images.

[0024] Various modifications to the preferred embodiments and the generic principles and features described herein will be readily apparent to those skilled in the art. Thus, the disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.Multi-Stage Processing

[0025] FIG. 1 illustrates an example auto content conform system 100 that implements multi-stage image processing operations. The auto content conform system 100 may be implemented with one or more computing devices and may be used to generate the output image / video content with the conforming visual appearance associated with the reference camera.

[0026] In the first stage or processing block 102 of the auto content conform system 100, the creative can provide user input to identify or specify the reference camera and associated working space including but not limited to a reference (e.g., HDR, SDR, etc.) color space. By way of example but not limitation, a first camera (e.g., the iPhone 14, etc.) may be selected as the reference camera, whereas the associated working space may be selected as the Hybrid Log Gamma (HLG) / Rec. 2020 space.

[0027] In the second stage or processing block 104 of the auto content conform system 100, specific video formats of the input image / video content. - e.g., video clips taken by different cameras - are interpreted or used to retrieve source image / video frames or source image data therein from the input image / video contents. The source image data in the source video frames of each of the input image / video contents can be subsequently linearized, if applicable. A video format may refer to a specific video file / container / signal (or its type) carrying an image (e.g., generated as output from the ISP of the camera, etc.) as described herein as well as image metadata relating to the image. Example image metadata may include, but is not necessarily limited to only, information on some or all of: color space, dynamic range, color gamut, bit depth, limited code space, full range code space, applicable video / image specification / standard, exposure, saturation, contrast, color compensation vector or scaling factors used by the camera.

[0028] In the third stage or processing block 106 of the auto content conform system 100, color space transformations may be performed on or applied to convert the (linearized) source image / video frames represented in (linearized) source color spaces - e.g., different from the reference color space associated with the working space - into corresponding color-space- converted (linearized) source image / video frames represented in the reference color space.

[0029] In the fourth stage or processing block 108 of the auto content conform system 100, content matching or conforming operations are applied to the (color-space-converted, linearized)source image / video frames to corresponding output image / video frames of the reference camera visual appearance.

[0030] The first three stages may include carrying out operations to parse or interpret input bitstream color and tone scale representations in the input image / video contents. In operational scenarios in which the source frames in the input image / video contents are represented in nonlinear color space different from the (e.g., linearized) reference color space, these source frames are converted from the non-linear color space such as YUV to the (color-space-converted, linearized) source image / video frames represented in the (linearized) reference color space such as a linear RGB color space or video format.

[0031] The color space conversion or transformation can be carried out based at least in part on the source or working color spaces or colorimetry respectively associated with the source or working color spaces. In some operational scenarios, one or more conversion matrixes such as those associated with YUV-to-RGB conversions may be applied to carry out the color space conversion or transformation.

[0032] FIG. 2A and FIG. 2B illustrate example reference and (linearized, color-space- converted) source images represented in the HLG Rec 2020 working (color) space after the previously mentioned first three stages or steps of image processing operations. The reference image in FIG. 2A may be originally in a first video clip captured by a reference camera such as iPhone 14 and represented in the HLG Rec 2020 working space, whereas the source image in FIG. 2A may be originally in a second video clip captured by a non-reference or source camera such as Sony A7 and represented in a non-reference working space such as Slog3 / SGamut3, which is subsequently converted to the HLG / Rec 2020 working space by the previously mentioned first three stages of image processing operations.

[0033] As illustrated, not only is the exposure different between the two cameras, but also the color balance, saturation, and overall contrast. These easily perceptible differences would incur significant effort on the part of a user to obtain a satisfactory conformance of image / video content acquired with the two cameras. For example, color editing / correction systems today do have some capability to automatically level content, but with mixed results. These color editing / grading systems or platforms - e.g., Adobe Premiere, Blackmagic Resolve, and Apple Final Cut Pro - still produce visibly perceptible rendering mismatches.

[0034] The problem is further complicated by the fact that for real-world applications, the cameras that are to be conformed may not have captured the same scene at the same time. The two video feeds may therefore have had zero or few common elements that could be used to conform.

[0035] FIG. 2C illustrates example reference image generated by a reference camera (e.g., iPhone 14 in the present example, etc.) and auto-leveled source (or non-reference) image generated by applying the auto-leveling feature of Adobe Premiere to a non-reference video frame or image in the Sony A7 Slog3 video clip. Although the image - which is on the right of FIG. 2C - generated by Adobe Premiere using the auto-leveling feature shows improvement relative to that of FIG. 2B, the exposure, saturation, and contrast are still significantly different from the reference image - which is on the left of FIG. 2C - acquired with iPhone 14. As a result, a user would still need to manually intervene and apply a relatively substantial amount of color grading to this clip in order to achieve a better similarity to the visual appearance in the video clip acquired with iPhone 14. This could potentially involve applying relatively complex and numerous secondary grading operations.

[0036] In contrast with these existing systems / platforms that are prone to produce visibly perceptible mismatches, an auto content conform system as described herein can be implemented to automatically perform content conforming operations on image / video contents acquired with different cameras to achieve a relatively satisfactory match in a relatively uniform visual appearance of the automatically conformed image / video contents, as if all the automatically conformed contents were derived from the same reference camera. These content conforming operations may be performed with no or little user modification and / or with no or little manual intervention. In some operational scenarios, these operations may be incorporated into the previously mentioned fourth stage (or step).

[0037] More specifically, after the initial ingest of input image / video contents acquired from non-reference or source cameras in the previously mentioned first three stages (or steps), the auto content conform system can apply image processing operations in the previously mentioned fourth stage (or step) to automatically conform the input image / video contents into corresponding conformed image / video contents with a visual appearance that is the same as or approximates the reference visual appearance of the reference camera in (reference) image / or content acquired with the reference camera.

[0038] The image processing operations by the auto content conform system may include tone editing / correction and / or color editing / correction on the intermediate output from the first three stages (or steps). The tone editing / correction and / or color editing / correction for contents acquired by each of the different cameras can be optimized for that camera. Hence, the system can perform robust image processing operations on contents generated by different cameras that are optimized for or conditioned on differences in the contents generated by these different cameras in relation to the reference camera, thereby realizing or generating a relatively uniform visual appearance associated with contents generated by the reference camera. As a result, the system can help effectively and efficiently solve the problem in real-world applications that the different cameras that are to be conformed may not have captured the same scene at the same time.

[0039] While multiple images or video feeds in the input image / video contents captured with different (sample or source) cameras may have had zero or few common elements that could be used to conform, all of these images or video feeds can be automatically conformed by the system to generate the conformed image / video feeds that have the relatively uniform visual appearance associated with contents generated by the reference camera, as if these conformed image / video feeds were captured by the reference camera in place of these different (sample or source) camera in the real-world applications or real-world scenes in which the input image / video contents were captured.Automatic Content Conforming

[0040] FIG. 4A illustrates an example process flow for auto content conforming. The process flow may be implemented or performed at least in part with one or more computing devices including but not limited to auto content conform systems, post-camera image / video processing systems, auto-conform tools installed on computing devices including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0041] Block 402 comprises reading or receiving reference image / video contents captured with the reference camera as well as reading or receiving non-reference image / video contents captured with a non-reference camera. By way of illustration but not limitation, the reference image / video contents may be referred to as a reference video stream. The reference camera maybe referred to as a hero camera. The non-reference camera may be referred to as a target or source or sample camera. The non-reference image / video contents may be referred to as a target or source or sample video stream.

[0042] Block 404 comprises linearizing the reference and non-reference (or sample / source) video streams according to their respective (input) inverse Optical-Electro Transfer Functions (OETFs). These inverse OETFs can be individually and specifically identified based in part or in whole on user input and / or image metadata received or decoded from the received reference and sample video streams.

[0043] Block 406 comprises applying a color correction matrix, if applicable (e.g., where the reference and sample cameras are of different camera models and / or makes, etc.), to the reference and sample video streams to convert video frames or images in these streams to the working (color) space of the reference camera. The working space can be specifically identified based in part or in whole on user input and / or image metadata received or decoded from the received reference video stream.

[0044] Block 408 comprises identifying common (visual or image) features between the reference and sample video streams or the video frames or images therein.

[0045] Block 410 comprises determining a codeword or tone mapping curve that maps luminances or luminance codewords of pixels in the common features in the sample video streams to luminances or luminance codewords of pixels in the corresponding common features in the reference video stream.

[0046] Block 412 comprises determining a color balancing matrix that maps colors or chrominances of the pixels in the common features in the sample video stream to colors or chrominances of the pixels in the corresponding common features in the reference video stream.

[0047] Block 414 comprises applying the codeword / tone mapping and color balancing matrixes to the linearized sample video stream to generate an adjusted sample video stream.

[0048] Block 416 comprises applying a specific OETF to the adjusted sample video stream to generate an auto content conformed sample video stream corresponding to the working space of the reference camera.

[0049] FIG. 4B illustrates an example process flow for determining tone and chrominance mapping curves (e.g., of FIG. 4A, etc.). The process flow may be implemented or performed at least in part with one or more computing devices including but not limited to auto contentconform systems, post-camera image / video processing systems, auto-conform tools installed on computing devices including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0050] As illustrated in FIG. 4A, the common features between the reference and sample video streams can be first identified before determining tone and chrominance mapping curves.

[0051] Block 422 comprises calculating a reference cumulative distribution function (CDF) of the luminances of the pixels in the common features in the reference video stream as well as a (non-reference) sample or source CDF of the luminances of the pixels in the common features in the sample or source video stream.

[0052] Block 424 comprises generating a codeword / tone mapping curve that maps the luminances (or luminance values) in the sample or source image data of the sample or source video stream to the reference-conformed luminances (or luminance values) of the autoconformed sample or source video stream.

[0053] Block 426 comprises, for each of the chroma channels, calculating a reference cumulative distribution function (CDF) of the chrominances of the pixels in the common features in the reference video stream as well as a (non-reference) sample or source CDF of the chrominances of the pixels in the common features in the sample or source video stream.

[0054] Block 428 comprises, for each of the chroma channels, generating a chrominance mapping curve that maps the chrominances (or chrominance values) in the sample or source image data of the sample or source video stream to the reference-conformed chrominances (or chrominance values) of the auto-conformed sample or source video stream.

[0055] FIG. 4C illustrates an example process flow for calculating spline curves to perform tone and chrominance mapping. The process flow may be implemented or performed at least in part with one or more computing devices including but not limited to auto content conform systems, post-camera image / video processing systems, auto-conform tools installed on computing devices including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0056] As illustrated in FIG. 4A and FIG. 4B, tone and chrominance mapping curves can be determined by first identifying common features between the reference and sample video streams.

[0057] Block 442 comprises calculating, based at least in part on the tone mapping curve and / or CDFs, a first spline curve that maps luminances (or luminance values) in the sample or source image data of the sample or source video stream to reference-conformed luminances (or luminance values) of the auto-conformed sample or source video stream. Additionally, optionally or alternatively, the first spline curve may be generated directly based on histograms of luminance values collected from the pixels of the common image features in the reference and non-reference images. The first spline curve may be made of B-spline functions joined at a set of (e.g., 11, etc.) nodes or knot points.

[0058] Block 444 comprises calculating, based at least in part on the chrominance mapping curve and / or CDFs, a second spline curve that maps chrominances (or chrominance values) in the sample or source image data of the sample or source video stream to reference-conformed chrominances (or chrominance values) of the auto-conformed sample or source video stream. Additionally, optionally or alternatively, the second spline curve may be generated directly based on histograms of chrominance values collected from the pixels of the common image features in the reference and non-reference images. The second spline curve may be made of B-spline functions joined at a set of (e.g., 11, etc.) nodes or knot points.

[0059] FIG. 4D illustrates an example process flow for training and applying an artificial neural network such as a convolutional neural network (CNN) predicting a tone (mapping) curve and a color balance matrix between the reference and sample cameras. The process flow may be implemented or performed at least in part with one or more computing devices including but not limited to auto content conform systems, post-camera image / video processing systems, autoconform tools installed on computing devices including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0060] Block 462 comprises using training reference and sample (or non-reference) image / video contents along with ground truth (e.g., known codeword / tone mapping curves and / or known color balance matrixes, etc.) to train the CNN in an offline training process / phase to learn or optimize operational parameters of the CNN for predicting codeword / tone mapping curves and color balance matrixes that have relatively small prediction errors. Input features may be extracted from the training reference and sample image / video contents and provided as input to the CNN for predicting or estimating tone mapping curves and / or color balance matrixes with minimized prediction errors in reference to the ground truth.

[0061] Block 464 comprises applying, in a runtime application process / phase, the trained CNN on video / image data in a non-training sample video stream to predict specific codeword / tone mapping curves and specific color balance matrixes using optimized operational parameters of the CNN that have been trained or optimized with the previously trained data. These codeword / tone mapping curves and specific color balance matrixes can be used to map video / image data in the non-training sample video stream to auto-conformed video / image data in an auto-conformed sample video stream.CDF Matching Overview

[0062] As noted, CDF matching operations in the process flow of FIG. 4B may be applied to luminance and / or chrominance values of pixels in common image features of reference image / video contents and non-reference (e.g., source, sample, etc.) image / video contents for the purpose of generating luminance (or tone) and / or chrominance mapping curves that can be used to map or convert the non-reference image / video contents into a reference visual appearance associated with the reference image / video contents.

[0063] In some operational scenarios, these image / video contents such as images or video clips may be captured by a reference camera and a non-reference camera with respect to the same physical environment or visual scene but from their respective different spatial camera locations in the physical environment or visual scene. For example, a reference image may be originally captured by the reference camera as a reference video frame in an iPhone 14 video clip as illustrated in FIG. 2A, whereas a non-reference image may be originally captured by the non- reference camera as a non-reference video frame in a Sony A7 video clip as illustrated in FIG. 2B. The reference and non-reference images can contain pixels in common image features as both images are captured (e.g., concurrently, close in time, close in light conditions, etc.) from the same environment or scene. However, the common image features may be geometrically or spatially displaced and hence located in respective spatial areas in the reference and non- reference images.

[0064] The reference and non-reference images - or pixel / codeword values therein - may be first linearized to generate linearized reference and non-reference images. A reference luminance image and a non-reference luminance image may be respectively created from the linearizedreference image (or video frame) and the linearized non-reference image (or video frame). For example, the linearized reference and non-reference images (or video frames) may be represented in a linearized RGB color space. An RGB-to-YCbCr (color space) conversion matrix may be applied to linearized RGB values in the linearized reference image (or video frame) and the linearized non-reference image (or video frame) to generate the reference luminance image and the non-reference image, respectively.

[0065] In some operational scenarios, luminance values of the pixels in the common image features of the reference and non-reference images (or their respective luminance images) can be used to establish correspondence relationships (e.g., feature points, corner points, eye corners, etc.) between the respective spatial areas of the common image features. These correspondence relationships may be used to generate or estimate an approximate spatial transform (or transformation). The spatial transform may operate in the luminance domain. For example, the spatial transform may be used to map between locations of non-reference luminance values in the non-reference (luminance) image originally from the non-reference camera and corresponding locations of reference luminance values in the non-reference (luminance) image originally from the reference image.

[0066] In some operational scenarios, reference and non-reference pixels belonging to the common image features in the reference and non-reference images may be identified, based at least in part on the spatial transform and / or the reference and non-reference luminance images respectively derived from the reference and non-reference luminance images.

[0067] For example, one or more computer vision techniques including but not limited to the oriented FAST and rotated BRIEF (ORB) algorithm, the speeded-up-robust-features (SURF) algorithm, scale-invariant-feature-transform (SIFT), a computer-implemented feature key point detection / matching algorithm, a computer-implemented feature detection / matching algorithm, may be used to match, identify or recognize common image features in both reference and non- reference luminance images. These common image features in the reference and non-reference luminance images - or in the reference and non-reference video frames giving rise to the luminance images - depict the same visual objects, characters, etc., albeit from the different spatial locations of the reference and non-reference cameras in the same physical environment or visual scene.

[0068] These common image features and spatial displacement (or disparity) information of pixels therein in the reference and non-reference luminance images can be used to estimate a projective transformation from a first image in the reference and non-reference luminance images to a second different image in these images. The projective transformation may be used to project or spatially transform pixels (or pixel locations) of the common image features in the first image (one of the reference and non-reference luminance images) to corresponding pixels (or pixel locations) of the common image features in the second image (the other of the reference and non-reference luminance images).

[0069] Histograms of luminance values in the pixels of the common image features in the reference and non-reference luminance images can be subsequently computed. These histograms or luminance value bins therein can be used by CDF matching operations to generate a luminance mapping curve - which may be referred to as a tone curve - that may be used to map or convert luminance values of the non-reference video frame into mapped / converted luminance values in an auto-conformed video frame corresponding to the non-reference video frame. In some operational scenarios, some or all of the luminance mapping curve may be represented as a lookup table.

[0070] Example codeword / tone mapping operations are described in U.S. Provisional Patent Application No. 62 / 356,087, fded on June 29, 2016, entitled "EFFICIENT HISTOGRAMBASED LUMA LOOK MATCHING" by Harshad Kadu et al. Example CDF operations are described in U.S. Provisional Patent Application No. 62 / 404,307, filed on October 5, 2016, entitled "INVERSE LUMA / CHROMA MAPPINGS WITH HISTOGRAM TRANSFER AND APPROXIMATION" by Bihan Wen et al. The above-mentioned patent applications are hereby incorporated by reference as if fully set forth herein.Example CDF Matching

[0071] By way of illustration but not limitation, while the reference and non-reference cameras are capturing visual elements such as visual objects, characters, background, etc., of the same physical environment or visual scene in the reference and non-reference images, each of the cameras may have a different field of view, different lens distortion, etc. Hence, there may be visual objects that appear in one image, but not in the other.

[0072] To identify the subset of some or all specific common or correlated pixels in each of the reference and non-reference image, common image features (e.g., specific texture, building corner, rear car window lower left corner, eye, nose, lips, chin, hand, table, car, etc.) between the two cameras may be identified, for example, using the SURF method / algorithm (or another computer vision technique such as a computer-implemented feature (key point) matching / detection algorithm). FIG. 3A illustrates example correlations or correspondence relationships - such as spatial displacements indicated by lines - between common image features of two images from the same physical environment or visual scene as identified using computer vision techniques including but not limited to the SURF method / algorithm.

[0073] In some operational scenarios, each of the reference and non-reference images can be linearized and represented in the same colorimetric working space or color space (e.g., RGB, etc ). The linearized reference and non-reference images in the working space can be subsequently converted to corresponding (reference and non-reference) single channel luminance images.

[0074] The SURF method / algorithm can then be used to detect common image (or visual) features in each of the reference and non-reference luminance images. Geometric / spatial displacements (e.g., indicated by lines in FIG. 3A) between pairs of specific spatial locations (e.g., specific pixel locations, eye corners, lips, etc.) of the common image features in the reference and non-reference luminance images can be estimated for those common image features that are common or correlated.

[0075] From the geometric / spatial displacements between the pairs of specific spatial locations as detected by the SURF method / algorithm, one or more projective (or spatial) transformations can be estimated and used to provide one or more geometric (or spatial) mappings from the coordinate system of the non-reference image to that of the reference image or vice versa.

[0076] In some operational scenarios, both forward and reverse projective (or spatial) transformations can be estimated and then used to identify valid pixel maps or pixels in the common image features in each of the reference and non-reference images. Source pixel coordinates of source pixels in the non-reference image are passed through the forward transformation to generate or identify corresponding reference pixels with corresponding reference pixel coordinates in the reference image. Any non-reference pixels that are forwardtransformed into reference pixel coordinates outside of the reference video frame used to contain the reference image may be excluded (as “invalid” source pixels) from the common image features.

[0077] The same or converse may be performed for reference pixels in the reference image using the reverse transformation to exclude any reference pixels - from the common image features - that are reverse transformed into non-reference pixel coordinates outside of the nonreference video frame used to contain the non-reference image.

[0078] FIG. 3B illustrates two example binary maps constructed from performing the forward and reverse transformations on the non-reference and reference images, respectively. In a first binary map on the left, white pixels indicate all valid (not excluded as “invalid”) reference pixels in the common image features of the reference image, whereas black pixels indicate all other pixels in the reference image that do not belong to the common image features. Likewise, in a second binary map on the right, white pixels indicate all valid (not excluded as “invalid”) non-reference pixels in the common image features of the non-reference image, whereas black pixels indicate all other pixels in the non-reference image that do not belong to the common image features.

[0079] An entire luminance (codeword) range or space for representing non-reference luminance values in the non-reference images may be partitioned (e.g., equally, etc.) into a first plurality of luminance (codeword) sub-ranges. Counts of (“valid” or not excluded as “invalid”) pixels - or total numbers of pixels in the common image features of the non-reference image - having non-reference luminance values in each sub-range in the first plurality of luminance subrange may be stored in a respective bin a plurality of bins of a non-reference histogram.

[0080] Likewise, an entire luminance (codeword) range or space for representing reference luminance values may be partitioned (e.g., equally, etc.) into a second plurality of luminance (codeword) sub-ranges. Counts of (“valid” or not excluded as “invalid”) pixels - or total numbers of pixels in the common image features of the reference image - having reference luminance values in each sub-range in the second plurality of luminance sub-range may be stored in a respective bin a plurality of bins of a reference histogram.

[0081] As both the reference and non-reference histograms are constructed from the same or correlated distribution of pixels of the common image features identified in the reference and non-reference images, these histograms can be used to compute or establish correlations orcorrespondences between reference and non-reference pixel or codeword values represented in the reference and non-reference pixels of the common image features.

[0082] For example, a reference cumulative density function denoted as CDFR(vR) - where vRdenotes reference pixel or codeword values (e.g., reference luminance values representing reference luminance sub-ranges, etc.) corresponding to bins of the reference histogram - can be obtained from the reference histogram by simply applying cumulative sums, along an incrementing direction of the reference pixel or codeword values, to the bins of the reference histogram.

[0083] Similarly, a non-reference cumulative density function denoted as CDFs(us') - where usdenotes non-reference pixel or codeword values (e.g., non-reference luminance values representing non-reference luminance sub-ranges, etc.) corresponding to bins of the non- reference histogram - can be obtained from the non-reference histogram by simply applying cumulative sums, along an incrementing direction of the non-reference pixel or codeword values, to the bins of the non-reference histogram.

[0084] Here the S and R subscripts refer to non-reference (or source) and reference pixel or codeword spaces, respectively. The non-reference pixel or codeword values us, and reference pixel or codeword values vR, can be respectively represented and indexed - e.g., with an index of i or / - in the non-reference (or source) and reference pixel or codeword scales / ranges.

[0085] For a given non-reference pixel or codeword value us(i), a corresponding reference pixel or codeword value vRmay be determined by enforcing or using an equality condition as follows: (1)

[0086] By way of illustration but not limitation, the correspondence or mapping of us(i) and vRmay be accomplished based on expression (1) above using simple linear interpolation. For example, CDFs(us(if) may be between CDFR(VR(J)) and CDFR(vR(jy). The corresponding reference pixel or codeword value vRcan be obtained by applying the simple linear interpolation to vR(j) and vR(j) based on respective distances or differences of CDFR(VR(J)) and

[0087] As a result, the reference and non-reference CDFs derived from the reference and non-reference histograms can be used to generate or derive luminance (or tone) and / or chrominance mapping curve(s) or LUT(s) - as illustrated in FIG. 3C - that map or convert non-reference pixel or codeword values usto their corresponding reference pixel or codeword values VR -

[0088] Additionally, optionally or alternatively, in some operational scenarios, the reference and non-reference CDFs derived from the reference and non-reference histograms can be likewise used to generate or derive inverse luminance (or tone) and / or chrominance mapping curve(s) or LUT(s) that map or convert reference pixel or codeword values vRto their corresponding non-reference pixel or codeword values us.

[0089] In some operational scenarios, a (forward and / or inverse) luminance and / or chrominance mapping curve or LUT may be represented or approximated by a set of B-spline curves with a relatively small number of nodes or knots (e.g., 11, etc.). This may help achieve the benefit of reducing the overall complexity in transforming images from one visual appearance to a different visual appearance (e.g., a reference visual appearance, etc.) as well as potentially producing a smoother function due to continuity properties of the B-spline functions.

[0090] In some operational scenarios, for a sequence of time-consecutive non-reference images (e.g., belonging to the same video clip, scene, group of pictures or GOPs, etc.), luminance and / or chrominance mapping curves may be stabilized temporally over a plurality of time points or time durations covered by the sequence of images. This helps relatively gracefully handle or avoid frame-specific errors in the feature detection and the subsequent projective transformation. For example, a relatively constant or fixed luminance transformation or mapping curve may be used across the entire sequence of images. Additionally, optionally or alternatively, a temporal low pass filter can be applied when applying mapping curves as described herein across multiple time-consecutive frames in the sequence.

[0091] Additionally, optionally or alternatively, instead of or in addition to using (framespecific, frame-pair-specific) valid pixel maps of a pair of reference and non-reference images to construct histograms and mapping curves for the pair of reference and non-reference images, in some operational scenarios, global valid pixel maps of multiple pairs of reference and non- reference images may be used to construct histograms and mapping curves as described herein for some or all of the multiple pairs of reference and non-reference images.Codeword or RGB Leveling

[0092] Subsequent image processing operations such as codeword or RGB leveling operations may be performed on the tone-mapped non-referenced image to generate a codeword or RGB leveled tone-mapped non-reference image. The codeword or RGB leveling operations may represent relatively minor adjustments to pixel or codeword values. Respective gain values in the non-RGB or RGB channels can be applied to adjust respective non-RGB or RGB pixel or codeword values in the tone-mapped non-reference image to generate the codeword or RGB leveled tone-mapped non-reference image. Along with the tone mapping, the leveling operations help provide or achieve a relatively high quality match between the leveled tone-mapped non- reference image (originally derived from the Sony A7 clip in the present example) and the reference image (originally derived from the iPhone 14 clip in the present example) in terms of luminance and color visual appearances. As shown, it is evident that the auto-conform approach as described herein produces a much better correspondence between the reference and sample images in comparison with other approaches.

[0093] By way of example but not limitation, the non-reference (or sample) image is represented in a linearized RGB color space, the tone curve or LUT generated from the CDF matching operations may be first applied equally to component pixel values in the RGB channels of the non-reference (or sample) image to generate a tone-mapped non-reference image.

[0094] Following this tone (or luminance-domain) mapping, RGB leveling operations can be performed to refine the visual appearance match with respect to the reference camera. For example, RGB pixel or codeword values in the individual RGB channels of the source image (or the tone-mapped non-reference image) may be adjusted by a respective ratio of the reference-to- source averages for each channel. For example, the component pixel or codewords denoted as 7?s(i) in the red channel may be adjusted to generate leveled or adjusted component pixel or codewords denoted as Rs' (i) in the red channel as follows:where (Rx) is the average over all valid pixels for image A; Xis S or R respectively denoting the source or reference images. In some operational scenarios, clipping operations may be performed after the scaling operations or adjustments using the reference-to- source ratios to constrain the adjusted values within a specific pixel or codeword value range. Additionally, optionally or alternatively, some or all codeword or RGB leveling operations as described herein may be applied using a white balancing matrix constructed based at least in part on the reference-to-source ratios computed for different color channels of the working color space.Saturation Leveling

[0095] In some operational scenarios, the CDF approach performs well in matching the luminance channel or appearance of the sample or non-reference image with that of the reference image, but the intensity of the colors may still show disparity. To deal with this problem, in the linear domain, after applying the tone mapping and (e.g., RGB, etc.) leveling operations such as gain value adjustments illustrated in expression (2) above, a saturation metric can be calculated for each of the sample and reference images for applying saturation leveling.

[0096] In various operational scenarios, different approaches may be used or implemented to calculate saturation metrics for applying saturation leveling operations. Aa a first example, an RGB reference or non-reference image for which saturation leveling is to be applied may be converted from the RGB color space to the CIE 1976 L*a*b* color space. As a second example, an RGB reference or non-reference image for which saturation leveling is to be applied may be converted from the RGB color space to the YCbCr color space (or model).

[0097] The saturation metric denoted as Sat can then be calculated using color (or nonluminance) channels in the CIE 1976 L*a*b* color space or the YCbCr color space, as follows:

[0098] The saturation metric may be used to calculate an overall saturation ratio of the mean saturation of the reference image to the mean saturation of the sample image for each pair of the reference and sample images. The sample image can then be (saturation) leveled by applying the (overall saturation) ratio (e.g., using a saturation balancing matrix, etc.) in a similar manner to that of expression (2) above.

[0099] FIG. 3D illustrates example pre- saturation-leveled and post-saturation-leveled tonemapped (e.g., using tone curve(s) or LUT(s) generated from CDF matching, etc.) non-reference images on the left and on the right, respectively. The original non-reference image giving rise to these images may be captured in a video clip by Sony A7.

[0100] In some operational scenarios, instead of or in addition to saturation leveling, saturation matching may be performed by way of a source-to-reference saturation mappinggenerated from CDFs of saturation metric histograms. For example, analogous to the tone (or luminance-domain) and / or chrominance mapping, one can calculate these CDFs of the saturation metrics for both reference and sample images using reference and non-reference histograms with bins storing count of valid pixels in different saturation metric sub-ranges in a plurality of saturation metric range. A saturation mapping curve or LUT may be computed or estimated by enforcing or using equality conditions between the CDFs of saturation metric histograms constructed from the reference and sample images. The saturation mapping curve or LUT may be applied to the RGB-leveled tone-mapped non-reference image to generate saturation-leveled RGB-leveled tone-mapped non-reference image with a relatively refined visual appearance matching the reference visual appearance.

[0101] Additionally, optionally, alternatively, an artificial neural network may be used and trained to perform some or all of: codeword or RGB leveling, saturation leveling and / or saturation matching.

[0102] For the purpose of illustration, it has been described that the leveling operations may be performed after luminance (or tone) and / or chrominance mapping operations. It should be noted that, in some other operational scenarios, some or all of the leveling operations such as RGB leveling (or white balancing) may be performed before luminance (or tone) and / or chrominance mapping operations. Additionally, optionally or alternatively, in some operational scenarios, some or all of the leveling operations such as RGB leveling (or white balancing) may be performed concurrently or in combination with luminance (or tone) and / or chrominance mapping operations.Noise Visibility

[0103] Techniques as described herein can be implemented or applied to address the problem of matching the look or rendering of image or video content captured by different cameras in different lighting conditions. As noted, a specific camera among a plurality of cameras used to capture the content can be selected as the reference or “hero” camera. Content portions - which may be considered as source or “test” image / video content portions to be auto-conformed into the look or rendering associated with the reference or hero camera - generated by other camera(s) can be modified into auto-conformed content portions that appear similar to the lookor rendering of the reference or hero camera. The effort for (human) creators of user generated content (UGC) to generate a seamless timeline from mixed content captured by multiple reference and / or non-reference cameras can be much reduced with the techniques as described herein.

[0104] FIG. 6A illustrates an example process flow, method or algorithm summarizing content auto-conform operations that have been described herein. This process flow may be used to auto-conform a source or non-reference look or rendering of a source or non-reference video / image stream into a reference look or rendering identical or similar to that of a reference video / image stream. The process flow may be implemented or performed at least in part with one or more computing devices including but not limited to auto content conform systems, postcamera image / video processing systems, auto-conform tools installed on computing devices including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0105] Block 602 comprises converting the source video / image stream and / or the reference video / image stream into the same (working) color space if the input color space of the source video / image stream is different from the color space of the reference video / image stream. Block 602 further comprises linearizing the reference and source video / image streams in the (working) color space.

[0106] Block 604 comprises, to generate an auto-conformed look or rendering (of the source video / image stream) that matches the look or rendering of the reference video / image stream, computing or deriving white balancing (e.g., codeword balancing, RGB balancing, etc.) gains based on codewords in images / frames in the reference and source video / image streams. For example, as illustrated in expression (2) above, average codewords (Rx) may be computed over valid pixels for source or reference images in the source or reference video / image streams, respectively, where the valid pixels are at pixel locations - in the reference and source video / image streams - at which pixel values are located near, at or within a specific configured or selected chromaticity neighborhood around a specific white point (e.g., Illuminant D65, etc.) of the reference color space in the CIE 1931 x, y chromaticity diagram. An average ratio or a white point (WP) gain may be determined based on these average codewords (Rx) from the reference and source video / image streams. In some operational scenarios, the reference andsource images are represented in an RGB color space, the average ratios or white point gains can be calculated individually for individual R, G, and B channels of the RGB color space.

[0107] Block 606 comprises multiplying the codewords of the source images in the source video / image stream with the average ratios or white point gains to perform white point leveling (or RGB / codeword leveling), as illustrated in expression (2) above. The application of the leveling operations to the source video / image stream produces a leveled source video / image stream.

[0108] Block 608 comprises generating and applying a tone map curve or LUT used to match the tone scale of the source video / image stream with the reference video / image stream. In some operational scenarios, this curve or LUT can then be (e.g., equally, individually, etc.) applied to the individual R, G, and B channels of the source video stream. The application of the tone map curve or LUT to the leveled source video / image stream produces a tone-mapped (or tone-matched) leveled source video / image stream.

[0109] Block 610 comprises applying a specific opto-electronic transfer function (OETF) designated for or associated with the reference video / image stream to the tone-matched leveled (linearized) source video / image stream to generate the auto-conformed source video / image stream.

[0110] FIG. 7A illustrates an example relatively bright reference image in a reference video / image stream (at the top of FIG. 7A), an example (relatively bright) source image in an original or pre-auto-conformed source video / image stream (in the middle of FIG. 7A), and an example auto conformed image (at the bottom of FIG. 7A) derived from the source image using the process flow of FIG. 6A. As shown, this process flow, method or algorithm of FIG. 6A produces relatively high quality results - e.g., with no or little noise visibility - in matching bright videos such as bright images shown in FIG. 7A.

[0111] FIG. 7B illustrates an example relatively dark source image in a source video / image stream (at the top of FIG. 7B) and an example auto conformed image derived from this source image using the process flow of FIG. 6A. As shown, in comparison with matching a bright source video / image with a bright reference video / image as illustrated in FIG. 7A using the process flow of FIG. 6A, when attempting to match a dark source video / image (at the middle of FIG. 7B) with a bright reference video / image (at the top of FIG. 7B) as illustrated in FIG. 7B using the same process flow of FIG. 6A, intrinsic noise introduced by the source camera in adark scene depicted in the source video / image causes objectionable noise visibility in the autoconformed or matched video / image (at the bottom of FIG. 7B).

[0112] This objectionable noise visibility in the adjusted or auto-conformed video / image is caused by performing the RGB / codeword leveling operations (also referred to as white point gains) using average ratios or white point gains computed in block 604 of FIG. 6A. In the case of the dark-to-bright adjustment, the average ratios or white point gains may be of relatively large values. As a result, the noise introduced by the source camera is amplified to a relatively significant extent to result in a relatively perceptible noise visibility in the adjusted or autoconformed video / image.

[0113] Under techniques as described herein, a noise analysis / characterization of the scenes captured in images can be performed or generated. Information obtained from this noise analysis / characterization may be used to control an amount of gain that is ultimately applied in the leveling operations.Noise Modulated Leveling

[0114] An ISO speed in a digital camera used to capture or acquire reference or nonreference source video / image content controls overall brightness of the captured video / image content. The ISO speed depends in whole or in part on (e.g., intrinsic, sensor-specific, etc.) light sensitivity characteristics of image sensor(s) of the camera as well as a gains (or gain settings) of image signal processor(s) (ISP(s)) of the camera. With an ISP gain, the video / image signal generated by the image sensor can be (e.g., electrically, etc.) amplified, including any noise present or inherent in the pre-amplified video / image signal.

[0115] In low-light situations or physical scenes, higher ISO speeds or higher gain values (e.g., user or system settable, etc.) can amplify a video / image signal - including any noise grains therein - generated by the image sensor and / ISP and result in a higher visibility of the noise grains in the captured video / image content generated by the image sensor and / or ISP.

[0116] FIG. 6B illustrates an example process flow, method or algorithm summarizing content auto-conform operations described herein, which is modified from the process flow of FIG. 6A. The process flow of FIG. 6B may be used (e.g., as an alternative or addition to that of FIG. 6A, with specific cameras, with specific image contents, with specific scene or brightnesschanges, etc.) to auto-conform a source or non-reference look or rendering of a source or nonreference video / image stream into a reference look or rendering identical or similar to that of a reference video / image stream. The process flow may be implemented or performed at least in part with one or more computing devices including but not limited to auto content conform systems, post-camera image / video processing systems, auto-conform tools installed on computing devices including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0117] As shown in FIG. 6A and FIG. 6B, block 606 of FIG. 6A is replaced by blocks 606-1 and 606-2 of FIG. 6B. Other blocks of FIG. 6B, other than blocks 606-1 and 602-2, may be performed in a manner same as or similar to those corresponding blocks of FIG. 6A. The average ratios or white point gains determined in block 604 of FIG. 6A or FIG. 6B may be used as initial white point gains to be further modified depending or based on noise analysis characterization.

[0118] More specifically, blocks 606-1 of FIG. 6B comprises applying or performing noise analysis or characterization to one or both of the reference and source video / image streams or images therein. Results of the noise analysis / characterization are applied or used to adjust the (initial) average ratios or white point gains to generate modified white point gains. The modified white point gains modulated or adjusted with the results of the noise analysis or characterization may be used in RGB / codeword leveling (or white point leveling) operations in place of the (preadjusted or pre-modified) average ratios or white point gains.

[0119] Block 606-2 of FIG. 6B comprises multiplying the codewords of the source images in the source video / image stream with the modified white point gains (as multiplicative factors) generated at least in part based on the noise analysis / characterization to perform a white point leveling (or RGB / codeword leveling). The application of the leveling operations to the source video / image stream produces a leveled source video / image stream that may be further processed by subsequent blocks 608 and 610 of FIG. 6A or FIG. 6B to generate a corresponding tone- matched or auto-conformed source video / image stream.Noise Analysis

[0120] FIG. 6C illustrates an example process flow for analyzing / characterizing noise in a video or image stream (e g., a source or reference video / image stream, an image therein, etc.).The process flow may be implemented or performed at least in part with one or more computing devices including but not limited to auto content conform systems, post-camera image / video processing systems, auto-conform tools installed on computing devices including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0121] Block 620 comprises applying an edge detector, algorithm, method, procedure and / or operator to the luminance channel of some or all images or frames of a video / image stream - e.g., after representing or converting the images or frames in a specific color space such as YUV that includes the luminance channel, etc. - to find or detect (e.g., visually perceptible, etc.) edges in the images / frames of the stream. Examples of the edge detector or detection algorithm, method, procedure and / or operator may include, but are not necessarily limited to only, Sobel, Canny, etc. Example edges may relate to visually perceptible structures of objects, characters, background, foreground, borders, frames, etc., visually depicted in the images or frames.

[0122] Block 622 comprises dividing each image or frame of the images / frames under noise analysis / characterization herein into multiple image blocks with block sizes such as 4x4, 8x8, and 16x16. Among these multiple image blocks, any image blocks including portions of the detected edges may be excluded from further noise analysis / characterization operations. Block 624 comprises determining, among these multiple image blocks, a specific set of image blocks to include some or all (good or valid) image blocks only with an absence of any of the detected edges. The specific set of (all good or valid) image blocks - excluding all image blocks with edge portions - is used to perform further noise analysis / characterization operations.

[0123] Block 626 comprises, for each image block in the specific of (all good or valid) image blocks, determining the block’s average pixel value and standard deviation. Further, a noise-to- signal ratio (NSR) value is derived by dividing the standard deviation of the block with the average pixel value (within the block).

[0124] Block 628 comprises computing a global or across-blocks average (denoted as LumaNSR) of the NSR values across all the image blocks in the specific of (all good or valid) image blocks in the luminance or luma channel - e.g., of the specific color space such as YUV in which the images or frames are represented, etc.

[0125] Block 628 further comprises computing a global or across-blocks average (denoted as ChromaNSR) of the NSR values across all the image blocks in the specific of (all good or valid)image blocks in each (e.g., U or V, etc.) of two chrominance or chroma channels - e.g., of the specific color space such as YUV in which the images or frames are represented, etc. The ChromaNSR value for the U channel or component of the specific color space may be specifically denoted as U_NSR, whereas the ChromaNSR value for the V channel or component of the specific color space may be specifically denoted as V_NSR.where b represents the total number of (good or valid) image blocks in the set; c refers to the standard deviation of an image block, p subscript refers to an average value of the image block.

[0126] An overall chroma noise-to-signal ratio may be computed as follows:ChromaNSR = (5)

[0127] Block 630 comprises computing or determining an overall or total noise (metric) using the noise-to-signal values computed for the color channels or components of the specific color space. In some operational scenarios, this noise metric may be computed as a sum of the luma and chroma noise-to-signal values using one or both of two options as follows:where a represents a user or system configurable numeric value between 0 and 1.

[0128] Expression (6-1) above may be applied to compute the noise metric in cases in which some or all noise sources are treated as correlated or mutually (or statistically) dependent, whereas expression (6-2) above may be applied to compute the noise metric in cases in which some or all noise sources are treated as uncorrelated or mutually (or statistically) independent. Additionally, optionally, alternatively, in various operational scenarios, one of the two options may be specifically selected as default for computing the noise metric. In an image sensor, pixels with higher exposure values may have less relative noise (e.g., noise in pixels versus luminance of the pixels, etc.) as compared with pixels with lower exposure values due to photon shot noise statistics in the image sensor - for example, noise generated in the signal generated by the image sensor may be proportional to square root of intensity or luminance (corresponding to expression(6-1) above) - or proportional to intensity or luminance (corresponding to expression (6-2) above) - of the signal.

[0129] By dividing noise with luminance in the noise-to-signal (NSR) values, the percentage of the impact of the noise relative to the luma value can be relatively adequately or accurately determined, estimated or accounted for. These ratios may be used to emphasize relative importance of noise in relatively dark image blocks or content portions.

[0130] By using these NSR values accounted for relative importance of noise in relatively dark image blocks or portions to compute or determine the overall noise metric, this overall noise metric may likewise emphasize the relative importance of the noise in the relatively dark image blocks or content portions. Hence, this overall metric can be used in leveling operations that reduce or prevent noise artifacts as illustrated in FIG. 7B.

[0131] A wide variety of other metrics can be used in addition to or in place of the NSR values. By way of illustration but not limitation, an ICtCp color representation format and / or color space - in place of the YUV color space - may be used to represent (e.g., intermediate, etc.) images or frames from which (e.g., white point, etc.) gains for leveling operations may be determined. This ICtCp color space or the gains computed therewith may provide better perceptual consistency with respect to the human visual system as compared with other color representation formats and / or color spaces.

[0132] Example operations and representation formats relating to the ICtCp color space are defined or described in Rec. ITU-R BT.2100, “Image parameter values for high dynamic range television for use in production and international programme exchange,” (06 / 2017), which are incorporated herein by reference in its entirety.

[0133] In some operational scenarios, an ICtCp color representation of images / frames in video / image content can be generated or derived by first converting input video / image content from a linear RGB color space to an (intermediate) LMS color space, then applying a nonlinear perceptually quantized (PQ) function to generate non-linear (PQ) video / image content from the linear video / image content in the LMS color space, and finally converting the nonlinear (PQ) video / image content (or signals) from the LMS color space to the ICtCp color space.

[0134] With the images / frames represented in the ICtCp color space, a noise metric may be derived from noise analysis / characterization using noise-to-signal values (or ratios) such as noise EI, noise CT / I and noise CP / I.

[0135] Here, I denotes an average value of I codewords in images or frames in the video / image content represented in the ICtCp color space; noise l denotes a noise value in the I component or channel of the ICtCp color space; noise CT denotes a noise value in the Ct component or channel of the ICtCp color space; and noise CP denotes a noise value in the Cp component or channel of the ICtCp color space.Noise Treatment

[0136] As previously noted, initial white point gains may be computed from RGB or codeword averages of images / frames. Where these initial white point gains are relatively small, applying these gains to scene portions relating to bright-to-dark, dark-to-dark or bright-to-bright scene matching or auto-conforming may produce no or relatively small impact or effect on noise perceptibility in the resultant auto-conformed visual / image content.

[0137] In comparison, where these initial white point gains are relatively large, applying these gains to scene portions involving dark-to-bright scene matching or auto-conforming may produce relatively large impact or effect on noise perceptibility in the resultant auto-conformed visual / image content. As the gains are increased, the visibility of noise amplification and perceptibility in the auto-conformed video / image content can become more and more readily perceptible and objectionable to a human viewer.

[0138] For example, the initial white point gain - which may be computed from RGB or codeword values in the reference image as illustrated at the top of FIG. 7B and the non-reference image at the middle of FIG. 7B - used to directly generate an auto-conformed image (at the bottom of FIG. 7B) may be relatively large such as ten (10). When the relatively large initial white point gain (10) is directly applied to the non-reference image without noise characterization or treatment, the noise in the non-reference image may be amplified substantially - as compared with relatively small gains - to produce relatively significant unpleasing visual result shown at the bottom of FIG. 7B.

[0139] One or more noise treatment approaches may be used or implemented to control or moderate these relatively large white point gains depending on intrinsic noise levels estimated, determined, or characterized from results of noise analysis performed on video / image content. Insome operational scenarios, the intrinsic noise levels may be estimated, determined or characterized using noise metrics as described herein.

[0140] In a first example approach, modified white point gains (denoted as “OutputWpGains”) may be generated by adjusting or modifying corresponding initial white point gains (denoted as “InputWpGains”) with an approximately (e.g., with a regulation factor or constant such as one (1), etc.) inverse relationship between a noise level as represented by a corresponding noise metric and the initial white point gains, as follows:OutputWpGains = (InputWpGains +p*noise) / (l+(p*noise)) (7) where noise may be the same as the total noise computed in expression (6-1) or (6-2) above; p denotes an adjustment or scaling parameter.

[0141] A specific value (e.g., one (1), three (3), etc.) of the adjustment or scaling parameter p can be experimentally derived by performing an objective and / or subjective evaluation of effects of noises on videos captured by different cameras with various intrinsic noise levels. These cameras may be from different camera or device manufacturers, different camera or device models, and so forth., and / or may use different image sensors with different intrinsic or ISP noise levels or characteristics. A relatively large value of p may be used for a given level of the intrinsic noise level to more aggressively adjust or limit any noise amplification impact or effect of the initial white point gains.

[0142] In a second example approach, a specific relationship between initial white point gains and modified white point gains may be defined, specified or tuned, for example, taking into account noise visibility reduction for a variety of image transitions or brightness changes. In some operational scenarios, - as illustrated in FIG. 8, this specific relationship may be represented as a spline with three different value regions - numeric constants in this relationship are for illustration purposes only - as follows: where InputWpGains in 0-4: OutputWpGains = InputWpGains (8-1) where InputWpGains in 4-15: OutputWpGains = 0.1818* InputWpGains + 4.27 (8-2) where InputWpGains in >15: OutputWpGains = 2 (8-3)

[0143] In a third example approach, a clip parameter / value or maximum modified white point gain may be determined based at least in part on noise level or metric. Modified white point gains are set to the initial or original white point gains where the gains are no more than the clip parameter / value. The modified white point gains are fixed or set to the clip parameter / valuewhere the initial or original white point gains are greater than the clip parameter / value. In some operational scenarios, this clip parameter / value may be set as follows: cZip=p / (s+ noise) (9) where P (e.g., 2, 6, 8, 11, etc.) and 8 (e.g., 1, etc.) are tuning parameters; noise may be total noise computed with expression (6-1) or (6-2) above.

[0144] These tuning parameters may be used to control the maximum modified white point for a given noise level or metric value such as a maximum noise value. One or both of these parameters may be determined, set or estimated experimentally by optimizing the clip parameter / value for the maximum noise value with no or little noise amplification artifacts in auto-conformed video / image content.

[0145] In some operational scenarios, in block 606-2 of FIG. 6B, the modified white point gains - generated from applying noise treatment or adjustment to initial white point gains - for each R, G or B channel in an RGB color space may be constrained to maintain the same interchannel ratios (e.g., between red and green or R / G, between green and blue or G / B, etc.) as those of the initial or original white point gains for the purpose of maintaining the same color balance between images generated from applying the modified white point gains and images generated from applying the initial white point gains.

[0146] FIG. 7C illustrates an example auto-conformed image generated from applying a modified white point gain to a source or non-reference image. As compared with the autoconformed image generated from applying an initial white point gain to the same source or nonreference image as illustrated at the bottom of FIG. 7B, the auto-conformed image of FIG. 7C is derived after performing noise-dependent gain adjustment, resulting in a relatively high quality image for the dark scene with no or little perceptible noise amplification or artifacts. Here, the modified white point gain is reduced from the initial white point gain of 10 to 1.5, which serves to prevent or lessen the amplification of noise present in the pre-auto-conformed source image and produces relatively pleasing visual look or results.Example Process Flows

[0147] FIG. 4E illustrates an example process flow according to an example embodiment. In some example embodiments, one or more computing devices or components (e.g., a videoediting system, a desktop computer, a video processing tool, etc.) implementing an auto-conform system as described herein may perform this process flow. In block 482, the system determines one or more reference images in a working color space and one or more source images in a selected working color space. The one or more reference images are originally captured by a reference camera. The one or more source images are originally captured by a source camera different from the reference camera.

[0148] In block 484, the system identifies one or more common image features between the one or more reference images and the one or more source images. The one or more common image features in the one or more reference images include a plurality of reference pixels. The one or more common image features in the one or more source images include a plurality of source pixels.

[0149] In block 486, the system uses a distribution of reference codeword values of the plurality of reference pixels and a distribution of source codeword values of the plurality of source pixels, as represented in the selected color space, to generate one or more source-to- reference codeword mappings.

[0150] In block 488, based at least in part on the one or more source-to-reference codeword mappings, the system converts a specific source image originally captured by the source camera to a corresponding auto-conformed source image conforming to a reference visual appearance associated with the one or more reference images originally captured by the reference camera.

[0151] In an embodiment, the reference camera differs from the source camera in one or more of: output image color spaces, output image dynamic ranges, output image color gamuts, output image bit depths, image signal processors, image signal processor configurations, device makers, device models, etc.

[0152] In an embodiment, the one or more reference images are originally captured by the reference camera in a high dynamic range format; the high dynamic range format is one of: a hybrid log gamma (HLG) format represented within a Rec. 2020 RGB color space or a perceptual quantizer (PQ) format represented within a Rec. 2020 RGB color space; the one or more reference images in the high dynamic range format are linearized to be represented in a linearized RGB color space representation using a non-linear inverse optical-to-electro transfer function (OETF).

[0153] In an embodiment, the common image features are identified based at least in part on one or more projective transformations constructed from image features detected using one or more computer vision techniques.

[0154] In an embodiment, the one or more source-to-reference codeword mappings are generated by one or more of: cumulative density function (CDF) matching operations, or one or more artificial neural networks.

[0155] In an embodiment, at least one of the one or more source-to-reference codeword mappings is represented by one of: a mapping function, a spline function, a piecewise function, a piecewise polynomial function, a lookup table, etc.

[0156] In an embodiment, the specific source image belongs to a sequence of consecutive source images of a specific visual scene; wherein the one or more source-to-reference codeword mappings are applied to each source image in the sequence of consecutive source images of the specific visual scene to generate a corresponding sequence of consecutive auto-conformed source images of the specific visual scene.

[0157] In an embodiment, the specific source image represents one of the one or more source images.

[0158] In an embodiment, the one or more reference images and the one or more source images are respectively captured by the reference camera and the source camera from a common physical environment; the reference camera and the source camera are situated at different spatial locations in the common physical environment.

[0159] In an embodiment, the specific source image is not one of the one or more source images.

[0160] In an embodiment, the one or more source images converted to one or more corresponding auto-conformed source images that match the reference visual appearance associated with the one or more reference images originally captured by the reference camera; at least a subset of the one or more corresponding auto-conformed source images is combined with at least a subset of the one or more reference images into a single sequence of time-consecutive images along a single playback timeline.

[0161] In an embodiment, the corresponding auto-conformed image is generated by performing RGB leveling operations on an intermediate image generated at least in part by applying the one or more source-to-reference codeword mappings to the specific source image.

[0162] In an embodiment, the corresponding auto-conformed image is generated by performing one of saturation matching or saturation leveling operations on an intermediate image generated at least in part by applying the one or more source-to-reference codeword mappings to the specific source image.

[0163] FIG. 4F illustrates an example process flow according to an example embodiment. In some example embodiments, one or more computing devices or components (e.g., a video editing system, a desktop computer, a video processing tool, etc.) implementing an auto-conform system as described herein may perform this process flow. In block 490, the system determines one or more reference images and one or more source images in a working color space. The one or more reference images are derived from a reference camera. The one or more source images are derived from a source camera different from the reference camera.

[0164] In block 492, the system derives initial gain values between the one or more reference images and the one or more source images. The initial gain values are determined based on reference codeword values of the one or more reference images and source codeword values of the one or more reference images.

[0165] In block 494, the system adjusts, based at least in part on one or more results of noise characterization performed with the one or more source images, the initial gain values into modified gain values.

[0166] In block 496, the system applies the modified gain values to the one or more source images to generate one or more leveled source images.

[0167] In block 498, based at least in part on one or more source-to-reference tone mappings, the system converts the one or more leveled source images to one or more tone-matched source images.

[0168] In an embodiment, the working color space represents a linear color space; the linear color space is one of: a linear RGB color space, a linear YUV color space, a linear YCbCr color space, or a different linear color space.

[0169] In an embodiment, the working color space represents a nonlinear color space; the nonlinear color space is one of: a perceptually quantized (PQ) color space, a hybrid log-gamma (HLG) color space, a nonlinear LMS color space, a nonlinear ICpCt color space, a different nonlinear color space, etc.

[0170] In an embodiment, the initial gain values are determined as average ratios based on average reference codeword values of the one or more reference images and average source codeword values of the one or more reference images.

[0171] In an embodiment, the system further performs: applying a specific opto-electronic transfer function (OETF) associated with the one or more reference images to the tone-matched source images to generate one or more auto-conformed source images conforming to a reference visual appearance associated with the reference camera.

[0172] In an embodiment, the results of noise characterization performed with the one or more source images are at least in part represented by a noise metric computed using individual noise-to-signal ratios between standard deviations of individual image blocks and intensities of the individual image blocks; the image blocks are partitioned from the one or more source images.

[0173] In an embodiment, the individual image blocks partitioned from the one or more source images and used to compute the individual noise-to-signal ratios constitute a set of image blocks; the set of image blocks excludes any image blocks, in the one or more source images, that contain visually perceptible edges detected with one or more edge detection methods.

[0174] In an embodiment, the individual noise-to-signal ratios are computed separately for each color channel of the working color space in which the one or more source images are represented.

[0175] In an embodiment, the noise metric is determined for a luminance channel of the working color space; two additional noise metrics are determined for chrominance channels of the working color space; a total noise is computed for the one or more source images from the noise metric and the two additional noise metric; the total noise is used to modify the initial gain values into the modified gain values.

[0176] In an embodiment, the total noise is used in an approximate inverse relationship to modify the initial gain values into the modified gain values.

[0177] In an embodiment, the total noise is used in a spline relationship to modify the initial gain values into the modified gain values.

[0178] In an embodiment, the total noise is used to apply maximum gain limits to the modified gain values.

[0179] In an embodiment, each and every pair of the modified gain values for two different color channels maintain a first ratio value same as a second ratio value maintained by a corresponding pair of the initial gain values for the different color channels.

[0180] In various example embodiments, an apparatus, a system, an apparatus, or one or more other computing devices performs any or a part of the foregoing methods as described. In an embodiment, a non-transitory computer readable storage medium stores software instructions, which when executed by one or more processors cause performance of a method as described herein.

[0181] Note that, although separate embodiments are discussed herein, any combination of embodiments and / or partial embodiments discussed herein may be combined to form further embodiments.Implementation Mechanisms - Hardware Overview

[0182] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hard-wired to perform the techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are persistently programmed to perform the techniques, or may include one or more general purpose hardware processors programmed to perform the techniques pursuant to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices may also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices or any other device that incorporates hard-wired and / or program logic to implement the techniques.

[0183] For example, FIG. 5 is a block diagram that illustrates a computer system 500 upon which an example embodiment of the invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled with bus 502 for processing information. Hardware processor 504 may be, for example, a general purpose microprocessor.

[0184] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored in non-transitory storage media accessible to processor 504, render computer system 500 into a special-purpose machine that is customized to perform the operations specified in the instructions.

[0185] Computer system 500 further includes a read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504.

[0186] A storage device 510, such as a magnetic disk or optical disk, solid state RAM, is provided and coupled to bus 502 for storing information and instructions.

[0187] Computer system 500 may be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, a trackball, or cursor direction keys for communicating direction information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane.

[0188] Computer system 500 may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which in combination with the computer system causes or programs computer system 500 to be a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.

[0189] The term “storage media” as used herein refers to any non-transitory media that storedata and / or instructions that cause a machine to operation in a specific fashion. Such storage media may comprise non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of storage media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, NVRAM, any other memory chip or cartridge.

[0190] Storage media is distinct from but may be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio-wave and infra-red data communications.

[0191] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infra-red transmitter to convert the data to an infra-red signal. An infra-red detector can receive the data carried in the infra-red signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504.

[0192] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an integrated services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 518 sends and receiveselectrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0193] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526. ISP 526 in turn provides data communication services through the world wide packet data communication network now commonly referred to as the “Internet” 528. Local network 522 and Internet 528 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518, which carry the digital data to and from computer system 500, are example forms of transmission media.

[0194] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518.

[0195] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution.Equivalents, Extensions, Alternatives and Miscellaneous

[0196] In the foregoing specification, example embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indicator of what is the invention, and is intended by the applicants to be the invention, is the set of claims that issue from this application, in the specific form in which such claims issue, including any subsequent correction. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Hence, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should limit the scope of such claim in any way. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.Enumerated Exemplary Embodiments

[0197] The invention may be embodied in any of the forms described herein, including, but not limited to the following Enumerated Example Embodiments (EEEs) which describe structure, features, and functionality of some portions of embodiments of the present invention.

[0198] EEE 1. A method compri sing : determining one or more reference images in a working color space and one or more source images in a selected working color space, wherein the one or more reference images are originally captured by a reference camera, wherein the one or more source images are originally captured by a source camera different from the reference camera; identifying one or more common image features between the one or more reference images and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, wherein the one or more common image features in the one or more source images include a plurality of source pixels; using a distribution of reference codeword values of the plurality of reference pixels and a distribution of source codeword values of the plurality of source pixels, as represented in the selected color space, to generate one or more source-to-reference codeword mappings; based at least in part on the one or more source-to-reference codeword mappings, converting a specific source image originally captured by the source camera to a corresponding autoconformed source image conforming to a reference visual appearance associated with the one or more reference images originally captured by the reference camera.

[0199] EEE2. The method of EEE1, wherein the reference camera differs from the source camera in one or more of: output image color spaces, output image dynamic ranges, output image color gamuts, output image bit depths, image signal processors, image signal processor configurations, device makers, or device models.

[0200] EEE3. The method of EEE 1 or EEE2, wherein the one or more reference images are originally captured by the reference camera in a high dynamic range format; wherein the high dynamic range format is one of: a hybrid log gamma (HLG) format represented within a Rec. 2020 RGB color space or a perceptual quantizer (PQ) format represented within a Rec. 2020 RGB color space; wherein the one or more reference images in the high dynamic range formatare linearized to be represented in a linearized RGB color space representation using a non-linear inverse optical-to-electro transfer function (OETF).

[0201] EEE4. The method of any of EEE1-EEE3, wherein the common image features are identified based at least in part on one or more geometric transformations constructed from image features detected using one or more computer vision techniques.

[0202] EEE5. The method of any of EEE1-EEE4, wherein the one or more source-to- reference codeword mappings are generated by one or more of: cumulative density function (CDF) matching operations, or one or more artificial neural networks.

[0203] EEE6. The method of any of EEE1-EEE5, wherein at least one of the one or more source-to-reference codeword mappings is represented by one of: a mapping function, a spline function, a piecewise function, a piecewise polynomial function, or a lookup table.

[0204] EEE7. The method of any of EEE1-EEE6, wherein the specific source image belongs to a sequence of consecutive source images of a specific visual scene; wherein the one or more source-to-reference codeword mappings are applied to each source image in the sequence of consecutive source images of the specific visual scene to generate a corresponding sequence of consecutive auto-conformed source images of the specific visual scene.

[0205] EEE8. The method of any of EEE1-EEE7, wherein the specific source image represents one of the one or more source images.

[0206] EEE9. The method of any of EEE1-EEE8, wherein the one or more reference images and the one or more source images are respectively captured by the reference camera and the source camera from a common physical environment; wherein the reference camera and the source camera are situated at different spatial locations in the common physical environment.

[0207] EEE10. The method of any of EEE1-EEE9, wherein the specific source image is not one of the one or more source images.

[0208] EEE11. The method of any of EEE 1 -EEE 10, wherein the one or more source images converted to one or more corresponding auto-conformed source images that match the reference visual appearance associated with the one or more reference images originally captured by the reference camera; wherein at least a subset of the one or more corresponding autoconformed source images is combined with at least a subset of the one or more reference images into a single sequence of time-consecutive images along a single playback timeline.

[0209] EEE12. The method of any of EEE1 -EEE11 , wherein the corresponding autoconformed image is generated by performing RGB leveling operations on an intermediate image generated at least in part by applying the one or more source-to-reference codeword mappings to the specific source image.

[0210] EEE13. The method of any of EEE 1 -EEE 12, wherein the corresponding autoconformed image is generated by performing one of saturation matching or saturation leveling operations on an intermediate image generated at least in part by applying the one or more source-to-reference codeword mappings to the specific source image.

[0211] EEE14. A non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of the method recited in any of EEE1-EEE13.

[0212] EEE15. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in any of EEE1-EEE13.

[0213] EEE16. An apparatus comprising: one or more processors; and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of: determining one or more reference images in a working color space and one or more source images in a selected working color space, wherein the one or more reference images are originally captured by a reference camera, wherein the one or more source images are originally captured by a source camera different from the reference camera; identifying one or more common image features between the one or more reference images and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, wherein the one or more common image features in the one or more source images include a plurality of source pixels; using a distribution of reference codeword values of the plurality of reference pixels and a distribution of source codeword values of the plurality of source pixels, as represented in the selected color space, to generate one or more source-to-reference codeword mappings; based at least in part on the one or more source-to-reference codeword mappings, converting a specific source image originally captured by the source camera to a corresponding auto-conformed source image conforming to a reference visual appearance associated with the one or more reference images originally captured by the reference camera.

[0214] EEE17. The apparatus of EEE16, wherein the reference camera differs from the source camera in one or more of: output image color spaces, output image dynamic ranges, output image color gamuts, output image bit depths, image signal processors, image signal processor configurations, device makers, or device models.

[0215] EEE 18. The apparatus of EEE 16 or EEE 17, wherein the one or more reference images are originally captured by the reference camera in a high dynamic range format; wherein the high dynamic range format is one of: a hybrid log gamma (HLG) format represented within a Rec. 2020 RGB color space or a perceptual quantizer (PQ) format represented within a Rec. 2020 RGB color space; wherein the one or more reference images in the high dynamic range format are linearized to be represented in a linearized RGB color space representation using a non-linear inverse optical-to-electro transfer function (OETF).

[0216] EEE19. The apparatus of any of EEE16 to EEE18, wherein the common image features are identified based at least in part on one or more geometric transformations constructed from image features detected using one or more computer vision techniques.

[0217] EEE20. The apparatus of any of EEE 16 to EEE19, wherein the one or more source-to-reference codeword mappings are generated by one or more of: cumulative density function (CDF) matching operations, or one or more artificial neural networks.

[0218] EEE21. The apparatus of any of EEE 16 to EEE20, wherein at least one of the one or more source-to-reference codeword mappings is represented by one of: a mapping function, a spline function, a piecewise function, a piecewise polynomial function, or a lookup table.

[0219] EEE22. The apparatus of any of EEE 16 to EEE21, wherein the specific source image belongs to a sequence of consecutive source images of a specific visual scene; wherein the one or more source-to-reference codeword mappings are applied to each source image in the sequence of consecutive source images of the specific visual scene to generate a corresponding sequence of consecutive auto-conformed source images of the specific visual scene.

[0220] EEE23. The apparatus of any of EEE 16 to EEE22, wherein the one or more source images converted to one or more corresponding auto-conformed source images that match the reference visual appearance associated with the one or more reference images originally captured by the reference camera; wherein at least a subset of the one or more corresponding auto-conformed source images is combined with at least a subset of the one or more reference images into a single sequence of time-consecutive images along a single playback timeline.

[0221] EEE24. The apparatus of any of EEE16 to EEE23, wherein the corresponding auto-conformed image is generated by performing one of RGB leveling operations, saturation matching or saturation leveling operations on an intermediate image generated at least in part by applying the one or more source-to-reference codeword mappings to the specific source image.

[0222] EEE25. A non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of: determining one or more reference images in a working color space and one or more source images in a selected working color space, wherein the one or more reference images are originally captured by a reference camera, wherein the one or more source images are originally captured by a source camera different from the reference camera; identifying one or more common image features between the one or more reference images and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, wherein the one or more common image features in the one or more source images include a plurality of source pixels; using a distribution of reference codeword values of the plurality of reference pixels and a distribution of source codeword values of the plurality of source pixels, as represented in the selected color space, to generate one or more source-to-reference codeword mappings; based at least in part on the one or more source-to-reference codeword mappings, converting a specific source image originally captured by the source camera to a corresponding autoconformed source image conforming to a reference visual appearance associated with the one or more reference images originally captured by the reference camera.

[0223] EEE26. The medium of EEE25, wherein the reference camera differs from the source camera in one or more of: output image color spaces, output image dynamic ranges, output image color gamuts, output image bit depths, image signal processors, image signal processor configurations, device makers, or device models.

[0224] EEE27. A method comprising: determining one or more reference images and one or more source images in a working color space, wherein the one or more reference images are derived from a reference camera, whereinthe one or more source images are derived from a source camera different from the reference camera; deriving initial gain values between the one or more reference images and the one or more source images, wherein the initial gain values are determined based on reference codeword values of the one or more reference images and source codeword values of the one or more reference images; adjusting, based at least in part on one or more results of noise characterization performed with the one or more source images, the initial gain values into modified gain values; applying the modified gain values to the one or more source images to generate one or more leveled source images; based at least in part on one or more source-to-reference tone mappings, converting the one or more leveled source images to one or more tone-matched source images.

[0225] EEE28. The method of EEE27, wherein the working color space represents a linear color space; wherein the linear color space is one of: a linear RGB color space, a linear YUV color space, a linear YCbCr color space, or a different linear color space.

[0226] EEE29. The method of EEE27, wherein the working color space represents a nonlinear color space; wherein the nonlinear color space is one of: a perceptually quantized (PQ) color space, a hybrid log-gamma (HLG) color space, a nonlinear LMS color space, a nonlinear ICpCt color space, or a different nonlinear color space.

[0227] EEE30. The method of EEE27, wherein the initial gain values are determined as average ratios based on average reference codeword values of the one or more reference images and average source codeword values of the one or more reference images.

[0228] EEE31. The method of EEE27, further comprising: applying a specific optoelectronic transfer function (OETF) associated with the one or more reference images to the tone-matched source images to generate one or more auto-conformed source images conforming to a reference visual appearance associated with the reference camera.

[0229] EEE32. The method of EEE27, wherein the results of noise characterization performed with the one or more source images are at least in part represented by a noise metric computed using individual noise-to-signal ratios between standard deviations of individual image blocks and intensities of the individual image blocks; wherein the image blocks are partitioned from the one or more source images.

[0230] EEE33. The method of EEE32, wherein the individual image blocks partitioned from the one or more source images and used to compute the individual noise-to-signal ratios constitute a set of image blocks; wherein the set of image blocks excludes any image blocks, in the one or more source images, that contain visually perceptible edges detected with one or more edge detection methods.

[0231] EEE34. The method of EEE32, wherein the individual noise-to-signal ratios are computed separately for each color channel of the working color space in which the one or more source images are represented.

[0232] EEE35. The method of EEE32, wherein the noise metric is determined for a luminance channel of the working color space; wherein two additional noise metrics are determined for chrominance channels of the working color space; wherein a total noise is computed for the one or more source images from the noise metric and the two additional noise metric; wherein the total noise is used to modify the initial gain values into the modified gain values.

[0233] EEE36. The method of EEE35, wherein the total noise is used in an approximate inverse relationship to modify the initial gain values into the modified gain values.

[0234] EEE37. The method of EEE35, wherein the total noise is used in a spline relationship to modify the initial gain values into the modified gain values.

[0235] EEE38. The method of EEE35, wherein the total noise is used to apply maximum gain limits to the modified gain values.

[0236] EEE39. The method of EEE27, wherein each and every pair of the modified gain values for two different color channels maintain a first ratio value same as a second ratio value maintained by a corresponding pair of the initial gain values for the different color channels.

[0237] EEE40. A non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of the method recited in any of EEE27-EEE39.

[0238] EEE41. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in any of EEE27-EEE39.

Claims

CLAIMS1. A method comprising: determining one or more reference images and one or more source images in a working color space, wherein the one or more reference images are derived from a reference camera, wherein the one or more source images are derived from a source camera different from the reference camera; deriving initial gain values between the one or more reference images and the one or more source images, wherein the initial gain values are determined based on reference codeword values of the one or more reference images and source codeword values of the one or more reference images; adjusting, based at least in part on one or more results of noise characterization performed with the one or more source images, the initial gain values into modified gain values; applying the modified gain values to the one or more source images to generate one or more leveled source images; based at least in part on one or more source-to-reference tone mappings, converting the one or more leveled source images to one or more tone-matched source images.

2. The method of claim 1, wherein the working color space represents a linear color space; wherein the linear color space is one of: a linear RGB color space, a linear YUV color space, a linear YCbCr color space, or a different linear color space.

3. The method of claim 1, wherein the working color space represents a nonlinear color space; wherein the nonlinear color space is one of: a perceptually quantized (PQ) color space, a hybrid log-gamma (HLG) color space, a nonlinear LMS color space, a nonlinear ICpCt color space, or a different nonlinear color space.

4. The method of any preceding claim, wherein the initial gain values are determined as average ratios based on average reference codeword values of the one or more reference images and average source codeword values of the one or more reference images.

5. The method of any preceding claim, further comprising: applying a specific optoelectronic transfer function (OETF) associated with the one or more reference images to the tone-matched source images to generate one or more auto-conformed source images conforming to a reference visual appearance associated with the reference camera.

6. The method of any preceding claim, wherein the results of noise characterization performed with the one or more source images are at least in part represented by a noise metric computed using individual noise-to-signal ratios between standard deviations of individual image blocks and intensities of the individual image blocks; wherein the image blocks are partitioned from the one or more source images.

7. The method of claim 6, wherein the individual image blocks partitioned from the one or more source images and used to compute the individual noise-to-signal ratios constitute a set of image blocks; wherein the set of image blocks excludes any image blocks, in the one or more source images, that contain visually perceptible edges detected with one or more edge detection methods.

8. The method of claim 6 or 7, wherein the individual noise-to-signal ratios are computed separately for each color channel of the working color space in which the one or more source images are represented.

9. The method of any one of claims 6 to 8, wherein the noise metric is determined for aluminance channel of the working color space; wherein two additional noise metrics are determined for chrominance channels of the working color space; wherein a total noise is computed for the one or more source images from the noise metric and the two additional noise metric; wherein the total noise is used to modify the initial gain values into the modified gain values.

10. The method of claim 9, wherein the total noise is used in an approximate inverse relationship to modify the initial gain values into the modified gain values.

11. The method of claim 9, wherein the total noise is used in a spline relationship to modify the initial gain values into the modified gain values.

12. The method of claim 9, wherein the total noise is used to apply maximum gain limits to the modified gain values.

13. The method of any preceding claim, wherein each and every pair of the modified gain values for two different color channels maintain a first ratio value same as a second ratio value maintained by a corresponding pair of the initial gain values for the different color channels.

14. A method compri sing : determining one or more reference images in a working color space and one or more source images in the same working color space, wherein the one or more reference images are originally captured by a reference camera, wherein the one or more source images are originally captured by a source camera different from the reference camera; identifying one or more common image features between the one or more referenceimages and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, wherein the one or more common image features in the one or more source images include a plurality of source pixels; using a distribution of reference codeword values of the plurality of reference pixels and a distribution of source codeword values of the plurality of source pixels to generate one or more source-to-reference codeword mappings; based at least in part on the one or more source-to-reference codeword mappings, converting a specific source image originally captured by the source camera to a corresponding auto-conformed source image conforming to a reference visual appearance associated with the one or more reference images originally captured by the reference camera.

15. The method of claim 14, wherein the reference camera differs from the source camera in one or more of: output image color spaces, output image dynamic ranges, output image color gamuts, output image bit depths, image signal processors, device makers, or device models.

16. The method of claim 14 or 15, wherein the one or more reference images are originally captured by the reference camera in a hybrid log gamma (HLG) Rec. 2020 RGB color space; wherein the one or more reference images in the HLG Rec. 2020 RGB color space are linearized to be represented in a linearized Rec. 2020 RGB color space using a nonlinear inverse optical-to-electro transfer function (OETF).

17. The method of any one of claims 14 to 16, wherein the common image features are identified based at least in part on one or more projective transformations constructed from image features detected using a speeded-up-robust-features (SURF) algorithm.

18. The method of any one of claims 14 to 17, wherein the one or more source-to-reference codeword mappings are generated by one or more of: cumulative density function (CDF) matching operations, or one or more artificial neural networks.

19. The method of any one of claims 14 to 18, wherein at least one of the one or more source-to-reference codeword mappings is represented by one of: a mapping function, a spline function, a piecewise function, a piecewise polynomial function, or a lookup table.

20. The method of any one of claims 14 to 19, wherein the specific source image represents one of the one or more source images.

21. The method of any one of claims 14 to 20, wherein the one or more reference images and the one or more source images are respectively captured by the reference camera and the source camera from a common physical environment; wherein the reference camera and the source camera are situated at different spatial locations in the common physical environment.

22. The method of any one of claims 14 to 21, wherein the one or more source images converted to one or more corresponding auto-conformed source images that match the reference visual appearance associated with the one or more reference images originally captured by the reference camera; wherein at least a subset of the one or more corresponding auto-conformed source images is combined with at least a subset of the one or more reference images into a single sequence of time-consecutive images along a single playback timeline.

23. The method of any one of claims 14 to 22, wherein the corresponding auto-conformed image is generated by performing RGB leveling operations on an intermediate imagegenerated at least in part by applying the one or more source-to-reference codeword mappings to the specific source image.

24. The method of any one of claims 14 to 22, wherein the corresponding auto-conformed image is generated by performing one of saturation matching or saturation leveling operations on an intermediate image generated at least in part by applying the one or more source-to-reference codeword mappings to the specific source image.

25. A non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of the method recited in any of claims 1 to 24.

26. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in any of claims 1 to 24.