Image or video content auto-adaptation techniques

The automatic content adaptation system solves the problem of inconsistencies in tonal scale and color space representation of images/videos captured by different cameras, achieving automatic adjustment to a consistent visual appearance and reducing manual operation.

CN121666601APending Publication Date: 2026-03-13DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480049385.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2024-07-16
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing image/video editing tools struggle to effectively manage the tonal scales and color space representations of different cameras, especially when mixing clips in high dynamic range and standard dynamic range formats, leading to visual inconsistencies.

Method used

The automatic content adaptation system utilizes the image processing pipeline of the reference camera to automatically adjust the image/video content captured by different cameras to match the visual appearance of the reference camera. This includes operations such as color space transformation, noise characteristic description and noise adjustment, and tone matching.

Benefits of technology

It enables automatic adaptation of images/videos captured by different cameras to a consistent visual appearance, reducing manual intervention and improving the consistency of visual effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121666601A_ABST
    Figure CN121666601A_ABST
Patent Text Reader

Abstract

A reference image and a source image in a working color space are determined. The reference image originates from a reference camera. The source image originates from a source camera. An initial gain value between the reference image and the source image is derived based on the reference image and a reference codeword value of the reference image. The initial gain value is adjusted to a modified gain value based on a noise feature description result performed with the source image. The modified gain value is applied to the source image to generate a leveled source image. Based on the source-to-reference tone mapping, the leveled source image is converted into a tone-matched source image.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This patent application claims the benefit of U.S. Provisional Application No. 63 / 515,642, filed July 26, 2023; European Patent Application No. 23190847.6, filed August 10, 2023; and U.S. Provisional Application No. 63 / 568,340, filed March 21, 2024, the entire contents of each of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates in its entirety to an image / video processing method, and in particular to an automatic image / video content adaptation technique. Background Technology

[0003] In recent years, prosumer content creation has surged, particularly among social media influencers. In these cases, multiple consumer-grade cameras can be used to capture content. Unlike professional workflows that work with raw image sequences, this modality involves extracting, combining, editing, and grading processed bitstreams, each of which may have already passed through an image processing pipeline. For content creation involving multiple segments, each camera's image signal processing pipeline will perceptibly render the scene differently—for example, one camera might choose to apply significantly more saturation and / or contrast than another. Editing / grading tools (such as commercially available Resolve, Final Cut Pro, or Adobe Premiere) allow the import of some common standard dynamic range (SDR) formats into their platforms, but these tools lack the ability to effectively manage high dynamic range formats with different tonal scales and color space representations specific to the underlying camera. Furthermore, this problem becomes even more complex and visually worse when editing a blend of SDR and high dynamic range (HDR) clips from different cameras into a combined video clip.

[0004] The methods described in this section are permissible but need not be methods previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any issues concerning one or more methods should not be considered to be in any prior art based on this section. Attached Figure Description

[0005] Embodiments of the invention are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the drawings: Figure 1 The diagram illustrates an example of an automatic content adaptation system; Figure 2A and Figure 2B The illustrations show example reference source images and non-reference source images; Figure 2C The illustration shows an example reference image generated by a reference camera and an auto-leveled non-reference source image; Figure 3A The illustration shows the instance space correlation or correspondence between common image features of two images; Figure 3B The illustration shows two example binary images constructed by performing a spatial transformation between two images; Figure 3C The example codeword mapping curve is illustrated. Figure 3D The illustration shows a non-reference image of the example pre-saturation leveling and the tone mapping after saturation leveling; Figures 4A to 4F The example process flow is illustrated; Figure 5 The illustration shows an example hardware platform on which a computer or computing device as described herein can be implemented; Figures 6A to 6C The example process flow is illustrated; Figures 7A to 7C The illustrations show example reference and source images, as well as the source image for tone matching; and Figure 8 The illustration shows an example spline relation used for noise-based control of white point gain. Detailed Implementation

[0006] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent that this disclosure may be practiced without these specific details. In other instances, well-known structures and devices have not been described in detail in order to avoid unnecessarily obscuring, obscuring, or confusing this disclosure.

[0007] Overview Input image / video content previously acquired or formed using different cameras or image acquisition / forming devices may need to be processed or converted into corresponding output image / video content with a relatively consistent or adapted visual appearance, for example, in a combination of video clips or image sequences on a common timeline. The relatively consistent or adapted visual appearance of the output image / video content may, but is not necessarily limited to, a Dolby Vision visual appearance with relatively high dynamic range and wide color gamut.

[0008] In this hybrid content scenario, creators (potentially consisting of one or more content creators or authors) can select a reference camera or "protagonist" camera from multiple different cameras. Some or all of these different cameras can be used to acquire or form the input image / video content. The selected reference camera generates image / video content with a reference visual appearance (also known as a reference camera style) that is specific to the reference camera having its particular image acquisition / processing settings.

[0009] After input image / video content with different visual appearances (or different camera styles) is formed using different image pipelines implemented by different cameras, additional image processing / conversion operations (referred to as auto-adaptation or auto-fitting operations) can be performed to modify or convert the input image / video content into corresponding output image / video content with the same reference visual appearance of a selected reference camera.

[0010] The input image / video content can be generated using a variety of cameras that implement different image forming pipelines and / or different image signal processors. Example cameras may include, but are not limited to, any, some, or all of the following: Canon C300, Apple iPhone 14, GoPro Hero 8, etc. The input image / video content to be used for auto-adaptation operations can be initially acquired or captured from the same and / or different physical environments or visual scenes where cameras are deployed.

[0011] The automatically adapted input image / video content can be further processed or incorporated into a combined editing or playback timeline, having the same relatively consistent reference camera visual appearance. It should be noted that some or all of the techniques described herein can be implemented or applied regardless of whether a selected reference camera can be deployed in the physical environment or visual scene to capture or generate any of the input image / video content.

[0012] Therefore, an automatic content adaptation system implementing some or all of the techniques described herein can automatically adapt source image / video content acquired using different cameras to a reference visual appearance associated with image / video content acquired using a reference camera, with little or no manual input or manipulation required to perform the automatic adaptation operation. The system achieves a relatively satisfactory match to the reference visual appearance, as if all automatically adapted source image / video content originated from the same reference camera or its corresponding image formation pipeline. The automatic content adaptation operation described herein can be performed without much or no subsequent manual or user modification of the final output image / video content, which has already achieved a visual appearance that is relatively consistent with the visual appearance of the selected reference camera.

[0013] In some operational scenarios, tone matching or auto-adaptation of the source content from a source camera with the appearance of a reference camera may involve scene (partial) changes from dark to light. To reduce noise visibility based on noise particles present or inherent in the amplified source content, additional noise characterization and noise-based adjustments can be performed to better control the amount of codeword or white point gain to be applied to the white (point) leveling operation. More specifically, prior to tone matching, noise characterization can be used to generate an estimate of the (overall) noise in the source content. Noise can be calculated using noise metrics generated for image patches within the source content. Noise adjustments can be used to adjust or limit the amount of codeword or white point gain. Therefore, noise visibility in the content to be tone matched or auto-adapted can be significantly reduced or prevented.

[0014] The example embodiments described herein relate to automatically adapting image content to a reference visual appearance of a reference camera. One or more reference images in a working color space and one or more source images in a selected working color space are determined. The one or more reference images are initially captured by a reference camera. The one or more source images are initially captured by source images different from those captured by the reference camera. One or more common image features are identified between the one or more reference images and the one or more source images. The one or more common image features in the one or more reference images include a plurality of reference pixels. The one or more common image features in the one or more source images include a plurality of source pixels. The distribution of reference codeword values ​​of the plurality of reference pixels and the distribution of source codeword values ​​of the plurality of source pixels are determined to generate one or more source-to-reference codeword maps. Based at least in part on the one or more source-to-reference codeword maps, a specific source image initially captured by a source camera is transformed into a corresponding automatically adapted source image adapted to the reference visual appearance associated with the one or more reference images initially captured by the reference camera.

[0015] The example embodiments described herein relate to automatically adapting image content to a reference visual appearance of a reference camera. One or more reference images and one or more source images are determined in a working color space. The one or more reference images originate from a reference camera. The one or more source images originate from a source camera other than the reference camera. An initial gain value is derived between the one or more reference images and the one or more source images. The initial gain value is determined based on reference codeword values ​​of the one or more reference images and source codeword values ​​of the one or more source images. The initial gain value is adjusted to a modified gain value, at least in part based on the results of one or more noise characteristic descriptions performed using the one or more source images. The modified gain value is applied to the one or more source images to generate one or more leveled source images. The one or more leveled source images are transformed into one or more tone-matched source images, at least in part based on one or more source-to-reference tone mappings.

[0016] Various modifications to the preferred embodiments, general principles, and features described herein will be readily apparent to those skilled in the art. Therefore, this disclosure is not intended to be limited to the illustrated embodiments, but rather to be accorded the maximum scope consistent with the principles and features described herein.

[0017] Multi-stage processing Figure 1 An example of an automatic content adaptation system 100 implementing multi-stage image processing operations is illustrated. The automatic content adaptation system 100 can be implemented using one or more computing devices and can be used to generate output image / video content with an adapted visual appearance associated with a reference camera.

[0018] In the first stage or processing frame 102 of the automatic content adaptation system 100, the creator can provide user input to identify or specify a reference camera and an associated workspace, which includes, but is not limited to, a reference color space (e.g., HDR, SDR, etc.). For example, but not limited to, a first camera (e.g., iPhone 14, etc.) can be selected as the reference camera, and the associated workspace can be selected as a Hybrid Log Gamma (HLG, Rec.2020) space.

[0019] In the second stage or processing block 104 of the automatic content adaptation system 100, a specific video format of the input image / video content (e.g., video clips obtained from different cameras) is interpreted or used to retrieve source image / video frames or source image data therein from the input image / video content. If applicable, the source image data in the source video frames of each input image / video content can then be linearized. The video format can refer to a specific video file / container / signal (or its type) carrying an image as described herein (e.g., generated as output by the camera's image signal processor (ISP)) and image metadata associated with the image. Example image metadata may include, but is not limited to, information about some or all of the following: the color space used by the camera, dynamic range, color gamut, bit depth, finite code space, full-range code space, applicable video / image specifications / standards, exposure, saturation, contrast, color compensation vector, or scaling factor.

[0020] In the third stage or processing block 106 of the automatic content adaptation system 100, a color space transformation can be performed or applied to convert a (linearized) source image / video frame represented in a (linearized) source color space (e.g., a reference color space different from the one associated with the working space) into a (linearized) source image / video frame represented in a corresponding color space transformation in the reference color space.

[0021] In the fourth stage or processing box 108 of the automatic content adaptation system 100, content matching or adaptation operations are applied to the source image / video frame (color space converted, linearized) to obtain a corresponding output source image / video frame with a reference camera visual appearance.

[0022] The first three stages may include performing operations to parse or interpret the input bitstream color and tone scale representation in the input image / video content. In operational scenarios where source frames in the input image / video content are represented in a non-linear color space different from (e.g., linearized) reference color space, these source frames are converted from non-linear color spaces (such as YUV) to (color space-converted, linearized) source image / video frames represented in (linearized) reference color spaces (such as linear RGB color space or video formats).

[0023] Color space conversion or transformation can be performed, at least in part, based on a source or working color space or colorimetry, which is associated with the source or working color space respectively. In some feasible scenarios, one or more transformation matrices (such as those associated with YUV to RGB conversion) can be applied to perform color space conversion or transformation.

[0024] Figure 2A and Figure 2B The illustration shows an example reference and a (linearized, color space-converted) source image represented in the HLGRec 2020 working (color) space after the first three stages or steps of the previously mentioned image processing operation. Figure 2A The reference image in the image can be initially a first video clip captured by a reference camera (such as an iPhone 14) and represented in the HLG Rec 2020 workspace, while Figure 2A The source image in the process can be initially a second video clip captured by a non-reference or source camera (such as a Sony A7) and represented in a non-reference workspace (such as Slog3 / SGamut3), which is then transformed into the HLG / Rec 2020 workspace by the first three stages of the previously mentioned image processing operations.

[0025] As shown in the image, not only are the exposures different between the two cameras, but the color balance, saturation, and overall contrast also differ. These easily perceptible differences will cause users to expend considerable effort to achieve a satisfactory fit between the images / video content captured using both cameras. For example, while current color editing / correction systems do have some ability to automatically level content, this often results in a blended effect. These color editing / grading systems or platforms (such as Adobe Premiere, Blackmagic Resolve, and Apple Final Cut Pro) still produce noticeable rendering mismatches.

[0026] This problem is further complicated by the fact that, in practical applications, the cameras to be adapted may not capture the same scene at the same time. Therefore, the two video feeds may have little or no common elements available for adaptation.

[0027] Figure 2C The illustration shows an example reference image generated by a reference camera (e.g., an iPhone 14 in this example, etc.) and an auto-leveled source (or non-reference) image generated by applying Adobe Premiere's auto-leveling feature to a non-reference video frame or image in a Sony A7 Slog3 video clip. Although the image generated by Adobe Premiere using the auto-leveling feature (its...) Figure 2C (Right side of the middle) shows relative to Figure 2B The image has been improved, but the exposure, saturation, and contrast are still significantly different from the reference image obtained using the iPhone 14 (which... Figure 2C (Left side of the middle). Therefore, users will still need to manually intervene and apply a relatively large amount of color grading to this segment in order to achieve a visual appearance more similar to that of the video clip captured using the iPhone 14. This could potentially involve relatively complex and numerous secondary grading operations.

[0028] Compared to existing systems / platforms that tend to produce noticeable mismatches, automated content adaptation systems as described herein can be implemented to automatically perform content adaptation operations on image / video content acquired using different cameras, thereby achieving a relatively satisfactory match in terms of the relatively consistent visual appearance of the automatically adapted image / video content, as if the automatically adapted image / video content originated from the same reference camera. These content adaptation operations can be performed with little or no user modification and / or little or no manual intervention. In some feasible scenarios, these operations can be combined with the previously mentioned fourth stage (or step).

[0029] More specifically, after the initial extraction of the input image / video content acquired from the non-reference or source camera in the first three stages (or steps) mentioned above, the automatic content adaptation system may apply image processing operations in the fourth stage (or step) mentioned above to automatically adapt the input image / video content into a corresponding adapted image / video content, which has the same or similar visual appearance as the reference visual appearance of the reference camera in the (reference) image / video content acquired using the reference camera.

[0030] Image processing operations performed by the automatic content adaptation system can include tone editing / correction and / or color editing / correction of the intermediate outputs from the first three stages (or steps). Tone editing / correction and / or color editing / correction of content acquired by each of the different cameras can be optimized for that camera. Therefore, the system can perform robust image processing operations on content generated by different cameras, optimizing and adjusting for differences in content generated by these different cameras relative to a reference camera, thereby achieving or generating a relatively consistent visual appearance associated with content generated by the reference camera. Thus, the system can help effectively and efficiently solve real-world applications where different cameras to be adapted may not capture the same scene at the same time.

[0031] While multiple image or video feeds from input image / video content captured using different (sample or source) cameras may have little or no common elements available for adaptation, all of these image or video feeds can be automatically adapted by the system to generate adapted image / video feeds with a relatively consistent visual appearance associated with the content generated by the reference camera, as if these adapted image / video feeds were captured by the reference camera rather than by these different (sample or source) cameras that captured the input image / video content in a real-world application or real-world scene.

[0032] Automatic content adaptation Figure 4A The illustration shows an example process flow for automatic content adaptation. This process flow can be implemented or executed at least in part using one or more computing devices, including but not limited to automatic content adaptation systems, post-camera image / video processing systems, computing devices installed on, but not limited to, image acquisition systems, mobile computing devices, automatic adaptation tools on non-mobile computing devices, cameras, etc.

[0033] Box 402 includes reading or receiving reference image / video content captured using a reference camera, and reading or receiving non-reference image / video content captured using a non-reference camera. The reference image / video content may be referred to as a reference video stream in an illustrated but not limited manner. The reference camera may be referred to as the main camera. The non-reference camera may be referred to as the target, source, or sample camera. The non-reference image / video content may be referred to as the target, source, or sample video stream.

[0034] Box 404 includes linearizing the reference and non-reference video streams according to the corresponding (input) inverse photoelectric conversion functions (OETFs) of the reference and non-reference (or sample / source) video streams. These inverse OETFs may be individually and specifically identified, in part or in whole, based on user input and / or image metadata received or decoded from the received reference and sample video streams.

[0035] Box 406 includes, if applicable (e.g., where the reference and sample cameras are different camera models and / or brands, etc.), applying a color correction matrix to the reference and sample video streams to convert video frames or images in these streams into the working (color) space of the reference camera. This working space may be specifically identified, in part or in whole, based on user input and / or image metadata received or decoded from the received reference video stream.

[0036] Box 408 includes identifying common (visual or image) features between the reference video stream and the sample video stream (or video frames or images therein).

[0037] Box 410 includes determining a codeword or tone mapping curve that maps the luminance or luminance codeword of a pixel in a common feature in a sample video stream to the luminance or luminance codeword of a pixel in a corresponding common feature in a reference video stream.

[0038] Box 412 includes determining a color balance matrix that maps the color or chromaticity of pixels in common features in the sample video stream to the color or chromaticity of pixels in corresponding common features in the reference video stream.

[0039] Box 414 includes applying codeword / tone mapping and color balance matrices to the linearized sample video stream to generate an adjusted sample video stream.

[0040] 416 includes applying a specific OETF to the adjusted sample video stream to generate a sample video stream corresponding to the automatic content adaptation of the workspace of the reference camera.

[0041] Figure 4B The diagram illustrates the method used to determine (e.g., Figure 4A Example process flow for hue and chroma mapping curves (e.g., etc.). This process flow can be implemented or executed at least in part using one or more computing devices, including but not limited to automatic content adaptation systems, such as rear camera image / video processing systems installed on computing devices, automatic adaptation tools including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0042] like Figure 4A As shown, before determining the hue and chroma mapping curves, common features between the reference video stream and the sample video stream can be identified first.

[0043] Box 422 includes a reference cumulative distribution function (CDF) for calculating pixel brightness in common features of a reference video stream and a (non-reference) sample or source CDF for pixel brightness in common features of a sample or source video stream.

[0044] Box 424 includes generating a codeword / tone mapping curve that maps the luminance (or luminance value) in the sample or source image data of the sample or source video stream to the automatically adapted luminance (or luminance value) of the sample or source video stream.

[0045] Box 426 includes, for each chroma channel in the chroma channel, calculating the reference cumulative distribution function (CDF) of pixel brightness in common features in the reference video stream and the (non-reference) sample or source CDF of pixel brightness in common features in the sample or source video stream.

[0046] Box 428 includes generating a chroma mapping curve for each chroma channel in the chroma channel, which maps the chroma (or chroma values) in the sample or source image data of the sample or source video stream to the automatically adapted chroma (or chroma values) of the sample or source video stream.

[0047] Figure 4C The illustration shows an example process flow for calculating spline curves used to perform hue and chroma mapping. This process flow can be implemented or performed, at least in part, using one or more computing devices, including but not limited to automatic content adaptation systems, such as rear-camera image / video processing systems installed on the computing devices, automatic adaptation tools including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0048] like Figure 4A and Figure 4B As shown, the hue and chroma mapping curves can be determined by the common features first identified between the reference video stream and the sample video stream.

[0049] Box 442 includes calculating a first spectrometer curve, at least in part, based on tone mapping curves and / or CDFs, which maps the luminance (or luminance values) in the sample or source image data of the sample or source video stream to automatically adapted luminance (or luminance values) of a reference in the sample or source video stream. Alternatively, optionally, or alternatively, the first spectrometer curve may be generated directly based on a histogram of luminance values ​​collected from pixels sharing common image features in both reference and non-reference images. The first spectrometer curve may consist of a set of nodes (e.g., 11, etc.) or B-spline functions connected at those nodes.

[0050] Box 444 includes a second spline curve computed at least in part based on tone mapping curves and / or CDFs, which maps chroma (or chroma values) in the sample or source image data of the sample or source video stream to automatically adapted chroma (or chroma values) of a reference in the sample or source video stream. Alternatively, the second spline curve may be generated directly based on a histogram of chroma values ​​collected from pixels sharing common image features in both reference and non-reference images. The second spline curve may consist of a set of nodes (e.g., 11, etc.) or B-spline functions connected at those nodes.

[0051] Figure 4D The illustration depicts an example process flow for training and applying an artificial neural network, such as a convolutional neural network (CNN), which predicts the tone (mapping) curves and color balance matrix between a reference camera and a sample camera. This process flow can be implemented or executed, at least in part, using one or more computing devices, including but not limited to automatic content adaptation systems, post-camera image / video processing systems, computing devices mounted on, but not limited to, image acquisition systems, mobile computing devices, automatic adaptation tools on non-mobile computing devices, cameras, etc.

[0052] Box 462 includes training a CNN in an offline training process / stage using training reference and sample (or non-reference) image / video content and a ground truth (e.g., known codeword / tone mapping curves and / or known color balance matrices, etc.) to learn or optimize operable parameters for predicting codeword / tone mapping curves and color balance matrices with relatively small prediction errors. Input features can be extracted from the training reference and sample image / video content and provided as input to the CNN for predicting or estimating tone mapping curves and / or color balance matrices with minimum prediction errors relative to the ground truth.

[0053] Box 464 includes, during the running application process / phase, applying a trained CNN to video / image data in a non-trained sample video stream to predict specific codeword / tone mapping curves and specific color balance matrices using optimized operational parameters of the CNN that has been trained or optimized using previously trained data. These codeword / tone mapping curves and specific color balance matrices can be used to map video / image data in a non-trained sample video stream to automatically adapted video / image data in an automatically adapted sample video stream.

[0054] CDF Matching Overview As pointed out, Figure 4BThe CDF matching operation in the process flow can be applied to the luminance and / or chrominance values ​​of pixels in the common image features of the reference image / video content and the non-reference (e.g., source, sample, etc.) image / video content, in order to generate luminance (or hue) and / or chrominance mapping curves that can be used to map / or transform the non-reference image / video content into a reference visual appearance associated with the reference image / video content.

[0055] In some operational scenarios, this image / video content (such as images or video clips) can be captured by a reference camera and non-reference cameras relative to the same physical or visual scene, but from correspondingly different spatial camera positions. For example, a reference image can be initially captured by a reference camera as... Figure 2A The reference video frame shown in the iPhone 14 video clip, rather than the reference image, can be initially captured by the non-reference camera as such. Figure 2B The image shows a non-reference video frame in a Sony A7 video clip. Reference and non-reference images can contain pixels with common image features because they were captured from the same environment or scene (e.g., concurrently, at similar times, under similar lighting conditions, etc.). However, common image features may be geometrically or spatially displaced and therefore located in corresponding spatial regions within the reference and non-reference images.

[0056] Reference and non-reference images (or their pixel / codeword values) can first be linearized to generate linearized reference and non-reference images. The reference and non-reference brightness images can be formed from the linearized reference image (or video frame) and the linearized non-reference image (or video frame), respectively. For example, the linearized reference and non-reference images (or video frames) can be represented in a linearized RGB color space. An RGB-to-YCbCr (color space) transformation matrix can be applied to the linearized RGB values ​​in the linearized reference image (or video frame) and the linearized non-reference image (or video frame) to generate the reference brightness image and non-reference image, respectively.

[0057] In some operational scenarios, the luminance values ​​of pixels in common image features of reference and non-reference images (or their corresponding luminance images) can be used to establish correspondences between corresponding spatial regions of the common image features (e.g., feature points, corner points, eye corners, etc.). These correspondences can be used to generate or estimate similar spatial transformations (or transitions). Spatial transformations can operate in the luminance domain. For example, a spatial transformation can be used to map the location of a non-reference luminance value in an initial non-reference (luminance) image from a non-reference image to the corresponding location of a reference luminance value in an initial non-reference (luminance) image from a reference image.

[0058] In some feasible scenarios, reference and non-reference pixels that belong to common image features in reference and non-reference images can be based at least in part on spatial transformations and / or reference and non-reference brightness images respectively derived from reference and non-reference brightness images.

[0059] For example, one or more computer vision techniques can be used to match, identify, or recognize common image features in both reference and non-reference images. These techniques include, but are not limited to, oriented FAST and rotated BRIEF (ORB) algorithms, speeded-up-robust-features (SURF) algorithms, scale-invariant-feature-transform (SIFT) algorithms, computer-implemented feature keypoint detection / matching algorithms, and computer-implemented feature detection / matching algorithms. These common image features in the reference and non-reference brightness images (or the reference and non-reference brightness video frames that generate the brightness images) depict the same visual objects, people, etc., even though these images originate from reference and non-reference cameras at different spatial locations within the same physical environment or visual scene.

[0060] The common image features in the reference and non-reference brightness images, and the spatial displacement (or difference) information of the pixels therein, can be used to estimate the projection transformation from a first image in the reference and non-reference brightness images to a second image that differs from these images. The projection transformation can be used to project or spatially transform the pixels (or pixel positions) of the common image features in the first image (one of the reference and non-reference brightness images) to the corresponding pixels (or pixel positions) of the common image features in the second image (the other of the reference and non-reference brightness images).

[0061] Histograms of luminance values ​​in pixels sharing common image features in both reference and non-reference luminance images can be subsequently calculated. These histograms, or the luminance value ranges therein, can be used by a CDF matching operation to generate luminance mapping curves (which may be referred to as hue curves). These luminance mapping curves can be used to map or convert the luminance values ​​of non-reference video frames to mapped / converted luminance values ​​in automatically adapted video frames corresponding to the non-reference video frames. In some feasible scenarios, some or all of the luminance mapping curves can be represented as lookup tables.

[0062] Example codeword / tone mapping operations are described in U.S. Provisional Patent Application No. 62 / 356,087, filed June 29, 2016, entitled “EFFICIENT HISTOGRAM-BASED LUMA LOOK MATCHING”, by Harshad Kadu et al. Example CDF operations are described in U.S. Provisional Patent Application No. 62 / 404,307, filed October 5, 2016, entitled “INVERSELUMA / CHROMA MAPPINGS WITH HISTOGRAM TRANSFER AND APPROXIMATION”, by Bihan Wen et al. The aforementioned patent applications are hereby incorporated by reference as if their entirety were shown herein.

[0063] Example CDF matching In an illustrative but unrestricted manner, although reference and non-reference cameras will capture visual elements (such as visual objects, people, backgrounds, etc.) of the same physical environment or visual scene in both reference and non-reference images, each of these cameras may have different viewpoints, different lens distortions, etc. Therefore, a visual object may appear in one image but not in another.

[0064] To identify some or all of a specific subgroup of common or associated pixels in each of the reference and non-reference images, common image features between the two cameras (e.g., specific textures, building corners, lower left corner of a rear window, eyes, nose, lips, chin, hands, tables, cars, etc.) can be identified, for example, using the SURF method / algorithm (or another computer vision technique such as a computer-implemented feature (keypoint) matching / detection algorithm). Figure 3A The illustration shows an example of correlation or correspondence (e.g., spatial displacement indicated by a line) between common image features of two images from the same physical environment or visual scene, identified using computer vision techniques including but not limited to the SURF method / algorithm.

[0065] In some feasible scenarios, each of the reference and non-reference images can be linearized and represented in a workspace or color space with the same chromaticity (e.g., RGB). The linearized reference and non-reference images in the workspace can then be converted into corresponding (reference and non-reference) single-channel luminance images.

[0066] The SURF method / algorithm can then be used to detect common image (or visual) features in each of the reference and non-reference brightness images. This involves geometric / spatial displacements (e.g., caused by specific pixel locations, corners of the eyes, lips, etc.) between multiple pairs of specific spatial locations (e.g., specific pixel locations, corners of the eyes, lips, etc.) of common image features in the reference and non-reference brightness images. Figure 3A The lines in the diagram indicate that common or related image features can be estimated.

[0067] Based on the geometric / spatial displacements between multiple pairs of specific spatial locations detected by the SURF method / algorithm, one or more projection (or spatial) transformations can be estimated and used to map one or more projection (or spatial) transformations from the coordinate system of a non-reference image to the coordinate system of a reference image, or vice versa.

[0068] In some operational scenarios, both forward and backward projection (or spatial) transformations can be estimated and subsequently used to identify pixels or valid pixel maps in common image features in both the reference and non-reference images. The source pixel coordinates of the source pixels in the non-reference image are passed through the forward transformation to generate or identify corresponding reference pixels in the reference image with corresponding reference pixel coordinates. Any non-reference pixels that can be excluded from the common image features (as “invalid” source pixels) are forward transformed to reference pixel coordinates located outside the reference video frame used to contain the reference image.

[0069] For reference pixels in a reference image, the same or opposite operations can be performed using an inverse transformation to exclude any reference pixels (from common image features) that are inversely transformed to non-reference pixel coordinates located outside of non-reference video frames containing non-reference images.

[0070] Figure 3B The illustration shows example binary images constructed by performing forward and inverse transformations on non-reference and reference images. In the first binary image on the left, white pixels indicate all valid (not excluded as "invalid") reference pixels in the common image features of the reference image, while black pixels indicate all other pixels in the reference image that do not belong to the common image features. Similarly, in the second binary image on the right, white pixels indicate all valid (not excluded as "invalid") reference pixels in the common image features of the non-reference image, while black pixels indicate all other pixels in the non-reference image that do not belong to the common image features.

[0071] The entire luminance (codeword) range or space used to represent non-reference luminance values ​​in a non-reference image can be (e.g., equally) divided into a first plurality of luminance (codeword) sub-ranges. The counts of ("valid" or not excluded as "invalid") pixels (or the total number of pixels in the common image features of the non-reference image) can be stored in corresponding intervals of a plurality of intervals of a non-reference histogram, these pixel counts having non-reference luminance values ​​in each of the first plurality of luminance sub-ranges.

[0072] Similarly, the entire luminance (codeword) range or space used to represent the reference luminance value can be (e.g., equally) divided into a second plurality of luminance (codeword) sub-ranges. The count of ("valid" or not excluded as "invalid") pixels (or the total number of pixels in the common image features of the reference image) can be stored in corresponding intervals of a plurality of intervals of the reference histogram, these pixel counts having non-reference luminance values ​​in each of the second plurality of luminance sub-ranges.

[0073] Since both reference and non-reference histograms are constructed based on the same and related distribution of pixels with common image features identified in reference and non-reference images, these histograms can be used to calculate or establish the correlation and correspondence between reference pixel or codeword values ​​and non-reference pixel or codeword values ​​represented in reference and non-reference pixels with common image features.

[0074] For example, the reference cumulative density function indicates as ,in The reference pixel or codeword value (e.g., a reference luminance value representing a reference luminance subrange) indicates the interval of the reference histogram. The reference cumulative density function can be obtained from the reference histogram by simply applying the cumulative sum along the direction of increasing reference pixel or codeword value to the interval of the reference histogram.

[0075] Similarly, the non-reference cumulative density function indicates as ,in The nonreference pixel or codeword value (e.g., nonreference luminance value representing a nonreference luminance subrange) indicates the interval of the nonreference histogram. The nonreference cumulative density function can be obtained from the nonreference histogram by simply applying the cumulative sum along the direction of increasing nonreference pixel or codeword value to the interval of the nonreference histogram.

[0076] Here, the subscripts S and R refer to the non-reference (or source) and reference pixel or codeword spaces, respectively. Non-reference pixel or codeword value and reference pixel or codeword value They can represent non-reference (or source) and reference pixel or codeword ranges / spaces respectively, and are indices of non-reference (or source) and reference pixel or codeword ranges / spaces (e.g., indexed by i or j).

[0077] For a given non-reference pixel or codeword value The corresponding reference pixel or codeword value This can be determined by imposing or using the following equal conditions: (1) In an illustrated but unrestricted manner, and The correspondence or mapping can be implemented based on the expression (1) above using simple linear interpolation. For example, It can be between and Between. The corresponding reference pixel or codeword value. It can be based on and arrive Applying simple linear interpolation to the corresponding distance or difference and To obtain.

[0078] Therefore, the reference and non-reference CDFs derived from the reference and non-reference histograms can be used to generate or derive non-reference image or codeword values. Mapping or transforming to its corresponding reference pixel or codeword value The brightness (or hue) and / or chromaticity mapping curve or LUT (such as...) Figure 3C shown).

[0079] Additionally, optionally or alternatively, in some feasible scenarios, the reference and non-reference CDFs derived from the reference and non-reference histograms can also be used to generate or derive reference image or codeword values. Mapping or transforming to its corresponding non-reference pixel or codeword value The inverse luminance (or hue) and / or chroma mapping curve or LUT.

[0080] In some operational scenarios, (forward and / or inverse) luminance and / or chrominance mapping curves or LUTs can be represented or approximated by a set of B-spline curves with a relatively small number of nodes or vertices (e.g., 11, etc.). This can help achieve the following benefits: reducing the overall complexity of transforming an image from one visual appearance to another (e.g., a reference visual appearance, etc.), and potentially producing a smoother function due to the continuity nature of B-spline functions.

[0081] In some feasible scenarios, for a temporally continuous sequence of non-reference images (e.g., belonging to the same video clip, scene, group of pictures, or GOP), the luminance and / or chrominance mapping curves can be temporally stable across multiple time points or time periods covered by the image sequence. This helps to handle or avoid errors in specific frames relatively properly in feature detection and subsequent projection transformations. For example, a relatively constant or fixed luminance transformation or mapping curve can be used throughout the image sequence. Additionally, optionally or alternatively, when applying mapping curves as described herein across multiple temporally continuous frames in the sequence, a temporal low-pass filter can be applied.

[0082] Additionally, optionally or alternatively, as a substitute or supplement to using effective pixel maps of a pair of reference and non-reference images (for a specific frame or a specific frame pair) to construct histograms and mapping curves for that pair of reference and non-reference images, in some feasible scenarios, globally effective pixel maps of multiple pairs of reference and non-reference images can be used to construct histograms and mapping curves for some and all pairs of reference and non-reference images as described herein.

[0083] Typography and RGB leveling Subsequent image processing operations (such as codeword or RGB leveling operations) can be performed on the tone-mapped non-reference image to generate a codeword- or RGB-leveled tone-mapped non-reference image. Codeword or RGB leveling operations can represent relatively small adjustments to image or codeword values. Corresponding gain values ​​in non-RGB or RGB channels can be used to adjust the corresponding non-RGB or RGB pixel or codeword values ​​in the tone-mapped non-reference image to generate a codeword- or RGB-leveled tone-mapped non-reference image. In terms of brightness and color visual appearance, leveling operations, along with tone mapping, help provide or achieve a relatively high quality match between the leveled tone-mapped non-reference image (initially derived from a Sony A7 clip in this example) and the reference image (initially derived from an iPhone 14 clip in this example). As shown in the figure, it is clear that the automatic adaptation method described in this paper produces a better correspondence between the reference image and the sample image compared to other methods.

[0084] By way of example, but not limitation, a non-reference (or sample) image is represented in a linearized RGB color space, and the tone curve or LUT generated by the CDF matching operation can first be applied equally to the component pixel values ​​in the RGB channels of the non-reference (or sample) image.

[0085] After this hue (or luminance) domain mapping, RGB leveling can be performed to match the visual appearance relative to a reference camera. For example, the RGB pixel or codeword values ​​in individual RGB channels of the source image (or the non-reference image in the hue mapping) can be adjusted by the corresponding ratio of the reference to the source's average value for each channel. For example, the values ​​in the red channel indicated by... The component pixels or codewords are used to generate the indicator in the red channel. The leveled or adjusted component pixels or codewords are as follows: (2) in It is an image X The average value of all valid pixels; X These respectively indicate the source image or the reference image. S or RIn some feasible scenarios, a clipping operation can be performed after scaling or adjustments using a reference-to-source ratio to constrain the adjustments to a specific range of pixel or codeword values. Additionally, optionally or alternatively, some or all of the codeword or RGB leveling operations described herein can be applied using a white balance matrix constructed, at least in part, based on reference-to-source ratios calculated for different color channels in the working color space.

[0086] Saturation leveling In some feasible scenarios, the CDF method performs well in matching the luminance channels or appearance of a sample or non-reference image with those of a reference image, but differences in color intensity may still exist. To address this issue, in the linear domain, after applying tone mapping and (e.g., RGB, etc.) leveling operations (such as gain adjustment as shown in expression (2) above), a saturation metric can be computed for either the sample or reference image to apply saturation leveling.

[0087] In various operational scenarios, different methods can be used or implemented to calculate saturation metrics for applying saturation leveling operations. As a first example, the RGB reference or non-reference image to which saturation leveling is to be applied can be converted from the RGB color space to CIE 1976 L. a b Color space. As a second example, to apply saturation leveling to an RGB reference or non-reference image, it can be converted from the RGB color space to the YCbCr color space (or mode).

[0088] Instructions are The saturation measure can then be used using CIE 1976 L a b The calculation is performed using the color (or non-luminance) channels in the color space or YCbCr color space, as follows: (3) The saturation metric can be used to calculate the overall saturation ratio of the average saturation of the reference image to the average saturation of the sample image for each pair of reference and sample images. The sample image can then be (saturation) leveled by applying the (overall saturation) ratio in a manner similar to the above expression (2) (e.g., using a saturation balance matrix).

[0089] Figure 3D The illustrations show pre-saturated leveled and saturated leveled tone maps (e.g., using tone curves or LUTs generated based on CDF matching) on ​​the left and right sides, respectively, of non-reference images. The initial non-reference images that produced these images could have been captured by a Sony A7 in a video clip.

[0090] In some feasible scenarios, saturation matching can be performed as an alternative to or complement to saturation leveling by generating a source-to-reference saturation mapping based on the CDF of the saturation metric histogram. For example, operations similar to hue (or lightness domain) and / or chroma mapping can use reference and non-reference histograms to compute these CDFs for both the reference and sample images, with intervals storing the effective pixel counts in different sub-ranges of saturation metrics across multiple saturation metric ranges. A saturation mapping curve or LUT can be computed or estimated by imposing or using equality conditions between the CDFs of the saturation metric histograms constructed from the reference and sample images. A saturation mapping curve or LUT can be applied to a non-reference image with RGB leveled tone mapping to generate a saturation-leveled RGB leveled tone mapping non-reference image with a relatively refined visual appearance that matches the reference visual appearance.

[0091] Additionally, optionally or alternatively, artificial neural networks can be used and trained to perform some or all of the following: codeword or RGB leveling, saturation leveling, and / or saturation matching.

[0092] For illustrative purposes, it is shown that the leveling operation can be performed after the luminance (or hue) and / or chroma mapping operation. It should be noted that in some other operational scenarios, such as RGB leveling (or white balance), some or all of the leveling operation can be performed before the luminance (or hue) and / or chroma mapping operation. Additionally, optionally or alternatively, in some operational scenarios, such as RGB leveling (or white balance), some or all of the leveling operation can be performed in parallel or in combination with the luminance (or hue) and / or chroma mapping operation.

[0093] Noise visibility The techniques described herein can be implemented or applied to address the problem of matching the style or rendering of image or video content captured by different cameras under varying lighting conditions. As shown, a specific camera among multiple cameras used to capture content can be selected as a reference or "protagonist" camera. Content portions generated by other cameras can be considered as source or "test" image / video content portions to be automatically adapted to the style or rendering associated with the reference or protagonist camera; these portions can be modified to automatically adapt to a style and rendering similar to that presented by the reference or protagonist camera. For user-generated content (UGC) creators, the techniques described herein can significantly reduce the workload of generating seamless timelines from mixed content captured by multiple reference and / or non-reference cameras.

[0094] Figure 6AThe illustration shows an example process flow, method, or algorithm for the automatic content adaptation operation described herein. This process flow can be used to automatically adapt the source or non-reference style or rendering of a source or non-reference video / image stream to a reference style or rendering that is the same as or similar to that of a reference video / image stream. This process flow can be implemented or performed, at least in part, using one or more computing devices, including but not limited to automatic content adaptation systems, such as post-camera image / video processing systems installed on the computing devices, automatic adaptation tools including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0095] Box 602 includes converting the source video / image stream and / or the reference video / image stream to the same (working) color space if the input color space of the source video / image stream is different from the color space of the reference video / image stream. Box 602 also includes linearizing the reference and source video / image streams in the (working) color space.

[0096] Box 604 includes calculating or deriving a white balance (e.g., codeword balance, RGB balance, etc.) gain based on the codewords in the images / frames of the reference and source video / image streams to generate an automatically adapted style or rendering that matches the reference video / image stream (of the source video / image stream). For example, the average codeword... The average ratio or white point (WP) gain can be calculated across the valid pixels of the source or reference image in the source or reference video / image stream, where the valid pixels are located at pixel positions (in the reference and source video / image streams) where the pixel value is located near, or within, a specific configuration or selected chromaticity neighborhood around, the specific white point (e.g., light source D65, etc.) in the reference color space of the CIE 1931 xy chromaticity diagram. The average ratio or white point (WP) gain can be based on these average codewords from the reference and source video / image streams. To determine this. In some operational scenarios, the reference and source images are represented in the RGB color space, and the average ratio or white point gain can be calculated separately for the individual R, G, and B channels in the RGB color space.

[0097] Box 606 includes multiplying the codewords of the source image in the source video / image stream by an average ratio or white point gain to perform white point leveling (or RGB / codeword leveling), as shown in expression (2) above. Applying the leveling operation to the source video / image stream produces a leveled source video / image stream.

[0098] Box 608 includes generating and applying a tone mapping curve or LUT to match the tonal range of a source video / image stream to the tonal range of a reference source video / image stream. In an operational scenario, this curve or LUT may subsequently (e.g., evenly, individually, etc.) be applied to the individual R, G, and B channels of the source video stream. Applying the tone mapping curve or LUT to a leveled source video / image stream produces a leveled source video / image stream with tone mapping (or tone matching).

[0099] Box 610 includes applying a specific photoelectric conversion function (OETF) to a tone-matched leveling (linearization) of a source video / image stream, specifying one for use with or associated with the reference video / image stream, to generate an automatically adapted source video / image stream.

[0100] Figure 7A The diagram illustrates the use of Figure 6A The example in the reference video / image stream of the process flow is a relatively bright reference image (in...) Figure 7A (at the top), example (relatively bright) source image in the initial or pre-auto-adapted source video / image stream (in) Figure 7A (in the middle) and example automatically adapted images derived from the source image (in Figure 7A (The bottom). As shown in the figure. Figure 6A The process, method, or algorithm will generate a matching bright video (e.g., Figure 7A The quality result of relative brightness in the bright image shown (e.g., no or almost no noise visibility).

[0101] Figure 7B The diagram illustrates the use of Figure 6A Example of a relatively dark source image in the source video / image stream of the process flow (in Figure 7B (at the top) and example automatically adapted images derived from that source image. As shown, compared to using Figure 6A The process flow will be as follows Figure 7A The illustrated bright source video / image is matched with a bright reference video / image when attempting to use... Figure 6A The same process flow will be as follows Figure 7B The dark source video / image shown in the illustration (in) Figure 7B (in the middle) with bright reference video / image (in) Figure 7B When the top of the image matches, the inherent noise introduced by the source camera in the dark scene depicted in the source video / image causes the video / image to automatically adapt or match (in the top of the image). Figure 7B The unpleasant noise visibility at the bottom.

[0102] Unpleasant noise visibility in adjusted or automatically adapted videos / images is achieved by using... Figure 6AThe average ratio or white point gain calculated in box 604 is caused by performing RGB / codeword leveling (referred to as white point gain). In the case of adjustment from dark to light, the average ratio or white point gain can be a relatively large value. Therefore, the noise introduced by the source camera is amplified to a relatively noticeable degree, resulting in a relatively perceptible noise visibility in the adjusted or automatically adapted video / image.

[0103] Using the techniques described herein, noise analysis / characteristic description of a scene captured in an image can be performed or generated. Information obtained from this noise analysis / characteristic description can be used to control the amount of gain ultimately applied to the leveling operation.

[0104] Noise modulation leveling The ISO speed of a digital camera used to capture or acquire reference or non-reference source video / image content controls the overall brightness of the captured video / image content. ISO speed depends in whole or in part on the light-sensitive characteristics of the camera's image sensor (e.g., inherent, sensor-specific, etc.) and the gain (or gain setting) of the camera's image signal processor (ISP). For ISP gain, the video / image signal generated by the image sensor can be amplified (e.g., electrically, etc.), including any noise present or inherent in the pre-amplified video / image signal.

[0105] In low light saturation or physical scenes, higher ISO speeds or higher gain values ​​(e.g., user- or system-configurable values) can amplify the video / image signals generated by the image sensor and ISP (including any noise particles therein), and cause higher visibility of noise particles in the captured video / image content generated by the image sensor and / or ISP.

[0106] Figure 6B The diagram illustrates the basis described in this article. Figure 6A Examples of process flows, methods, or algorithms that automatically adapt to changes in the process flow. Figure 6B The process flow can be used (for example, as) Figure 6A This process, which is an alternative or supplement to the existing workflow, automatically adapts the source or non-reference style or rendering of a source or non-reference video / image stream to a reference style or rendering that is the same as or similar to that of the reference video / image stream, taking into account specific cameras, specific image content, specific scenes, or changes in brightness. This process can be implemented or executed, at least in part, using one or more computing devices, including but not limited to automatic content adaptation systems. These devices include post-camera image / video processing systems installed on the computing devices, automatic adaptation tools including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, and cameras.

[0107] like Figure 6A and Figure 6BAs shown, Figure 6A Box 606 was Figure 6B Replace boxes 606-1 and 606-2. Figure 6B All boxes except boxes 606-1 and 602-2 can be used with Figure 6A The corresponding box is executed in the same or similar manner. Figure 6A or Figure 6B The average ratio or white point gain determined in box 604 can be used as the initial white gain to be further modified based on or based on the noise analysis characteristics.

[0108] More specifically, Figure 6B Box 606-1 includes applying or performing noise analysis or characterization to one or both of a reference and a source video / image stream or images therein. The results of the noise analysis / characterization are applied or used to adjust the (initial) average ratio or white point gain to generate a modified white point gain. Instead of the (pre-adjusted or pre-modified) average ratio or white point gain, the modified white point gain modulated or adjusted using the results of the noise analysis / characterization can be used for RGB / codeword leveling (or white point leveling) operations.

[0109] Figure 6B Box 606-2 includes multiplying the codewords of the source image in the source video / image stream by a modified white point gain (as a multiplication factor) generated at least in part based on noise analysis / characterization, to perform white point leveling (or RGB / codeword leveling). Applying the leveling operation to the source video / image stream produces a leveled source video / image stream, which can be further processed by... Figure 6A or Figure 6B The subsequent boxes 608 and 610 are processed to generate corresponding tone-matched or automatically adapted source video / image streams.

[0110] Noise Analysis Figure 6C The illustration depicts an example process flow for noise analysis / characterization in a video or image stream (e.g., a source or reference video / image stream, images therein, etc.). This process flow can be implemented or performed, at least in part, using one or more computing devices, including but not limited to automatic content adaptation systems, such as post-camera image / video processing systems mounted on computing devices, automatic adaptation tools including but not limited to image acquisition systems, mobile computing devices, non-mobile computing devices, cameras, etc.

[0111] Box 620 includes applying edge detectors, algorithms, methods, procedures, and / or operators to the luminance channels of some or all images or frames in a video / image stream—for example, after representing or transforming the images or frames in a specific color space (such as YUV, which includes luminance channels)—to discover or detect (e.g., visually perceptible) edges in the images / frames of the stream. Examples of edge detectors or detection algorithms, methods, procedures, and / or operators may include, but are not limited to, Sobel, Canny, etc. Example edges may involve visually perceptible structures such as objects, people, backgrounds, foregrounds, boundaries, borders, etc., that are visually depicted in an image or frame.

[0112] Box 622 includes dividing each image or frame in the noise analysis / characterization process herein into multiple image blocks with block sizes such as 4×4, 8×8, and 16×16. Any image block containing detected edges within these multiple image blocks can be excluded from further noise analysis / characterization operations. Box 624 includes identifying a specific group of image blocks within these multiple image blocks to include some or all (all good or valid) image blocks that do not have any detected edges. The specific group of (all good or valid) image blocks (excluding all image blocks with edge portions) is used for further noise analysis / characterization operations.

[0113] Box 626 includes determining the average pixel value and standard deviation for each image block in a specific (all good or valid) image block. Additionally, the signal-to-noise ratio (NSR) value is obtained by dividing the standard deviation of the block by the average pixel value (within that block).

[0114] Box 628 includes calculating the global or cross-block average (indicated as LumaNSR) of the NSR value across all image blocks in a specific (all good or valid) image block across (e.g., representing a specific color space (such as YUV) of an image or frame) luminance or luminance channel.

[0115] Box 628 also includes calculating the global or cross-block average (indicated as ChromaNSR) of the NSR value across all image blocks in a specific (all good or valid) image block across (e.g., representing a specific color space (such as YUV) of an image or frame) two chroma or chroma channels, each (e.g., U or V, etc.). The ChromaNSR value of the U channel or component of a specific color space can be specifically indicated as U_NSR, while the ChromaNSR value of the V channel or component of a specific color space can be specifically indicated as V_NSR.

[0116] (4-1) (4-2) (4-3) Where σ represents the total number of (good or valid) image patches in the group; σ is the standard deviation of the image patches, and μ is the average value of the image patches.

[0117] The overall colorimetric signal-to-noise ratio can be calculated as follows: (5) Box 630 includes calculating or determining the overall or total noise (metric) using signal-to-noise values ​​calculated for color channels or components of a specific color space. In some operational scenarios, this noise metric can be calculated as the sum of luminance and chrominance signal-to-noise values ​​using one or both of the following options: Total Noise = α LumaNSR + (1 α) ChromaNSR (6-1) Total noise = sqrt(α) LumaNSR^2 + (1 α) ChromaNSR^2) (6-2) Where α represents a configurable value for the user or system, between 0 and 1.

[0118] The above expression (6-1) can be used to calculate the noise metric when some or all noise sources are considered correlated or mutually (or data-wise) related, and the above expression (6-2) can be used to calculate the noise metric when some or all noise sources are considered uncorrelated or mutually (or data-wise) unrelated. Additionally, optionally, alternatively, in various operational scenarios, one of the two options can be specifically selected as the default value for calculating the noise metric. In image sensors, due to the photon shot noise statistics in the image sensor (e.g., the noise generated in the signal generated by the image sensor can be proportional to the square root of the signal intensity or brightness (corresponding to the above expression (6-1)), or proportional to the signal intensity or brightness (corresponding to the above expression (6-2)), pixels with higher exposure values ​​can have lower relative noise compared to pixels with lower exposure values ​​(e.g., noise in a pixel compared to the brightness of a pixel, etc.).

[0119] By dividing the noise by the brightness in the signal-to-noise (NSR) value, the proportion of noise's influence relative to the brightness value can be determined, estimated, or accounted for relatively sufficiently or accurately. These ratios can be used to emphasize the relative importance of noise in relatively dark image blocks or content areas.

[0120] By calculating or determining the overall noise metric using NSR values ​​that take into account the relative importance of noise in relatively dark image patches or portions, this overall noise metric can also emphasize the relative importance of noise in relatively dark image patches or portions of content. Therefore, this overall noise metric can be used for leveling operations to reduce or prevent noise from entering the image. Figure 7B The illustration shows noise artifacts.

[0121] Several other metrics can be used to supplement or replace NSR values. In an illustrative but unrestricted manner, the ICtCp color representation format and / or color space (instead of the YUV color space) can be used to represent images or frames from which gains (e.g., intermediate values, etc.) can be determined for leveling operations (e.g., white point, etc.). Compared to other color representation formats and / or color spaces, the ICtCp color space, or the gains calculated using it, can provide better visual consistency relative to human vision systems.

[0122] Example operations and representation formats related to the ICtCp color space are defined or described in Rec. ITU-R BT.2100 “Picture parameter values ​​for the production and international program exchange of high dynamic range television” (06 / 2017), the entire text of which is incorporated herein by reference.

[0123] In some feasible scenarios, the ICtCp color representation of an image / frame in video / image content can be generated or derived by first converting the input video / image content from the linear RGB color space to the (intermediate) LMS color space, then applying a nonlinear perceptual quantization (PQ) function to generate nonlinear (PQ) video / image content from the linear video / image content in the LMS color space, and finally converting the nonlinear (PQ) video / image content (or signal) from the LMS color space to the ICtCp color space.

[0124] Images / frames are represented in the ICtCp color space, and noise metrics can be derived from noise analysis / characteristic descriptions using signal-to-noise ratios (or ratios) such as noise_I / I, noise_CT / I, and noise_CP / I.

[0125] Here, I indicates the average value of the I codewords in the image / frame of the video / image content represented in the ICtCp color space; noise_I indicates the noise value in the I component or channel of the ICtCp color space; noise_CT indicates the noise value in the Ct component or channel of the ICtCp color space; and noise_CP indicates the noise value in the Cp component or channel of the ICtCp color space.

[0126] Noise processing As previously indicated, the initial white point gain can be calculated based on the RGB or codeword average of the image / frame. With these initial white point gains being relatively small, applying these gains to the scene portion related to light-to-dark, dark-to-dark, or light-to-light scene matching or autofit may have little or no impact on the perceived noise in the resulting autofitted visual / image content.

[0127] In contrast, when these initial white point gains are relatively large, applying these gains to scene matching or autofitting involving dark to light scenes can have a relatively large impact on the perceived noise in the resulting autofitted visual / image content. As the gain increases, the amplification of noise and the visibility of perceived noise in the autofitted visual / image content can become increasingly perceptible and unpleasant to human observers.

[0128] For example, used to directly generate automatically adapted images (in) Figure 7B The initial white point gain (at the bottom) can be determined as follows: Figure 7B The reference image shown at the top and Figure 7B The initial white point gain can be relatively large, such as 10 (10), calculated from the RGB or codeword values ​​in the non-reference image in the middle. When a relatively large initial white point gain (10) is applied directly to the non-reference image without noise characterization or processing, the noise in the non-reference image can be significantly amplified (compared to a relatively small gain) to produce noise in the non-reference image. Figure 7B The bottom section shows a relatively unpleasant visual result.

[0129] One or more noise processing methods can be used or implemented to control or adjust these relatively large white point gains based on an inherent noise level estimated, determined, or characterized from the results of noise analysis performed on the video / image content. In some operational scenarios, the inherent noise level can be estimated, determined, or characterized using noise metrics as described herein.

[0130] In the first example method, the modified white point gain (indicated as "OutputWpGains") can be generated by adjusting or modifying the corresponding initial white point gain (indicated as "InputWpGains") by utilizing the inverse relationship between the noise level, as represented by the corresponding noise metric, and the initial white point gain (e.g., with an adjustment factor or constant such as -1, etc.), as follows: OutputWpGains = (InputWpGains +p (noise) / (1+(p) Noise) (7) The noise can be the same as the total noise calculated in the above expressions (6-1) or (6-2); p indicates the adjustment or scaling parameter.

[0131] A specific value for adjusting or scaling parameter p (e.g., one (1), three (3), etc.) can be experimentally derived by performing objective and / or subjective noise effect assessments on videos captured by different cameras with various inherent noise levels. These cameras may come from different camera or device manufacturers, different camera or device models, etc., and / or may use different image sensors with different inherent or ISP noise levels or characteristics. For a given inherent noise level, a relatively large value of p can be used to more aggressively adjust or limit any noise amplification effects or influences of the initial white point gain.

[0132] In the second example method, for example, considering the reduced visibility due to noise during various image transitions or brightness variations, a specific relationship can be defined, specified, or adjusted between the initial white point gain and the modified white point gain. In some operable scenarios—such as… Figure 8 The specific relationship illustrated can be represented as a spline with three distinct ranges—the numerical constants in this relationship are for illustrative purposes only—as follows: Where InputWpGains is between 0 and 4: OutputWpGains = InputWpGains (8-1) Where InputWpGains is between 4 and 15: OutputWpGains = 0.1818 InputWpGains + 4.27 (8-2) Where InputWpGains>15: OutputWpGains = 2 (8-3) In the third example method, the fragment parameter / value or the maximum modified white point gain can be determined at least in part based on the noise level or metric. If the initial or original white point gain does not exceed the fragment parameter / value, the modified white point gain is set to the initial or original white point gain. If the initial or original white point gain is greater than the fragment parameter / value, the modified white point gain is fixed or set to the fragment parameter / value. In some operational scenarios, the fragment parameter / value can be set as follows: Excerpt =β / (ε+ noise)(9) Where β (e.g., 2, 6, 8, 11, etc.) and ε (e.g., 1, etc.) are adjustment parameters; the noise can be the total noise calculated using the above expressions (6-1) or (6-2).

[0133] These adjustment parameters can be used to control the maximum white point modification for a given noise level or metric (such as the maximum noise value). One or two of these parameters can be experimentally determined, set, or estimated by optimizing the segment parameters / values ​​for the maximum noise value with little or no automatic adaptation to noise amplification artifacts in video / image content.

[0134] In some feasible scenarios, Figure 6B In box 606-2, the modified white point gain for each R, G, or B channel in the RGB color space (generated by applying noise processing or adjustment to the initial white point gain) can be constrained to maintain the same inter-channel ratio (e.g., between red and green or R / G, between green and blue or G / B, etc.) as the initial or original white point gain, so as to maintain the same color balance between the image generated by applying the modified white point gain and the image generated by applying the initial white point gain.

[0135] Figure 7C The illustration shows an example of an automatically adapted image generated by applying a modified white point gain to a source or non-reference image. This is in contrast to an image generated by applying an initial white point gain as shown in the example. Figure 7B Compared to the automatically adapted image generated from the same source or non-reference image illustrated at the bottom, the image obtained after performing noise-related noise adjustment is shown below. Figure 7C The automatically adapted image results in a relatively high-quality image for dark scenes with little or no perceptible noise amplification or artifacts. Here, the modified white point gain is reduced from 10 to 1.5 of the initial white point gain, which helps to prevent or reduce noise amplification present in the pre-automatically adapted source image and produces a relatively satisfactory visual style or result.

[0136] Example process flow Figure 4E An example process flow according to an example embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., video editing systems, desktop computers, video processing tools, etc.) implementing the automatic adaptation system as described herein can perform this process flow. In block 482, the system determines one or more reference images in a working color space and one or more source images in a selected working color space. The one or more reference images are initially captured by a reference camera. The one or more source images are initially captured by a source camera different from the reference camera.

[0137] In box 484, the system identifies one or more common image features between one or more reference images and one or more source images. The one or more common image features in the one or more reference images include multiple reference pixels. The one or more common image features in the one or more source images include multiple source pixels.

[0138] In box 486, the system uses the distribution of reference codeword values ​​of the plurality of reference pixels as represented in the selected color space and the distribution of source codeword values ​​of the plurality of source pixels to generate one or more source-to-reference codeword maps.

[0139] In box 488, based at least in part on one or more source-to-reference codeword mappings, the system transforms a specific source image initially captured by a source camera into a corresponding auto-adapted source image that is adapted to a reference visual appearance associated with one or more reference images initially captured by a reference camera.

[0140] In the embodiments, the reference camera differs from the source camera in one or more of the following: output image color space, output image dynamic range, output image color gamut, output image bit depth, image signal processor, image signal processor configuration, device manufacturer, device model, etc.

[0141] In an embodiment, one or more reference images are initially captured by a reference camera in a high dynamic range format; the high dynamic range format is one of the following: a mixed log-gamma (HLG) format represented in the Rec.2020 RGB color space, or a perceptual quantizer (PQ) format represented in the Rec.2020 RGB color space; the one or more reference images in the high dynamic range format are linearized to be represented in a linearized RGB color space representation using a nonlinear inverse photoelectric conversion function (OETF).

[0142] In an embodiment, common image features are identified at least in part based on one or more projection transformations constructed from image features detected using one or more computer vision techniques.

[0143] In an embodiment, one or more source-to-reference codeword mappings are generated by one or more of the following: cumulative density function (CDF) matching operation, or one or more artificial neural networks.

[0144] In an embodiment, at least one source-to-reference codeword mapping in one or more source-to-reference codeword mappings is represented by one of the following: mapping function, spline function, piecewise function, piecewise polynomial function, lookup table, etc.

[0145] In one embodiment, a particular source image belongs to a sequence of consecutive source images of a particular visual scene; wherein one or more source-to-reference codeword mappings are used for each source image in the sequence of consecutive source images of a particular visual scene to generate a corresponding sequence of consecutively auto-adapted source images of the particular visual scene.

[0146] In an embodiment, a specific source image refers to one of one or more source images.

[0147] In one embodiment, one or more reference images and one or more source images are captured by a reference camera and a source camera from a common physical environment, respectively; the reference camera and the source camera are located at different spatial locations in the common physical environment.

[0148] In this embodiment, a specific source image is not one of one or more source images.

[0149] In one embodiment, one or more source images are converted into one or more corresponding autofit source images, which are matched with a reference visual appearance associated with one or more reference images initially captured by a reference camera; at least a subgroup of the one or more corresponding autofit source images is combined with at least a subgroup of the one or more reference images to form a single temporally continuous image sequence along a single playback timeline.

[0150] In an embodiment, the corresponding auto-adapted image is generated by performing an RGB leveling operation on an intermediate image, which is generated at least in part by applying one or more source-to-reference codeword mappings to a specific source image.

[0151] In an embodiment, the corresponding auto-adapted image is generated by performing either a saturation matching operation or a saturation leveling operation on an intermediate image, which is generated at least in part by applying one or more source-to-reference codeword mappings to a particular source image.

[0152] Figure 4F The illustration depicts an example process flow according to an example embodiment. In some embodiments, one or more computing devices or components (e.g., video editing systems, desktop computers, video processing tools, etc.) implementing the automatic adaptation system described herein can perform this process flow. In block 490, the system determines one or more reference images and one or more source images in a working color space. The one or more reference images are derived from a reference camera. The one or more source images are derived from a source camera other than the reference camera.

[0153] In box 492, the system derives an initial gain value between one or more reference images and one or more source images. The initial gain value is determined based on the reference codeword values ​​of one or more reference images and the source codeword values ​​of one or more source images.

[0154] In box 494, the system adjusts an initial gain value to a modified gain value based at least in part on one or more results of a noise characteristic description performed using one or more source images.

[0155] In box 496, the system applies the modified gain value to one or more source images to generate one or more leveled source images.

[0156] In box 498, the system transforms one or more leveled source images into one or more tone-matched source images, based at least in part on one or more source-to-reference tone maps.

[0157] In this embodiment, the working color space represents a linear color space; a linear color space is one of the following: linear RGB color space, linear YUV color space, linear YCbCr color space, or a different linear color space.

[0158] In this embodiment, the working color space represents a non-linear color space; the non-linear color space is one of the following: perceptual quantization (PQ) color space, mixed log-gamma (HLG) color space, non-linear LMS color space, non-linear ICpCt color space, and various non-linear color spaces.

[0159] In an embodiment, the initial gain value is determined as an average ratio based on the average reference codeword value of one or more reference images and the average source codeword value of one or more reference images.

[0160] In one embodiment, the system further performs the following: applying a specific photoelectric conversion function (OETF) associated with one or more reference images to a tone-matched source image to generate one or more automatically adapted source images that are adapted to a reference visual appearance associated with a reference camera.

[0161] In an embodiment, the noise characterization performed using one or more source images is represented at least in part by a noise metric calculated using the standard deviation of individual image patches and the signal-to-noise ratio between the intensity of individual image patches; these image patches are partitioned from one or more source images.

[0162] In an embodiment, the individual image blocks that are segmented from one or more source images and used to calculate individual signal-to-noise ratios constitute a set of image blocks; the set of image blocks does not include any image blocks in one or more source images that contain visually perceptible edges detected using one or more edge detection methods.

[0163] In the embodiments, the individual signal-to-noise ratio is calculated for each color channel of the working color space representing one or more source images.

[0164] In this embodiment, a noise metric is determined for the luminance channel of the working color space; two additional noise metrics are determined for the chrominance channels of the working color space; total noise is calculated for one or more source images based on the noise metric and the two additional noise metrics; the total noise is used to modify the initial gain value to the modified gain value.

[0165] In an embodiment, total noise is used to modify the initial gain value to a modified gain value in approximating the inverse relation.

[0166] In an embodiment, total noise is used to modify the initial gain value to a modified gain value in the spline relation.

[0167] In an embodiment, total noise is used to apply the maximum gain limit to the modified gain value.

[0168] In an embodiment, for each pair of modified gain values ​​for two different color channels, a first ratio value is maintained that is the same as a second ratio value maintained by the corresponding initial gain values ​​for the different color channels.

[0169] In various example embodiments, an apparatus, system, device, or one or more other computing devices performs any or a portion of the methods described herein. In embodiments, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause the methods described herein to be performed.

[0170] Note that although individual embodiments are discussed herein, any combination of the embodiments and / or some of the embodiments discussed herein can be combined to form further embodiments.

[0171] Implementation Mechanism – Hardware Overview According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to perform these techniques, or may include digital electronic devices persistently programmed to perform these techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or combinations thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device incorporating hardwired and / or program logic to implement the techniques.

[0172] For example, Figure 5 This is a block diagram illustrating a computer system 500 on which an example embodiment of the invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor.

[0173] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When stored in a non-transitory storage medium accessible to processor 504, such instructions enable computer system 500 to become a dedicated machine defined to perform the operations specified in the instructions.

[0174] The computer system 500 further includes a read-only memory (ROM) 508 or other static storage device coupled to the bus 502 for storing static information and instructions of the processor 504.

[0175] Storage devices 510, such as disks or optical discs, solid-state RAM, etc., are provided and coupled to bus 502 for storing information and instructions.

[0176] Computer system 500 can be coupled to display 512, such as an LCD, via bus 502 for displaying information to a computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or cursor arrow keys, for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Typically, this input device has two degrees of freedom on two axes (a first axis (e.g., x-axis) and a second axis (e.g., y-axis)), allowing the device to specify a position in a plane.

[0177] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic. These custom hardwired logics, one or more ASICs or FPGAs, firmware, and / or program logic, combined with the computer system, enable computer system 500 to be a dedicated machine or programmed to be a special-purpose machine. According to one embodiment, the techniques described herein are executed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium (such as storage device 510). Execution of the sequence of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0178] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, floppy hard disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, any other memory chips or memory cartridges.

[0179] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0180] Various forms of media can involve loading one or more sequences of one or more instructions to processor 504 for execution. For example, instructions may initially be loaded onto a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 loads the data into main memory 506, from which processor 504 retrieves and executes the instructions. Instructions received in main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.

[0181] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which connects to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity with a corresponding type of telephone line. As another example, communication interface 518 may be a Local Area Network (LAN) card for providing data communication connectivity with a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying streams of digital data representing various types of information.

[0182] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection via local network 522 to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526. ISP 526, in turn, provides data communication services via a global packet data communication network now commonly referred to as the “Internet” 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.

[0183] Computer system 500 can send messages and receive data, including program code, through multiple networks, network links 520, and communication interfaces 518. In the Internet example, server 530 can transmit application request codes through the Internet 528, ISP 526, local network 522, and communication interface 518.

[0184] The received code may be executed by processor 504 upon receipt and / or stored in storage device 510 or other non-volatile storage device for later execution.

[0185] Equivalents, extensions, alternatives and miscellaneous In the foregoing description, exemplary embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the set of claims issued in the specific form of such claims, including any subsequent amendments, indicates the scope of the invention and is intended by the applicant as the sole and exclusive embodiment of the invention. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and the drawings should be viewed in an illustrative rather than restrictive sense.

[0186] Exemplary examples of enumeration This invention may be practiced in any of the forms described herein, including but not limited to the enumerated example embodiments (EEE) that describe some parts of the structure, features and functions of embodiments of the invention.

[0187] EEE1. A method comprising: Identify one or more reference images in the working color space and one or more source images in the selected working color space, wherein the one or more reference images are initially captured by a reference camera and the one or more source images are initially captured by a source camera different from the reference camera; Identify one or more common image features between the one or more reference images and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, and wherein the one or more common image features in the one or more source images include a plurality of source pixels; One or more source-to-reference codeword maps are generated using the distribution of reference codeword values ​​of the plurality of reference pixels and the distribution of source codeword values ​​of the plurality of source pixels as represented in the selected color space. Based at least in part on the codeword mapping from one or more sources to references, a specific source image initially captured by the source camera is transformed into a corresponding automatically adapted source image that is adapted to the reference visual appearance associated with one or more reference images initially captured by the reference camera.

[0188] EEE2. The method as described in EEE1, wherein the reference camera differs from the source camera in one or more of the following: output image color space, output image dynamic range, output image color gamut, output image bit depth, image signal processor, image signal processor configuration, device manufacturer, or device model.

[0189] EEE3. The method as described in EEE1 or EEE2, wherein the one or more reference images are initially captured by the reference camera in a high dynamic range format; wherein the high dynamic range format is one of the following: a mixed log-gamma (HLG) format represented in the Rec.2020 RGB color space, or a perceptual quantizer (PQ) format represented in the Rec.2020 RGB color space; wherein the one or more reference images in the high dynamic range format are linearized to be represented in a linearized RGB color space representation using a nonlinear inverse photoelectric conversion function (OETF).

[0190] EEE4. The method of any one of EEE1 to EEE3, wherein the common image features are identified at least in part based on one or more geometric transformations constructed from image features detected using one or more computer vision techniques.

[0191] EEE5. The method of any one of EEE1 to EEE4, wherein the one or more source-to-reference codeword mappings are generated by one or more of the following: cumulative density function (CDF) matching operation, or one or more artificial neural networks.

[0192] EEE6. The method of any one of EEE1 to EEE5, wherein at least one of the one or more source-to-reference codeword mappings is represented by one of the following: a mapping function, a spline function, a piecewise function, a piecewise polynomial function, or a lookup table.

[0193] EEE7. The method of any one of EEE1 to EEE6, wherein the particular source image belongs to a continuous source image sequence of a particular visual scene; wherein the one or more source-to-reference codeword mappings are applied to each source image in the continuous source image sequence of the particular visual scene to generate a corresponding continuous auto-adapted source image sequence for the particular visual scene.

[0194] EEE8. The method of any one of EEE1 to EEE7, wherein the particular source image represents one of the one or more source images.

[0195] EEE9. The method of any one of EEE1 to EEE8, wherein the one or more reference images and the one or more source images are captured by the reference camera and the source camera respectively from a common physical environment; wherein the reference camera and the source camera are located at different spatial locations in the common physical environment.

[0196] EEE10. The method of any one of EEE1 to EEE9, wherein the particular source image represents one of the one or more source images.

[0197] EEE11. The method of any one of EEE1 to EEE10, wherein the one or more source images are converted into one or more corresponding autofit source images, the one or more corresponding autofit source images being matched with the reference visual appearance associated with the one or more reference images initially captured by the reference camera; wherein at least a subset of the one or more corresponding autofit source images is combined with at least a subset of the one or more reference images to form a single temporally continuous image sequence along a single playback timeline.

[0198] EEE12. The method of any one of EEE1 to EEE11, wherein the corresponding auto-adapted image is generated by performing an RGB leveling operation on an intermediate image, the intermediate image being generated at least in part by applying the one or more source-to-reference codeword mappings to the particular source image.

[0199] EEE13. The method of any one of EEE1 to EEE12, wherein the corresponding autofit image is generated by performing either a saturation matching operation or a saturation leveling operation on an intermediate image, the intermediate image being generated at least in part by applying one or more source-to-reference codeword mappings to the particular source image.

[0200] EEE14. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause to perform the method as described in any one of EEE1 to EEE13.

[0201] EEE15. A computing device comprising one or more processors and one or more storage media storing an instruction set that, when executed by the one or more processors, causes to perform the method as described in any one of EEE1 to EEE13.

[0202] EEE16. An apparatus comprising: One or more processors; and One or more storage media storing an instruction set that, when executed by one or more processors, causes the following to be executed: Identify one or more reference images in the working color space and one or more source images in the selected working color space, wherein the one or more reference images are initially captured by a reference camera and the one or more source images are initially captured by a source camera different from the reference camera; Identify one or more common image features between the one or more reference images and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, and wherein the one or more common image features in the one or more source images include a plurality of source pixels; One or more source-to-reference codeword maps are generated using the distribution of reference codeword values ​​of the plurality of reference pixels and the distribution of source codeword values ​​of the plurality of source pixels as represented in the selected color space. Based at least in part on the codeword mapping from one or more sources to references, a specific source image initially captured by the source camera is transformed into a corresponding automatically adapted source image that is adapted to the reference visual appearance associated with one or more reference images initially captured by the reference camera.

[0203] EEE17. The apparatus as described in EEE16, wherein the reference camera differs from the source camera in one or more of the following: output image color space, output image dynamic range, output image color gamut, output image bit depth, image signal processor, image signal processor configuration, device manufacturer, or device model.

[0204] EEE18. An apparatus as described in EEE16 or EEE17, wherein the one or more reference images are initially captured by the reference camera in a high dynamic range format; wherein the high dynamic range format is one of the following: a mixed log-gamma (HLG) format represented in the Rec.2020 RGB color space, or a perceptual quantizer (PQ) format represented in the Rec.2020 RGB color space; wherein the one or more reference images in the high dynamic range format are linearized to be represented in a linearized RGB color space representation using a nonlinear inverse photoelectric conversion function (OETF).

[0205] EEE19. The apparatus of any one of EEE16 to EEE18, wherein the common image feature is identified at least in part based on one or more geometric transformations constructed from image features detected using one or more computer vision techniques.

[0206] EEE20. The apparatus of any one of EEE16 to EEE19, wherein the one or more source-to-reference codeword mappings are generated by one or more of the following: cumulative density function (CDF) matching operation, or one or more artificial neural networks.

[0207] EEE21. The apparatus of any one of EEE16 to EEE20, wherein at least one of the one or more source-to-reference codeword mappings is represented by one of the following: a mapping function, a spline function, a piecewise function, a piecewise polynomial function, or a lookup table.

[0208] EEE22. The apparatus of any one of EEE16 to EEE21, wherein the particular source image belongs to a continuous source image sequence of a particular visual scene; wherein the one or more source-to-reference codeword mappings are applied to each source image in the continuous source image sequence of the particular visual scene to generate a corresponding continuous auto-adaptive source image sequence for the particular visual scene.

[0209] EEE23. The apparatus of any one of EEE16 to EEE22, wherein the one or more source images are converted into one or more corresponding auto-adapted source images, the one or more corresponding auto-adapted source images being matched with the reference visual appearance associated with the one or more reference images initially captured by the reference camera; wherein at least a subset of the one or more corresponding auto-adapted source images is combined with at least a subset of the one or more reference images to form a single temporally continuous image sequence along a single playback timeline.

[0210] EEE24. The apparatus of any one of EEE16 to EEE23, wherein the corresponding auto-adapted image is generated by performing one of an RGB leveling operation, a saturation matching operation, or a saturation leveling operation on an intermediate image, the intermediate image being generated at least in part by applying the one or more source-to-reference codeword mappings to the particular source image.

[0211] EEE25. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause the following to be performed: Identify one or more reference images in the working color space and one or more source images in the selected working color space, wherein the one or more reference images are initially captured by a reference camera and the one or more source images are initially captured by a source camera different from the reference camera; Identify one or more common image features between the one or more reference images and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, and wherein the one or more common image features in the one or more source images include a plurality of source pixels; One or more source-to-reference codeword maps are generated using the distribution of reference codeword values ​​of the plurality of reference pixels and the distribution of source codeword values ​​of the plurality of source pixels as represented in the selected color space. Based at least in part on the codeword mapping from one or more sources to references, a specific source image initially captured by the source camera is transformed into a corresponding automatically adapted source image that is adapted to the reference visual appearance associated with one or more reference images initially captured by the reference camera.

[0212] EEE26. The medium as described in EEE25, wherein the reference camera differs from the source camera in one or more of the following: output image color space, output image dynamic range, output image color gamut, output image bit depth, image signal processor, image signal processor configuration, device manufacturer, or device model.

[0213] EEE27. A method comprising: Identify one or more reference images and one or more source images in the working color space, wherein the one or more reference images are derived from a reference camera, and wherein the one or more source images are derived from a source camera different from the reference camera; The initial gain values ​​between the one or more reference images and the one or more source images are obtained, wherein these initial gain values ​​are determined based on the reference codeword values ​​of the one or more reference images and the source codeword values ​​of the one or more reference images; These initial gain values ​​are adjusted to modified gain values ​​based at least in part on the results of one or more noise characteristic descriptions performed using the one or more source images; These modified gain values ​​are applied to the one or more source images to generate one or more leveled source images; The one or more leveled source images are converted into one or more tone-matched source images, based at least in part on one or more source-to-reference tone maps.

[0214] EEE28. The method as described in EEE27, wherein the working color space represents a linear color space; wherein the linear color space is one of the following: linear RGB color space, linear YUV color space, linear YCbCr color space, or a different linear color space.

[0215] EEE29. The method as described in EEE27, wherein the working color space represents a non-linear color space; wherein the non-linear color space is one of the following: perceptual quantization (PQ) color space, mixed log-gamma (HLG) color space, non-linear LMS color space, non-linear ICpCt color space, or a different non-linear color space.

[0216] EEE30. The method as described in EEE27, wherein these initial gain values ​​are determined as an average ratio based on the average reference codeword value of the one or more reference images and the average source codeword value of the one or more reference images.

[0217] EEE31. The method as described in EEE27, further comprising: applying a specific photoelectric conversion function (OETF) associated with the one or more reference images to these tone-matched source images to generate one or more automatically adapted source images that are adapted to a reference visual appearance associated with the reference camera.

[0218] EEE32. The method as described in EEE27, wherein the results of these noise characteristic descriptions performed using the one or more source images are represented at least in part by a noise metric calculated using a separate signal-to-noise ratio between the standard deviation of a separate image patch and the intensity of the separate image patch; wherein the image patches are partitioned from the one or more source images.

[0219] EEE33. The method as described in EEE32, wherein the individual image patches segmented from the one or more source images and used to calculate the individual signal-to-noise ratios constitute a set of image patches; wherein the set of image patches does not include any image patches in the one or more source images that contain visually perceptible edges detected using one or more edge detection methods.

[0220] EEE34. The method as described in EEE32, wherein these individual signal-to-noise ratios are calculated for each color channel of the working color space representing the one or more source images.

[0221] EEE35. The method as described in EEE32, wherein the noise metric is determined for the luminance channel of the working color space; wherein two additional noise metrics are determined for the chrominance channel of the working color space; wherein the total noise is calculated for the one or more source images based on the noise metric and the two additional noise metrics; wherein the total noise is used to modify the initial gain values ​​to the modified gain values.

[0222] EEE36. The method as described in EEE35, wherein the total noise is used to modify these initial gain values ​​to the modified gain values ​​in the approximation inverse relation.

[0223] EEE37. The method as described in EEE35, wherein the total noise is used to modify these initial gain values ​​to the modified gain values ​​in the spline relation.

[0224] EEE38. The method as described in EEE35, wherein the total noise is used to apply the maximum gain limit to the modified gain value.

[0225] EEE39. The method as described in EEE27, wherein for each pair of modified gain values ​​for two different color channels, a first ratio value is maintained that is the same as a second ratio value maintained by the corresponding initial gain values ​​for the different color channels.

[0226] EEE40. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause to perform the method as described in any one of EEE27 to EEE39.

[0227] EEE41. A computing device comprising one or more processors and one or more storage media storing an instruction set that, when executed by the one or more processors, causes to perform a method as described in any one of EEE27 to EEE39.

Claims

1. A method comprising: Identify one or more reference images and one or more source images in the working color space, wherein the one or more reference images are derived from a reference camera, and wherein the one or more source images are derived from a source camera different from the reference camera; An initial gain value is obtained between the one or more reference images and the one or more source images, wherein the initial gain value is determined based on the reference codeword values ​​of the one or more reference images and the source codeword values ​​of the one or more reference images; The initial gain value is adjusted to a modified gain value based at least in part on one or more results of a noise characteristic description performed using the one or more source images; The modified gain value is applied to the one or more source images to generate one or more leveled source images; The one or more leveled source images are converted into one or more tone-matched source images, at least in part based on one or more source-to-reference tone mappings.

2. The method of claim 1, wherein the working color space represents a linear color space; wherein the linear color space is one of the following: linear RGB color space, linear YUV color space, linear YCbCr color space, or different linear color spaces.

3. The method of claim 1, wherein the working color space represents a nonlinear color space; wherein the nonlinear color space is one of the following: perceptual quantization (PQ) color space, mixed log-gamma (HLG) color space, nonlinear LMS color space, nonlinear ICpCt color space, or different nonlinear color spaces.

4. The method of any of the preceding claims, wherein the initial gain value is determined as an average ratio based on the average reference codeword value of the one or more reference images and the average source codeword value of the one or more reference images.

5. The method of any of the preceding claims, further comprising: A specific photoelectric conversion function (OETF) associated with the one or more reference images is applied to the tone-matched source image to generate one or more automatically adapted source images that are adapted to the reference visual appearance associated with the reference camera.

6. The method of any preceding claim, wherein the result of the noise characteristic description performed using the one or more source images is represented at least in part by a noise metric calculated using a separate signal-to-noise ratio between the standard deviation of a separate image patch and the intensity of the separate image patch; wherein the image patch is partitioned from the one or more source images.

7. The method of claim 6, wherein the individual image blocks segmented from the one or more source images and used to calculate the individual signal-to-noise ratio constitute a set of image blocks; wherein the set of image blocks does not include any image blocks in the one or more source images that contain visually perceptible edges detected using one or more edge detection methods.

8. The method of claim 6 or 7, wherein the individual signal-to-noise ratio is calculated for each color channel of the working color space representing the one or more source images.

9. The method of any one of claims 6 to 8, wherein the noise metric is determined for the luminance channel of the working color space; wherein two additional noise metrics are determined for the chroma channel of the working color space; wherein the total noise is calculated for the one or more source images based on the noise metric and the two additional noise metrics; wherein the total noise is used to modify these initial gain values ​​to the modified gain values.

10. The method of claim 9, wherein the total noise is used to modify the initial gain value to the modified gain value in the approximation inverse relation.

11. The method of claim 9, wherein the total noise is used to modify the initial gain value to the modified gain value in the spline relation.

12. The method of claim 9, wherein the total noise is used to apply the maximum gain limit to the modified gain value.

13. The method of any of the preceding claims, wherein each pair of modified gain values ​​for two different color channels maintains a first ratio value, the first ratio value being the same as a second ratio value maintained by the corresponding initial gain values ​​for the different color channels.

14. A method comprising: Identify one or more reference images in the working color space and one or more source images in the same working color space, wherein the one or more reference images are initially captured by a reference camera and wherein the one or more source images are initially captured by a source camera different from the reference camera; Identify one or more common image features between the one or more reference images and the one or more source images, wherein the one or more common image features in the one or more reference images include a plurality of reference pixels, and wherein the one or more common image features in the one or more source images include a plurality of source pixels; One or more source-to-reference codeword mappings are generated using the distribution of reference codeword values ​​of the plurality of reference pixels and the distribution of source codeword values ​​of the plurality of source pixels; Based at least in part on the codeword mapping from one or more sources to references, a specific source image initially captured by the source camera is transformed into a corresponding automatically adapted source image that is adapted to a reference visual appearance associated with one or more reference images initially captured by the reference camera.

15. The method of claim 14, wherein the reference camera differs from the source camera in one or more of the following: output image color space, output image dynamic range, output image color gamut, output image bit depth, image signal processor, device manufacturer, or device model.

16. The method of claim 14 or 15, wherein the one or more reference images are initially captured by the reference camera in a hybrid log-gamma (HLG) Rec.2020 RGB color space; wherein the one or more reference images in the HLG Rec.2020 RGB color space are linearized to be represented in the linearized Rec.2020 RGB color space using a nonlinear inverse photoelectric conversion function (OETF).

17. The method of any one of claims 14 to 16, wherein the common image features are identified at least in part based on one or more projection transformations constructed from image features detected using the Accelerated Robust Features (SURF) algorithm.

18. The method of any one of claims 14 to 17, wherein the one or more source-to-reference codeword mappings are generated by one or more of the following: cumulative density function (CDF) matching operation, or one or more artificial neural networks.

19. The method of any one of claims 14 to 18, wherein at least one of the one or more source-to-reference codeword mappings is represented by one of the following: a mapping function, a spline function, a piecewise function, a piecewise polynomial function, or a lookup table.

20. The method of any one of claims 14 to 19, wherein the specific source image represents one of the one or more source images.

21. The method of any one of claims 14 to 20, wherein the one or more reference images and the one or more source images are captured by the reference camera and the source camera, respectively, from a common physical environment; wherein the reference camera and the source camera are located at different spatial locations within the common physical environment.

22. The method of any one of claims 14 to 21, wherein the one or more source images are converted into one or more corresponding auto-adapted source images, the one or more corresponding auto-adapted source images being matched with the reference visual appearance associated with the one or more reference images initially captured by the reference camera; wherein at least a subset of the one or more corresponding auto-adapted source images is combined with at least a subset of the one or more reference images to form a single sequence of temporally continuous images along a single playback timeline.

23. The method of any one of claims 14 to 22, wherein the corresponding auto-adapted image is generated by performing an RGB leveling operation on an intermediate image, the intermediate image being generated at least in part by applying the one or more source-to-reference codeword mappings to the particular source image.

24. The method of any one of claims 14 to 22, wherein the corresponding autofit image is generated by performing either a saturation matching operation or a saturation leveling operation on an intermediate image, the intermediate image being generated at least in part by applying the one or more source-to-reference codeword mappings to the particular source image.

25. A non-transitory computer-readable storage medium storing software instructions that, when executed by one or more processors, cause to perform the method as described in any one of claims 1 to 24.

26. A computing device comprising one or more processors and one or more storage media storing an instruction set, the instruction set causing, when executed by the one or more processors, to perform the method as claimed in any one of claims 1 to 24.