User-adjustable semantic tone mapping for photographic capture and editing

The system allows user-adjustable semantic tone mapping through a tone gain map and semantic masks, addressing the loss of original light ratios in tone mapping and providing flexible post-processing options.

WO2026055691A1PCT designated stage Publication Date: 2026-03-12APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing tone mapping techniques in photography lose the inherent light ratio of the captured scene, making it impossible to later render images with the original light conditions, and there is a lack of user-adjustable and reversible local tone mapping with semantic information integration.

Method used

A system that allows for user-adjustable semantic tone mapping by generating a tone gain map from a linear image and combining it with semantic masks and user inputs to produce a locally tone-mapped image, preserving the original light ratios and enabling flexible post-processing.

Benefits of technology

Enables users to adjust tone mapping in real-time and post-processing without increasing file size, allowing for both aesthetic and artistic flexibility while maintaining the original light conditions of the captured scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045592_12032026_PF_FP_ABST
    Figure US2025045592_12032026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of photography, videography and digital graphics. More particularly, but not by way of limitation, it relates to a camera control system and image processing system, which can take an input asset (e.g., a still image, a video, or a still image with an associated video) and produce a user-adjustable semantically, locally tone-mapped version of the input asset. Some embodiments include obtaining a "linear" version of the input asset (i.e., a scene-referred version that is not tone mapped), producing a tone gain map that maps between an (already tone-mapped) input asset and the linear version of the input asset, as well as one or more semantic masks for the input asset and user inputs related thereto. Then, the input asset may be modified according to the tone gain map in user-customizable ways, e.g., according to the one or more semantic masks and inputs related thereto.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT COOPERATION TREATY (PCT) APPLICATIONTitle USER- AD JUST ABLE SEMANTIC TONE MAPPING FOR PHOTOGRAPHICCAPTURE AND EDITINGInventors JOHNSON ET AL.Docket No. : P69531WO1 (119-2207WO1) Customer No : 61947TECHNICAL FIELD

[0001] This disclosure relates generally to the field of photography, videography and digital graphics. More particularly, but not by way of limitation, it relates to a camera control system and image processing system, which can take an input asset (e.g., a still image, a video, or a still image with an associated video) and produce a user-adjustable, semantically tonemapped version of the input asset.BACKGROUND

[0002] In the field of photography, tone mapping is a common technique used to compress the large dynamic range (i.e., the ratio of the lightest areas in the scene to the darkest areas in the scene) into the smaller dynamic range that is typically available in an image file format or output display. These techniques usually come in a combination of two forms: global tone mapping and local tone mapping (referred to collectively herein as “LTM”). The mathematics behind tone mapping operation is typically quite complex, and is usually left to the designers of the camera system to specify.

[0003] The end result of these typical LTM operators is essentially to locally adjust the image to fall into the available dynamic range of either a typical file format or output display. For example, an LTM operator may have the role of mapping a real word scene that has a dynamic range of say, 1-10,000 nits, into the much smaller range of values available in an image file format (e.g., a JPEG or HEIC). For example, the values representing the scene brightness levels may need to be compressed into a range of 0..255, 0.. 1,024, or the like. To do this, the LTM will locally adjust the image — typically by lifting the darkest shadow areas in the scene and bringing down (i.e., lowering) the brightest highlight regions of the image. Often, these operators are designed to mimic the adaptive behavior of the of the human visual system (HVS) itself, and they generally work quite well for typical artistic intents and displays.Customer No. 61947

[0004] While this dynamic range compression is often necessary to squeeze as much information into the final image itself, by nature, an LTM operator destroys the inherent ratio of light that was in the actual scene when captured. For example, once an LTM operation has lifted the shadows and brought down the highlights in a captured image, the information of what regions in the captured scene were actually was a shadow (and / or how dark those shadows were) in lost. Thus, if the end user wishes to later render their image more like how the light was originally in the captured scene, it is impossible to do so.

[0005] Modem day image processing for still images often relies on semantic information, such as scene classification and segmentation. The segmentation masks can be used to apply different processing algorithms or parameters to various components of the scene and allow for dedicated processing of, e.g., person regions, sky regions, skin tone regions, or the like. Modem machine learning (ML)-based techniques provide fairly reliable image segmentation of various object types (e.g., persons, sky, skin, etc.), and the inferred masks can be aligned with the image content using matting algorithms. However, high image quality segmentation is not presently utilized in conjunction with user-adjustable and reversible local tone mapping techniques.

[0006] Over time and / or during an editing process, a user's preference for the look of processed images or videos might change, and users might want to go back to the look of the light as it was captured in the ‘Taw-” (i.e., unstyled) image(s), i.e., without semantic effects applied or with different LTM effects applied. Thus, what is further needed is an approach to allow semantically-based image and video LTM operation with reversibility and ongoing modification capabilities.SUMMARY

[0007] Devices, methods, and non-transitory program storage devices (PSDs) are disclosed herein to obtain an input asset (e.g.. a still image, a video, or a still image with an associated video) that is tone-mapped according to “stock” or “default” processing operations, and then allow- for user- adjustable tone mapping (e.g., using on a tone gain map produced for the image based upon a scene-referred linear light version of the image) according to customizable semantic masks (and user inputs related thereto) to and output a locally tone-mapped image asset that is rendered with a particular aesthetic style to match a particular artistic intent (e.g.. balancing betw-een recovering the light levels more closely to how they originally appeared inCustomer No. 61947 the captured scene and fully recovering details related to a subject or other region of interest in the image).

[0008] According to one embodiment, a device is disclosed, comprising: a memory; and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to: obtain a first linear image of a scene; obtain, from the image capture device, a first tone-mapped image of the scene having a first resolution, wherein the first tone-mapped image is a tone-mapped version of the first linear image of the scene; divide the first linear image by the first tonemapped image to produce a tone gain map for the first tone-mapped image; obtain one or more semantic masks for the first tone-mapped image; obtain user input related to the one or more semantic masks; and combine the first tone-mapped image and the tone gain map according to the one or more semantic masks and the respective user input related to the one or more semantic masks to produce a semantically, locally tone-mapped output image.

[0009] According to some embodiments, the first linear image comprises an image with scene-referred light values.

[0010] According to other embodiments, the first linear image has the first resolution or a resolution that is smaller than (or larger than) the first resolution.

[0011] According to other embodiments, the first linear image comprises a single exposure image or a fusion of multiple bracketed image captures.

[0012] According to other embodiments, the instructions further comprise instructions causing the one or more processors to: obtain a global tone curve for the first linear image and / or an exposure metric for the first linear image. According to some such embodiments, the global tone curve and / or the exposure metric may be applied to the first linear image before the dividing occurs.

[0013] According to other embodiments, the exposure metric comprises one or more of: an aperture setting, an ISO setting, or an integration time.

[0014] According to other embodiments, at least one of the one or more user inputs comprises: a slider value, a dial value, a directional pad value, a numerical value, or a qualitative value.Customer No. 61947

[0015] According to other embodiments, the instructions to combine the first tone-mapped image and the tone gain map according to the one or more semantic masks and the respective user input related to the one or more semantic masks further comprise instructions causing the one or more processors to: perform a pixel-wise multiplication operation between the first tonemapped image and the tone gain map. According to some such embodiments, the pixel-wise multiplication operation between the first tone-mapped image and the tone gain map may be performed subject to corresponding pixel values in the one or more semantic masks and the respective user input related to the one or more semantic masks.

[0016] According to other embodiments, at least one of the one or more semantic masks comprises: a face mask, a skin tone mask, a subject mask, a person mask, an image foreground mask, an image background mask, a pet mask, a sky mask, a water mask, a teeth mask, a lips mask, an eyes mask, an eye glasses mask, a plated food mask, a Sun mask, a light mask, or a hair mask.

[0017] According to other embodiments, at least one of the one or more semantic masks comprises a machine learning (ML)-based semantic mask.

[0018] According to other embodiments, the first tone-mapped image may be obtained from the image capture device during an image preview mode of operation.

[0019] According to still other embodiments, the first tone-mapped image may be an image that is already stored in the memory.

[0020] Various other device, non -transitory program storage device (PSD) and method embodiments are also disclosed herein. Such PSDs are readable by one or more processors. Instructions may be stored on the PSDs for causing the one or more processors to perform any of the embodiments disclosed herein. Various electronic devices are also disclosed herein, e.g., comprising memory, one or more processors, one or more image capture devices, displays and / or other electronic components (e.g., IMUs, microphones, etc.), and programmed to perform in accordance with the various method and PSD embodiments disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1A illustrates an exemplary scene to be captured by an image capture device, according to one or more embodiments.Customer No. 61947

[0022] Figure IB illustrates an exemplary scene-referred linear light image version of the captured scene, according to one or more embodiments.

[0023] Figure 1C illustrates an exemplary “stock” tone-mapped image version of the captured scene, according to one or more embodiments.

[0024] Figure 2 illustrates an exemplary tone gain map generation operation, according to one or more embodiments.

[0025] Figure 3 illustrates an exemplar}' user-adjustable semantic local tone mapping operation, according to one or more embodiments.

[0026] Figure 4 illustrates an exemplary set of semantically tone mapped images, according to one or more embodiments.

[0027] Figure 5 is a flow diagram illustrating a method of performing user-adjustable semantic tone mapping, according to various embodiments.

[0028] Figure 6 is a block diagram illustrating a programmable electronic computing device, in which one or more of the techniques disclosed herein may be implemented.DETAILED DESCRIPTION

[0029] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the inventions disclosed herein. It will be apparent, however, to one skilled in the art that the inventions may be practiced without these specific details. In other instances, structure and devices are shown in block diagram form in order to avoid obscuring the inventions. References to numbers without subscripts or suffixes are understood to reference all instance of subscripts and suffixes corresponding to the referenced number. Moreover, the language used in this disclosure has been principally selected for readability and instructional purposes and may not have been selected to delineate or circumscribe the inventive subject matter, and, thus, resort to the claims may be necessary’ to determine such inventive subject matter. Reference in the specification to “one embodiment” or to “an embodiment” (or similar) means that a particular feature, structure, or characteristic described in connection with the embodiments is included in at least one embodiment of one of the inventions, and multiple references to “one embodiment” or “an embodiment” should not be understood as necessarily all referring to the same embodiment.

[0030] This present disclosure relates to a camera control and asset processing system that allows for an asset to be rendered with a particular style, e.g., through modification of the toneCustomer No. 61947 and / or color content in the asset. Such systems allow for the continuous restyling of the asset through the ability to add, change, or remove various semantically-based local tone mapping operations — without losing the ability to return the original un-styled or ‘’stock” input.

[0031] The system may also consist of a display device that presents a preview of the semantically-based locally tone mapped asset that will be processed as the image(s) are being captured by a user with an image capture device of an electronic device. In addition, the techniques described herein are also able to store metadata representations of the necessary scene-referred light levels, semantic information, global tone curves, and / or exposure metrics, which allow the original input image (and / or differently tone-mapped versions of the original input image) to be recovered or recreated on-demand.

[0032] According to some embodiments, the user-adjustable semantic tone mapping systems described herein may be enabled by a particular system and / or set of software- implemented algorithms, as well as novel additions to standardized image and video file formats, which novel file format additions are used to encapsulate the necessary metadata to enable the adjustable semantic local tone mapping systems described herein.

[0033] Exemplary Non-Tone-Mapped and Tone-Mapped Images

[0034] As described above, the dynamic range compression of ty pical LTM operations removes or deletes from the image data the inherent ratio of light that was in the actual scene when captured. Thus, if the end user wishes to later render their image more like how the light was originally in the captured scene, it is impossible to do so. This is shown in greater detail in Figures 1A-1C, below.

[0035] Turning now to Figure 1A, an exemplary scene 100 to be captured by an image capture device is shown, according to one or more embodiments. Scene 100 includes the Sun 106, and a human subject 102. who is standing under a tree 104, while being photographed by an exemplary image capture device, i.e., camera 108. The exemplary scene 100 may represent a bright and sunny day, yet the subject 102 may be partially (or completely) placed into shadows created by tree 104. Thus, exemplary scene 100 is a ty pical example of scene that is difficult for traditional tone mapping algorithms to handle, without imposing a particular stylistic preference that may or may not match the user’s photographic or artistic intent. For example, in order to prevent highlight blowouts in the background of the scene, the image highlights may be lowered, causing already dark / shadowy areas, such as on subject 102’s face to appear even more dark. Alternatively, by raising the shadows to bring out additional detailCustomer No. 61947 on subject 102 ’s face, a tone mapping algorithm may risk further exacerbating highlight blowouts in the already bright background of the sunny scene. Another aspect of this type of scene that may prove difficult for traditional tone mapping algorithms to handle is so called “dappled” light 112 (i.e., a pattern of bright spots and shaded areas) appearing on the subject 102’s face, e.g., as caused by some of the light on the subject 102’s face being blocked by the branches and leaves from tree 104. In some cases, an end user might only be interested in seeing the face of the subject 102, so the tone mapping will remove the shadows of the dappled light to reveal the face more fully. However, in other cases, an end user may instead have actually taken that photo specifically because of the unique qualities and the interplay of light and shadows on the subject 102’s face. In those types of cases, the end user may actually want to darken the shadows even more, e.g., to show off the uniqueness of the lighting in the scene. The same end user might also want to share one version of the image with the lifted shadows to their friends and family, while keeping a version emphasizing the “dappled” lighting effects for themselves. Thus, the techniques disclosed herein allow for both potential renderings of the image, i.e.. those where the viewer is “seeing the person” — and those where the viewer is “seeing the light” in the scene.

[0036] Turning now to Figure IB, an exemplary' scene-referred linear light image version of the captured scene 120 is shown, according to one or more embodiments. In the scene- referred linear light image 120, also referred to herein as a “non-tone-mapped” image, both the rendition of the subject 102A and the tree 104A reflect the dark shadows and dappled light patterns 112 that were present in the scene as originally captured, as well as the bright and sunny background 110A of the scene as originally captured. The linear light image 120 may comprise a single image capture, or it may be produced as a fusion of multiple bracketed captures (i.e., a high dynamic range (HDR) image), i.e., in order to represent more of the original light in the scene as captured in the image data. According to some implementations, a global tone curve and / or an exposure metric, e.g., representative of the average or “baseline” exposure of the image, may be stored along with linear light image 120. This baseline exposure may represent the overall exposure gain that is typically realized by the local tone mapping operator, i.e., the amount of gain that would move the linear image into the same exposure range as the “stock” tone-mapped image, as will be described below' with reference to Figure 1C

[0037] Turning now to Figure 1C, an exemplary “stock” tone-mapped image version of the captured scene 140 is shown, according to one or more embodiments. In the “stock” tone-Customer No. 61947 mapped image 140, the rendition of the subject 102B and the tree 104B reflect that the dark shadows that were present in the scene as originally captured have been brightened up to reveal more detail, e.g., in the face of subject 102B. (It is also possible that the sunny background of the image HOB may experience even more exposure highlight blowouts, based on the stock tone mapping operations.)

[0038] In many cases, e.g.. when auser is only interested in seeing the detail in the subject's face well, the tone mapping treatment detailed in Figure 1C may very well be desired. However, if the user is interested in reproducing the quality of light in the original scene, e.g., the interplay of light and shadows underneath the tree, that information is lost. In other words, after shadow detail recovery, it is impossible for the end user to manipulate the locally tonemapped image to obtain the interesting quality' of light that was in the original scene.

[0039] Image rendering and, in particular, tone mapping, is ultimately a user-preference, e.g., based upon both the aesthetic taste of the user, as well as the light and subject of the original photographed scene itself. Thus, the techniques disclosed herein present systems and methods that allow' a user to control the global and local tone mapping operations both at capture time, and after an image has been stored — allowing the user to change their mind and adjust the tone mapping further in an editing session. The techniques disclosed herein may also advantageously include additional semantic information with an image than is typically available to a camera at capture time, including scene type, and one or more ML-based segmentation masks (typically, but not limited to, a face mask, a skin tone mask, a subject mask, a person mask, an image foreground mask, an image background mask, a pet mask, a sky mask, a water mask, a hair mask, or the like).

[0040] Some prior techniques have attempted to solve the problem by storing the scene’s original light ratios (known as scene-referred linear light) in a “raw” image file. These raw image files may also contain a grid of local tone mapping curves, along with a global tone curve. Then, when the user chooses to open such a file, e g., in an image editing application, the application must take the linear light representation of the scene and apply the grid of local tone mapping curves and the global tone curve before it can be displayed. It is difficult to semantically manipulate this grid of tone curves later to achieve a different style of tone mapping. Additionally, these types of raw image files will end up being considerably larger than typical HEIC / JPEG files and typically require some level of image processing before they can be view ed or shared, w hich is not desirableCustomer No. 61947

[0041] Thus, the approaches to user-adjustable local tone mapping disclosed herein may allow for the semantically-based manipulation of local tone mapping curves in real-time and / or post-processing operations, without substantially increasing the overall fde size of the resultant output image. For example, it is possible (although not necessary) to store the “secondary” asset of the resultant output image (i.e., the asset that may be used to semantically guide how the raw image is developed at display time) at something less than full resolution size. For example, a linear light version of the original input image (or, alternatively, a tone-mapped version of the original input image) may be stored at a “thumbnail” resolution or as an otherwise heavily-compressed version, so that its impact on the resultant output image's file size is not substantial.

[0042] Tone Gain Map Generation

[0043] Turning now to Figure 2, an exemplary' tone gain map generation operation 200 is illustrated, according to one or more embodiments. According to some embodiments, the first stage in a tone gain map generation operation is to take the linear light image 120 (optionally gained up by the baseline exposure metric and / or with the global tone curve applied), and divide (205) that linear light image 120 by the “stock” tone-mapped image 140. In some implementations, the baseline exposure metric and amount of global tone to be applied to the image may also be exposed to for control by the end user, i.e., assuming the values are stored as metadata in the image. The result of the division operation at 205 is atone gain map 210.

[0044] According to some embodiments, the tone gain map 210 specifies how each pixel in the stock tone-mapped image 140 has been altered by the local tone mapping operation. For example, according to some implementations, when the value in the tone gain map 210 is above 1 .0, it may indicate that the corresponding pixel(s) in the linear light image 120 has been tone mapped to be lighter, while tone gain map 210 values below 1.0 indicate that the corresponding pixels(s) in the linear light image 120 have been mapped to be darker. In some embodiments, tone gain map 210 may specify distinct gain values at each pixel location for each color channel used in the particular image format (e.g., a red (R) gain value, a green (G) gain value, and a blue (B) gain value, in the case of RGB-encoded image file formats).

[0045] This tone gain map 210 is thus a useful and powerful piece of metadata that gives an end user much more control than simply manipulating the individual LTM curves themselves. For instance, an end user might want to maintain the highlight recovers’ of the stock tone-mapped image 140, but they may not want to keep the shadow lifting that it imposedCustomer No. 61947 on the linear light image 120. In such a situation, because the tone gain map 210 advantageously specifies precisely which pixels have been recovered and which have been lifted (and by how much), it is possible to provide the end user with that degree of customizable choice for post-processing adjustment.

[0046] Creation of User-Adjustable, Semantically Locally Tone-Mapped Images

[0047] Turning now to Figure 3. an exemplary user-adjustable semantic local tone mapping operation 300 is illustrated, according to one or more embodiments. In some embodiments, to create a new version of a tone-mapped image, the operation 300 may being with the stock tonemapped image 140, the tone gain map 210, one or more semantic masks 220, and user input from one or more controls 225 that suggest the user’s intent with respect to the desired tone mapping treatment to be applied to the various semantically -defined regions of the image. For example, a user input control 225A may be used to allow the user to input their preference with respect to a first “Semantic Mask A” 220A (which, in this case, is a semantic mask over an identified foreground subject in the scene), and a user input control 225B may be used to allow the user to input their preference with respect to a second “Semantic Mask A” 220B (which, in this case, is a semantic mask over identified skin tone regions in the scene). It is to be understood that these semantic masks 220A / 220B are merely exemplary and that, greater, fewer, and / or different semantic masks may be utilized in a given implementation. Further, semantic masks may be applied in a serial fashion or after a ‘union’ or other morphological operation has been applied to any masks that will be utilized in the given implementation. For example, in the case of combining “skin” and “person” semantic masks, the person mask may be inclusive of all the “skin” regions in the image, but the end user may want to apply some effect at 100% to the skin regions of the image — but only at 50% strength to the person regions of the image. In this case, assuming both masks have values ranging from 0 to 1 (e.g., where ‘ 1 ’ is indicative of full confidence in a pixel being a “skin” pixel or a “person” pixel), the masks may simply be added together. Then, the resultant combined mask (with values ranging from 0 to 2) may be scaled by dividing all values by 2, resulting in a mask that is 1.0 where skin and person overlap and 0.5 where there is a person — but not skin. In this example, the desired effect will be applied at 100% strength to skin regions and at 50% strength to person (but not skin) regions.

[0048] According to some embodiments, the corresponding values in each of the stock tone-mapped image 140 and tone gain map 210 may be multiplied together in a pixelwise fashion to obtain the pixel values for the tone mapped output image, subject to the values inCustomer No. 61947 the one or more semantic masks 220 and user input values from respective one or more controls 225. In other words, the tone mapping preferences specified by user input control 225A may be applied only in the “white” areas of semantic mask A 220A, and the tone mapping preferences specified by user input control 225B may be applied only in the “white” areas of semantic mask B 220B. It is to be understood, that although binary semantic masks 220 are illustrated in Figure 3. it is also possible for semantic masks to use (or convert to the use of) continuous values (e.g., between 0..1) at each pixel location, e.g., based on the confidence the semantic segmentation algorithm may have in a given pixel actually being part of the relevant semantic class. In such cases, the tone mapping preferences specified by the relevant user input control 225 may be applied in accordance with the non-binary values in the respective semantic mask 220 (e.g., via a further multiplication operation with the corresponding pixel value in the respective non-binary semantic mask).

[0049] Turning now to Figure 4, an exemplary set 400 of semantically tone mapped images 410A and 410B are shown, according to one or more embodiments. For one example, an end user may want the shadow details on the subject’s face brought back down to the original lighting conditions (e.g., with an interesting interplay between the sunlight and the heavy shadows caused by the tree that were cast on the subject’s face during the original image capture). In another examples, a user may prefer that the shadows in the background are brought down even further, e.g., to bring more emphasis to the face of the subject. In still another example, a user may actually want the highlights on the subject’s face to be rendered even brighter, i.e., to give a dramatic SLR-like look to the resultant tone-mapped output image.

[0050] As may now be appreciated, there are many possible ways that an end user may wish to tone map (and re-tone map) an originally-captured image. The tools and techniques disclosed herein can give the end user a large amount of flexibility in terms of how and when they choose to change the tone mapping of an image, either at capture time or in postproduction.

[0051] In accordance with some embodiments, the user’s input with respect to particular semantic masks and / or respective tone mapping strength decisions may be mapped to simple sliders or other graphical user interface features, such as dials, directional pads, qualitative terms, etc. For instance, looking at exemplary sliders 425A / 425B in Figure 4, placing an indicator at one end of the slider (e.g., the bottom, as illustrated in slider 425B of Figure 4) could indicate the user’s desire to return the tonality (i.e.. light levels) of the tone mapped image 410 all the way back towards the original scene-referred linear light levels (e.g., such as linearCustomer No. 61947 light image 120, described above), subject to any corresponding semantic masks (e.g., such as semantic mask 420B of Figure 4, which would, in this example, limit any tonality modifications to the skin tone portions of the original tone mapped image 410B). This is reflected in exemplary semantically tone mapped image B (430B), which returns some of the original shadow detail (including the ’’dappled" light 112) onto the subject's face, but only in those portions of the image that are predicted to contain skin tones. On the other hand, placing the indicator at the other end of the slider (e g., the top, as illustrated in slider 425A of Figure 4) could indicate the user’s desire to increase the tonality (i.e., light levels) of the original tone mapped image 410A to recover as much of the subject detail (and remove as much shadow) as is possible. This is reflected in exemplary semantically tone mapped image A (430A), which highlights the appearance of the subject tin the image, but is only applied in those portions of the image that are predicted to contain the subject. Continuing with the simple slider example shown in Figure 4, it follow s that placing the indicator near the middle of the slider could strike a tonality balance between the original scene-referred linear light levels and attempting to fully recover as much detail as possible in the image. In still other embodiments, differently-mapped sliders could also be employed, if so desired. For example, moving the indicator towards either extreme of the slider could actually serve as an indication to apply even more “tone mapping” or more lifting of shadows in the image. Sliders could also be mapped such that moving the indicator towards either extreme of the slider could serve as an indication to apply even more contrast to the image, e.g., such that shadows can actually go darker than they were in the linear image (e.g., for artistic reasons, or otherwise).

[0052] In some embodiments, the tone gain map 210 may be further analyzed when making adjustments to highlight (and / or shadow) regions of a captured image. For example, in some implementations, tone gain map values greater than 1.0 indicate regions of the image where the local tone mapping is lifting shadows, whereas tone gain map values less than 1.0 indicate regions of the image where the local tone mapping is lowering highlights. If the end user likes the image processing that is being applied to the highlights regions, all the gain values less than 1.0 may be maintained, while the image can revert to its values from the “scene-referred” linear light image for regions of the image with gain values greater than 1.0. Additionally, because the linear image is “scene referred,” a “highlight” map and ’’shadow" map may be generated based on the luminance of the linear image (i.e., since it is known that the linear image maps directly to how much light was really in the captured scene). It is more difficult to create suchCustomer No. 61947 a highlight and / or shadow map with an already tone-mapped image, since the tone mapping itself could lift shadow regions even above highlight regions, in some renderings.

[0053] Exemplary Methods of Performing User-Adjustable Semantic Tone Mapping

[0054] Turning next to Figure 5, a flow diagram is shown, illustrating a method 500 of performing user-adjustable semantic tone mapping. First, at Step 502, the method 500 may obtain a first linear image of a scene, i.e.. a version of the scene with scene-referred linear light. Next, at Step 504, the method 500 may obtain a first tone-mapped “stock’’ image of the scene, wherein the first tone-mapped image is a tone-mapped version of the first linear image of the scene. In some embodiments, the first linear image may have the same full-resolution as the stock image. In other embodiments, the first linear image may have a larger resolution (or a smaller resolution) than the stock image, depending on the desires of a given implementation. In still other embodiments, the linear image may comprise a single image capture. In other such embodiments, the linear image may more likely be produced as a fusion of multiple bracketed captures (i.e., a high dynamic range (HDR) image).

[0055] Next, at Step 506, the method 500 may optionally obtain a global tone curve and / or an exposure metric for the first linear image. The exposure metric may, e.g., comprise one or more of: an aperture setting, an ISO setting, or an integration time. Next, at Step 508, the method 500 may divide the first linear image (optionally, as gained up by the global tone curve and / or with the exposure metric from Step 506 applied) by the first tone-mapped stock image to produce a tone gain map for the first tone-mapped stock image. For example, in some embodiments, it may be desirable for the “overall brightness” to be about the same between the linear image and the first tone-mapped image, thus the exposure metric may comprise the result of calculating, on average, how much “brighter” the local tone mapping made the first tone-mapped image, and then applying that same amount of gain to the linear image, in the form of an “exposure multiplier” metric.

[0056] Next, at Step 510, the method 500 may obtain one or more semantic masks for the first tone-mapped stock image. At Step 512. the method 500 may also obtain user input related to the one or more semantic masks, e.g., in the form of a specified numerical value, a slider value, a dial value, a 2-D directional pad position, or the like. In some embodiments, the user input may be normalized, e.g., to a scale of 0..1, wherein, e.g.: a value of 0 implies the user’s desire not to apply the tone gain map to the first tone-mapped stock image (e.g., at least within the bounds of the particular semantic mask that the user input applies to); a value of 0.5 implies the user’s desire to apply a balance between the first tone-mapped stock image and the firstCustomer No. 61947 linear image, and a value of 1 implies the user’s desire to fully recover as much detail as possible in the output image, i.e., recreating the look of the original scene-referred light in the image. As mentioned above, however, other user input mappings are possible, as well, such as those that allow users to implement tonality changes even beyond what was visible in the originally captured linear light image.

[0057] Finally, at Step 514, the method 500 may combine: (1) the first tone-mapped stock image; (2) the tone gain map; (3) the one or more semantic masks; and (4) the respective user input related to the one or more semantic masks to produce a semantically, locally tone mapped output image. For example, in some embodiments, the combining may comprise a pixel-wise multiplication operation between the elements (1) through (4). above, wherein it is to be understood that tone mapping for a particular semantic region does not occur outside the “mask” region for the respective semantic region.

[0058] The various methods and techniques described herein, e.g., with reference to Figures 1-5 may be performed by an electronic device, e.g., via being initiated by an application (or “App”) executing on the device and / or the device’s native operating system (OS). For example, an App executing on the device could initiate or implement all of the steps in a method, or at least a portion of the steps in the method, while making calls to the device’s OS to perform other steps in the method. Similarly, a device’s OS can receive API calls from an App or elsewhere and process / perform the calls to cause the method to be performed by the device(s). In some implementations, one or more of the processing steps may also be performed by a device that is remote to the electronic device, e.g., on a smartphone, laptop or other electronic device associated with the user, and / or on a server device accessible to the electronic device via a network connection (which server device may, e.g., have greater processing capacity than a wearable electronic device).

[0059] Exemplary Electronic Computing Devices

[0060] Referring now7to Figure 6, a simplified functional block diagram of illustrative programmable electronic computing device 600 is shown according to one embodiment. Electronic device 600 could be, for example, a mobile telephone, personal media device, portable camera, or a tablet, notebook or desktop computer system. As shown, electronic device 600 may include processor 605, display 610, user interface 615, graphics hardware 620, device sensors 625 (e.g., proximity sensor / ambient light sensor, accelerometer, inertial measurement unit, and / or gyroscope), microphone 630, audio codec(s) 635, speaker(s) 640, communications circuitry 645, image capture device 650, which may, e.g., comprise multipleCustomer No. 61947 camera units / optical image sensors having different characteristics or abilities (e.g., Still Image Stabilization (SIS), HDR, OIS systems, optical zoom, digital zoom, etc.), video codec(s) 655, memory 660, storage 665, and communications bus 670.

[0061] Processor 605 may execute instructions necessary' to cany' out or control the operation of many functions performed by electronic device 600 (e.g., such as the generation, processing, and / or streaming of image and video data, in accordance with the various embodiments described herein). Processor 605 may, for instance, drive display 610 and receive user input from user interface 615. User interface 615 can take a variety of forms, such as a button, keypad, dial, a click wheel, keyboard, display screen and / or a touch screen. User interface 615 could, for example, be the conduit through which a user may view a captured video stream and / or indicate particular image frame(s) that the user would like to capture (e.g., by clicking on a physical or virtual button at the moment the desired image frame is being displayed on the device’s display screen). In one embodiment, display 610 may display a video stream as it is captured while processor 605 and / or graphics hardware 620 and / or image capture circuitry contemporaneously generate and store the video stream in memory 660 and / or storage 665. Processor 605 may be a system-on-chip (SOC) such as those found in mobile devices and include one or more dedicated graphics processing units (GPUs). Processor 605 may be based on reduced instruction-set computer (RISC) or complex instruction-set computer (CISC) architectures or any other suitable architecture and may include one or more processing cores. Graphics hardware 620 may be special purpose computational hardware for processing graphics and / or assisting processor 605 perform computational tasks. In one embodiment, graphics hardware 620 may include one or more programmable graphics processing units (GPUs) and / or one or more specialized SOCs, e.g.. an SOC specially designed to implement neural network and machine learning operations (e.g., convolutions) in a more energy-efficient manner than either the main device central processing unit (CPU) or a typical GPU, such as Apple’s Neural Engine processing cores.

[0062] Image capture device 650 may comprise one or more camera units configured to capture images, e.g.. images which may be processed to generate cropped, augmented, and / or semantically tone-mapped versions of said captured images, e.g., in accordance with this disclosure. Image capture device(s) 650 may include two (or more) lens assemblies 680A and 680B, where each lens assembly may have a separate focal length. For example, lens assembly 680A may have a shorter focal length relative to the focal length of lens assembly 680B. Each lens assembly may have a separate associated sensor element, e.g., sensor elementsCustomer No. 61947690A / 690B. Alternatively, two or more lens assemblies may share a common sensor element. Image capture device(s) 650 may capture still and / or video images. Output from image capture device 650 may be processed, at least in part, by video codec(s) 655 and / or processor 605 and / or graphics hardware 620, and / or a dedicated image processing unit or image signal processor incorporated within image capture device 650. Images so captured may be stored in memory 660 and / or storage 665.

[0063] Memory 660 may include one or more different types of media used by processor 605, graphics hardware 620, and image capture device 650 to perform device functions. For example, memory 660 may include memory cache, read-only memory (ROM), and / or random access memory (RAM). Storage 665 may store media (e.g., audio, image and video files), computer program instructions or software, preference information, device profile information, and any other suitable data. Storage 665 may include one more non-transitory storage mediums including, for example, magnetic disks (fixed, floppy, and removable) and tape, optical media such as CD-ROMs and digital video disks (DVDs), and semiconductor memory devices such as Electrically Programmable Read-Only Memory (EPROM), and Electrically Erasable Programmable Read-Only Memory (EEPROM). Memory 660 and storage 665 may be used to retain computer program instructions or code organized into one or more modules and written in any desired computer programming language. When executed by, for example, processor 605. such computer program code may implement one or more of the methods or processes described herein. Power source 675 may comprise a rechargeable battery (e.g., a lithium-ion battery, or the like) or other electrical connection to a power supply, e.g., to a mains power source, that is used to manage and / or provide electrical power to the electronic components and associated circuitry of electronic device 600.

[0064] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described embodiments may be used in combination with each other. Many other embodiments will be apparent to those of skill in the art upon reviewing the above description. The scope of the invention therefore should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

Customer No. 61947CLAIMSWhat is claimed is:

1. A device, comprising: an image capture device; a memory; and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to: obtain a first linear image of a scene; obtain a first tone-mapped image of the scene having a first resolution, wherein the first tone-mapped image is a tone-mapped version of the first linear image of the scene; divide the first linear image by the first tone-mapped image to produce a tone gain map for the first tone-mapped image; obtain one or more semantic masks for the first tone-mapped image; obtain user input related to the one or more semantic masks; and combine the first tone-mapped image and the tone gain map according to the one or more semantic masks and the respective user input related to the one or more semantic masks to produce a semantically, locally tonemapped output image.

2. The device of claim 1, wherein the first linear image comprises an image with scene- referred light values.

3. The device of claim 1, wherein the first linear image has the first resolution.

4. The device of claim 1, wherein the first linear image has a resolution that is smaller than the first resolution or larger than the first resolution.

5. The device of claim 1, wherein the first linear image comprises a single exposure image.Customer No. 619476. The device of claim 1, wherein the first linear image comprises a fusion of multiple bracketed image captures.

7. The device of claim 1, wherein the instructions further comprise instructions causing the one or more processors to: obtain a global tone curve for the first linear image.

8. The device of claim 7, wherein the instructions further comprise instructions causing the one or more processors to: obtain an exposure metric for the first linear image.

9. The device of claim 8, wherein the exposure metric comprises one or more of: an aperture setting, an ISO setting, or an integration time.

10. The device of claim 7, wherein the instructions to divide the first linear image by the first tone-mapped image to produce a tone gain map for the first tone-mapped image further comprise instructions causing the one or more processors to: apply the global tone curve to the first linear image before the dividing.

11. The device of claim 8, wherein the instructions to divide the first linear image by the first tone-mapped image to produce a tone gain map for the first tone-mapped image further comprise instructions causing the one or more processors to: apply the global tone curve and the exposure metric to the first linear image before the dividing.

12. The device of claim 1, wherein at least one of the one or more user inputs comprises: a slider value, a dial value, a directional pad value, a numerical value, or a qualitative value.

13. The device of claim 1, wherein the instructions to combine the first tone-mapped image and the tone gain map according to the one or more semantic masks and the respective user input related to the one or more semantic masks further comprise instructions causing the one or more processors to:Customer No. 61947 perform a pixel-wise multiplication operation between the first tone-mapped image and the tone gain map.

14. The device of claim 13, wherein the instructions to perform a pixel-wise multiplication operation between the first tone-mapped image and the tone gain map further comprise instructions causing the one or more processors to: perform a pixel-wise multiplication operation between the first tone-mapped image and the tone gain map, subject to corresponding pixel values in the one or more semantic masks and the respective user input related to the one or more semantic masks.

15. The device of claim 1, wherein at least one of the one or more semantic masks comprises: a face mask, a skin tone mask, a subject mask, a person mask, an image foreground mask, an image background mask, a pet mask, a sky mask, a water mask, a teeth mask, a lips mask, an eyes mask, an eye glasses mask, a plated food mask, a Sun mask, a light mask, or a hair mask.

16. The device of claim 1, wherein at least one of the one or more semantic masks comprises a machine learning (ML)-based semantic mask.

17. The device of claim 1, wherein the instructions to obtain a first tone-mapped image further comprise instructions causing the one or more processors to: obtain the first tone-mapped image from the image capture device during an image preview mode of operation.

18. The device of claim 1, wherein the instructions to obtain a first tone-mapped image further comprise instructions causing the one or more processors to: obtain the first tone-mapped image from the memory.

19. A non-transitory program storage device, comprising instructions stored thereon, to cause one or more processors to: perform operations in accordance with the operations performed by the one or more processors of any of claims 1-18.Customer No. 6194720. An image processing method, comprising: performing operations in accordance with the operations performed by the one or more processors of any of claims 1-18.

Citation Information

Patent Citations

  • Automatically Segmenting and Adjusting Images

    US20220230323A1

  • Machine learning segmentation-based tone mapping in high noise and high dynamic range environments or other environments

    US20240257324A1