Method and apparatus for processing image frames

By using machine learning-based segmentation technology in high-dynamic range environments, the problem of image quality degradation and darkening of important areas in the prior art is solved, and a higher quality image output is achieved.

CN120092257APending Publication Date: 2025-06-03SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380075502.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-23
Filing Date
2023-12-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In high noise and high dynamic range environments, it is difficult for the prior art to effectively perform tone mapping, resulting in image quality degradation, important areas become darker, object boundaries are unclear, and undesirable artifacts are easily generated.

Method used

Using machine learning-based segmentation technology, a high-dynamic range hybrid image is generated by generating semantic incremental weight maps to tone fusion operations to generate a fusion image. The method includes multiple steps: generating a high dynamic range mixed image, synthesizing a plurality of low dynamic range images, generating an initial weight map and a filter weight map, performing image decomposition and weighting mixing to achieve tone fusion.

Benefits of technology

It effectively reduces the darkening of important areas in the image, maintains the clarity of object boundaries, reduces the appearance of undesired artifacts, and improves image quality and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092257A_ABST
    Figure CN120092257A_ABST
Patent Text Reader

Abstract

A method for processing image frames by an electronic device includes obtaining a plurality of input image frames; generating a high dynamic range (HDR) hybrid image based on the input image frame; and performing a hue fusion operation on the HDR mixed image based on the semantic increment weight map to generate a fused image. Performing the tone fusion operation includes: synthesizing a plurality of low dynamic range (LDR) images based on the HDR mixed image, and generating an initial weight map based on the LDR images; generating a filtering weight map based on the initial weight map, the semantic increment weight map and a guide filter; and generating a fused image based on the filtered weight map and the decomposed version of the LDR image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to imaging systems. More specifically, the present disclosure relates to machine learning segmentation-based tone mapping in high-noise and high-dynamic range environments or other environments. Background Art

[0002] Mobile electronic devices such as smart phones and tablet computers have become the most common device types for capturing, uploading, and sharing digital images. "Computational photography" refers to image generation techniques commonly used in these and other types of devices, where one or more digital images of a scene are captured and processed digitally in order to produce a desired effect within one or more final images of the scene. In the past few years, computational photography has evolved to significantly narrow the gap with traditional digital single-lens reflex (DSLR) cameras. For example, computational photography techniques have been developed to effectively support functions such as zooming, low-light photography, and the use of under-screen cameras. Summary of the Invention

[0003] Solution to the Problem

[0004] The present disclosure relates to machine learning segmentation-based tone mapping in high-noise and high-dynamic range environments or other environments.

[0005] According to an embodiment of the present disclosure, a method for an electronic device to process an image frame is provided. The method may include: obtaining a plurality of input image frames. The method may include: generating a high-dynamic range (HDR) hybrid image based on the plurality of input image frames. The method may include: performing a tone fusion operation on the HDR hybrid image based on a semantic incremental weight map to generate a fused image. Performing the tone fusion operation may include: synthesizing a plurality of low-dynamic range (LDR) images based on the HDR hybrid image. Performing the tone fusion operation may include: generating an initial weight map based on the LDR images. Performing the tone fusion operation may include: generating a filtered weight map based on the initial weight map, the semantic incremental weight map, and a guided filter. Performing the tone fusion operation may include: generating a fused image based on the filtered weight map and a decomposed version of the LDR images.

[0006] Generating a fused image based on a filtered weight map and a decomposed version of an LDR image may include: performing image decomposition of each LDR image to generate a base component and a detail component of each LDR image. Generating a fused image based on a filtered weight map and a decomposed version of an LDR image may include: performing a base blending operation based on the filtered weight map and the base component of the LDR image. Generating a fused image based on a filtered weight map and a decomposed version of an LDR image may include: performing a detail blending operation based on the filtered weight map and the detail component of the LDR image. Generating a fused image based on a filtered weight map and a decomposed version of an LDR image may include: combining the results of the base blending operation and the detail blending operation to generate a fused image.

[0007] Synthesizing an LDR image based on an HDR blended image may include: identifying an image histogram of the HDR blended image. Synthesizing an LDR image based on an HDR blended image may include: determining a plurality of fusion scales based on the image histogram. Synthesizing an LDR image based on an HDR blended image may include: multiplying the image data of the HDR blended image by the fusion scale and cropping the resulting image data to generate cropped image data. Synthesizing an LDR image based on an HDR blended image may include: applying an image signal processing (ISP) conversion to the cropped image data so as to generate a YUV image.

[0008] Generating an initial weight map based on an LDR image may include: generating a saliency metric for the YUV image. Generating an initial weight map based on an LDR image may include: generating a color saturation metric for the YUV image using a first look-up table. Generating an initial weight map based on an LDR image may include: generating a good exposure metric for the YUV image using a second look-up table. Generating an initial weight map based on an LDR image may include: combining the saliency metric, the color saturation metric, and the good exposure metric for each YUV image and normalizing the combined metrics to generate an initial weight map for each YUV image.

[0009] Generating a filtered weight map may include: generating a modified weight map based on the initial weight map and a semantic incremental weight map. Generating a filtered weight map may include: performing removal using a guided filter.

[0010] The method may include generating a semantic incremental weight map. Generating the semantic incremental weight map may include: generating a lower-resolution HDR mixed image based on the HDR mixed image. Generating the semantic incremental weight map may include: generating a lower-resolution LDR mixed image based on the lower-resolution HDR mixed image. Generating the semantic incremental weight map may include: generating a lower-resolution LDR mixed YUV image based on the lower-resolution LDR mixed image. Generating the semantic incremental weight map may include: generating a semantic segmentation mask based on the lower-resolution LDR mixed YUV image. Generating the semantic incremental weight map may include: generating the semantic incremental weight map based on the semantic segmentation mask using a mapping.

[0011] The semantic segmentation mask may be generated using a trained machine learning model that processes the lower-resolution LDR mixed YUV image.

[0012] A mapping may be used to transform different values of different semantic classes in the semantic segmentation mask into corresponding values in the semantic incremental weight map.

[0013] According to an embodiment of the present disclosure, an electronic device is provided. The electronic device may include a memory that stores instructions, and at least one processor operably coupled to the memory. When the at least one processor executes the instructions, the at least one processor causes the electronic device to perform operations. The operations may include: obtaining a plurality of input image frames. The operations may include: generating a high dynamic range (HDR) mixed image based on the plurality of input image frames. The operations may include: performing a tone fusion operation on the HDR mixed image based on the semantic incremental weight map to generate a fused image. Performing the tone fusion operation may include: synthesizing a plurality of low dynamic range LDR images based on the HDR mixed image. Performing the tone fusion operation may include: generating an initial weight map based on the LDR images. Performing the tone fusion operation may include: generating a filtered weight map based on the initial weight map, the semantic incremental weight map, and a guidance filter. Performing the tone fusion operation may include: generating a fused image based on the filtered weight map and a decomposed version of the LDR images.

[0014] According to an embodiment of the present disclosure, a non-transitory computer-readable medium storing instructions is provided. When the instructions are executed by at least one processor of an electronic device (101), the electronic device is caused to perform operations. The operations may include: obtaining a plurality of input image frames. The operations may include: generating a high dynamic range (HDR) hybrid image based on the plurality of input image frames. The operations may include: performing a tone fusion operation on the HDR hybrid image based on a semantic incremental weight map to generate a fused image. Performing the tone fusion operation may include: synthesizing a plurality of low dynamic range (LDR) images based on the HDR hybrid image. Performing the tone fusion operation may include: generating an initial weight map based on the LDR images. Performing the tone fusion operation may include: generating a filtered weight map based on the initial weight map, the semantic incremental weight map, and a guidance filter. Performing the tone fusion operation may include: generating a fused image based on the filtered weight map and a decomposed version of the LDR images. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more fully understand the present disclosure and its advantages, the following description is now made in conjunction with the accompanying drawings, in which like reference numerals represent like components:

[0016] Figure 1 An example network configuration including an electronic device according to the present disclosure is shown;

[0017] Figure 2 An example architecture supporting tone mapping based on machine learning segmentation according to the present disclosure is shown;

[0018] Figure 3 An example multi-exposure analysis operation in the architecture according to the present disclosure is shown Figure 2 of;

[0019] Figure 4 An example multi-exposure blending operation in the architecture according to the present disclosure is shown Figure 2 of;

[0020] Figure 5 An example tone fusion operation in the architecture according to the present disclosure is shown Figure 2 of;

[0021] Figure 6 An example image histogram for use in a tone fusion operation according to the present disclosure is shown for Figure 5 ;

[0022] Figure 7 and Figure 8 An example look-up table for use in a tone fusion operation according to the present disclosure is shown for Figure 5 ;

[0023] Figure 9 An example in the architecture according to the present disclosure is shown Figure 2Exemplary semantic-based machine learning model segmentation operations in the architecture;

[0024] Figure 10 Illustrates an example lookup table for use in semantic-based machine learning model segmentation operations according to the present disclosure; Figure 9 for

[0025] Figure 11 Illustrates an example mapping for use in noise-aware segmentation to incremental weight conversion operations in the architecture according to the present disclosure; and Figure 2 for

[0026] Figure 12 Illustrates an example method for tone mapping based on machine learning segmentation according to the present disclosure. Detailed Description

[0027] Before proceeding with the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "send," "receive," and "communicate," and derivatives thereof, include both direct and indirect communication. The terms "include" and "comprise," and derivatives thereof, mean including but not limited to. The term "or" is inclusive and means and / or. The phrase "associated with," and derivatives thereof, means including, being included within, interconnecting with, containing, being contained within, connected to or coupling with, coupled to or coupling with, capable of communicating with, cooperating with, interlacing, juxtaposing, proximate to, bound to or binding with, having, having the attribute of, having a relationship to or having a relationship with, and the like.

[0028] Furthermore, the various functions described below may be implemented or supported by one or more computer programs, each computer program formed from computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof, suitable for implementation in appropriate computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory. A "non-transitory" computer-readable medium does not include a wired, wireless, optical, or other communication link that transmits transitory electrical or other signals. Non-transitory computer-readable media include media that can permanently store data and media that can store data and then rewrite the data (e.g., rewritable compact discs or erasable memory devices).

[0029] As used herein, terms and phrases such as "having", "may have", "including", or "may include" features (such as numbers, functions, operations, or components such as parts) indicate the presence of the features and do not exclude the presence of other features. In addition, as used herein, the phrases "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may indicate any of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate different user devices from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, the first component may be represented as the second component, and vice versa.

[0030] It will be understood that when an element (e.g., a first element) is referred to as "coupled (or communicatively coupled) to" or "coupled to" another element (e.g., a second element), or "connected (or communicatively connected) to" another element, the element may be coupled or connected to the other element directly or via a third element. Conversely, it will be understood that when an element (e.g., a first element) is referred to as "directly coupled to" or "directly connected to" another element (e.g., a second element), there is no other element (e.g., a third element) between the element and the other element.

[0031] As used herein, the phrase "configured (or set) to" may be used interchangeably with the phrases "suitable for", "capable of", "designed to", "adapted to", "manufactured to", or "able to" as the context requires. The phrase "configured (or set) to" does not essentially mean "specially designed in hardware to". Instead, the phrase "configured to" may mean that a device can perform operations together with another device or component. For example, the phrase "a processor configured (or set) to perform A, B, and C" may mean a general-purpose processor (e.g., a CPU or an application processor) that can perform operations by executing one or more software programs stored in a storage device, or a dedicated processor (e.g., an embedded processor) for performing the operations.

[0032] The terms and phrases used herein are provided only to describe some embodiments of the present disclosure and are not intended to limit the scope of other embodiments of the present disclosure. It should be understood that, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" include plural references. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of the present disclosure pertain. It will be further understood that terms and phrases defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In some instances, the terms and phrases defined herein may be interpreted so as to exclude embodiments of the present disclosure.

[0033] Examples of an "electronic device" according to an embodiment of the present disclosure may include at least one of the following: a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (e.g., smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, electronic accessories, an electronic tattoo, a smart mirror, or a smart watch). Other examples of electronic devices include smart home appliances. Examples of smart home appliances may include at least one of the following: a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a cleaner, an oven, a microwave oven, a washing machine, a dryer, an air cleaner, a set-top box, a home automation control panel, a security control panel, a TV box (e.g., SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker, or a speaker with an integrated digital assistant (e.g., SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (e.g., XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camera, or an electronic photo frame. Other examples of electronic devices include at least one of the following: various medical devices (e.g., various portable medical measurement devices such as a blood glucose measurement device, a heart rate measurement device, or a body temperature measurement device, a magnetic resource angiography (MRA) device, a magnetic resource imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an in-vehicle infotainment device, marine electronic devices (e.g., marine navigation devices or gyro compasses), avionics, security devices, a vehicle head unit, industrial or home robots, an automated teller machine (ATM), a point of sale (POS) device, or an Internet of Things (IoT) device (e.g., a light bulb, various sensors, a water meter, or a gas meter, a sprinkler, a fire alarm, a thermostat, a street lamp, a toaster, a fitness device, a hot water tank, a heater, or a boiler). Other examples of electronic devices include a piece of furniture or at least a part of a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measurement devices (e.g., devices for measuring water, electricity, gas, or electromagnetic waves). Note that according to various embodiments of the present disclosure, an electronic device may be one or a combination of the devices listed above. According to some embodiments of the present disclosure, an electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to the devices listed above and may include new electronic devices according to the development of technology

[0034] In the following description, according to various embodiments of the present disclosure, an electronic device is described with reference to the accompanying drawings. As used herein, the term "user" may refer to a person using the electronic device or another device (e.g., an artificial intelligence electronic device).

[0035] Throughout this patent document, definitions may be provided for certain other words and phrases. Those of ordinary skill in the art should understand that in many, if not most, instances, such definitions apply to the prior as well as future use of the words and phrases so defined.

[0036] The description in this application should not be construed as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the patent subject matter is defined only by the claims. The applicant understands that any other terms used in the claims, including but not limited to "mechanism", "module", "device", "unit", "component", "element", "member", "apparatus", "machine", "system", "processor", or "controller", refer to structures known to those skilled in the relevant art.

[0037] The following discussion is described with reference to the accompanying drawings Figures 1 to 12 and various embodiments of the present disclosure. However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalents or substitutes thereof also fall within the scope of the present disclosure.

[0038] As mentioned above, mobile electronic devices such as smart phones and tablet computers have become the most common types of devices for capturing, uploading, and sharing digital images. "Computational photography" refers to image generation techniques commonly used in these and other types of devices, in which one or more digital images of a scene are captured and processed digitally in order to produce a desired effect within one or more final images of the scene. In the past few years, computational photography has evolved to significantly narrow the gap with traditional digital single-lens reflex (DSLR) cameras. For example, computational photography techniques have been developed to effectively support functions such as zooming, low-light photography, and the use of under-display cameras.

[0039] An example image processing function typically performed by a mobile electronic device is contrast enhancement representing a tone mapping operation, which generally involves adjusting the contrast within an image to improve image features while attempting to keep dark regions dark and bright regions bright. Among other things, this can help reduce ambiguity or other issues in the captured image. Contrast enhancement generally takes one of two forms, namely, global contrast enhancement and local contrast enhancement. Global contrast enhancement generally involves performing tone mapping over the entire image, while local contrast enhancement generally involves performing tone mapping within local regions of the image. Unfortunately, global contrast enhancement generally darkens important regions within the image (e.g., a face). Additionally, local contrast enhancement generally does not consider object boundaries and may create brighter or darker spots within an object (e.g., when a brighter halo or a darker blotch or other shape appears within the sky or other image regions). In both cases, these issues may create undesirable artifacts in the image, which reduces the quality of the image and thus user satisfaction.

[0040] The present disclosure provides various techniques related to machine learning-based segmentation tone mapping in high-noise and high-dynamic range environments or other environments. As described in more detail below, multiple input image frames (e.g., multiple image frames of a scene obtained by a mobile electronic device or other electronic device) can be obtained. A high-dynamic range (HDR) blended image can be generated based on the input image frames, and the HDR blended image can have a higher dynamic range than each of the input image frames. A tone fusion operation can be performed on the HDR blended image based on a semantic incremental weight map to generate a fused image. Performing the tone fusion operation can include: synthesizing multiple low-dynamic range (LDR) images based on the HDR blended image, and generating an initial weight map based on the LDR images. Performing the tone fusion operation can further include: generating a filtered weight map based on the initial weight map, the semantic incremental weight map, and a guided filter. Performing the tone fusion operation can further include: generating a fused image based on the filtered weight map and a decomposed version of the LDR images.

[0041] In some cases, a fused image can be generated in the following manner: perform image decomposition on each LDR image to generate a base component and a detail component of each LDR image, perform a base blending operation based on a filtering weight map and the base component of the LDR image, perform a detail blending operation based on the filtering weight map and the detail component of the LDR image, and combine the results of the base blending operation and the detail blending operation to generate a fused image. Additionally, in some cases, a semantic incremental weight map can be obtained in the following manner: generate a lower-resolution HDR blended image based on an HDR blended image, generate a lower-resolution LDR blended image based on the lower-resolution HDR blended image, generate a lower-resolution LDR blended YUV image based on the lower-resolution LDR blended image, generate a semantic segmentation mask based on the lower-resolution LDR blended YUV image, and generate a semantic incremental weight map based on the semantic segmentation mask using a mapping.

[0042] In this way, the disclosed techniques support the use of semantics-based tone mapping, which can be performed using tone fusion operations guided by semantics or a segmentation map or mask associated with the content within an input image frame. For example, the tone fusion operation can obtain an HDR blended image and a semantic incremental weight map (which can be generated using a trained machine learning model), where the semantic incremental weight map is used to process the HDR blended image to perform tone fusion. These techniques allow for more effective tone mapping operations, for example, during contrast enhancement operations. Additionally, these techniques can reduce or avoid darkening of important regions within an image, respect object boundaries while performing tone mapping, and reduce or avoid creating unwanted brighter or darker spots within an object. Overall, these techniques can be used to produce scene images of higher quality, which can improve user satisfaction.

[0043] Note that while some of the embodiments discussed below are described in the context of using a consumer electronic device (e.g., a smart phone), this is merely an example. It will be understood that the principles of the present disclosure can be implemented in any number of other suitable contexts and can use any suitable one or more devices. Also note that while some of the embodiments discussed below are described under the assumption that a machine learning model trained on one device (e.g., a server) is deployed to one or more other devices (e.g., one or more consumer electronic devices), this is also merely an example. It will be understood that the principles of the present disclosure can be implemented using any number of devices (including a single device that both trains and uses the machine learning model). Generally, the present disclosure is not limited to use with any particular type of device.

[0044] Figure 1 An example network configuration 100 including an electronic device in accordance with the present disclosure is shown. Figure 1The illustrated embodiment of the network configuration 100 is for illustrative purposes only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.

[0045] According to an embodiment of the present disclosure, the electronic device 101 is included in the network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or a sensor 180. In some embodiments, the electronic device 101 may not include at least one of these components, or may add at least one other component. The bus 110 includes circuitry for connecting the components 120-180 to each other and for transmitting communications (e.g., control messages and / or data) between the components.

[0046] The processor 120 includes one or more processing devices (e.g., one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs)). In some embodiments, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), or a graphics processing unit (GPU). The processor 120 is capable of performing control over at least one of the other components of the electronic device 101, and / or performing operations or data processing related to communication or other functions. As described in more detail below, the processor 120 may perform various operations related to machine learning-based segmentation tone mapping in a high-noise and high-dynamic range environment or other environments.

[0047] The memory 130 may include volatile and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one of the other components of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or a program 140. The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or “app”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).

[0048] The kernel 141 may control or manage system resources (e.g., the bus 110, the processor 120, or the memory 130) for performing operations or functions implemented in other programs (e.g., the middleware 143, the API 145, or the application 147). The kernel 141 provides an interface that allows the middleware 143, the API 145, or the application 147 to access the components of the electronic device 101 to control or manage the system resources. The application 147 may support various functions related to machine learning-based segmentation tone mapping. These functions may be performed by a single application or by multiple applications, with each application performing one or more of these functions. For example, the middleware 143 may act as a relay to allow the API 145 or the application 147 to transfer data to and from the kernel 141. Multiple applications 147 may be provided. The middleware 143 is capable of controlling work requests received from the application 147, for example, by assigning priorities to the use of system resources (such as the bus 110, the processor 120, or the memory 130) of the electronic device 101 to at least one of the multiple applications 147. The API 145 is an interface that allows the application 147 to control functions provided by the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (e.g., a command) for archive control, window control, image processing, or text control.

[0049] The I / O interface 150 serves as an interface that can transfer commands or data input from a user or other external device, for example, to other components of the electronic device 101. The I / O interface 150 may also output commands or data received from other components of the electronic device 101 to the user or other external device.

[0050] The display 160 includes, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, a microelectromechanical systems (MEMS) display, or an electronic paper display. The display 160 may also be a depth perception display (e.g., a multi-focus display). The display 160 is capable of displaying various contents (e.g., text, images, videos, icons, or symbols) to a user. The display 160 may include a touch screen and may receive touch, gesture, proximity, or hover inputs using an electronic pen or a user's body part.

[0051] The communication interface 170 is capable of establishing communication, for example, between the electronic device 101 and an external electronic device (e.g., the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 may be connected to the network 162 or 164 through wireless or wired communication to communicate with the external electronic device. The communication interface 170 may be a wired or wireless transceiver, or any other component for transmitting and receiving signals.

[0052] Wireless communication can use at least one of the following as a communication protocol: for example, WiFi, Long Term Evolution (LTE), Long Term Evolution - Advanced (LTE - A), 5th Generation Wireless System (5G), millimeter - wave or 60GHz wireless communication, Wireless USB, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM). The wired connection can include at least one of, for example, Universal Serial Bus (USB), High - Definition Multimedia Interface (HDMI), Recommended Standard 232 (RS - 232), or Plain Old Telephone Service (POTS). Network 162 or 164 includes at least one communication network, for example, a computer network (such as a Local Area Network (LAN) or a Wide Area Network (WAN)), the Internet, or a telephone network.

[0053] The electronic device 101 also includes one or more sensors 180, which can measure a physical quantity or detect the activation state of the electronic device 101 and convert the measured or detected information into an electrical signal. For example, one or more sensors 180 can include one or more cameras or other imaging sensors for capturing a scene image. The sensor 180 can also include one or more buttons for touch input, a gesture sensor, a gyroscope or gyroscope sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (e.g., a Red - Green - Blue (RGB) sensor), a biophysical sensor, a temperature sensor, a humidity sensor, an illuminance sensor, an Ultraviolet (UV) sensor, an Electromyogram (EMG) sensor, an Electroencephalogram (EEG) sensor, an Electrocardiogram (ECG) sensor, an Infrared (IR) sensor, an ultrasonic sensor, an iris sensor, or a fingerprint sensor. The sensor 180 can also include an Inertial Measurement Unit, which can include one or more accelerometers, gyroscopes, and other components. Additionally, the sensor 180 can include a control circuit for controlling at least one of the sensors included herein. Any of these sensors 180 can be located within the electronic device 101.

[0054] In some embodiments, the first external electronic device 102 or the second external electronic device 104 can be a wearable device or a wearable device installable on an electronic device (e.g., an HMD). When the electronic device 101 is installed in the electronic device 102 (e.g., an HMD), the electronic device 101 can communicate with the electronic device 102 through the communication interface 170. The electronic device 101 can be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network. The electronic device 101 can also be an augmented reality wearable device (e.g., glasses) including one or more imaging sensors.

[0055] The first external electronic device 102, the second external electronic device 104, and the server 106 can each be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. Additionally, according to certain embodiments of the present disclosure, all or some of the operations performed on the electronic device 101 can be performed on another or multiple other electronic devices (e.g., electronic devices 102 and 104, or the server 106). Further, according to certain embodiments of the present disclosure, when the electronic device 101 is to automatically or upon request perform some functions or services, the electronic device 101 can request another device (e.g., electronic devices 102 and 104, or the server 106) to perform at least some functions associated therewith, rather than performing the functions or services itself or additionally. Another electronic device (e.g., electronic devices 102 and 104, or the server 106) is capable of performing the requested function or additional functions and transmitting the execution result to the electronic device 101. The electronic device 101 can provide the requested function or service by processing the received result as is or additionally. To this end, for example, cloud computing, distributed computing, or client-server computing technologies can be used. Although Figure 1 it is shown that the electronic device 101 includes a communication interface 170 to communicate with the external electronic device 104 or the server 106 via the network 162 or 164, according to some embodiments of the present disclosure, the electronic device 101 can operate independently without a separate communication function.

[0056] The server 106 can include components 110 - 180 (or a suitable subset thereof) that are the same as or similar to those of the electronic device 101. The server 106 can support driving the electronic device 101 by performing at least one operation (or function) implemented on the electronic device 101. For example, the server 106 can include a processing module or a processor that can support the processor 120 implemented in the electronic device 101. As described in more detail below, the server 106 can perform various operations related to machine learning-based segmentation tone mapping in a high-noise and high-dynamic range environment or other environments.

[0057] Although Figure 1 it shows an example of the network configuration 100 including the electronic device 101, various changes can be made to Figure 1 it. For example, the network configuration 100 can include any number of each component arranged in any suitable manner. Generally, computing and communication systems have a wide variety of configurations, and Figure 1 the scope of the present disclosure is not limited to any specific configuration. Additionally, although Figure 1 it shows an operating environment in which various features disclosed in this patent document can be used, these features can be used in any other suitable system.

[0058] Figure 2 An example architecture 200 for supporting tone mapping based on machine learning segmentation according to the present disclosure is shown. For ease of explanation, Figure 2 The illustrated architecture 200 is described as consisting of Figure 1 The illustrated network configuration 100 is used by the electronic device 101. However, the architecture 200 may be used by any other suitable device (eg, the server 106) and in any other suitable system.

[0059] like Figure 2 As shown, the architecture 200 generally receives and processes a plurality of input image frames 202. Each input image frame 202 represents an image frame of a captured scene. The input image frames 202 may be obtained from any suitable source, such as when the input image frames 202 are generated using at least one camera or other imaging sensor 180 of the electronic device 101 during an image capture operation. According to an embodiment, the input image frames 202 may be captured simultaneously using different cameras or other imaging sensors 180 of the electronic device 101, or may be captured sequentially (e.g., in rapid succession) using one or more cameras or other imaging sensors 180 of the electronic device 101. In some cases, the input image frames 202 may be captured in response to a capture event (e.g., when the processor 120 detects that a user initiates image capture by pressing a hard button or soft button of the electronic device 101). The input image frames 202 may have any suitable resolution, and the resolution of each input image frame 202 may depend on the capabilities of the imaging sensor 180 in the electronic device 101, and may depend on one or more user settings that affect the resolution. In some embodiments, the input image frames 202 may represent raw image frames, RGB image frames, or image frames in any other suitable image data space. In addition, in some embodiments, each input image frame 202 may include twelve-bit image data values, although data values ​​of other bit lengths may be used.

[0060] The input image frames 202 may include image frames captured using different exposure levels, such as when the input image frames 202 include one or more shorter exposure image frames and one or more longer exposure image frames. As a specific example, the input image frames 202 may include one or more image frames captured at an EV-0 exposure level (also commonly referred to as an "auto" exposure level), one or more image frames captured at an EV-2 exposure level, and one or more image frames captured at an EV-4 exposure level. Note, however, that these exposure levels are for illustration only, and the input image frames 202 may be captured at any other or additional exposure levels (e.g., EV-1, EV-3, EV-5, EV-6, or EV+1 exposure levels).

[0061] In this example, an input image frame 202 is provided to a multi-exposure (ME) analysis operation 204 that processes the input image frame 202 to generate various outputs 206. The outputs 206 include information that can be used to support blending of the input image frame 202. For example, the ME analysis operation 204 can process the input image frame 202 to generate a shadow map, a saturation map, and an ME motion map for at least some of the input image frames 202 in the input image frame 202. A shadow map generally represents a set of maps or other indicators identifying where shadows are present in the input image frame 202. A saturation map generally represents a set of maps or other indicators identifying the saturation level within the input image frame 202, e.g., by identifying regions of the input image frame 202 that are over-saturated or under-saturated. An ME motion map generally represents a set of maps or other indicators identifying locations where motion is detected between different input image frames 202. The ME analysis operation 204 can use any suitable techniques for generating the shadow map, the saturation map, and the ME motion map. An example implementation of the ME analysis operation 204 is described in the Figure 3 shown below.

[0062] The input image frame 202 and the outputs 206 from the ME analysis operation 204 are provided to an ME blending operation 208 that represents a multi-frame blending operation for combining the input image frame 202 based on the outputs 206 to generate a high dynamic range (HDR) blended image 210. Since the HDR blended image 210 generally has a higher dynamic range than any single one of the input image frames 202, the HDR blended image 210 is referred to as a "high dynamic range" image. The ME blending operation 208 can use any suitable techniques for blending the input image frame 202. For example, in some embodiments, the ME blending operation 208 can blend the red (R), green (G), and blue (B) channels (and possibly one or more additional channels) of the image data included in the input image frame 202 independently. Blending of the image data can be controlled based on the outputs 206 of the ME analysis operation 204, such as when the ME blending operation 208 performs blending to restore regions of the input image frame 202 that are over-saturated or under-saturated and not associated with motion. For regions of the input image frame 202 that are not over-saturated or under-saturated or are associated with motion, the ME blending operation 208 can use the image data from one of the input image frames 202 (typically a "reference" image frame) without blending or with very little blending. An example implementation of the ME blending operation 208 is described in the Figure 4 shown below.

[0063] An HDR hybrid image 210 is provided to a tone fusion operation 212 that processes the HDR hybrid image 210 to modify the dynamic range within the HDR hybrid image 210 and generate a fused image 214. This allows the fused image 214 to be displayed or otherwise presented in a form with a smaller dynamic range. For example, this can be done when an HDR image is to be presented on a display device having a smaller dynamic range than the HDR image itself. For example, the tone fusion operation 212 can generate multiple low dynamic range (LDR) images from the HDR hybrid image 210 and decompose the LDR images. The tone fusion operation 212 can also generate, modify, and filter an initial weight map to generate a filtered weight map. The filtered weight map can be used during weighted blending to fuse the decomposed versions of the LDR images to produce the fused image 214. Ideally, these capabilities allow the tone fusion operation 212 to weight the image data differently to achieve the desired tone modification of the HDR hybrid image 210 and generate the fused image 214. An example implementation of the tone fusion operation 212 is described in the Figure 5 shown below.

[0064] The fused image 214 here can be subjected to a tone mapping operation 216 that can include local tone mapping (LTM) and / or global tone mapping (GTM). The tone mapping operation 216 can generally be used as a post-processing operation to fine-tune the tone. In some cases, this can allow the fused image 214 to have higher or lower contrast and / or color saturation, or can make the fused image 214 brighter or darker. The tone mapping operation 216 can use any suitable technique to perform the tone mapping. The result of the tone mapping operation 216 is an output image 218 that can be stored, output, or used in any suitable manner as needed or desired.

[0065] The HDR mixed image 210 is also provided to an HDR semantics-based machine learning model (MLM) segmentation operation 220 that processes the HDR mixed image 210 to generate a semantic segmentation mask 222 associated with the HDR mixed image 210. The semantic segmentation mask 222 represents a graph or other set of indicators identifying one or more objects within the HDR mixed image 210 and which pixels of the HDR mixed image 210 are associated with each object. For example, the semantic segmentation mask 222 can identify the pixels of the HDR mixed image 210 associated with people, animals, trees or other plants, buildings / walls / floors / windows / other man-made structures, ground, grass, water, sky, or background within the HDR mixed image 210. In some embodiments, the HDR semantics-based MLM segmentation operation 220 can include a machine learning model (e.g., a convolutional neural network (CNN) or other neural network) that has been trained to perform image segmentation. An example implementation of the HDR semantics-based MLM segmentation operation 220 is described below in Figure 9 as shown.

[0066] The semantic segmentation mask 222 is provided to a noise-aware segmentation-to-incremental weight conversion operation 224 that processes the semantic segmentation mask 222 to generate a semantic incremental weight map 226. The semantic incremental weight map 226 identifies how the initial weight map generated by the tone fusion operation 212 can be modified to support proper object segmentation when performing tone fusion to generate the fused image 214. For example, the segmentation-to-incremental weight conversion operation 224 can use a mapping to transform the segmentation mask values from the semantic segmentation mask 222 into corresponding incremental weight map values in the semantic incremental weight map 226. The tone fusion operation 212 can modify the initial weight map based on the incremental weight map values in the semantic incremental weight map 226. As described below, in some cases, the mapping used by the segmentation-to-incremental weight conversion operation 224 can be based on metadata associated with the HDR mixed image 210 (e.g., ISO, exposure time, and / or luminance) or the image data contained within the HDR mixed image 210. Since factors such as ISO, exposure time, and luminance can affect the amount of noise contained within the HDR mixed image 210, this can allow the segmentation-to-incremental weight conversion operation 224 to be “noise-aware”. The noise-aware segmentation-to-incremental weight conversion operation 224 can use any suitable technique to perform the segmentation-to-incremental weight conversion. Details of an example method used by the noise-aware segmentation-to-incremental weight conversion operation 224 are described below in Figure 11 as shown.

[0067] As can be seen here, the architecture 200 supports the execution of semantic-based tone mapping. This is accomplished using a tone fusion operation 212 that operates based on a semantic delta weight map 226. Since the semantic delta weight map 226 is generated using a segmentation of the HDR mixed image 210, the semantic delta weight map 226 can be used to guide the fusion of image data through the tone fusion operation 212, which allows the tone fusion to be guided based on the semantic content of the HDR mixed image 210 defined by the semantic segmentation mask 222. In addition, the HDR semantic-based MLM segmentation operation 220 and the noise-aware segmentation to delta weight conversion operation 224 can help provide robust and reliable handling of semantic map inaccuracies (e.g., inaccuracies that typically occur in the presence of high noise and HDR scenes). Among other aspects, this can be achieved by processing the HDR mixed image 210 instead of each input image frame 202.

[0068] although Figure 2 One example of an architecture 200 that supports tone mapping based on machine learning segmentation is shown, but can be used for Figure 2 Make various changes. For example, they may be combined, further subdivided, rearranged, duplicated, or omitted Figure 2 Various components or functions are shown, and additional components may be added according to specific needs.

[0069] Figure 3 It shows that according to the present disclosure Figure 2 Example multi-exposure (ME) analysis operation 204 in the architecture 200. Figure 3 As shown, the ME analysis operation 204 receives as input the input image frames 202a-202c and the alignment maps 302, 304. The input image frames 202a-202c may represent input image frames 202 captured using different exposure levels. In this particular example, one or more input image frames 202a may represent one or more EV-0 image frames, one or more input image frames 202b may represent one or more EV-2 image frames, and one or more input image frames 202c may represent one or more EV-4 image frames. However, as described above, input image frames 202 with other or additional exposure levels may be used.

[0070] Each alignment map 302, 304 represents the alignment difference between two input image frames among the input image frames 202a - 202c. For example, alignment map 302 may represent the alignment difference between input image frames 202a and 202b, so alignment map 302 may be referred to as an EV - 2 alignment map. Similarly, alignment map 304 may represent the alignment difference between input image frames 202a and 202c, so alignment map 304 may be referred to as an EV - 4 alignment map. Here, input image frame 202a can be regarded as a reference frame, and an alignment map can be generated or otherwise obtained for each of the other input image frames 202b - 202c (which are referred to as non - reference image frames). Note that although alignment maps 302, 304 are shown here as being input to the ME analysis operation 204, the ME analysis operation 204 itself can determine alignment maps 302, 304 based on input image frames 202a - 202c.

[0071] The electronic device 101 performs multiplication operations 306 and 308 on input image frame 202b and input image frame 202c respectively, so as to substantially equalize the brightness levels among input image frames 202a - 202c. For example, multiplication operations 306 and 308 can be per - pixel multiplication operations, where each pixel of input image frames 202b - 202c is multiplied by a multiplication value. In this particular example, the pixel values in input image frame 202b are multiplied by four in multiplication operation 306, and the pixel values in input image frame 202c are multiplied by sixteen in multiplication operation 308. The multiplication values four and sixteen are based on the use of EV - 2 and EV - 4 exposure levels, so different multiplication values can be used if different exposure levels are used.

[0072] The electronic device 101 also performs linear red - green - blue (RGB) conversion operations 310a - 310c on input image frame 202a and the equalized versions of input image frames 202b - 202c. In some embodiments, input image frames 202a - 202c may be received in the Bayer or other raw image domain, and the electronic device 101 performs linear RGB conversion operations 310a - 310c to convert input image frames 202a - 202c (or their equalized versions) to the RGB domain. For example, in some cases, each of the linear RGB conversion operations 310a - 310c may perform the following operations.

[0073] (1)

[0074] (2)

[0075] (3)

[0076] Here, R, G, and B respectively represent the red value, green value, and blue value of each pixel in the RGB domain, and , , , and respectively represent the red value, first green value, second green value, and blue value of each pixel in the Bayer or other raw domain. Of course, these operations are merely examples, and other suitable conversion operations can be used. Note that although Figure 3 shows three instances of the linear RGB conversion operations 310a - 310c (which can operate in parallel), one or two linear RGB conversion operations can be used, and different image data can be processed sequentially.

[0077] The electronic device 101 performs an RGB warping operation 312 on an equalized version of the input image frame 202b converted to the RGB domain. For example, the RGB warping operation 312 can use the alignment map 302 to warp the equalized converted version of the input image frame 202b so that the equalized converted version of the input image frame 202b is aligned with the converted version of the input image frame 202a. The electronic device 101 can use any suitable technique to warp the image data during the RGB warping operation 312. In some embodiments, for example, the electronic device 101 independently performs bilinear image warping for each of the red, green, and blue channels.

[0078] The electronic device 101 performs a luminance - chrominance (YUV) conversion operation 314a on the converted input image frame 202a, and performs a YUV conversion operation 314b on the equalized, converted, and warped input image frame 202b. Using the YUV conversion operations 314a - 314b, the electronic device 101 converts these versions of the input image frames 202a and 202b from the RGB domain to the YUV domain. In some embodiments, each of the YUV conversion operations 314a - 314b can be a twelve - bit operation (e.g., twelve - bit depth per pixel), although this is not required. In some cases, each of the YUV conversion operations 314a - 314b can perform the following operations.

[0079] (4)

[0080] (5)

[0081] (6)

[0082] (7)

[0083] (8)

[0084] (9)

[0085] Here, R, G, and B respectively represent the red value, green value, and blue value of each pixel in the RGB domain; b represents the number of bits of the depth of each pixel (e.g., twelve); , and respectively represent the intermediate values obtained after gamma operation; and Y, U, and V respectively represent the Y (luminance) value, U (blue chrominance) value, and V (red chrominance) value of each pixel in the YUV domain. Of course, these operations are only examples, and other suitable conversion operations can be used. Note that although Figure 3 shows two instances of the YUV conversion operations 314a - 314b (which can operate in parallel), one YUV conversion operation can be used, and different image data can be processed sequentially.

[0086] Once the input image frames 202a - 202b have been processed using the above operations, the electronic device 101 performs a shadow map operation 316 to generate an EV - 0 shadow map 322, performs a saturation map operation 318 to generate an EV - 0 saturation map 324, and performs an ME motion map operation 320 to generate an EV - 0 ME motion map 326. For example, during the shadow map operation 316, the electronic device 101 can create a base map from the converted input image frame 202a by using, for example, the following operations .

[0087] (10)

[0088] Here, R, G, and B respectively represent the red value, green value, and blue value of each pixel in the RGB domain. The electronic device 101 can filter the base map using, for example, a filter to generate the final EV - 0 shadow map 322 . In some cases, this can be expressed as follows.

[0089] (11)

[0090] During the saturation map operation 318, the electronic device 101 can create a base map from the converted input image frame 202a by using, for example, the following operations .

[0091] (12)

[0092] The electronic device 101 can also filter the base map using, for example, a filter Perform filtering to generate a final EV-0 saturation map 324 . In some cases, this can be expressed as follows.

[0093] (13)

[0094] During the ME motion map operation 320, the electronic device 101 can determine the motion represented between the input image frames 202a - 202b processed and converted in the YUV domain. In some embodiments, the ME motion map operation 320 is a per-pixel operation that results in the generation of an EV-0 ME motion map 326. The ME motion map operation 320 can use any suitable technique for generating a motion map.

[0095] The electronic device 101 performs an RGB ME blending operation 328 using a converted version of the input image frames 202a - 202b in the RGB domain and the outputs from the saturation map operation 318 and the ME motion map operation 320 to generate a blended image 330. The RGB ME blending operation 328 is a multi-exposure blending operation that can be performed independently for each of the red, green, and blue channels. In other words, the RGB ME blending operation 328 can blend the red image data in the converted version of the input image frames 202a - 202b, blend the green image data in the converted version of the input image frames 202a - 202b, and blend the blue image data in the converted version of the input image frames 202a - 202b. Here, the RGB ME blending operation 328 can operate to combine the image data in each color channel of the converted version of the input image frames 202a - 202b based on the EV-0 saturation map 324 and the EV-0 motion map 326. This results in the generation of a blended image 330 that represents a combination of the image data of the input image frames 202a - 202b converted to the RGB domain.

[0096] The electronic device 101 performs an RGB warping operation 332 on an equalized version of the input image frame 202c converted to the RGB domain. For example, similar to the RGB warping operation 312, the RGB warping operation 332 can use the alignment map 304 to warp the equalized converted version of the input image frame 202c so that the equalized converted version of the input image frame 202c is aligned with the converted version of the input image frame 202a. The electronic device 101 can use any suitable technique to warp the image data during the RGB warping operation 332. In some embodiments, for example, the electronic device 101 performs bilinear image warping independently for each of the red, green, and blue channels. Note that the RGB warping operation 332 can represent the same functional component as the RGB warping operation 312 but applied to different image data, or the RGB warping operation 332 can be implemented independently.

[0097] The electronic device 101 performs YUV conversion operations 334a - 334b on the hybrid image 330 and the warped version of the converted input image frame 202c, respectively. For example, using the YUV conversion operations 334a - 334b, the electronic device 101 can convert the hybrid image 330 and the warped version of the converted input image frame 202c from the RGB domain to the YUV domain. The YUV conversion operations 334a - 334b can be the same as or similar to the above-mentioned YUV conversion operations 314a - 314b. However, in some cases, the YUV conversion operations 334a - 334b can represent a fourteen-bit operation instead of a twelve-bit operation. Note that although Figure 3 two instances of the YUV conversion operations 334a - 334b (which can operate in parallel) are shown, one YUV conversion operation can be used, and different image data can be processed sequentially.

[0098] The electronic device 101 performs a shadow map operation 336 to generate an EV-2 shadow map 342, a saturation map operation 338 to generate an EV-2 saturation map 344, and a ME motion map operation 340 to generate an EV-2 ME motion map 346. Except that the input and output are based on the EV-2 input image frame 202b and the EV-4 input image frame 202c instead of the EV-0 input image frame 202a and the EV-2 input image frame 202b, these operations 336, 338, 340 can be the same as or similar to the corresponding operations 316, 318, 320 discussed previously. Note that the operations 336, 338, 340 can represent functional components that are the same as those of the operations 316, 318, 320 but applied to different image data, or the operations 336, 338, 340 can be implemented independently.

[0099] Regarding Figure 3 additional details of the specific example implementation of the ME analysis operation 204 shown can be found in U.S. Patent Application No. 18 / 149,714, filed on January 4, 2023 (the entire content of which is incorporated herein by reference). Note that although Figure 3Assume that the input image frames 202a - 202c are converted to the RGB domain and then to the YUV domain, but other embodiments of the ME analysis operation 204 can be used. For example, the input image frames 202a - 202c can be directly converted to the YUV domain, and U.S. Patent Application No. 18 / 149,714 provides additional details on how the ME analysis operation 204 can be implemented in this way. Additionally, any other or additional techniques can be used to generate the shadow map, saturation map, ME motion map, or other output 206. Generally, the present disclosure is not limited to any specific technique for generating the shadow map, saturation map, ME motion map, or other output 206 used for blending purposes, and the image data can be used in any suitable domain to generate the output 206.

[0100] Although Figure 3 illustrates Figure 2 an example of the ME analysis operation 204 in the architecture 200, various changes can be made to Figure 3 it. For example, the various components or functions shown can be combined, further subdivided, rearranged, duplicated, or omitted, and additional components can be added according to specific needs. As a specific example, as described above, the various components in Figure 3 can perform the same or similar functions, and these components can be shared across different image data or implemented independently for different image data. Additionally, although Figure 3 it is assumed that the EV - 0 image frame 202a is selected as the reference frame, this is not necessary. Another image frame 202b or 202c can be selected as the reference frame, in which case, the positions of at least some of the image frames 202a - 202c can be swapped in Figure 3 and suitable multiplication or other scaling operations can be used to substantially equalize the image frames 202a - 202c. Figure 3

[0101] Figure 4 illustrates an example multi - exposure (ME) blending operation 208 in the architecture 200 according to the present disclosure. Among other aspects, the ME blending operation 208 can be used to recover saturation from each color channel of the various input image frames 202a - 202c. As Figure 2 shown, the EV - 0 saturation map 324 and the EV - 0 ME motion map 326 are combined using a multiplication operation 402, which generates a combined map 404. The combined map 404 can represent the collective content of the saturation map 324 and the motion map 326. For example, the combined map 404 can represent the pixels of the image frames to be included during blending based on the content of the saturation map 324 and the motion map 326. Figure 4

[0102] Using multiplication operation 406, combine the composite graph 404 with an equalized version of the input image frame 202b, which is identified as input image frame 202b' and represented as (since it can be equalized with the input image frame 202a). Also provide the composite graph 404 to the complement operation 408, which performs a uniform inversion operation (e.g., 1 - x) on the composite graph 404. Here, the composite graph 404 represents the weight to be applied to the input image frame 202b' during blending, while its complement represents the weight to be applied to the input image frame 202a during blending. Combine the input image frame 202a (represented as I EV-0 ) with the output of the complement operation 408 using multiplication operation 410. Combine the resulting products generated and output by the multiplication operations 406 and 410 using the summation operation 412, which results in the generation of the blended image frame 414. The blended image frame 414 represents a weighted combination of the pixels from the input image frames 202a and 202b'.

[0103] Note that these operations combine the input image frame 202a and the input image frame 202b'. In some cases, the input image frame 202a can be combined with multiple input image frames 202b', in which case the same or similar process shown Figure 4 can be repeated or extended to use at least one additional equalized input image frame 202b' and its associated graphs 324, 326. The same or similar process shown Figure 4 can also be used to combine the input image frame 202a or the blended image frame 414 with at least one equalized version of at least one input image frame 202c (which can be represented as I 16×EV-4 ), such as when replacing the EV - 0 input image frame 202a in Figure 4 with the blended image frame 414, replacing the equalized version of the EV - 2 input image frame 202b' in Figure 4 with an equalized version of the EV - 4 input image frame 202c, and repeating the process shown in Figure 4 when replacing the graphs 324, 326 with the graphs 344, 346. Additionally, note that Figure 4 the same or similar process shown Figure 3 can be used to implement the RGB ME blending operation 328 shown in

[0104] In this case, the RGB ME blending operation 328 can combine the image data from the equalized and processed versions of the input image frames 202a - 202b using the graphs 324, 326. Figure 4Additional details of the specific example implementation of the ME mixing operation 208 shown can be found in U.S. Patent Application No. 18 / 149,714, which is incorporated herein by reference. Additionally, any other or additional techniques can be used to mix image frames in the ME mixing operation 208 and / or the RGB ME mixing operation 328, and each of these mixing operations can be implemented in any other suitable manner. Generally, the present disclosure is not limited to any specific technique for mixing image frames.

[0105] Although Figure 4 an example of the ME mixing operation 208 in the architecture 200 is shown, Figure 2 various changes can be made to Figure 4 it. For example, the various components or functions shown can be combined, further subdivided, rearranged, copied, or omitted, Figure 4 and additional components can be added according to specific needs. As a specific example, the various components in Figure 4 can perform the same or similar functions, and these components can be shared across different image data or implemented independently for different image data. Additionally, although Figure 4 the EV-0 image frame 202a is assumed to be selected as the reference frame, this is not required, and another image frame 202b or 202c can be selected as the reference frame (in which case, the positions of certain image frames can be swapped).

[0106] Figure 5 An example tone mapping operation 212 in the architecture 200 according to the present disclosure is shown. As Figure 2 shown, the tone mapping operation 212 receives the HDR mixed image 210 and provides it to the LDR image synthesis operation 502. The LDR image synthesis operation 502 uses the HDR mixed image 210 to generate a plurality of LDR images 504, and each LDR image 504 can represent image data from the HDR mixed image 210 but having a lower dynamic range than the HDR mixed image 210. The LDR image synthesis operation 502 can use any suitable technique for generating the LDR images 504 based on the HDR mixed image 210. Details of an example method used by the LDR image synthesis operation 502 are shown in Figure 5 described below. In some cases, each LDR image 504 can represent a YUV image. Figure 6

[0107] The LDR image 504 is provided to a weight map generation operation 506 that processes the LDR image 504 to generate an initial weight map 508 associated with the LDR image 504. The initial weight map 508 represents an initial map or other set of indicators identifying how the image data from the LDR image 504 can be combined to produce the fused image 214. For example, each initial weight map 508 can include weights to be applied to the image data from one of the LDR images 504 during blending. The weight map generation operation 506 can use any suitable technique for generating the initial weight map 508 for the LDR image 504. For example, the weight map generation operation 506 can generate saliency measures, color saturation measures, and good exposure measures for each LDR image 504. The saliency measures, color saturation measures, and good exposure measures can be combined for each LDR image 504, and the combined measures can be normalized to generate the initial weight map 508 for each LDR image 504. Details of one example method used by the weight map generation operation 506 are described in Figure 7 and Figure 8 shown in

[0108] The initial weight map 508 is provided to a weight map modification operation 510 that modifies one or more of the initial weight maps 508 based on the semantic incremental weight map 226 to generate a modified weight map 512. As described above, the semantic incremental weight map 226 is based on the semantic content of the HDR blended image 210, and the semantic incremental weight map 226 can be used to modify one or more of the initial weight maps 508 to adjust the initial weight map 508 based on the semantic content of the HDR blended image 210. As a specific example, the weight map modification operation 510 can scale, increase, or decrease the values in one or more of the initial weight maps 508 based on the values in the semantic incremental weight map 226, such as by adding the value based on the semantic incremental weight map 226 to the value in the initial weight map 508, or subtracting the value based on the semantic incremental weight map 226 from the value in the initial weight map 508. The weight map modification operation 510 can use any suitable technique for modifying the initial weight map 508 and generate the modified weight map 512.

[0109] The modified weight map 512 is provided to a guided filtering operation 514, which processes the modified weight map 512 to generate a filtered weight map 516. The guided filtering operation 514 can filter the modified weight map 512 to provide any desired effect in the filtered weight map 516. As a specific example, the guided filtering operation 514 can process the modified weight map 512 to remove noise from the modified weight map 512 while preserving the edges in the modified weight map 512. The guided filtering operation 514 can use any suitable technique for filtering the weight map. In some embodiments, for example, the guided filtering operation 514 can filter the weight map using the method described by He et al. in "Guided image filtering" in IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 35, Issue 6, pp. 1397-1409, in June 2013, the entire content of which is incorporated herein by reference.

[0110] Here, the LDR image 504 is also provided to a decomposition operation 518, which decomposes each LDR image 504 into a base layer 520 containing the base components of the LDR image 504 and a detail layer 522 containing the detail components of the LDR image 504. The base layer 520 of each LDR image 504 can include large-scale variations in the intensity of the LDR image 504, and the detail layer 522 of each LDR image 504 can capture the fine-scale details of the LDR image 504. The decomposition operation 518 can use any suitable technique for decomposing the LDR image 504 into the base layer 520 and the detail layer 522. In some cases, for example, the decomposition operation 518 can use box filtering. As a specific instance, using the luminance channel (Y) as an example, let 1, ……, N represent the luminance channels of each LDR image 504. The base layer of each LDR image 504 can be defined as follows.

[0111] (14)

[0112] Here, θ represents the size of the box filter. The detail layer of each LDR image 504 can be obtained by subtracting the associated base layer 520 from the corresponding LDR image 504, which can be defined as follows.

[0113] (15)

[0114] In addition to first subtracting a specified value from the chromaticity values (e.g., 0.5 if the chromaticity values are between 0 and 1, or 128 if the chromaticity values are between 0 and 255), the decomposition for the chromaticity (U and V) channels can occur in a similar manner. Note that the choice of the value for θ may be important here. If the value of θ is too large, each detail layer 522 may include more low-frequency information and may produce unnatural transition boundaries in the resulting image. If the value of θ is too small, the details may not be well preserved.

[0115] The base blend operation 524 is used to blend the base layer 520 of the LDR image 504, and the detail blend operation 526 is used to blend the detail layers 522 of the LDR image 504. The blend operations 524 and 526 use the filter weight maps 516 here to blend the base layer 520 and the detail layers 522 of the LDR image 504, respectively. For example, each filter weight map 516 can identify the weights to be applied to one of the LDR images 504, and each of the blend operations 524 and 526 can perform a weighted combination of the image data in the base layer 520 or the image data in the detail layers 522 (where the values in each of the base layer 520 or the detail layers 522 are scaled or otherwise weighted using the values in the associated filter weight map 516). Thus, a combined base layer 528 can be generated using the base blend operation 524, and the combined base layer 528 represents a weighted combination of the base layer 520 of the LDR image 504. Similarly, a combined detail layer 530 can be generated using the detail blend operation 526, and the combined detail layer 530 represents a weighted combination of the detail layers 522 of the LDR image 504. Each of the blend operations 524 and 526 can use any suitable technique for performing the weighted blend of the LDR image 504.

[0116] The combined base layer 528 and the combined detail layer 530 are provided to a combining operation 532, and the combining operation 532 combines the combined base layer 528 and the combined detail layer 530 to generate a fused image 214. The combining operation 532 can use any suitable technique to combine the base image data and the detail image data to generate the fused image 214. In some cases, for example, the combining operation 532 can add the values of the combined base layer 528 and the values of the combined detail layer 530 to generate the fused image 214.

[0117] As described above, the LDR image synthesis operation 502 can generate the LDR image 504 based on the HDR blended image 210. In some embodiments, the LDR image synthesis operation 502 can generate the LDR image 504 by calculating a fusion scale based on the histogram of the image data included in the HDR blended image 210 and using the HDR blended image 210 and the fusion scale to generate the LDR image 504.Figure 6 illustrates an example image histogram 600 used for the tone blending operation 212 according to the present disclosure. As can be seen in Figure 5 , the image histogram 600 identifies the number of pixels from the HDR blended image 210 within each of a plurality of bins 602, where each bin 602 is associated with a different range of pixel values. In some cases, the image histogram 600 can be generated based on the luminance associated with each pixel of the HDR blended image 210. In this particular example, the image histogram 600 is shown to be generated in the log Figure 6 domain, and the bins 602 can be associated with unequal numbers of pixel values. However, any other suitable image histogram 600 can be generated and used here. 2 domain, and the bins 602 can be associated with unequal numbers of pixel values. However, any other suitable image histogram 600 can be generated and used here.

[0118] Based on the image histogram 600, a plurality of blending scales can be generated based on the number of pixels in at least some of the bins 602 of the image histogram 600. In some embodiments, for example, five blending scales can be determined as follows.

[0119] (16)

[0120] (17)

[0121] (18)

[0122] (19)

[0123] (20)

[0124] In equations (16)-(20), represents a blending scale, represents the number of pixels in the i-th bin 602, and represents the j-th adjustable parameter. Each adjustable parameter can be used to adjust the weighting between two adjacent bins 602 of the image histogram 600. The LDR image composition operation 502 can multiply the image data of the HDR blended image 210 by each blending scale , the scaled image data obtained by cropping, and apply any desired image signal processing (ISP) conversion operations to the cropped and scaled image data. For example, the ISP conversion operations can include demosaicing, dynamic range control (DRC), color correction using a color correction matrix (CCM), gamma correction, and RGB to YUV conversion. Details of example implementations of specific ISP conversion operations can be found in U.S. Patent No. 11,388,348, the entire content of which is incorporated herein by reference. The result of the ISP conversion operations represents the LDR image 504. Note that in this particular example, the LDR image synthesis operation 502 can generate five LDR images 504 based on five fusion scales, although other numbers of fusion scales and LDR images 504 can also be generated.

[0125] As described above, the weight map generation operation 506 processes the LDR image 504 to generate an initial weight map 508 associated with the LDR image 504. In some embodiments, the weight map generation operation 506 can generate the initial weight map 508 by calculating a saliency metric, a color saturation metric, and a good exposure metric for each LDR image 504, combining the metrics for each LDR image 504, and normalizing the combined metrics to generate the initial weight map 508 for each LDR image 504. In some embodiments, the saliency metric, color saturation metric, and good exposure metric for each LDR image 504 can be determined as follows. To generate the saliency metric for each LDR image 504, the weight map generation operation 506 can perform Laplacian filtering on the LDR image 504, obtain the absolute value of the filtered output generated by the Laplacian filtering, and apply Gaussian filtering to the absolute value. In some cases, this can generate the saliency metric for each pixel of each LDR image 504.

[0126] To generate the color saturation metric for each LDR image 504, the weight map generation operation 506 can apply a color saturation lookup table to the image chrominance data (U and V channels) of the LDR image 504 and sum the values from the color saturation lookup table. To generate the good exposure metric for each LDR image 504, the weight map generation operation 506 can apply a good exposure lookup table to the image data of the LDR image 504. In some cases, this can generate the color saturation metric and the good exposure metric for each pixel of each LDR image 504. Each lookup table can be used to transform the pixel values from the LDR image 504 into corresponding values for generating the color saturation metric or the good exposure metric.

[0127] Figure 7 and Figure 8 illustrates for Figure 5Examples of the look-up tables 700 and 800 used in the hue fusion operation 212. More specifically, the look-up table 700 represents an example implementation of a color saturation look-up table, where the look-up table 700 is used to transform the image chromaticity data of each LDR image 504 plotted along the horizontal axis into corresponding values plotted along the vertical axis. The corresponding values for each pixel in the U and V image data of each LDR image 504 can be added and used as a color saturation measure for that pixel in the LDR image 504. The look-up table 800 represents an example implementation of a good exposure look-up table, where the look-up table 800 is used to transform the image data of each LDR image 504 plotted along the horizontal axis into corresponding values plotted along the vertical axis. The corresponding values for each pixel in the image data of each LDR image 504 can be used as a good exposure measure for that pixel in the LDR image 504.

[0128] These methods can generate a saliency measure, a color saturation measure, and a good exposure measure for each pixel of each LDR image 504. The saliency measure, color saturation measure, and good exposure measure for each pixel of each LDR image 504 can be combined and normalized to generate an initial weight map 508 for each LDR image 504. For example, in some embodiments, the saliency measure, color saturation measure, and good exposure measure determined for each LDR image 504 can be combined as follows.

[0129] (21)

[0130] Here, represents the combined measure for each pixel of the j-th LDR image 504, represents the saliency measure for each pixel of the j-th LDR image 504, represents the color saturation measure for each pixel of the j-th LDR image 504, and represents the good exposure measure for each pixel of the j-th LDR image 504. The variable j here is assumed to have five possible values due to the generation of five LDR images 504 based on five fusion scales, although other numbers of LDR images 504 can be generated as described above. The normalization of the combined measure can be performed as follows.

[0131] (22)

[0132] Here, represents the normalized combined measure for each pixel of the j-th LDR image 504. The normalized combined measures for all pixels of each LDR image 504 can be used to form the initial weight map 508 for that LDR image 504.

[0133] As described above, one or more in the initial weight map 508 are modified by the weight map modification operation 510 to generate a modified weight map 512, and the guided filtering operation 514 filters the modified weight map 512 to generate a filtered weight map 516. Among other things, the guided filtering operation 514 can be used to remove noise from the modified weight map 512 while preserving the edges identified in the modified weight map 512. Tone fusion can be achieved by weighted mixing of the base layer 520 and the detail layer 522 of the LDR image 504 (which is based on the filtered weight map 516) and combining the resulting combined base layer 528 and combined detail layer 530. In some cases, the operations for generating the fused image 214 can be represented as follows.

[0134] (23)

[0135] (24)

[0136] Here, represents the combined base layer 528 or the combined detail layer 530, represents the base layer 520 or the detail layer 522 of the j-th LDR image 504, and I represents the fused image 214. Also, since five LDR images 504 are generated, the variable j is assumed to have five possible values here, although other numbers of LDR images 504 can be generated as described above.

[0137] Although Figures 5 to 8 shows Figure 2 an example and related details of the tone fusion operation 212 in the architecture 200 of Figures 5 to 8 various changes can be made to Figure 5 For example, various components or functions shown in Figure 6 The content of the histogram 600 in can vary according to the situation (e.g., based on the content of the input image frame 202 actually being processed). Additionally, Figure 7 and Figure 8 The content of the look-up tables 700, 800 in are only examples and can vary according to the implementation.

[0138] Figure 9 shows an exemplary semantic-based machine learning model (MLM) segmentation operation 220 in the architecture 200 according to the present disclosure. As Figure 2 in Figure 9As shown, the semantics-based MLM segmentation operation 220 receives the HDR mixed image 210 and provides it to the down-conversion operation 902. The down-conversion operation 902 operates to obtain a higher-resolution image and generate a lower-resolution version of the image. Here, the down-conversion operation 902 converts the HDR mixed image 210 into a lower-resolution HDR mixed image 904. The down-conversion operation 902 can use any suitable technique to down-convert the resolution of the image or otherwise reduce the resolution of the image.

[0139] The lower-resolution HDR mixed image 904 is provided to the HDR gamma correction operation 906, which processes the lower-resolution HDR mixed image 904 to reduce the dynamic range of the lower-resolution HDR mixed image 904 and generate a lower-resolution LDR mixed image 908. The HDR gamma correction operation 906 can use any suitable technique for performing gamma correction and reducing the dynamic range of the image. In some embodiments, for example, the HDR gamma correction operation 906 uses a gamma correction look-up table to perform gamma correction. Figure 10 An example look-up table 1000 used by the semantics-based machine learning model segmentation operation 220 according to the present disclosure is shown. Here, the look-up table 1000 represents a gamma correction look-up table and can be used to convert values associated with the lower-resolution HDR mixed image 904 plotted along the horizontal axis into corresponding values associated with the lower-resolution LDR mixed image 908 plotted along the vertical axis. Figure 9

[0140] The lower-resolution LDR mixed image 908 is provided to the raw-to-YUV conversion operation 910, which converts the lower-resolution LDR mixed image 908 from the raw image domain to the YUV image domain. Here, it is assumed that the HDR mixed image 210 (and thus the lower-resolution HDR mixed image 904 and the lower-resolution LDR mixed image 908) is in the raw image domain, so the raw-to-YUV conversion operation 910 can be used to convert the lower-resolution LDR mixed image 908 into a lower-resolution LDR mixed YUV image 912. The raw-to-YUV conversion operation 910 can use any suitable technique to convert the raw image data into YUV image data. In some cases, for example, the raw-to-YUV conversion operation 910 can use the above equations (1)-(3) and (7)-(9) to convert the image data of the lower-resolution LDR mixed image 908 into the corresponding image data of the lower-resolution LDR mixed YUV image 912.

[0141] A lower resolution LDR hybrid YUV image 912 is provided to a trained machine learning model 914, which processes the lower resolution LDR hybrid YUV image 912 to generate a semantic segmentation mask 222 associated with the HDR hybrid image 210. The trained machine learning model 914 can represent any suitable machine learning model structure (e.g., a convolutional neural network (CNN) or other neural network). In some embodiments, for example, the trained machine learning model 914 can represent a trained U-Net architecture, which is a CNN-based architecture designed to perform image segmentation. The trained machine learning model 914 can be trained to process the lower resolution LDR hybrid YUV image 912, identify objects within the lower resolution LDR hybrid YUV image 912, and generate the semantic segmentation mask 222, which identifies which pixels in the lower resolution LDR hybrid YUV image 912 are associated with different objects identified within the lower resolution LDR hybrid YUV image 912.

[0142] In some cases, the machine learning model 914 can be trained by providing training images to the machine learning model 914 and comparing the semantic segmentation mask 222 generated by the machine learning model 914 with a ground truth semantic segmentation mask. The ground truth semantic segmentation mask represents the desired output of the machine learning model 914, and comparing the semantic segmentation mask 222 actually generated by the machine learning model 914 with the ground truth semantic segmentation mask can produce a loss value that identifies the accuracy of the machine learning model 914. When the loss value exceeds a threshold loss value, the weights or other parameters of the machine learning model 914 can be adjusted, and another training iteration can occur. This can be repeated any number of times, ideally until the machine learning model 914 generates a semantic segmentation mask 222 that is the same or substantially similar to the ground truth semantic segmentation mask (at least within the desired accuracy defined by the threshold loss value).

[0143] Note that the input to the semantic-based MLM segmentation operation 220 here is the HDR hybrid image 210 generated based on mixing multiple input image frames 202. The semantic-based MLM segmentation operation 220 does not simply receive and process a single input image frame 202 with a lower dynamic range than the HDR hybrid image 210. Generating the semantic segmentation mask 222 using the HDR hybrid image 210 can help improve the accuracy and robustness of the semantic segmentation mask 222 compared to a semantic segmentation mask that might be generated using only one of the input image frames 202.

[0144] Although Figure 9 and Figure 10 shows Figure 2An example and related details of the semantic-based MLM segmentation operation 220 in the architecture 200 of FIG. 2 are provided, but may be used for Figure 9 and Figure 10 Make various changes. For example, they may be combined, further subdivided, rearranged, duplicated, or omitted Figure 9 Various components or functions are shown, and additional components can be added according to specific needs. Figure 10 The contents of the lookup table 1000 in are merely examples and may vary depending on the implementation.

[0145] Figure 11 The present disclosure shows Figure 2 10. As described above, the noise-aware segmentation to delta-weight conversion operation 224 operates to convert the contents of the semantic segmentation mask 222 into a semantic delta-weight map 226, where the semantic delta-weight map 226 is used to modify the initial weight map 508 (thereby allowing for semantic-based guidance for the hue blending operation 212). In some embodiments, the noise-aware segmentation to delta-weight conversion operation 224 may map different values ​​from the semantic segmentation mask 222 to corresponding values ​​in the semantic delta-weight map 226. The different values ​​in the semantic segmentation mask 222 represent different types or categories of semantic content in the HDR hybrid image 210, and the different values ​​in the semantic delta-weight map 226 represent different modifications to be made to at least one initial weight map 508. In certain embodiments, the noise-aware segmentation to delta-weight conversion operation 224 may map the values ​​differently based on metadata associated with the HDR hybrid image 210 (e.g., ISO, exposure time, and brightness), or image data contained in the HDR hybrid image 210.

[0146] like Figure 11 As shown, the map 1100 here associates ISO values ​​plotted along the horizontal axis with incremental weight values ​​plotted along the vertical axis. In this particular example, an incremental weight value of 128 is used to indicate that no adjustment is to be made to the values ​​in the at least one initial weight map 508. Incremental weight values ​​above 128 may be used to indicate that the values ​​in the at least one initial weight map 508 are to be increased, and incremental weight values ​​below 128 may be used to indicate that the values ​​in the at least one initial weight map 508 are to be decreased.

[0147] In this example, the mapping 1100 varies based on which semantic class or classes of content are identified in the semantic segmentation mask 222. For example, line 1102 can be used to represent a mapping for transforming the value for the "background" semantic class in the semantic segmentation mask 222 into a corresponding incremental weight value. Line 1104 can be used to represent a mapping for transforming the value for the "sky" semantic class in the semantic segmentation mask 222 into a corresponding incremental weight value. Line 1106 can be used to represent a mapping for transforming the value for the "person" semantic class in the semantic segmentation mask 222 into a corresponding incremental weight value. Line 1108 can be used to represent a mapping for transforming the value for the "grass" semantic class in the semantic segmentation mask 222 into a corresponding incremental weight value. Line 1110 can be used to represent a mapping for transforming the value for the "ground" semantic class in the semantic segmentation mask 222 into a corresponding incremental weight value. Line 1112 can be used to represent a mapping for transforming the value for the "building" semantic class in the semantic segmentation mask 222 into a corresponding incremental weight value.

[0148] Based on this, if the noise-aware segmentation to incremental weight conversion operation 224 determines that the semantic segmentation mask 222 identifies pixels in the HDR composite image 210 associated with the background, the conversion operation 224 can use the ISO value associated with the HDR composite image 210 to determine the incremental weight correction value to use along line 1102 in the semantic incremental weight map 226, and this value can be used in the semantic incremental weight map 226 for each pixel associated with the background. If the noise-aware segmentation to incremental weight conversion operation 224 determines that the semantic segmentation mask 222 identifies pixels in the HDR composite image 210 associated with the sky, the conversion operation 224 can use the ISO value associated with the HDR composite image 210 to determine the incremental weight correction value to use along line 1104 in the semantic incremental weight map 226, and this value can be used in the semantic incremental weight map 226 for each pixel associated with the sky. This can occur for each semantic class included in the semantic segmentation mask 222, thus transforming the values in the semantic segmentation mask 222 (each value identifying a specific semantic class) into corresponding values in the semantic incremental weight map 226.

[0149] As described above, the weight map modification operation 510 modifies one or more of the initial weight maps 508 based on the semantic incremental weight map 226 to generate a modified weight map 512. In some embodiments, and without loss of generality (since the weights are normalized for blending and changing one weight means changing all), the weight map modification operation 510 can select the initial weight map 508 associated with the brightest LDR image 504 (meaning the LDR image 504 associated with the highest fusion scale). The selected initial weight map 508 can undergo incremental weight modification, for example, in the following manner.

[0150] (25)

[0151] Here, represents each pixel value in the selected initial weight map 508, and represents each pixel value in the corresponding modified weight map 512. Additionally, δ represents the incremental weight value from the semantic incremental weight map 226 applied to each pixel value in the selected initial weight map 508. Note that if each incremental weight value from the semantic incremental weight map 226 could alternatively have a range from 0 to 1, the expression "δ / 256" in equation (25) could be replaced with "δ". Similar operations can occur for other initial weight maps 508 in order to generate additional modified weight maps 512, or other initial weight maps 508 can be output unchanged and used as additional modified weight maps 512.

[0152] Although Figure 11 shows an example of the mapping 1100 used for noise-aware segmentation to incremental weight conversion operation 224 in the architecture 200 provided for Figure 2 , various changes can be made to Figure 11 . For example, other or additional mappings can be based on factors different from or in addition to the ISO value (e.g., exposure time and / or brightness). Additionally, Figure 11 the specific lines shown and their associated semantic classes are for illustration only. The mapping 1100 can include any number of lines, and each line can be associated with any suitable one or more semantic classes.

[0153] It should be noted that Figures 2 to 11 the functions shown or described above can be implemented in any suitable manner in the electronic devices 101, 102, 104, the server 106, or other devices. For example, in some embodiments, one or more software applications or other software instructions executed by the processor 120 of the electronic devices 101, 102, 104, the server 106, or other devices can be used to implement or support Figures 2 to 11 at least some of the functions shown or described above. In other embodiments, dedicated hardware components can be used to implement or support Figures 2 to 11 at least some of the functions shown or described above. Generally, any suitable hardware, or any suitable combination of hardware and software / firmware instructions can be used to perform Figures 2 to 11 the functions shown or described above. Additionally, the functions can be performed by a single device or by multiple devices Figures 2 to 11The functions shown or described above. For example, the server 106 can be used to train the machine learning model 914, and the server 106 can deploy the trained machine learning model 914 to one or more other devices (e.g., the electronic device 101) for use.

[0154] Figure 12 FIG. 1200 illustrates an example method for tone mapping based on machine learning segmentation according to the present disclosure. For ease of explanation, method 1200 is described as being performed by Figure 1 the electronic device 101 in the network configuration 100 shown Figure 2 using the architecture 200 shown. However, method 1200 can be performed using any other suitable device that supports any other suitable architecture and in any other suitable system.

[0155] As Figure 12 shown, at step 1202, a plurality of input image frames of a scene are obtained. This can include, for example, the processor 120 of the electronic device 101 obtaining the input image frames 202 using one or more imaging sensors 180 of the electronic device 101. At step 1204, an HDR mixed image is generated based on the input image frames. This can include, for example, the processor 120 of the electronic device 101 performing ME analysis operations 204 to generate a shadow map, a saturation map, an ME motion map, or other outputs 206. This can also include the processor 120 of the electronic device 101 performing ME mixing operations 208 to mix the input image frames 202 based on the outputs 206, thereby generating an HDR mixed image 210. The HDR mixed image 210 can have a higher dynamic range than any single one of the input image frames 202.

[0156] Tone mapping can be performed on the HDR hybrid image to generate a tone-mapped image. For example, at step 1206, a semantic segmentation mask is generated based on the HDR hybrid image. This can include, for example, the processor 120 of the electronic device 101 performing an HDR-semantics-based MLM segmentation operation 220 to generate a semantic segmentation mask 222 based on the content of the HDR hybrid image 210. As a specific example, this can include the processor 120 of the electronic device 101 generating a lower-resolution HDR hybrid image 904 based on the HDR hybrid image 210, generating a lower-resolution LDR hybrid image 908 based on the lower-resolution HDR hybrid image 904, generating a lower-resolution LDR hybrid YUV image 912 based on the lower-resolution LDR hybrid image 908, and (possibly by using a trained machine learning model 914) generating a semantic segmentation mask 222 based on the lower-resolution LDR hybrid YUV image 912. At step 1208, a semantic delta weight map is generated based on the semantic segmentation mask. This can include, for example, the processor 120 of the electronic device 101 performing a noise-aware segmentation-to-delta weight conversion operation 224 to convert the semantic segmentation mask 222 into a semantic delta weight map 226. As a specific example, this can include the processor 120 of the electronic device 101 using a mapping 1100 to transform different values of different semantic classes in the semantic segmentation mask 222 into corresponding values in the semantic delta weight map 226, which can be done based on metadata associated with the HDR hybrid image 210 (e.g., ISO, exposure time, and / or luminance) or image data contained in the HDR hybrid image 210.

[0157] At step 1210, multiple LDR images are synthesized based on the HDR hybrid image. This can include, for example, the processor 120 of the electronic device 101 performing an LDR image synthesis operation 502 to generate multiple LDR images 504 based on the HDR hybrid image 210. As a specific example, this can include the processor 120 of the electronic device 101 identifying an image histogram 600 of the HDR hybrid image 210, determining multiple fusion scales based on the image histogram 600, multiplying the image data of the HDR hybrid image 210 by the fusion scales and cropping the resulting image data to generate cropped image data, and applying an ISP conversion (e.g., demosaicing, DRC, color correction using CCM, gamma correction, and RGB-to-YUV conversion) to the cropped image data to generate a YUV image.

[0158] At step 1212, an initial weight map for the LDR image is generated. This can include, for example, the processor 120 of the electronic device 101 performing a weight map generation operation 506 to process the LDR image 504 and generate an initial weight map 508 based on the image content of the LDR image 504. As a specific example, this can include the processor 120 of the electronic device 101 generating a saliency measure, a color saturation measure, and a good exposure measure for the LDR image 504. The color saturation measure and the good exposure measure can be generated using a first look-up table 700 and a second look-up table 800 respectively. This can also include the processor 120 of the electronic device 101 combining the saliency measure, the color saturation measure, and the good exposure measure for each LDR image 504 and normalizing the combined measures to generate an initial weight map 508 for the LDR image 504. At step 1214, the initial weight map is modified and filtered to generate a filtered weight map. This can include, for example, the processor 120 of the electronic device 101 performing a weight map modification operation 510 to modify one or more initial weight maps 508 based on the semantic incremental weight map 226 and generate a modified weight map 512. This can also include the processor 120 of the electronic device 101 performing a guided filtering operation 514 to filter the modified weight map 512 using a guided filter and generate a filtered weight map 516.

[0159] At step 1216, the LDR image is decomposed into a base layer and a detail layer. This can include, for example, the processor 120 of the electronic device 101 performing a decomposition operation 518 to decompose each LDR image 504 into a base layer 520 and a detail layer 522. At step 1218, based on the filtered weight map, the base layers of the LDR images are mixed with each other, and the detail layers of the LDR images are mixed with each other. This can include, for example, the processor 120 of the electronic device 101 performing a base mixing operation 524 to mix the base layers 520 of the LDR images 504 and generate a combined base layer 528. This can also include the processor 120 of the electronic device 101 performing a detail mixing operation 526 to mix the detail layers 522 of the LDR images 504 and generate a combined detail layer 530. At step 1220, an output image is generated based on the mixed base layer and the mixed detail layer. This can include, for example, the processor 120 of the electronic device 101 performing a combining operation 532 that combines the combined base layer 528 and the combined detail layer 530 to generate a fused image 214. The fused image 214 can optionally undergo a tone mapping operation 216 or other post-processing to generate an output image 218.

[0160] The output image can be used in any suitable manner. For example, at step 1222, the output image can be stored, output, or used. This can include, for example, the processor 120 of the electronic device 101 presenting the output image 218 on the display 160 of the electronic device 101, saving the output image 218 to a camera roll stored in the memory 130 of the electronic device 101, or attaching the output image 218 to a text message, email, or other communication to be sent from the electronic device 101. However, note that the output image 218 can be used in any other or additional manner.

[0161] Although Figure 12 an example of a method 1200 for tone mapping based on machine learning segmentation is shown, various changes can be made to Figure 12 it. For example, although shown as a series of steps, Figure 12 the individual steps in

[0162] it can overlap, occur in parallel, occur in a different order, or occur any number of times (including zero times). Although the present disclosure has been described with reference to various example embodiments, various changes and modifications can be made by those skilled in the art. The present disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims.

Claims

1. A method for processing image frames by an electronic device (101), comprising: obtaining (202) a plurality of input image frames; generating (208) a high dynamic range HDR hybrid image based on the plurality of input image frames; and performing (212) a tone fusion operation on the HDR hybrid image based on a semantic incremental weight map to generate a fused image; wherein performing the tone fusion operation comprises: synthesizing a plurality of low dynamic range LDR images based on the HDR hybrid image; generating an initial weight map based on the plurality of LDR images; generating a filtered weight map based on the initial weight map, the semantic incremental weight map, and a guidance filter; and generating the fused image based on the filtered weight map and a decomposed version of the plurality of LDR images.

2. The method according to claim 1, wherein generating the fused image based on the filtered weight map and the decomposed version of the plurality of LDR images comprises: performing image decomposition on each LDR image of the plurality of LDR images to generate a base component and a detail component for each LDR image of the plurality of LDR images; performing a base blending operation based on the filtered weight map and the base components of the plurality of LDR images; performing a detail blending operation based on the filtered weight map and the detail components of the plurality of LDR images; and combining the results of the base blending operation and the detail blending operation to generate the fused image.

3. The method according to claim 1, wherein synthesizing the plurality of LDR images based on the HDR hybrid image comprises: identifying an image histogram of the HDR hybrid image; determining a plurality of fusion scales based on the image histogram; multiplying the image data of the HDR hybrid image by the fusion scales and cropping the resulting image data to generate cropped image data; and applying an image signal processing ISP conversion to the cropped image data so as to generate a YUV image.

4. The method according to claim 3, wherein generating the initial weight map based on the plurality of LDR images comprises: generating a saliency measure for the YUV image; generating a color saturation measure for the YUV image using a first look-up table; generating a good exposure measure for the YUV image using a second look-up table; and combining the saliency measure, the color saturation measure, and the good exposure measure for each YUV image and normalizing the combined measures to generate an initial weight map for each YUV image.

5. The method according to claim 1, wherein generating the filtered weight map comprises: generating a modified weight map based on the initial weight map and the semantic incremental weight map; and removing noise from the modified weight map using the guidance filter while preserving edges in the modified weight map.

6. The method according to claim 1, further comprising: generating the semantic incremental weight map, wherein generating the semantic incremental weight map comprises: generating a lower resolution HDR hybrid image based on the HDR hybrid image; Generate a lower-resolution LDR mixed image based on the lower-resolution HDR mixed image; Generate a lower-resolution LDR mixed YUV image based on the lower-resolution LDR mixed image; Generate a semantic segmentation mask based on the lower-resolution LDR mixed YUV image; and Generate the semantic incremental weight map based on the semantic segmentation mask using a mapping.

7. The method according to claim 6, wherein: the semantic segmentation mask is generated using a trained machine learning model that processes the lower-resolution LDR mixed YUV image; and the mapping is used to transform different values of different semantic classes in the semantic segmentation mask into corresponding values in the semantic incremental weight map.

8. An electronic device (101), comprising: a memory (130) storing instructions; and at least one processor (120) operatively coupled to the memory, wherein when the at least one processor (120) executes the instructions, the at least one processor (120) causes the electronic device to perform operations, the operations including: obtain (202) a plurality of input image frames; generate (208) a high dynamic range HDR mixed image based on the plurality of input image frames, the HDR mixed image having a higher dynamic range than each of the input image frames in the plurality of input image frames; and perform (212) a tone fusion operation on the HDR mixed image based on a semantic incremental weight map to generate a fused image; wherein performing the tone fusion operation includes: synthesizing a plurality of low dynamic range LDR images based on the HDR mixed image; generating an initial weight map based on the plurality of LDR images; generating a filtered weight map based on the initial weight map, the semantic incremental weight map, and a guided filter; and generating the fused image based on the filtered weight map and a decomposed version of the plurality of LDR images.

9. The electronic device according to claim 8, wherein, generating the fused image based on the filtered weight map and the decomposed version of the plurality of LDR images includes: performing image decomposition on each of the plurality of LDR images to generate a base component and a detail component of each of the plurality of LDR images; performing a base mixing operation based on the filtered weight map and the base components of the plurality of LDR images; performing a detail mixing operation based on the filtered weight map and the detail components of the plurality of LDR images; and combining the results of the base mixing operation and the detail mixing operation to generate the fused image.

10. The electronic device according to claim 8, wherein, synthesizing the plurality of LDR images based on the HDR mixed image includes: identifying an image histogram of the HDR mixed image; determining a plurality of fusion scales based on the image histogram; multiplying image data of the HDR mixed image by the fusion scales and cropping the resulting image data to generate cropped image data; and applying an image signal processing ISP conversion to the cropped image data so as to generate a YUV image.

11. The electronic device according to claim 10, wherein, generating the initial weight map based on the plurality of LDR images includes: generating a saliency measure for the YUV image; generating a color saturation measure for the YUV image using a first look-up table; generating a good exposure measure for the YUV image using a second look-up table; and combining the saliency measure, the color saturation measure, and the good exposure measure for each YUV image, and normalizing the combined measures to generate an initial weight map for each YUV image.

12. The electronic device according to claim 8, wherein, generating the filtered weight map includes: generating a modified weight map based on the initial weight map and the semantic incremental weight map; and using the guided filter to remove noise from the modified weight map while retaining edges in the modified weight map.

13. The electronic device according to claim 8, wherein, the operation further includes: generating the semantic incremental weight map, and wherein generating the semantic incremental weight map includes: generating a lower-resolution HDR hybrid image based on the HDR hybrid image; generating a lower-resolution LDR hybrid image based on the lower-resolution HDR hybrid image; generating a lower-resolution LDR hybrid YUV image based on the lower-resolution LDR hybrid image; generating a semantic segmentation mask based on the lower-resolution LDR hybrid YUV image; and generating the semantic incremental weight map using a mapping based on the semantic segmentation mask.

14. The electronic device according to claim 13, wherein: the semantic segmentation mask is generated using a trained machine learning model that processes the lower-resolution LDR hybrid YUV image; and using the mapping to transform different values of different semantic classes in the semantic segmentation mask into corresponding values in the semantic incremental weight map.

15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor (120) of an electronic device (101), cause the electronic device (101) to perform the operations of the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Systems and methods for dynamic range compression in multi-frame processing

    US11388348B2

  • System and method for scene-adaptive denoise scheduling and efficient deghosting

    US20240221130A1