Improved fusion of image frames for high dynamic range content

By receiving image frames of different exposure times and calculating multiple motion maps, combined with time filtering processing, the distortion and noise problems caused by object movement in conventional image sensors in high dynamic range image processing are solved, and image quality and clarity are improved.

CN120266153APending Publication Date: 2025-07-04QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081661.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-05
Filing Date
2023-11-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When conventional image sensors capture image frames, it is difficult to balance the details of the bright and dark parts, resulting in distortion and motion artifacts caused by object movement. Especially in high dynamic range image processing, the difference in noise level leads to difficulty in adjusting the motion threshold.

Method used

By receiving image frames of different exposure times, multiple motion maps are determined and combined to generate high dynamic range image frames, time filtering is applied to correct motion distortion, and errors caused by noise are reduced.

Benefits of technology

Improve image quality, reduce motion artifacts and noise, achieve more accurate motion threshold adjustment, and enhance sharpness and detail retention of high dynamic range images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266153A_ABST
    Figure CN120266153A_ABST
Patent Text Reader

Abstract

This disclosure provides systems, methods, and apparatus for image signal processing detection and correction of motion during HDR fusion. In a first aspect, a method is provided that includes receiving a first image frame, a second image frame, and a third image frame. The second image frame may be captured at a short exposure time, and the third image frame may be captured at a long exposure time. The method may include determining a first motion map based on the first image frame and the third image frame and determining a second motion map based on the second image frame and the third image frame. A third motion map may be determined based on the first motion map and the second motion map, and used to determine a fourth image frame by combining the second image frame and the third image frame. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Patent Application No. 18 / 061,856, filed on December 5, 2022, entitled "IMPROVED FUSION OF IMAGE FRAMES FOR HIGH DYNAMIC RANGE CONTENT", which is hereby incorporated by reference in its entirety. Technical Field

[0003] Aspects of the present disclosure generally relate to image processing, and more particularly to high - dynamic - range (HDR) images. Some features enable and provide improved image processing, including improved techniques for detecting and correcting movement in image frames fused to form high - dynamic - range content. Background Art

[0004] An image - capture device is a device that can capture one or more digital images (either a still image for a photograph or a sequence of images for a video). The capture device can be incorporated into various devices. By way of example, an image - capture device can include a standalone digital camera or digital video camera, a wireless communication device phone equipped with a camera (such as a mobile phone, cellular or satellite radiotelephone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, a video surveillance camera), or other devices having digital imaging or video capabilities.

[0005] When capturing a representation of a scene with a wide color gamut using an image - capture device, dynamic range can be important for image quality. Conventional image sensors have a limited dynamic range, which may be less than the dynamic range of the human eye. Dynamic range can refer to the light range between the bright part of an image and the dark part of the image. Conventional image sensors can increase the exposure time to improve details in the dark part of the image, at the cost of saturating the bright part of the image. Alternatively, conventional image sensors can decrease the exposure time to improve details in the bright part of the image, at the cost of losing details in the dark part of the image. Thus, image - capture devices typically balance conflicting desires by adjusting the exposure time, thereby preserving details in either the bright or dark part of the image. High - dynamic - range (HDR) photography improves photography using these conventional image sensors by combining multiple recorded representations of a scene from an image sensor.

[0006] Movement of an object when capturing an image frame (i.e., an image frame for synthesis into a single still image, an image frame used as part of a video sequence) can introduce various distortions within the image frame. For example, movement of one or more objects within the image frame can blur and / or blend these objects together or may leave motion artifacts within the captured image frame. Summary of the Invention

[0007] Some aspects of the present disclosure are summarized below to provide a basic understanding of the technologies being discussed. This Summary of the Invention is not an exhaustive overview of all the intended features of the present disclosure, and is neither intended to identify key or important elements of all aspects of the present disclosure, nor to delineate the scope of any or all aspects of the present disclosure. The sole purpose of this Summary of the Invention is to present some concepts of one or more aspects of the present disclosure in a generalized form as a prelude to the more detailed embodiments that are presented later.

[0008] In some aspects, HDR processing of received image frames may be performed based on multiple motion maps determined for different sets of image frames. A first motion map may be determined based on a previous reference frame and a current image frame having a long exposure time. A second motion map may be determined based on a current image frame having a short exposure time and the current image frame having a long exposure time. Then, the first motion map and the second motion map may be combined to form a third motion map that excludes motion regions caused by noise within the first motion map and the second motion map. Then, the current image frame may be blended based on the third motion map to generate a fourth image frame (such as an HDR image frame for the current image frame). In certain instances, the improved third motion map may also be used to apply TF processing and generate a fifth image frame to correct for motion between the previous reference image frame and the fourth image frame.

[0009] In some aspects, the techniques described herein relate to an apparatus that includes: a memory that stores processor-readable code; and at least one processor coupled to the memory, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to perform operations including: receiving a first image frame, a second image frame, and a third image frame, where the second image frame is captured at a first exposure time and the third image frame is captured at a second exposure time that is longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0010] In some aspects, the techniques described herein relate to a method that includes: receiving a first image frame, a second image frame, and a third image frame, where the second image frame is captured at a first exposure time and the third image frame is captured at a second exposure time that is longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0011] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations including: receiving a first image frame, a second image frame, and a third image frame, where the second image frame is captured at a first exposure time and the third image frame is captured at a second exposure time that is longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0012] In some aspects, the techniques described herein relate to an image capture device that includes: an image sensor; a memory that stores processor-readable code; and at least one processor coupled to the memory and the image sensor, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to perform: receiving a first image frame, a second image frame, and a third image frame, where the second image frame is captured at a first exposure time and the third image frame is captured at a second exposure time that is longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0013] The image processing methods described herein may be performed by an image capture device and / or on image data captured by one or more image capture devices. An image capture device (a device that can capture one or more digital images, whether a still image photograph or an image sequence of a video) may be incorporated into various devices. By way of example, an image capture device may include a standalone digital camera or digital video camera, a wireless communication device phone equipped with a camera (such as a mobile phone, cellular or satellite radiotelephone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, a video surveillance camera), or other devices having digital imaging or video capabilities.

[0014] The image processing techniques described herein may relate to a digital camera having an image sensor and processing circuitry (e.g., an application specific integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), or a central processing unit (CPU)). An image signal processor (ISP) may include one or more of these processing circuitries and be configured to perform operations to obtain image data for processing according to the image processing techniques described herein and / or involved in the image processing techniques described herein. The ISP may be configured to control the capture of image frames from one or more image sensors and determine one or more image frames from the one or more image sensors to generate a view of a scene in an output image frame. The output image frame may be part of a sequence of image frames forming a video sequence. The video sequence may include other image frames received from the image sensor or other image sensors.

[0015] In an example application, an image signal processor (ISP) may receive instructions to capture a sequence of image frames in response to the loading of software (such as a camera application) to produce a preview display from the image capture device. The image signal processor may be configured to generate a single output image frame stream based on the image frames received from one or more image sensors. The single output image frame stream may include raw image data from the image sensor, merged image data from the image sensor, or corrected image data processed by one or more algorithms within the image signal processor. For example, image frames may be processed by an image post - processing engine (IPE) and / or other image processing circuitry to process the image frames obtained from the image sensor (the image frames may have had some processing of the data performed on them before being output to the image signal processor) so as to perform one or more of tone mapping, portrait illumination, contrast enhancement, gamma correction, etc. The output image frames from the ISP may be stored in a memory and retrieved by an application processor executing the camera application, which may perform further processing on the output image frames to adjust the appearance of the output image frames and reproduce the output image frames on a display for viewing by a user.

[0016] After an output frame representing a scene is determined by an image signal processor and / or an application processor (such as through the image processing techniques described in various embodiments herein), the output image frame can be displayed on a device display as a single still image and / or as part of a video sequence, saved to a storage device as a picture or video sequence, sent over a network, and / or printed to an output medium. For example, an image signal processor (ISP) can be configured to obtain an input frame of image data (e.g., pixel values) from one or more image sensors and, in turn, generate a corresponding output image frame (e.g., a preview display frame, a still image capture, a frame for video, a frame for object tracking, etc.). In other examples, the image signal processor can output the image frame to various output devices and / or camera modules for further processing, such as for 3A parameter synchronization (e.g., autofocus (AF), auto white balance (AWB), and automatic exposure control (AEC)), generating a video file via the output frame, configuring the frame for display, configuring the frame for storage, sending the frame over a network connection, etc. Generally, an image signal processor (ISP) can obtain incoming frames from one or more image sensors, generate an output frame stream, and output the output frame stream to various output destinations.

[0017] In some aspects, the output image frame can be generated by combining aspects of the image correction of the present disclosure with other computational photography techniques such as high dynamic range (HDR) photography or multi-frame noise reduction (MFNR). In the case of HDR photography, a first image frame and a second image frame are captured using different exposure times, different apertures, different lenses, and / or other characteristics that can result in an improved dynamic range of the fused image when combining the two image frames. In some aspects, the method can be performed for MFNR photography, where the first image frame and the second image frame are captured using the same or different exposure times, and the first image frame and the second image frame are fused to generate a corrected first image frame that has reduced noise compared to the captured first image frame.

[0018] In some aspects, the device can include an image signal processor or a processor (e.g., an application processor) that includes specific functionality for camera control and / or processing, such as enabling or disabling a merging module or otherwise controlling aspects of the image correction. The methods and techniques described herein can be performed entirely by the image signal processor or the processor, or various operations can be split between the image signal processor and the processor and, in some aspects, across additional processors.

[0019] The device may include one, two, or more image sensors, such as a first image sensor. When there are multiple image sensors, the configurations of these image sensors may be different. For example, the first image sensor may have a larger field of view (FOV) than the second image sensor, or the first image sensor may have different sensitivity or different dynamic range from the second image sensor. In one example, the first image sensor may be a wide-angle image sensor, and the second image sensor may be a tele image sensor. In another example, the first sensor is configured to obtain an image through a first lens having a first optical axis, and the second sensor is configured to obtain an image through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens may have a first magnification, and the second lens may have a second magnification different from the first magnification. Any of these or other configurations may be part of a lens cluster on a mobile device, such as where multiple image sensors and associated lenses are located at offset positions on the front or rear side of the mobile device. Additional image sensors with larger, smaller, or the same field of view may be included. The image processing techniques described herein may be applied to image frames captured from any of the image sensors in a multi-sensor device.

[0020] In an additional aspect of the present disclosure, a device configured for image processing and / or image capture is disclosed. The device includes components for capturing image frames. The device also includes one or more components for capturing data representative of a scene, such as image sensors (including charge-coupled devices (CCDs), Bayer filter sensors, infrared (IR) detectors, ultraviolet (UV) detectors, complementary metal oxide semiconductor (CMOS) sensors), and time-of-flight detectors. The device may also include one or more components for focusing and / or concentrating light onto one or more image sensors (including simple lenses, compound lenses, spherical lenses, and aspherical lenses). These components may be controlled to capture a first image frame and / or a second image frame input to the image processing techniques described herein.

[0021] For those of ordinary skill in the art, other aspects, features, and specific implementations will become apparent when the following description of specific exemplary aspects is reviewed in conjunction with the accompanying drawings. Although the features may be discussed below with respect to certain aspects and drawings, each aspect may include one or more of the advantageous features discussed herein. In other words, although one or more aspects may be discussed as having certain advantageous features, one or more of such features may also be used according to each aspect. In a similar manner, although the exemplary aspects may be discussed below as device, system, or method aspects, the exemplary aspects may be implemented in various devices, systems, and methods.

[0022] The method can be embedded as computer program code in a computer-readable medium, the computer program code including instructions that cause a processor to perform the steps of the method. In some embodiments, the processor can be part of a mobile device that includes: a first network adapter configured to send data, such as an image or video as recorded data or as streaming data, over a first network connection of a plurality of network connections; and a processor coupled to the first network adapter and a memory. The processor can cause the output image frames described herein to be sent over a wireless communication network, such as a 5G NR communication network.

[0023] The features and technical advantages of examples in accordance with the present disclosure have been outlined above rather broadly in order that the detailed description that follows may be better understood. Additional features and advantages will be described below. The disclosed concepts and specific examples can be readily used as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein, both as to their organization and method of operation, as well as associated advantages, will be better understood when considered in connection with the accompanying drawings. Each of the drawings provided is for the purpose of illustration and description and not as a definition of the limits of the claims.

[0024] While aspects and specific implementations are described herein by way of some examples, those skilled in the art will understand that additional specific implementations and use cases may arise in many different arrangements and scenarios. The innovations described herein can be implemented across many different platform types, devices, systems, shapes, sizes, and packaging arrangements. For example, aspects and / or uses can be implemented via integrated chips and other non-module-component-based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchase devices, medical devices, artificial intelligence (AI)-enabled devices, etc.). Although some examples may or may not specifically refer to use cases or applications, a wide variety of applicability of the described innovations may occur. The scope of specific implementations can range from chip-level or modular components to non-modular, non-chip-level implementations, and further to the scope of aggregated, distributed, or original equipment manufacturer (OEM) devices or systems that incorporate one or more aspects of the described innovations. In some practical environments, devices that incorporate the described aspects and features may also necessarily include additional components and features for implementing and practicing the claimed and described aspects. For example, the transmission and reception of wireless signals necessarily includes multiple components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders / summers, etc.). The innovations described herein are intended to be practiced in a variety of devices, chip-level components, systems, distributed arrangements, end-user devices, etc. having different sizes, shapes, and configurations. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] A further understanding of the nature and advantages of the present disclosure can be realized by reference to the following drawings. In the drawings, like components or features may have the same reference numeral. Additionally, various components of the same type can be distinguished by adding a dash and a second numeral used to differentiate between like components after the reference numeral. If only the first reference numeral is used in the specification, the description applies to any one of the like components having the same first reference numeral, regardless of the second reference numeral.

[0026] Figure 1 A block diagram illustrating an example device for performing image capture from one or more image sensors is shown.

[0027] Figure 2 is a block diagram illustrating an example data flow path for image data processing in an image capture device in accordance with one or more embodiments of the present disclosure.

[0028] Figure 3A Depicts a system for correcting motion in HDR image frame fusion in accordance with some embodiments of the present disclosure.

[0029] Figure 3B depicts a frame timing sequence according to an exemplary embodiment of the present disclosure.

[0030] Figure 4 shows a flowchart of an example method for processing image data to correct for movement in HDR image frame fusion according to some embodiments of the present disclosure.

[0031] Figure 5 is a block diagram illustrating an example processor configuration for image data processing in an image capture device according to one or more embodiments of the present disclosure.

[0032] Like reference numerals and names in different figures represent like elements. Detailed Description

[0033] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to limit the scope of the present disclosure. On the contrary, the detailed description includes specific details for providing a thorough understanding of the subject matter of the present invention. It will be apparent to those skilled in the art that these specific details are not required in every instance and that in some instances, for the sake of clarity, well-known structures and components are shown in block diagram form.

[0034] In some aspects of the present disclosure, temporal filtering (TF) may be applied or not applied to portions of an image frame based on motion detected within the image frame. Temporal filtering may be applied to remove distortions within an image frame caused by movement of an object when capturing image frames for synthesis into a single still image or as part of a video sequence. For example, temporal filtering applied to a region of an image frame may reduce distortions caused by blurring and / or movement of one or more objects blended together within the image frame. In some aspects, the temporal filtering applied to correct an image frame may be motion-compensated temporal filtering (MCTF). MCTF reduces noise and / or motion artifacts in a video by filtering motion regions based on global and / or local motion of the current frame relative to the previous frame.

[0035] HDR processing can combine multiple image frames with different exposure times to generate an HDR image frame. However, movement of an object between or during these image frames (or previous HDR image frames) can cause distortions similar to those discussed above with respect to TF processing. Motion maps for these image frames are often inaccurate due to different noise levels in the image frames being blended. Specifically, an image frame with a shorter exposure time can have a higher noise level than an image frame with a longer exposure time. In such instances, the different noise levels can cause differences between the image frames that are interpreted as motion between the image frames. This can cause additional regions within the short-exposure image frames to be unnecessarily blended within the final HDR image frame, thereby increasing the noise level and degrading the image quality. Reducing such errors requires adjusting the motion threshold of the HDR fusion process, which can be tricky in practice. Specifically, if the motion threshold is too high, ghosting artifacts may be introduced because regions with actual motion are excluded. Additionally, the ideal motion threshold is different for different shooting conditions (such as at night, during the day, indoors, outdoors). Therefore, it may not be feasible to use the motion threshold alone to properly identify motion regions within individual image frames.

[0036] One solution to this problem can be to calculate multiple motion maps based on different sets of image frames. Specifically, a first motion map can be determined based on a previous reference frame and a current image frame with a long exposure time. A second motion map can be determined based on a current image frame with a short exposure time and a current image frame with a long exposure time. The first motion map and the second motion map can then be combined such that motion regions caused by noise within each motion map are eliminated. For example, the first motion map can be subtracted from the second motion map to generate a third motion map. Then, the current image frame can be blended based on the third motion map to generate a fourth image frame (such as an HDR image frame for the current image frame). In some instances, the improved third motion map can also be used to apply TF processing and generate a fifth image frame to correct for motion between the previous reference image frame and the fourth image frame.

[0037] The disadvantages mentioned here are merely representative and are included to emphasize the problems that the inventors have identified and sought to improve with respect to existing devices. Aspects of the devices described below can address some or all of these disadvantages as well as other disadvantages known in the art. Aspects of the improved devices described herein can present other benefits different from those described above and can be used in other applications different from those described above.

[0038] Example devices for capturing image frames using one or more image sensors, such as smartphones, may include a configuration of one, two, three, four, or more cameras on the rear side of the device (e.g., the side opposite the main user display) and / or the front side (e.g., the side same as the main user display). These devices may include one or more image signal processors (ISPs), computer vision processors (CVPs) (e.g., AI engines), or other suitable circuitry for processing the images captured by the image sensors. The one or more image signal processors (ISPs) may store the output image frames in memory and / or otherwise provide the output image frames to the processing circuitry (such as via a bus). The processing circuitry may perform further processing, such as encoding, storing, transmitting, or other manipulation of the output image frames.

[0039] As used herein, an image sensor may refer to the image sensor itself and any particular other components coupled to the image sensor for generating image frames for processing by an image signal processor or other logic circuitry or for storage in memory (whether a short-term buffer or long-term non-volatile memory). For example, an image sensor may include other components of a camera, including a shutter, buffer, or other readout circuitry for accessing the individual pixels of the image sensor. An image sensor may also refer to an analog front end or other circuitry for converting an analog signal to a digital representation of an image frame, which is provided to digital circuitry coupled to the image sensor.

[0040] In the description of the embodiments herein, numerous specific details (such as examples of specific components, circuits, and processes) are set forth to provide a thorough understanding of the present disclosure. As used herein, the term "coupled" means directly connected or connected through one or more intermediate components or circuits. Additionally, in the following description and for purposes of explanation, specific terms are set forth to provide a thorough understanding of the present disclosure. However, those skilled in the art will appreciate that implementing the teachings disclosed herein may not require these specific details. In other instances, well-known circuits and devices are shown in block diagram form to avoid obscuring the teachings of the present disclosure.

[0041] Certain portions of the detailed description that follows are presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. In the present disclosure, procedures, logic blocks, processes, etc. are conceived as self-consistent sequences of steps or instructions leading to a desired result. These steps are those requiring physical manipulation of physical quantities. Although not necessarily, typically, these physical quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system.

[0042] In the figures, a single box may be described as performing one or more functions. The one or more functions performed by the box may be performed in a single component or across multiple components, and / or may be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are described below in terms of their functionality. Implementing such functionality as hardware or software depends on the particular application and the design constraints imposed on the overall system. A person skilled in the art may implement the described functionality in different ways for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of the present disclosure. Moreover, an example device may include components other than those shown, including well-known components such as processors, memories, and the like.

[0043] Aspects of the present disclosure are applicable to any electronic device that includes, is coupled to, or otherwise processes data from one, two, or more image sensors capable of capturing image frames (or "frames"). The terms "output image frame" and "corrected image frame" may refer to an image frame that has been processed by any of the techniques discussed herein. Additionally, aspects of the present disclosure may be implemented in image sensors or devices coupled to the image sensors that have the same or different capabilities and characteristics, such as resolution, shutter speed, sensor type, and the like. Further, aspects of the present disclosure may be implemented in a device for processing image frames, whether or not the device includes or is coupled to an image sensor, such as a processing device that can retrieve stored images for processing, including processing devices present in a cloud computing system.

[0044] Unless otherwise specifically stated, it should be understood that throughout this application, discussions using terms such as "access," "receive," "transmit," "use," "select," "determine," "normalize," "multiply," "average," "monitor," "compare," "apply," "update," "measure," "derive," "set," "generate," etc., refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the registers, memories, or other such information storage, transmission, or display devices of the computer system.

[0045] The terms "device" and "apparatus" are not limited to one or a specific number of physical objects (such as a smart phone, a camera controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that can implement at least some parts of the present disclosure. Although the description and examples herein use the term "device" to describe various aspects of the present disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. As used herein, an apparatus can include a device or a part of a device for performing the described operations.

[0046] Some components in a device or apparatus described as "components for accessing", "components for receiving", "components for transmitting", "components for using", "components for selecting", "components for determining", "components for normalizing", "components for multiplying", or other similarly named terms referring to one or more operations on data (such as image data) can refer to processing circuitry (e.g., application specific integrated circuit (ASIC), digital signal processor (DSP), graphics processing unit (GPU), central processing unit (CPU)) configured to perform the functions by a combination of hardware, software, or hardware configured by software.

[0047] Figure 1 A block diagram of an example device 100 for performing image capture from one or more image sensors is shown. Device 100 can include or otherwise be coupled to an image signal processor 112 for processing image frames from one or more image sensors (such as a first image sensor 101, a second image sensor 102, and a depth sensor 140). In some specific embodiments, device 100 further includes or is coupled to a processor 104 and a memory 106 storing instructions 108. Device 100 can also include or be coupled to a display 114 and input / output (I / O) components 116. The I / O components 116 can be used to interact with a user, such as a touch screen interface and / or physical buttons.

[0048] The I / O components 116 can also include a network interface for communicating with other devices including a wide area network (WAN) adapter 152, a local area network (LAN) adapter 153, and / or a personal area network (PAN) adapter 154. An example WAN adapter is a 4G LTE or 5G NR wireless network adapter. An example LAN adapter 153 is an IEEE 802.11 WiFi wireless network adapter. An example PAN adapter 154 is a Bluetooth wireless network adapter. Each of the adapters 152, 153, and / or 154 can be coupled to an antenna, which includes multiple antennas configured for main set reception and diversity reception and / or configured for receiving a specific frequency band.

[0049] Device 100 may also include or be coupled to a power source 118 for device 100, such as a battery or a component that couples device 100 to an energy source. Device 100 may also include or be coupled to Figure 1 additional feature parts or components not shown in Figure 1 . In one example, a wireless interface that may include multiple transceivers and a baseband processor may be coupled to or included in WAN adapter 152 for a wireless communication device. In another example, an analog front end (AFE) for converting analog image frame data to digital image frame data may be coupled between image sensors 101 and 102 and image signal processor 112.

[0050] The device may include or be coupled to a sensor hub 150 that is used to interface with sensors to receive data about the movement of device 100, data about the environment around device 100, and / or other non-camera sensor data. One example of a non-camera sensor is a gyroscope, i.e., a device configured to measure rotation, orientation, and / or angular velocity to generate motion data. Another example of a non-camera sensor is an accelerometer, i.e., a device configured to measure acceleration, which can also be used to determine velocity and distance traveled by appropriately integrating the measured acceleration, and one or more of acceleration, velocity, and / or distance may be included in the generated motion data. In some aspects, the gyroscope in an electronic image stabilization system (EIS) may be coupled to the sensor hub or directly coupled to image signal processor 112. In another example, the non-camera sensor may be a global positioning system (GPS) receiver.

[0051] Image signal processor 112 may receive image data such as for forming an image frame. In one implementation, a local bus connection couples image signal processor 112 to image sensors 101 and 102 of first camera 103 and second camera 105, respectively. In another implementation, a wire interface couples image signal processor 112 to an external image sensor. In yet another implementation, a wireless interface couples image signal processor 112 to image sensors 101, 102.

[0052] First camera 103 may include a first image sensor 101 and a corresponding first lens 131. The second camera may include a second image sensor 102 and a corresponding second lens 132. Each of lenses 131 and 132 may be controlled by an associated autofocus (AF) algorithm 133 executed in ISP 112, and the autofocus (AF) algorithm adjusts lenses 131 and 132 to focus on a specific focal plane at a certain scene depth from image sensors 101 and 102. AF algorithm 133 may be assisted by depth sensor 140.

[0053] The first image sensor 101 and the second image sensor 102 are configured to capture one or more image frames. The lenses 131 and 132 focus light onto the image sensors 101 and 102 respectively through one or more apertures for receiving light, one or more shutters for blocking light when outside the exposure window, one or more color filter arrays (CFAs) for filtering light outside a specific frequency range, one or more analog front ends for converting analog measurements to digital information, and / or other suitable components for imaging. The first lens 131 and the second lens 132 may have different fields of view to capture different representations of a scene. For example, the first lens 131 may be an ultra-wide (UW) lens, and the second lens 132 may be a wide (W) lens. The plurality of image sensors may include a combination of ultra-wide (high field of view (FOV)) sensors, wide sensors, tele sensors, and ultra-tele (low FOV) sensors.

[0054] That is, each image sensor can be configured, through hardware configuration and / or software settings, to obtain different but overlapping fields of view. In one configuration, the image sensor is configured with different lenses having different magnifications, which results in different fields of view. The sensors can be configured such that the UW sensor has a larger FOV than the W sensor, the W sensor has a larger FOV than the T sensor, and the T sensor has a larger FOV than the UT sensor. For example, a sensor configured for a wide FOV can capture a field of view in the range of 64 degrees to 84 degrees, a sensor configured for an ultra-side FOV can capture a field of view in the range of 100 degrees to 140 degrees, a sensor configured for a tele FOV can capture a field of view in the range of 10 degrees to 30 degrees, and a sensor configured for an ultra-tele FOV can capture a field of view in the range of 1 degree to 8 degrees.

[0055] The camera 103 may have a variable aperture (VA) camera, where the aperture can be controlled to a specific size. Example aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. A larger aperture value corresponds to a smaller aperture size, and a smaller aperture value corresponds to a larger aperture size. The camera 103 may have different characteristics based on the current aperture size, such as different depths of field (DOF) at different aperture sizes.

[0056] The image signal processor 112 processes the image frames captured by the image sensors 101 and 102. Although Figure 1Illustrated is that device 100 includes two image sensors 101 and 102 coupled to image signal processor 112, but any number (e.g., one, two, three, four, five, six, etc.) of image sensors may be coupled to image signal processor 112. In some aspects, a depth sensor such as depth sensor 140 may be coupled to image signal processor 112 and process the output from the depth sensor in a manner similar to that of image sensors 101 and 102. Example depth sensors include active sensors, including one or more of indirect time of flight (iToF), direct time of flight (dToF), light detection and ranging (LiDAR), mmWave, radio detection and ranging (radar), and / or hybrid depth sensors (such as structured light). In embodiments without depth sensor 140, similar information regarding the depth or depth map of an object can be generated passively from the disparity between two image sensors (e.g., using disparity depth measurement or stereo depth measurement), phase detection autofocus (PDAF) sensors, etc. Additionally, there may be any number of additional image sensors or image signal processors for device 100.

[0057] In some embodiments, image signal processor 112 may execute instructions from a memory, such as instructions 108 from memory 106, instructions stored in a separate memory coupled to or included in image signal processor 112, or instructions provided by processor 104. Additionally or alternatively, image signal processor 112 may include specific hardware (such as one or more integrated circuits (ICs)) configured to perform one or more operations described in this disclosure. For example, image signal processor 112 may include one or more image front ends (IFE) 135, one or more image post - processing engines 136 (IPE), one or more automatic exposure compensation (AEC) 134 engines, and / or one or more video analysis engines (EVA). AF 133, AEC 134, IFE 135, IPE 136, and EVA 137 may each include dedicated circuitry, may be embodied as software code executed by ISP 112, and / or a combination of hardware and software code executed on ISP 112.

[0058] In some embodiments, the memory 106 may include a non-transitory or non-volatile computer-readable medium storing computer-executable instructions 108 to perform all or a portion of one or more operations described in the present disclosure. In some embodiments, the instructions 108 include a camera application (or other suitable application) for generating images or videos to be executed by the device 100. The instructions 108 may also include other applications or programs executed by the device 100, such as an operating system and specific applications other than those for image or video generation. A camera application, such as executed by the processor 104, may cause the device 100 to generate images using the image sensors 101 and 102 and the image signal processor 112. The memory 106 may also be accessed by the image signal processor 112 to store the processed frames or may be accessed by the processor 104 to obtain the processed frames. In some embodiments, the device 100 does not include the memory 106. For example, the device 100 may be a circuit including the image signal processor 112, and the memory may be external to the device 100. The device 100 may be coupled to an external memory and configured to access the memory to write output frames for display or long-term storage. In some embodiments, the device 100 is a system-on-chip (SoC) that combines the image signal processor 112, the processor 104, the sensor hub 150, the memory 106, and the input / output component 116 into a single package.

[0059] In some embodiments, at least one of the image signal processor 112 or the processor 104 executes instructions to perform the various operations described herein, including improved motion and noise correction operations for HDR fusion. For example, the execution of the instructions may direct the image signal processor 112 to start or end capturing an image frame or a sequence of image frames, where the capture includes motion and noise correction operations for HDR fusion as described in the embodiments herein. In some embodiments, the processor 104 may include one or more general-purpose processor cores 104A capable of executing a script or instructions of one or more software programs (such as the instructions 108 stored in the memory 106). For example, the processor 104 may include one or more application processors configured to execute a camera application (or other suitable application for generating images or videos) stored in the memory 106.

[0060] When executing a camera application, the processor 104 may be configured to instruct the image signal processor 112 to perform one or more operations with reference to the image sensors 101 or 102. For example, a camera application executed on the processor 104 may receive a user command to start a video preview display, and upon receiving the user command, capture and process a video including a sequence of image frames from one or more of the image sensors 101 or 102 via the image signal processor 112. Image processing for generating an "output" or "corrected" image frame, such as according to the techniques described herein, may be applied to one or more of the image frames in the sequence. Execution of instructions 108 by the processor 104 external to the camera application may also cause the device 100 to perform any number of functions or operations. In some embodiments, the processor 104 may include an IC or other hardware (e.g., an artificial intelligence (AI) engine 124 or other coprocessor) to offload certain tasks from the core 104A. The AI engine 124 may be used to offload tasks related to, for example, face detection and / or object recognition. In some other embodiments, the device 100 does not include the processor 104, such as when all of the described functionality is configured in the image signal processor 112.

[0061] In some embodiments, the display 114 may include one or more suitable displays or screens that allow user interaction and / or present items to the user, such as a preview of image frames captured by the image sensors 101 and 102. In some embodiments, the display 114 is a touch-sensitive display. The I / O component 116 may be or include any suitable mechanism, interface, or device to receive input (such as commands) from the user and provide output to the user via the display 114. For example, the I / O component 116 may include (but is not limited to) a graphical user interface (GUI), a keyboard, a mouse, a microphone, a speaker, a squeezable bezel, one or more buttons (such as a power button), a slider, a switch, etc.

[0062] Although shown as being coupled to each other via the processor 104, components such as the processor 104, the memory 106, the image signal processor 112, the display 114, and the I / O component 116 may be coupled to each other in various other arrangements, such as via one or more local buses, which are not shown for simplicity. Although the image signal processor 112 is illustrated as being separate from the processor 104, the image signal processor 112 may be a core of the processor 104, which is an application processor unit (APU), included in a system-on-chip (SoC), or otherwise included in the processor 104. Although reference is made herein to the device 100 to perform aspects of the present disclosure, some device components may not be Figure 1shown to prevent obscuring aspects of the present disclosure. Additionally, other components, the number of components, or combinations of components may be included in a suitable device for performing aspects of the present disclosure. Accordingly, the present disclosure is not limited to a particular device or component configuration, including device 100.

[0063] Operable Figure 1 exemplary image capture device is to obtain improved images through improved motion detection and correction while combining multiple image frames to form an HDR image or image frame. Figure 2 shown and described below is an example method of operating one or more cameras, such as camera 103.

[0064] Figure 2 is a block diagram illustrating an example data flow path for image data processing in an image capture device according to one or more embodiments of the present disclosure. The processor 104 of the system 200 may communicate with an image signal processor (ISP) 112 via a bi - directional bus and / or separate control and data lines. The processor 104 may control the camera 103 via the camera control 210, such as to configure the camera 103 via a driver executed on the processor 104. The camera control 210 may be managed by a camera application 204 executed on the processor 104, which provides user - accessible settings such that the user may specify individual camera settings or select a profile with corresponding camera settings. The camera control 210 communicates with the camera 103 to configure the camera 103 according to commands received from the camera application 204. The camera application 204 may be, for example, a photography application, a document scanning application, a messaging application, or other application that processes image data obtained from the camera 103.

[0065] The camera configuration may be parameters that specify, for example, frame rate, image resolution, readout duration, exposure level, aspect ratio, aperture size, etc. The camera 103 may obtain image data based on the camera configuration. For example, the processor 104 may execute the camera application 204 to instruct the camera 103 via the camera control 210 to set a first image sensor configuration of the camera 103, obtain first image data from the camera 103 operating in the first camera configuration, instruct the camera 103 to set a second camera configuration of the camera 103, and obtain second image data from the camera 103 operating in the second camera configuration.

[0066] In some embodiments in which camera 103 is a variable aperture (VA) camera system, processor 104 may execute camera application 204 to instruct camera 103 to be configured to a first aperture size, obtain first image data from camera 103, instruct camera 103 to be configured to a second aperture size, and obtain second image data from camera 103. Reconfiguration of the aperture and the obtaining of the first and second image data may occur when there is little or no change in the scene that can be captured at the first aperture size and at the second aperture size. Example aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. Larger aperture values correspond to smaller aperture sizes, and smaller aperture values correspond to larger aperture sizes. That is, f / 2.0 is a larger aperture size than f / 8.0.

[0067] The image data received from camera 103 may be processed in one or more blocks of ISP 112 to form image frame 230 stored in memory 106 and / or provided to processor 104. Processor 104 may further process the image data to apply effects to image frame 230. The effects may include Bokeh, illumination, color cast, and / or high dynamic range (HDR) merging. In some embodiments, the functionality may be embedded in different components, such as ISP 112, DSP, ASIC, or other custom logic circuitry for performing additional image processing.

[0068] For example, Figure 3A System 300 for correcting motion in HDR image frame fusion according to some embodiments of the present disclosure is depicted. System 300 includes first image frame 302, second image frame 304, third image frame 306, ISP 112, and fourth image frame 308. ISP 112 includes first motion map 310, second motion map 312, third motion map 314, HDR processing 316, and temporal filtering (TF) processing 318. HDR processing 316 includes fusion map 320, and TF processing 318 includes transform matrix 322.

[0069] The ISP 112 can be configured to receive a first image frame 302, a second image frame 304, and a third image frame 306. Specifically, the ISP 112 can be configured to generate an output image frame, such as a fourth image frame 308, based on the image frames 302, 304, 306. In certain embodiments, the first image frame 302 can be a reference frame, the second image frame 304 can be captured with a short exposure time, the third image frame 306 can be captured with a long exposure time, and combinations thereof. In certain embodiments, the short exposure time can be less than 2 ms and the long exposure time can be greater than or equal to 2 ms. In certain embodiments, the first image frame 302 can be captured before the second image frame 304 and the third image frame 306. For example, the second image frame 304 and the third image frame 306 can be current image frames (such as current image frames for the image processing pipeline of the image signal processor 112). The first image frame 302 can be a previous reference frame (such as a previous reference frame for the image processing pipeline of the ISP 112). In certain embodiments, the second image frame 304 can be captured before the third image frame 306, or vice versa. In certain embodiments, the first image frame 302 can be a reference frame for a previous image processing operation. For example, the first image frame 302 can be a reference frame for TF processing (such as TF processing 318). In certain embodiments, the second image frame 304 can be subjected to spatial denoising processing before further processing. In certain embodiments, the third image frame 306 can be processed to correct for a global movement (such as the movement of an image capture device) between the capture of the first image frame 302 and at least one of the second image frame 304 and the third image frame 306. In certain embodiments, the image frames 302, 304, 306 can be captured by the same image sensor (such as the image sensor 101). In additional or alternative embodiments, the image frames 302, 304, 306 can be captured by different image sensors (such as simultaneously or successively by two or more image sensors).

[0070] The ISP 112 can be configured to determine a first motion map 310 based on the first image frame 302 and the third image frame 306. The ISP 112 can also be configured to determine a second motion map 312 based on the second image frame 304 and the third image frame 306. Motion maps 310, 312 can be generated to indicate the motion within one image frame relative to another image frame. For example, the first motion map 310 can be generated to indicate the motion within the third image frame 306 relative to the first image frame 302. As another example, the motion map 312 can be generated to indicate the motion within the third image frame 306 relative to the second image frame 304. In some embodiments, the motion maps 310, 312 can be calculated based on the differences between the image frames 302, 304, 306. For example, determining the first motion map 310 can include determining the difference between the third image frame 306 and the first image frame 302. Similarly, determining the second motion map 312 can include determining the difference between the second image frame 304 and the first image frame 302. Specifically, the local motion estimation can be determined by comparing the image frames 302, 304, 306. For example, the local motion estimation can be determined by comparing the positions of one or more objects within each of the image frames 302, 304, 306. The local motion estimation can be calculated as the difference between the image frames 302, 304, 306. As another example, the local motion estimation can be calculated based on texture processing using Harris corner detection and related techniques. Then the local motion estimation and / or the global motion estimation can be combined to generate the motion maps 310, 312.

[0071] In additional or alternative embodiments, the motion maps 310, 312 can be calculated based on sensor data (such as motion sensor data, depth sensor data, or a combination thereof) from a device (such as device 100) that captures the image frames 302, 304, 306. For example, the ISP 112 can be configured to calculate one or more global motion estimations that reflect the movement of the device and local motion estimations that reflect the movement of one or more objects depicted within the image frames 302, 304, 306. In some instances, the global motion estimation can be calculated based on sensor data (such as gyroscope or accelerometer data indicating the movement of the image capture device). The global motion estimation can include one or both of the magnitude and direction of the movement of the image capture device.

[0072] In some specific implementations, the motion maps 310, 312 can be implemented separately from the corresponding image frames (e.g., implemented as a separate data structure corresponding to the corresponding image frames). In additional or alternative specific implementations, the motion maps 310, 312 can be implemented as part of the corresponding image frames (e.g., implemented as a data layer or metadata layer of the corresponding image frames). For example, the first motion map 310 can be implemented as a metadata layer for the third image frame 306, and the motion map 312 can be implemented as a metadata layer for the second image frame 304. Additionally, the content of the motion maps 310, 312 can correspond to specific portions of the corresponding image frames. For example, each pixel of the corresponding image frame can have a corresponding entry in the motion maps 310, 312. As another example, each entry in the motion maps 310, 312 can correspond to multiple pixels (e.g., 4 pixels, 9 pixels, 16 pixels, or more). The content of the motion maps 310, 312 can indicate the movement within the corresponding portions of the corresponding image frames. For example, the entries in the motion maps 310, 312 can indicate the magnitude of the movement of the objects or features depicted within the corresponding portions of the corresponding image frames (e.g., relative to the first image frame 302). As a specific example, Figure 3A includes an exemplary motion map 324, which can be calculated as the motion map 312. The exemplary motion map 324 identifies regions 328, 330 where motion may occur (such as in the third image frame relative to the first image frame). In various specific implementations, the magnitude of the movement can be indicated as one or more of the number of pixels moved, the distance moved, the speed of movement, etc. In some specific implementations, the entry can also indicate the direction of the movement.

[0073] In one specific implementation, the motion maps 310, 312 can be calculated by the ISP 112 by comparing the corresponding image frames to generate an estimate of the motion vectors between the image frames 302, 304, 306. The motion vectors and the sensor data can then be analyzed together to determine the alignment of the image capture device (e.g., to separate the global motion of the image capture device from the local motion of the objects depicted within the image frames 302, 304, 306). The ISP 112 can then perform a matching process based on the alignment and the image frames 302, 304, 306 to generate the motion maps 310, 312. In some specific implementations, the matching process can be performed as a semi-global matching (SGM) process.

[0074] The ISP 112 can be configured to determine a third motion map 314 based on a first motion map 310 and a second motion map 312. In some embodiments, the third motion map 314 can be determined to remove a high-noise region that is incorrectly identified as a movement within the second motion map 312. For example, as described above, the difference in noise values between an image frame 304 having a short exposure time and an image frame 306 having a longer exposure time can result in the incorrect detection of one or more motion regions (such as where there is no motion). As a specific example, the exemplary motion map 324 includes a region 328 that is identified as containing motion but does not contain any actual movement of the objects depicted in the image frames 304, 306. The third motion map 314 can be determined to remove the region 328 from the motion map 324 and can thus include only a region 332 corresponding to the correct region 330 within the exemplary motion map 324. In some embodiments, determining the third motion map 314 can include combining the first motion map 310 and the second motion map 312 (such as by subtracting the first motion map 310 from the second motion map 312). In some embodiments, the difference between corresponding values of the first motion map 310 and the second motion map 312 can be stored as a movement value within the third motion map 314.

[0075] In some embodiments, correctly calculating the third motion map 314 may require that the image frame 304 having a short exposure time be captured before the image frame 306 having a longer exposure time. For example, Figure 3BDepicts a frame timing sequence 350 according to an exemplary embodiment of the present disclosure. The frame timing sequence 350 includes frame timings 352, 354, 356, 358. The durations of frame timings 352, 356 are shorter and may represent the exposure times of image frames captured with short exposure times. The durations of frame timings 354, 358 are longer and may represent the exposure times of image frames captured with longer exposure times. Specifically, the durations of the shorter frame timings 352, 356 (such as the time difference between T1 and T2 of frame timing 352 and the time difference between T4 and T5 of frame timing 356) may be less than 2 ms (such as 1 ms). The durations of the longer frame timings 354, 358 (such as the time difference between T2 and T3 of frame timing 354 and the time difference between T5 and T6 of frame timing 358) may be longer than 2 ms (such as 5 ms). In some specific implementations, the time difference between sets of image frames (such as the time difference between T3 and T4) may be greater than 5 ms (such as 33 ms) and may vary depending on the frame rate of the captured video. In some specific implementations, image frames with long exposure times may be captured simultaneously with image frames with short exposure times. For example, data for image frames with long exposure times may be captured during the frame times 352, 356 of image frames with short exposure. As a specific example, the frame times 352, 356 may be 1 ms long and the frame times 354, 358 may be 4 ms long, thus allowing both a 1 ms exposure time for image frames with short exposure and a 5 ms exposure time for image frames with long exposure.

[0076] Note that the frame timing sequence 350 includes the frame times 352, 356 of the short exposure image frames before the frame times 354, 358 of the long exposure image frames. This can be different from the conventional frame timing of an image sensor with HDR capabilities, which typically may capture the long exposure image frame before the short exposure image frame. Such frame timing may be necessary for correctly determining the third motion map 314 such that the third motion map 314 removes the noise regions from the motion map 312. For example, the third image frame 306 may typically have the lowest noise among the image frames 302, 304, 306 (such as when the first image frame 302 is a previous frame captured with a short exposure time). In such instances, generating the two motion maps 310, 312 based on the third image frame 306 can ensure that the two motion maps 310, 312 have a similar noise level, but the moving regions caused by the noise in the motion maps 310, 312 are not relevant. Additionally, the time difference between the time of capturing the first image frame 302 and the time of capturing the third image frame 306 (such as the time difference from T1 - T6) can ensure that the local motion of the objects within the image frames 302, 304, 306 is captured in the two motion maps 310, 312. Specifically, the movement that occurs between T4 - T6 is captured in both the first motion map 310 (determined based on the image frames 302, 306 captured from T1 - T6) and the second motion map 312 (determined based on the image frames 304, 306 captured from T4 - T6). This can help ensure that when the first motion map 310 is subtracted from the second motion map 312, the motion is conservatively excluded from the third motion map 314, thereby reducing the likelihood that the desired regions 330, 332 are removed from the third motion map 314, and can help reduce motion artifacts in the regions of the image frames 304, 306 that are highlighted and saturated due to increased motion coverage.

[0077] The ISP 112 can be configured to determine a fourth image frame 308 by combining a second image frame 304 and a third image frame 306. For example, the ISP 112 can combine the second image frame 304 and the third image frame 306 according to a third motion map 314. In some specific implementations, the image frames 304, 306 can be combined to generate an HDR image frame. For example, the image frames 304, 306 can be combined according to an HDR process 316. The HDR process 316 can generate a fused map 320 based on the third motion map 314. The fused map 320 can include different weights assigned when blending the second image frame 304 with the third image frame 306 to generate an HDR image frame. The fused map 320 can be generated based on various aspects (such as brightness values, noise values, etc.) of the image frames 304, 306 according to the HDR process 316. The HDR process 316 can further generate or update the fused map 320 based on the third motion map 314. Specifically, the blending weight of the second image frame 304 can be increased within a region 332 indicating motion in the third motion map 314 and can be decreased in a portion of the third motion map 314 that does not indicate motion. Then, the second image frame 304 and the third image frame 306 can be blended by applying the fused map 320 to generate the fourth image frame 308.

[0078] The ISP 112 can be further configured to determine a fifth image frame by applying the TF processing 318 to correct the motion within the fourth image frame 308 based on the third motion map 314. In such instances, the fifth image frame can be used as the output image frame from the ISP 112. In some specific implementations, the fifth image frame can be generated according to the TF processing 318 (such as MCTF processing), which can receive the first image frame 302, the fourth image frame 308, the third motion map 314, or a combination thereof. In such instances, the TF processing 318 can be used to generate a transformation matrix 322 based on the third motion map 314. In some specific implementations, the ISP 112 can then apply the transformation matrix 322 to the first image frame 302, the fourth image frame 308, or a combination thereof to generate the fifth image frame. Specifically, the transformation matrix 322 can be generated to correct the distortion or other errors caused by the movement within the fourth image frame 308 relative to the first image frame 302. For example, the transformation matrix 322 can be generated to indicate the image transformation that should be applied to the first image frame 302 to correct the motion distortion within the fourth image frame 308. The transformation matrix 322 can include an image transformation based on the individual pixels within the fourth image frame 308 and / or one or more adjacent pixels within the image frame 308. Additionally or alternatively, according to the TF processing 318, the transformation matrix 322 can include a transformation based on other image frames (e.g., the image frame 302 captured before the image frames 304, 306 and / or the image frames captured after the image frames 304, 306). In various specific implementations, the transformation matrix 322 can include the same or similar transformation for each pixel and / or part of the image frames 302, 308. In additional or alternative specific implementations, the transformation matrix 322 can indicate different transformations for different pixels and / or different parts of the image frames 302, 308.

[0079] In some specific implementations, all or part of the above functionality can be implemented as one or more hardware blocks. For example, the ISP 112 can include an HDR hardware block, a TF processing hardware block, and combinations thereof, which can be configured to perform any of the above functionality. In further specific implementations, all or part of the above functionality can be implemented as software configured to execute on a general-purpose processor (such as the processor 104).

[0080] The systems 200, 300 can be configured to perform the operations described in reference Figure 5 to determine the output image frame 230. Figure 4 A flowchart of an example method 400 for processing image data to correct movement in HDR image frame fusion according to some embodiments of the present disclosure is shown. Figure 5 The capture in can obtain an improved digital representation of the scene, thereby producing a photo or video with higher image quality (IQ).

[0081] Method 400 includes receiving a first image frame, a second image frame, and a third image frame (block 402). For example, ISP 112 may receive the first image frame 302, the second image frame 304, and the third image frame 306. As described above, the first image frame 302 may be a reference frame, the second image frame 304 may be captured with a short exposure time, and the third image frame 306 may be captured with a long exposure time. The image frames 302, 304, 306 may be received at ISP 112, processed by the image front end (IFE) and / or the image post-processing engine (IPE) of ISP 112, and stored in a memory. In some embodiments, the capture of the image frames 302, 304, 306 may be initiated by a camera application executing on processor 104, which causes camera control 210 to activate camera 103 to capture the image frames 302, 304, 306 and causes the image frames 302, 304, 306 to be provided to a processor, such as processor 104 or ISP 112.

[0082] Method 400 includes determining a first motion map based on the first image frame and the third image frame (block 404). For example, ISP 112 may determine the first motion map 310 based on the first image frame 302 and the third image frame 306. Method 400 also includes determining a second motion map 312 based on the second image frame 304 and the third image frame 306 (block 406). For example, ISP 112 may determine the second motion map 312 based on the second image frame 304 and the third image frame 306. In certain implementations, the motion maps 310, 312 may be calculated based on the differences between the image frames 302, 304, 306. For example, determining the first motion map 310 may include determining the difference between the third image frame 306 and the first image frame 302. Similarly, determining the second motion map 312 may include determining the difference between the second image frame 304 and the first image frame 302. Specifically, local motion estimates may be determined by comparing the image frames 302, 304, 306. Global motion estimates may also be determined, as described above. Then the local motion estimates, the global motion estimates, and combinations thereof may be combined to generate the motion maps 310, 312.

[0083] Method 400 includes determining a third motion map based on the first motion map and the second motion map (block 408). For example, ISP 112 may determine the third motion map 314 based on the first motion map 310 and the second motion map 312. In certain implementations, the third motion map 314 may be determined to remove high-noise regions that are misidentified as moving within the second motion map 312. For example, an image frame with a short exposure time may be captured before an image frame with a long exposure time. In such instances, determining the third motion map 314 may include subtracting the first motion map 310 from the second motion map 312.

[0084] Method 400 includes determining a fourth image frame (block 410) by combining a second image frame and a third image frame according to a third motion map. For example, ISP 112 may determine fourth image frame 308 by combining second image frame 304 and third image frame 306 according to third motion map 314. In some embodiments, second image frame 304 and third image frame 306 may be combined according to HDR processing 316. Specifically, blend map 320 of HDR processing 316 may be determined based at least in part on third motion map 314, and second image frame 304 may be combined with third image frame 306 according to blend map 320. In some embodiments, as described above, a fifth image frame may further be corrected for motion within fourth image frame 308 based on third motion map 314 (such as according to TF processing 318) by applying temporal filtering processing. An output image frame may also be determined based on fourth image frame 308, the fifth image frame, or a combination thereof. Image frame 230 may be determined by processor 104 or ISP 112 and stored in memory 106. The stored image frame may be read by processor 104 and used to form a preview display on a display of device 100 and / or processed to form a photo for storage in memory 106 and / or sent to another device.

[0085] Figure 5 is a block diagram illustrating an example processor configuration for image data processing in an image capture device according to one or more embodiments of the present disclosure. Processor 104 or other processing circuitry may be configured to operate on image data to perform Figure 4 one or more operations of the method. Image data may be processed to determine one or more output image frames 512 based on one or more input image frames 510. Processor 104 includes image receiving logic 502, motion map determining logic 504, motion map combining logic 506, and image fusion logic 508.

[0086] Image receiving logic 502 may be configured to receive image frames 510. Image frames 510 may include first image frame 302, second image frame 304, and third image frame 306. As described above, first image frame 302 may be a reference frame, second image frame 304 may be captured with a short exposure time, and third image frame 306 may be captured with a long exposure time.

[0087] The motion map determination logic 504 may be configured to determine a first motion map 310 based on the first image frame 302 and the third image frame 306, and determine a second motion map 312 based on the second image frame 304 and the third image frame 306. In some embodiments, the motion maps 310, 312 may be calculated based on the differences between the image frames 510. For example, determining the first motion map 310 may include determining the difference between the third image frame 306 and the first image frame 302. Similarly, determining the second motion map 312 may include determining the difference between the second image frame 304 and the first image frame 302. Specifically, the local motion estimation may be determined by comparing the image frames 302, 304, 306. The global motion estimation may also be determined as described above. Then the local motion estimation, the global motion estimation, and their combination may be combined to generate the motion maps 310, 312.

[0088] The motion map combination logic 506 may be configured to determine a third motion map 314 based on the first motion map 310 and the second motion map 312. In some embodiments, the third motion map 314 may be determined to remove high-noise regions that are incorrectly identified as moving within the second motion map 312, and may be determined to reduce motion artifacts in the highlighted saturation regions. For example, an image frame with a short exposure time may be captured before an image frame with a long exposure time. In such instances, determining the third motion map 314 may include subtracting the first motion map 310 from the second motion map 312.

[0089] The image fusion logic 508 may be configured to determine a fourth image frame 308 by combining the second image frame 304 and the third image frame 306 according to the third motion map 314. In some embodiments, the second image frame 304 and the third image frame 306 may be combined according to the HDR processing 316. Specifically, the fusion map 320 of the HDR processing 316 may be determined at least in part based on the third motion map 314, and the second image frame 304 may be combined with the third image frame 306 according to the fusion map 320. In some embodiments, as described above, the fifth image frame may further correct the motion within the fourth image frame 308 based on the third motion map 314 (such as according to the TF processing 318) by applying a temporal filtering process. The output image frame 512 may also be determined based on the fourth image frame 308, the fifth image frame, or their combination. The image frame 512 may be determined by the image fusion logic 508 and stored in the memory 106. The stored image frame 512 may be read by the processor 104 and used to form a preview display on the display of the device 100 and / or processed to form a photo for storage in the memory 106 and / or sent to another device.

[0090] In one or more aspects, techniques for supporting image processing may include additional aspects, such as any single aspect or any combination of aspects described below or in connection with one or more other processes or devices described elsewhere herein. In a first aspect, the techniques described herein relate to an apparatus that includes: a memory that stores processor-readable code; and at least one processor coupled to the memory, the at least one processor configured to execute the processor-readable code to cause the at least one processor to perform operations including: receiving a first image frame, a second image frame, and a third image frame, where the second image frame is captured at a first exposure time and the third image frame is captured at a second exposure time that is longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0091] Additionally, the apparatus may perform or operate according to one or more aspects described below. In some specific implementations, the apparatus includes a wireless device, such as a UE. In some specific implementations, the apparatus includes a remote server (such as a cloud-based computing solution) that receives image data for processing to determine an output image frame. In some specific implementations, the apparatus may include at least one processor and a memory coupled to the processor. The processor may be configured to perform the operations described herein with respect to the apparatus. In some other specific implementations, the apparatus may include a non-transitory computer-readable medium having program code recorded thereon, and the program code may be executable by a computer to cause the computer to perform the operations described herein with reference to the apparatus. In some specific implementations, the apparatus may include one or more components configured to perform the operations described herein. In some specific implementations, a method of wireless communication may include one or more operations described herein with respect to the apparatus.

[0092] In a second aspect according to the first aspect, determining the fourth image frame includes: determining a fusion map for HDR processing based on the third motion map; and determining the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

[0093] In a third aspect according to at least one of the first aspect and the second aspect, the operations further include determining a fifth image frame by applying temporal filtering processing to correct motion within the fourth image frame based on the third motion map.

[0094] In a fourth aspect according to the third aspect, the first image frame is a reference frame from a previous application of the temporal filtering process and has a third exposure time shorter than the second exposure time, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

[0095] In a fifth aspect according to at least one of the first to fourth aspects, the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

[0096] In a sixth aspect according to at least one of the first to fifth aspects, the second image frame is captured before the third image frame.

[0097] In a seventh aspect according to at least one of the first to sixth aspects, determining the third motion map includes combining the first motion map and the second motion map.

[0098] In an eighth aspect according to at least one of the first to seventh aspects, determining the first motion map includes determining the difference between the third image frame and the first image frame, and wherein determining the second motion map includes determining the difference between the second image frame and the first image frame.

[0099] In a ninth aspect, the techniques described herein relate to a method that includes: receiving a first image frame, a second image frame, and a third image frame, wherein the second image frame is captured at a first exposure time and the third image frame is captured at a second exposure time longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0100] In a tenth aspect according to the ninth aspect, determining the fourth image frame includes: determining a fusion map for HDR processing based on the third motion map; and determining the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

[0101] In an eleventh aspect according to at least one of the ninth and tenth aspects, the method further includes determining a fifth image frame by applying a temporal filtering process to correct motion within the fourth image frame based on the third motion map.

[0102] In a twelfth aspect according to the eleventh aspect, the first image frame is a reference frame from a previous application of the temporal filtering process and has a third exposure time that is shorter than the second exposure time, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

[0103] In a thirteenth aspect according to at least one of the ninth to twelfth aspects, the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

[0104] In a fourteenth aspect according to at least one of the ninth to thirteenth aspects, the second image frame is captured before the third image frame.

[0105] In a fifteenth aspect according to at least one of the ninth to fourteenth aspects, determining the third motion map includes combining the first motion map and the second motion map.

[0106] In a sixteenth aspect according to at least one of the ninth to fifteenth aspects, determining the first motion map includes determining the difference between the third image frame and the first image frame, and wherein determining the second motion map includes determining the difference between the second image frame and the first image frame.

[0107] In a seventeenth aspect, the techniques described herein relate to a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations including: receiving a first image frame, a second image frame, and a third image frame, wherein the second image frame is captured at a first exposure time and the third image frame is captured at a second exposure time that is longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0108] In an eighteenth aspect according to the seventeenth aspect, determining the fourth image frame includes: determining a fusion map for HDR processing based on the third motion map; and determining the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

[0109] In a nineteenth aspect according to at least one of the seventeenth and eighteenth aspects, the operations further include determining a fifth image frame by applying a temporal filtering process to correct motion within the fourth image frame based on the third motion map.

[0110] In a twentieth aspect according to the nineteenth aspect, the first image frame is a reference frame from a previous application of the temporal filtering process and has a third exposure time shorter than the second exposure time, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

[0111] In a twenty-first aspect according to at least one of the seventeenth aspect to the twentieth aspect, the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

[0112] In a twenty-second aspect according to at least one of the seventeenth aspect to the twenty-first aspect, the second image frame is captured before the third image frame.

[0113] In a twenty-third aspect according to at least one of the seventeenth aspect to the twenty-second aspect, determining the third motion map includes combining the first motion map and the second motion map.

[0114] In a twenty-fourth aspect, the techniques described herein relate to an image capture device including: an image sensor; a memory storing processor-readable code; and at least one processor coupled to the memory and the image sensor, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to perform: receiving a first image frame, a second image frame, and a third image frame, wherein the second image frame is captured with a first exposure time and the third image frame is captured with a second exposure time longer than the first exposure time; determining a first motion map based on the first image frame and the third image frame; determining a second motion map based on the second image frame and the third image frame; determining a third motion map based on the first motion map and the second motion map; and determining a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

[0115] In a twenty-fifth aspect according to the twenty-fourth aspect, determining the fourth image frame includes: determining a fusion map for HDR processing based on the third motion map; and determining the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

[0116] In a twenty-sixth aspect according to at least one of the twenty-fourth and twenty-fifth aspects, the processor-readable code further causes the processor to determine a fifth image frame by applying temporal filtering processing to correct motion within the fourth image frame based on the third motion map.

[0117] In a twenty-seventh aspect according to the twenty-sixth aspect, the first image frame is a reference frame from a previous application of the temporal filtering process and has a third exposure time shorter than the second exposure time, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

[0118] In a twenty-eighth aspect according to at least one of the twenty-fourth aspect to the twenty-seventh aspect, the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

[0119] In a twenty-ninth aspect according to at least one of the twenty-fourth aspect to the twenty-eighth aspect, the second image frame is captured before the third image frame.

[0120] In a thirtieth aspect according to at least one of the twenty-fourth aspect to the twenty-ninth aspect, determining the third motion map includes combining the first motion map and the second motion map.

[0121] Those skilled in the art should understand that: Any one of a variety of different techniques and arts can be used to represent information and signals. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.

[0122] This article is relative to Figure 1 The components, functional blocks, and modules described with respect to FIGS. 1 to 3 include a processor, an electronic device, a hardware device, an electronic component, a logic circuit, a memory, software code, firmware code, etc., or any combination thereof. Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, procedures, and / or functions, etc., regardless of whether it is called software, firmware, middleware, microcode, hardware description language, or other terms. In addition, the features discussed herein can be implemented via dedicated processor circuitry, via executable instructions, or a combination thereof.

[0123] Those skilled in the art should understand that: With reference to Figure 4 and Figure 5 One or more of the boxes (or operations) described can be combined with one or more of the boxes (or operations) described with reference to another figure in the reference drawings. For example, Figure 4 One or more of the boxes (or operations) of Figure 1 can be combined with one or more of the boxes (or operations) of FIGS. 1 to 3. As another example, one or more of the boxes associated with Figure 5 can be combined with those associated with Figure 1One or more combinations of boxes (or operations) associated with FIG. 3.

[0124] Those of ordinary skill in the art should also recognize that: all of the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. The skilled person may implement the described functionality in a different manner for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of this disclosure. The skilled person will also readily recognize that the order or combination of the components, methods, or interactions described herein are merely examples, and the components, methods, or interactions of the various aspects of this disclosure can be combined or performed in ways other than those illustrated and described herein.

[0125] The various illustrative logics, logical blocks, modules, circuits, and algorithmic processes described in connection with the specific implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. The interchangeability of hardware and software has been generally described in terms of functionality and illustrated in the various illustrative components, blocks, modules, circuits, and processes described above. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system.

[0126] A general purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein can be used to implement or perform the hardware and data processing apparatus for implementing the various illustrative logics, logical blocks, modules, and circuits described in connection with the various aspects disclosed herein. The general purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some specific implementations, the processor may be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. In some specific implementations, specific processes and methods may be performed by circuitry specific to a given function.

[0127] In one or more aspects, the described functionality may be implemented in hardware, digital electronic circuitry, computer software, firmware, including the structures disclosed in this specification and their structural equivalents, or any combination thereof. Specific implementations of the subject matter described in this specification may also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus.

[0128] If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code. The processes of the methods or algorithms disclosed herein may be implemented in a processor-executable software module that may reside on a computer-readable medium. Computer-readable media includes both computer storage media and communication media including any medium that can be implemented to transfer a computer program from one place to another. The storage media may be any available media that is accessible by a computer. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection may be properly termed a computer-readable medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, operations of a method or algorithm may be as a code and instruction set, or any combination of a code and instruction set, residing on a machine-readable medium and a computer-readable medium, which may be incorporated into a computer program product.

[0129] Various modifications to the specific implementations described in this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to some other specific implementations without departing from the spirit or scope of this disclosure. Thus, the claims are not intended to be limited to the specific implementations shown herein but are to be accorded the widest scope consistent with this disclosure, the principles disclosed herein, and the novel features.

[0130] In addition, those of ordinary skill in the art will readily recognize that, for the sake of convenience in describing the figures, terms such as "upper" and "lower" or "front" and "back" or "top" and "bottom" may sometimes be used, and indicate relative positions corresponding to the orientation of the figures on a correctly oriented page, and may not reflect the correct orientation of any device as implemented.

[0131] Certain features that are described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may have been described above as acting in certain combinations and even initially claimed as such, one or more features from the claimed combination can in some cases be excluded from the combination, and the claimed combination can be directed to a sub-combination or variations of a sub-combination.

[0132] Similarly, although operations are depicted in the figures in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. Additionally, the figures may schematically depict one or more example processes in the form of a flowchart. However, other operations not depicted can be incorporated into the example processes that are schematically illustrated. For example, one or more additional operations can be performed before, after, simultaneously with, or between any of the illustrated operations. In certain environments, multitasking and parallel processing are advantageous. Moreover, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other embodiments also fall within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result.

[0133] As used herein (including in the claims), the term "or" as used in a list of two or more items means that any one of the listed items can be employed alone, or any combination of two or more of the listed items can be employed. For example, if a composition is described as containing components A, B, or C, the composition can contain A alone; B alone; C alone; a combination of A and B; a combination of A and C; a combination of B and C; or a combination of A, B, and C. Additionally, as used herein (including in the claims), "or" as used in a list starting with "at least one" indicates a disjunctive list, such that a list of "at least one of A, B, or C" means A or B or C or AB or AC or BC or ABC (i.e., A and B and C) or any combination of any of these items.

[0134] The term "substantially" is defined as being largely but not necessarily wholly that which is specified (and includes that which is specified; for example, substantially 90 degrees includes 90 degrees, and substantially parallel includes parallel), as understood by one of ordinary skill in the art. In any of the specific embodiments disclosed, the term "substantially" may be replaced by "[percentage] within" that which is specified, where the percentage includes 0.1%, 1%, 5%, or 10%.

[0135] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An apparatus, the apparatus comprising: a memory that stores processor-readable code; and at least one processor coupled to the memory, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to perform operations including the following: receive a first image frame, a second image frame, and a third image frame, wherein the second image frame is captured at a first exposure time, and the third image frame is captured at a second exposure time longer than the first exposure time; determine a first motion map based on the first image frame and the third image frame; determine a second motion map based on the second image frame and the third image frame; determine a third motion map based on the first motion map and the second motion map; and determine a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

2. The apparatus according to claim 1, wherein determining the fourth image frame comprises: determine a fusion map for HDR processing based on the third motion map; and determine the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

3. The apparatus according to claim 1, wherein the operations further comprise determining a fifth image frame by applying temporal filtering processing to correct motion within the fourth image frame based on the third motion map.

4. The apparatus according to claim 3, wherein the first image frame is a reference frame from a previous application of the temporal filtering processing, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

5. The apparatus according to claim 1, wherein the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

6. The apparatus according to claim 1, wherein the second image frame is captured before the third image frame.

7. The apparatus according to claim 1, wherein determining the third motion map comprises combining the first motion map and the second motion map.

8. The apparatus according to claim 1, wherein determining the first motion map comprises determining the difference between the third image frame and the first image frame, and wherein determining the second motion map comprises determining the difference between the second image frame and the first image frame.

9. A method, the method comprising: receive a first image frame, a second image frame, and a third image frame, wherein the second image frame is captured at a first exposure time, and the third image frame is captured at a second exposure time longer than the first exposure time; determine a first motion map based on the first image frame and the third image frame; determine a second motion map based on the second image frame and the third image frame; determine a third motion map based on the first motion map and the second motion map; and determine a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

10. The method according to claim 9, wherein determining the fourth image frame comprises: Determine a fusion map for HDR processing based on the third motion map; and Determine the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

11. The method according to claim 9, wherein the method further comprises determining a fifth image frame by applying temporal filtering processing to correct motion within the fourth image frame based on the third motion map.

12. The method according to claim 11, wherein the first image frame is a reference frame from a previous application of the temporal filtering processing, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

13. The method according to claim 9, wherein the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

14. The method according to claim 9, wherein the second image frame is captured before the third image frame.

15. The method according to claim 9, wherein determining the third motion map comprises combining the first motion map and the second motion map.

16. The method according to claim 9, wherein determining the first motion map comprises determining a difference between the third image frame and the first image frame, and wherein determining the second motion map comprises determining a difference between the second image frame and the first image frame.

17. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations comprising: Receive a first image frame, a second image frame, and a third image frame, wherein the second image frame is captured at a first exposure time, and the third image frame is captured at a second exposure time longer than the first exposure time; Determine a first motion map based on the first image frame and the third image frame; Determine a second motion map based on the second image frame and the third image frame; Determine a third motion map based on the first motion map and the second motion map; and Determine a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

18. The non-transitory computer-readable medium according to claim 17, wherein determining the fourth image frame comprises: Determine a fusion map for HDR processing based on the third motion map; and Determine the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

19. The non-transitory computer-readable medium according to claim 17, wherein the operations further comprise determining a fifth image frame by applying temporal filtering processing to correct motion within the fourth image frame based on the third motion map.

20. The non-transitory computer-readable medium according to claim 19, wherein the first image frame is a reference frame from a previous application of the temporal filtering processing and has a third exposure time shorter than the second exposure time, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

21. The non-transitory computer-readable medium according to claim 17, wherein the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

22. The non-transitory computer-readable medium according to claim 17, wherein the second image frame is captured before the third image frame.

23. The non-transitory computer-readable medium according to claim 17, wherein determining the third motion map includes combining the first motion map and the second motion map.

24. An image capture device, the image capture device comprising: an image sensor; a memory that stores processor-readable code; and at least one processor coupled to the memory and the image sensor, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to: receive a first image frame, a second image frame, and a third image frame, wherein the second image frame is captured with a first exposure time, and the third image frame is captured with a second exposure time that is longer than the first exposure time; determine a first motion map based on the first image frame and the third image frame; determine a second motion map based on the second image frame and the third image frame; determine a third motion map based on the first motion map and the second motion map; and determine a fourth image frame by combining the second image frame and the third image frame according to the third motion map.

25. The image capture device according to claim 24, wherein determining the fourth image frame includes: determining a fusion map for HDR processing based on the third motion map; and determining the fourth image frame by mixing the second image frame and the third image frame according to the fusion map.

26. The image capture device according to claim 24, wherein the processor-readable code further causes the processor to determine a fifth image frame by applying temporal filtering processing to correct motion within the fourth image frame based on the third motion map.

27. The image capture device according to claim 26, wherein the first image frame is a reference frame from a previous application of the temporal filtering processing, and wherein the reference frame has been motion-aligned and processed to reduce spatial noise.

28. The image capture device according to claim 24, wherein the first exposure time is less than 2 ms, and the second exposure time is greater than or equal to 2 ms.

29. The image capture device according to claim 24, wherein the second image frame is captured before the third image frame.

30. The image capture device according to claim 24, wherein determining the third motion map includes combining the first motion map and the second motion map.