Frame processing and / or capture instruction systems and techniques

By dynamically adjusting frame capture settings based on motion information, the system addresses the challenge of capturing clear images in low-illumination scenes, achieving reduced noise and blur with efficient image processing.

JP7783879B2Active Publication Date: 2025-12-10QUALCOMM INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023520217
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-26
Filing Date
2021-10-01
Publication Date
2025-12-10
Estimated Expiration
2041-10-01

AI Technical Summary

Technical Problem

Existing camera systems struggle to capture clear, well-exposed images in low-illumination conditions without introducing motion blur or noise, particularly in scenes with weak illumination such as nighttime or indoor settings, due to challenges in synchronizing image sensor and image signal processor settings.

Method used

The system dynamically determines the number of frames and exposure duration based on motion information in the scene, using a motion map to adjust settings like gain, exposure, and number of frames to minimize blur and noise, employing techniques like multi-frame noise reduction and high dynamic range processing.

Benefits of technology

This approach generates well-exposed, sharp images with accurate color and low noise, maintaining dynamic range while reducing processing time and controlling artifacts, even in low-light conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007783879000003
    Figure 0007783879000003
  • Figure 0007783879000004
    Figure 0007783879000004
  • Figure 0007783879000005
    Figure 0007783879000005
Patent Text Reader

Abstract

Techniques and systems are provided for processing one or more frames or images. For example, a process for determining an exposure for one or more frames includes obtaining a motion map for the one or more frames. The process includes determining motion associated with one or more frames of a scene based on the motion map. The motion corresponds to movement of one or more objects in the scene relative to a camera used to capture the one or more frames. The process includes determining a number of frames and an exposure for capturing the number of frames based on the determined motion. The process further includes sending a request to capture the number of frames using the determined exposure duration.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to frame processing and / or capture instruction systems and techniques. [Background technology]

[0002] Cameras can be configured with various image capture and image processing settings to change the appearance of an image. Some camera settings, such as ISO, exposure time, aperture size, F-stop, shutter speed, focus, and gain, are determined and applied before or during photograph capture. Other camera settings, such as contrast, brightness, saturation, sharpness, levels, curves, or color modification, can configure post-processing of a photograph. A host processor (HP) can be used to configure the image sensor settings and / or the image signal processor (ISP) settings. To ensure that images are processed correctly, the configuration of settings between the image sensor and the ISP should be synchronized. Summary of the Invention [Means for solving the problem]

[0003]

[0006] Systems and techniques for performing image or frame processing and / or capture instruction configurations are described herein. According to an illustrative example, a method for determining an exposure duration and a number of frames is provided. The method includes obtaining a motion map for one or more frames, determining motion associated with one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects in the scene relative to a camera used to capture the one or more frames, determining the number of frames and an exposure duration for capturing the number of frames based on the determined motion, and sending a request to capture the number of frames using the determined exposure duration.

[0004] In another example, an apparatus for determining exposure durations for a number of frames is provided, the apparatus including: a memory configured to store at least one frame; and one or more processors (e.g., implemented in circuitry) coupled to the memory. The one or more processors are configured to and capable of: obtaining a motion map for the one or more frames; determining motion associated with the one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects in the scene relative to a camera used to capture the one or more frames; determining a number of frames and an exposure duration for capturing the number of frames based on the determined motion; and sending a request to capture the number of frames using the determined exposure duration.

[0005] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to obtain a motion map for one or more frames; determine motion associated with one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects in the scene relative to a camera used to capture the one or more frames; determine a number of frames and an exposure duration for capturing the number of frames based on the determined motion; and send a request to capture the number of frames using the determined exposure duration.

[0006] In another example, an apparatus for determining exposure durations for a number of frames is provided, the apparatus including: means for obtaining a motion map for one or more frames; means for determining motion associated with one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects in the scene relative to a camera used to capture the one or more frames; means for determining a number of frames and an exposure duration for capturing the number of frames based on the determined motion; and means for sending a request to capture the number of frames using the determined exposure duration.

[0007] In some aspects, one or more frames are acquired before a capture command to capture that number of frames is received.

[0008] In some aspects, the methods, apparatus, and computer-readable media described above further comprise performing temporal blending on the number of frames captured using the determined exposure duration to generate a temporally blended frame.

[0009] In some aspects, the methods, apparatus, and computer-readable media described above further comprise performing spatial processing on the temporally blended frames using a machine learning-based image signal processor, hi some aspects, the machine learning-based image signal processor uses the motion map as an input for performing spatial processing on the temporally blended frames.

[0010] In some aspects, the determined exposure duration is based on the gain.

[0011] In some aspects, the motion map includes an image, and each pixel of the image includes a value indicative of at least one of an amount of motion for each pixel and a confidence value associated with the amount of motion.

[0012] In some aspects, the above-described methods, devices, and computer-readable media further comprise determining a global motion associated with the camera based on one or more sensor measurements. In some cases, the number of frames and the exposure duration for capturing the number of frames are determined based on the determined motion and the global motion. For example, one or more processors of the device may be configured to determine the number of frames and the exposure duration for capturing the number of frames based on the determined motion and the global motion.

[0013] In some aspects, the above-described methods, devices, and computer-readable media further comprise determining a final motion instruction based on the determined motion and the global motion. In some cases, the number of frames and the exposure duration for capturing the number of frames are determined based on the final motion instruction. For example, one or more processors of the device may be configured to determine the number of frames and the exposure duration for capturing the number of frames based on the final motion instruction.

[0014] In some aspects, the final motion instruction is based on a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion. For example, to determine the final motion instruction based on the determined motion and the global motion, the one or more processors may be configured to determine a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion.

[0015] In some aspects, the methods, apparatus, and computer-readable media described above further comprise determining, based on the final motion indication, that an amount of motion in one or more frames is less than a motion threshold, and, based on the amount of motion in the one or more frames being less than the motion threshold, decreasing the number of frames for that number of frames and increasing the amount of exposure duration for the determined exposure duration.

[0016] In some aspects, the methods, apparatus, and computer-readable media described above further comprise determining, based on the final motion indication, that an amount of motion in one or more frames is greater than a motion threshold, and increasing the number of frames for that number of frames and decreasing the amount of exposure duration for the determined exposure duration based on the amount of motion in the one or more frames being greater than the motion threshold.

[0017] According to at least one other illustrative example, a method for performing temporal blending on one or more frames is provided, the method including: obtaining a raw frame, the raw frame including a single color component for each pixel of the raw frame; dividing the raw frame into a first color component, a second color component, and a third color component; generating a plurality of frames, at least in part, by adding at least a first chrominance value to the first color component, at least a second chrominance value to the second color component, and at least a third chrominance value to the third color component; and performing temporal blending on the plurality of frames.

[0018] In another example, an apparatus for performing temporal blending on one or more frames is provided, the apparatus including: a memory configured to store at least one image; and one or more processors (e.g., implemented in circuitry) coupled to the memory. The one or more processors are configured to: obtain a raw frame, the raw frame including a single color component for each pixel of the raw frame; divide the raw frame into a first color component, a second color component, and a third color component; generate a plurality of frames, at least in part, by adding at least a first chrominance value to the first color component, at least a second chrominance value to the second color component, and at least a third chrominance value to the third color component; and perform temporal blending on the plurality of frames.

[0019] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided that, when executed by one or more processors, cause the one or more processors to obtain a raw frame, the raw frame including a single color component for each pixel of the raw frame; divide the raw frame into a first color component, a second color component, and a third color component; generate a plurality of frames, at least in part, by adding at least the first chrominance value to the first color component, the at least the second chrominance value to the second color component, and the at least the third chrominance value to the third color component; and perform temporal blending on the plurality of frames.

[0020] In another example, an apparatus for performing temporal blending on one or more frames is provided, the apparatus including: means for obtaining a raw frame, the raw frame including a single color component for each pixel of the raw frame; means for dividing the raw frame into a first color component, a second color component, and a third color component; means for generating a plurality of frames, at least in part, by adding at least a first chrominance value to the first color component, at least a second chrominance value to the second color component, and at least a third chrominance value to the third color component; and means for performing temporal blending on the plurality of frames.

[0021] In some embodiments, the raw frame includes a color filter array (CFA) pattern.

[0022] In some embodiments, the first color component comprises a red color component, the second color component comprises a green color component, and the third color component comprises a blue color component.

[0023] In some embodiments, the first color component includes all red pixels of the raw frame, the second color component includes all green pixels of the raw frame, and the third color component includes all blue pixels of the raw frame.

[0024] In some aspects, generating the plurality of frames includes generating a first frame, at least in part, by adding at least a first chrominance value to a first color component, generating a second frame, at least in part, by adding at least a second chrominance value to a second color component, and generating a third frame, at least in part, by adding at least a third chrominance value to a third color component.

[0025] In some aspects, generating the first frame includes adding the first chrominance value and the second chrominance value to a first color component. In some aspects, generating the second frame includes adding the first chrominance value and the second chrominance value to a second color component. In some aspects, generating the third frame includes adding the first chrominance value and the second chrominance value to a third color component.

[0026] In some embodiments, the first chrominance value and the second chrominance value are the same value.

[0027] In some aspects, performing temporal blending on the plurality of frames includes temporally blending a first frame of the plurality of frames with one or more additional frames having a first color component, temporally blending a second frame of the plurality of frames with one or more additional frames having a second color component, and temporally blending a third frame of the plurality of frames with one or more additional frames having a third color component.

[0028] In some aspects, the device is, is part of, and / or includes a camera, a mobile device (e.g., a mobile phone, or so-called “smartphone,” or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, a vehicle or a computing device or component of a vehicle, or other device. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the devices described above may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).

[0029] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0030] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0031] Exemplary embodiments of the present application are described in detail below with reference to the following figures: [Brief explanation of the drawings]

[0032] [Figure 1]1 is a block diagram illustrating an example architecture of a frame processing and / or capture instruction system, according to some examples. [Figure 2] FIG. 10 illustrates various lux values ​​for different exemplary scenarios. [Figure 3] FIG. 1 illustrates an image captured during very low lighting conditions. [Figure 4] FIG. 1 is a block diagram illustrating an example of a frame processing and / or capture instruction system, according to some examples. [Figure 5] 5 is a timing diagram illustrating an example of the timing of various processes performed by the frame processing and / or capture instruction system of FIG. 4, according to some examples. [Figure 6] 5 illustrates an example implementation of the frame processing and / or capture instruction system of FIG. 4, according to some examples. [Figure 7] 5 illustrates an example implementation of the frame processing and / or capture instruction system of FIG. 4, according to some examples. [Figure 8] 5 illustrates an example implementation of the frame processing and / or capture instruction system of FIG. 4, according to some examples. [Figure 9] 5 illustrates an example implementation of the frame processing and / or capture instruction system of FIG. 4, according to some examples. [Figure 10] 5 illustrates an example implementation of the frame processing and / or capture instruction system of FIG. 4, according to some examples. [Figure 11] 5 illustrates an example implementation of the frame processing and / or capture instruction system of FIG. 4, according to some examples. [Figure 12] 1 is a graph plotting motion versus exposure duration (or exposure time) and frame number, according to some examples. [Figure 13] 1A-1C illustrate an image and a temporal filter indication (TFI) image, according to some examples. [Figure 14A] FIG. 1 illustrates another example of a frame processing and / or capture instruction system, according to some examples. [Figure 14B] FIG. 1 illustrates another example of a frame processing and / or capture instruction system, according to some examples. [Figure 15] FIG. 15B illustrates an example of a machine learning image signal processor (ML ISP) of the frame processing and / or capture instruction system of FIG. 15A, according to some examples. [Figure 16] FIG. 15C illustrates an example neural network of the ML ISP of FIGS. 15A and 15B, according to some examples. [Figure 17A] 15B illustrates the frame processing and / or capture command system of FIG. 15A with additional processing used for white balance refinement, according to some examples. [Figure 17B] 15B illustrates the frame processing and / or capture command system of FIG. 15A with additional processing used for white balance refinement, according to some examples. [Figure 18] FIG. 1 illustrates an example process for gradually displaying an image, according to some examples. [Figure 19] FIG. 10 illustrates an example of raw temporal blending based on chrominance (U and V) channels, according to some examples. [Figure 20] FIG. 10 illustrates an example of raw temporal blending based on chrominance (U and V) channels, according to some examples. [Figure 21] 10A-10C are diagrams including images resulting from raw temporal blending and from using standard YUV images, according to some examples. [Figure 22] FIG. 10 is a flow diagram illustrating an example of a process for determining exposure duration for several frames, according to some examples. [Figure 23] FIG. 10 is a flow diagram illustrating another example of a process for performing temporal blending, according to some examples. [Figure 24]FIG. 1 is a block diagram illustrating an example of a neural network, in accordance with some examples. [Figure 25] FIG. 1 is a block diagram illustrating an example of a convolutional neural network (CNN), in accordance with some examples. [Figure 26] FIG. 1 illustrates an example of a system for implementing some aspects described herein. DETAILED DESCRIPTION OF THE INVENTION

[0033] Some aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for purposes of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments can be practiced without these specific details. The figures and description are not intended to be limiting.

[0034] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with an enabling description for practicing the exemplary embodiments. It will be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application, as set forth in the appended claims.

[0035] A camera is a device that uses an image sensor to receive light and capture image frames, such as still images or video frames. The terms “image,” “image frame,” and “frame” are used interchangeably herein. A camera system may include a processor (e.g., an image signal processor (ISP)) that can receive and process one or more image frames. For example, raw image frames captured by a camera sensor may be processed by the ISP to generate a final image. The ISP may perform processing by applying multiple filters or processing blocks to the captured image frames, such as noise removal or filtering, edge enhancement, color balancing, contrast, intensity adjustment (such as darkening or lightening), and tone adjustment, among others. Image processing blocks or modules may include lens / sensor noise correction, Bayer filtering, demosaicing, color conversion, image attribute correction or enhancement / suppression, noise removal filters, and sharpening filters, among others.

[0036] In many camera systems, a host processor (HP), sometimes called an application processor (AP), is used to dynamically configure the image sensor with new parameter settings. The HP is also used to dynamically configure the parameter settings of the ISP pipeline to match the settings of the input image sensor frame so that the image data can be processed correctly.

[0037] Cameras can be configured with a variety of image capture and image processing settings. The application of different settings can result in frames or images with different appearances. Some camera settings, such as ISO, exposure time (also called exposure duration), aperture size, f-stop, shutter speed, focus, and gain, are determined and applied before or during the capture of a photograph. Other camera settings, such as contrast, brightness, saturation, sharpness, levels, curves, or color modifications, can constitute post-processing of the photograph.

[0038] Challenges exist, particularly when attempting to capture frames or images in scenes with weak illumination, such as nighttime scenes or indoor scenes with weak or low illumination. For example, weakly lit scenes are typically dark with saturated bright areas (if any exist). Images of scenes with weak illumination are referred to herein as low-illumination images. Low-illumination images are typically dark, noisy, and colorless. For example, low-illumination images typically have dark pixels along with overly bright areas relative to the bright areas of the scene. Furthermore, the signal-to-noise ratio (SNR) in low-illumination images is extremely small. Noise in low-illumination images is a manifestation of random fluctuations in brightness and / or color information caused by low-lighting conditions. The result of noise is that low-illumination images appear grainy. In some cases, the signal of a low-illumination image must be amplified due to a small SNR. For example, signal amplification may introduce more noise and inaccurate white balance. In some cases, the exposure duration / time for the camera can be increased to help increase the amount of light exposed to the image sensor. However, increased exposure duration can introduce motion blur artifacts that result in blurry images due to more light hitting the camera sensor during shutter operation.

[0039] Described herein are systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to herein as “systems and techniques”) for providing improved image processing techniques. In some cases, the systems and techniques described herein can be used to provide improved low-lighting frames or images. However, the systems and techniques can also be applied to frames or images captured in other lighting conditions. For example, as described in more detail below, the systems and techniques can be used to generate well-exposed (e.g., correctly exposed with little blur), sharp (e.g., having a high texture level maintained with low noise) frames with accurate color, high texture, and low noise. In some cases, the systems and techniques can control artifacts, preserve most or all of the dynamic range of the captured frame, and / or provide good quality inter-take and processing times. For example, using the systems and techniques described herein, frames can be generated while maintaining comparable frame quality with reduced processing time (compared to other image capture systems). In some examples, the systems and techniques can generate an interactive preview (eg, by gradually displaying frames as new frames are buffered and / or processed).

[0040] In some aspects, the systems and techniques may leverage the use of existing hardware components, such as multi-frame noise reduction (MFNR) hardware components, multi-frame high dynamic range (MFHDR) hardware components, massive multi-frame (MMF) hardware components, staggered HDR (sHDR) hardware components (which may be a subset of MFHDR), any combination thereof, and / or other hardware components. In some aspects, the systems and techniques may use longer exposure times and greater gain to capture dark frames. In some examples, the systems and techniques may use adaptive dynamic range control (ADRC) and / or multi-frame high dynamic range (MFHDR), sHDR, and / or MMF for highly saturated (or clipped or misshapen) brightest parts. ADRC may be used to achieve high dynamic range (HDR) from a single image or frame. For example, ADRC can underexpose frames to preserve the brightest parts, and then apply later gain to compensate for shadows and mid-tones. In some embodiments, systems and techniques can use MFxR, smaller gain with longer exposure durations, and possibly machine learning systems for frames with high noise and / or low texture. The term MFxR can refer to multi-frame noise reduction (MFNR) and / or multi-frame super resolution (MFSR).Also, as used herein, when describing MFxR (e.g., MFNR, MFSR, etc.), the same or similar techniques may be performed using MMF in addition to or as an alternative to MFxR. In MFNR, MFSR, MMF, or other related techniques, a final frame may be generated by blending two or more frames.

[0041] In some examples, for frames with motion blur, systems and techniques can utilize dynamic determination of a combination of the number of frames and exposure duration (and / or gain) according to motion information associated with one or more preview frames (e.g., effective motion / movement in the scene, global motion of image capture determined using sensor measurements, or a combination thereof). For example, using motion information associated with one or more preview frames, systems and techniques can determine the number of frames to capture (which may later be combined using MFNR, MMF, etc.) and thereby the exposure duration for capturing that number of frames. In some cases, systems and techniques (e.g., a low-light engine) can obtain various decisions and statistics from a preview pipeline (e.g., generating preview frames). Systems and techniques can output capturing instructions for an offline pipeline (e.g., capturing output frames). For example, systems and techniques can calculate an exposure duration (e.g., longest single-frame exposure duration) for an optimal balance between motion blur and SNR. In some cases, the SNR variation in this case corresponds to the actual sensor gain to be applied and is a product of a target exposure divided by the exposure duration / time. In some cases, the systems and techniques can calculate the number of frames in a multi-frame (the number of frames described above) to satisfy a required inter-take time or duration. The inter-take time can refer to the duration between two consecutive user-initiated frame captures (e.g., between activation of a shutter or capture option, such as selecting the shutter / capture button). The required inter-take duration can be the product of the number of frames multiplied by the single-frame exposure duration (in addition to a default pipeline latency).

[0042] In some aspects, to find an accurate white balance, systems and techniques can calculate auto-white balance (AWB) statistics from longer exposed frames (also referred to herein as long-exposure frames). The long-exposure frames are captured using an exposure time longer than the standard exposure time used to capture several frames (e.g., in a scene that does not have low-lighting conditions, such as the exemplary standard lighting conditions shown in FIG. 2). In some cases, systems and techniques can calculate AWB statistics from a longer exposed, integrated frame (e.g., by combining multiple longer-exposed frames). In some examples, to address processing latency, systems and techniques can process short-exposure frames (e.g., captured using a standard exposure time) while capturing longer-exposed frames (referred to herein as long-exposure frames). In some examples, systems and techniques can continue to process short-exposure frames in the background after capture has finished. For example, systems and techniques can collect a queue of short-exposure frames and / or long-exposure frames and process frames while collecting a subsequent set of short-exposure frames and / or long-exposure frames. In some aspects, the systems and techniques can continually provide "improving" longer exposed frames to a preview engine while still capturing, and the preview engine can output frames as previews (e.g., before a shutter button or shutter option is activated and / or while the capture process based on the shutter button or shutter option is still occurring).

[0043] In some examples, such as addressing quantization issues (e.g., chroma staining), systems and techniques can use a post-image processing engine (post-IPE) after MFHDR in the camera pipeline and / or can use a machine learning system. The term “chroma staining” is a visual term for chroma quantization and is sometimes referred to as “chroma banding” or “chroma contour.” Chroma staining can occur with frames that have insufficient color depth and undergo additional processes to smooth (e.g., noise reduction) and enhance color. The results of chroma staining can include contours or steps on flat (nearly gray) areas. The post-IPE is an additional hardware image processing engine (IPE) instance that can be used to further smooth the generated contours. For example, the post-IPE can be located at the end of the frame processing and / or capture instruction pipeline (e.g., when the frame has its final tone).

[0044] In some examples, the systems and techniques can activate or deactivate some low-light processing based on illuminance (lux) metering. Using illuminance metering, the systems and techniques can be dynamically enabled based on lighting conditions. For example, an image signal processor (ISP) or other processor of a frame processing and / or capture instruction system can measure the amount of light. Based on the amount of light (e.g., low lighting, standard lighting, very low lighting, etc., such as that shown in FIG. 2), the ISP can determine whether to activate one or more of the techniques described herein.

[0045] In some examples, long-exposure frames have significantly higher light sensitivity than short-exposure frames. Light sensitivity may also be referred to as “exposure,” “image exposure,” or “image sensitivity” and may be defined as follows: light sensitivity = gain * exposure_time. Exposure time may also be referred to as exposure duration. Furthermore, the term exposure refers to exposure duration or exposure time. The exposure scale used to capture short-exposure and long-exposure frames may vary. Short-exposure and long-exposure frames may span the entire available gain range in some cases. In some examples, short-exposure frames may be captured using an exposure of 33 milliseconds (ms) or 16.7 ms. In some cases, short-exposure frames may be used as candidates for supporting standard preview (e.g., previewed in a device's display before a shutter button or shutter option is activated and / or while a capture process based on the shutter button or shutter option is still running). In some cases, the exposure time for short-exposure frames may be extremely short (e.g., 0.01 seconds), such as to meet flicker-free conditions, or even shorter if a flicker-free, direct light source is detected. In some cases, to reach maximum light sensitivity for a particular frame rate (e.g., a user-defined frame rate), the exposure time for a short-exposure frame may be 1 / frame_rate seconds. In some examples, the exposure of a short-exposure frame may vary within a range of [0.01, 0.08] seconds. A long-exposure frame or image may be captured using any exposure greater than the exposure used to capture the short-exposure frame. For example, the exposure of a long-exposure frame may vary within a range of [0.33, 1] seconds. For example, in one illustrative example using a tripod without in-scene movement, the exposure duration for a long-exposure frame may be approximately 1 second.In some cases, the exposure used for the long exposure frame may reach a longer duration (eg, 3 seconds or other duration), but is no "shorter" than the short exposure frame.

[0046] As used herein, the terms short-term, medium-term, safe, and long-term refer to relative characterizations between a first setting and a second setting and do not necessarily correspond to defined ranges for a particular setting. That is, a long-term exposure (or a long-term exposure duration or a long-term exposure frame or image) simply refers to an exposure time that is longer than a second exposure (e.g., a short-term exposure or a medium-term exposure). In another example, a short-term exposure (or a short-term exposure duration or a short-term exposure frame) refers to an exposure time that is shorter than a second exposure (e.g., a long-term exposure or a medium-term exposure).

[0047] Various aspects of the application are described with reference to the figures. FIG. 1 is a block diagram illustrating the architecture of a frame capture and processing system 100. The frame capture and processing system 100 includes various components used to capture and process frames of a scene (e.g., frames of a scene 110). The frame capture and processing system 100 can capture standalone frames (or photographs) and / or can capture videos that include multiple frames (or video frames) in a particular sequence. A lens 115 of the system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the light toward an image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0048] The one or more controls 120 may control exposure, focus, and / or zoom based on information from image sensor 130 and / or based on information from image processor 150. The one or more controls 120 may include multiple mechanisms and components, for example, the control 120 may include one or more exposure controls 125A, one or more focus controls 125B, and / or one or more zoom controls 125C. The one or more controls 120 may also include additional controls beyond those shown, such as controls for analog gain, flash, HDR, depth of field, and / or other image capture characteristics.

[0049] The focus control mechanism 125B of the control mechanism 120 can obtain the focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or servo, thereby adjusting the focus. In some cases, additional lenses, such as one or more microlenses above each photodiode of the image sensor 130, can be included in the system 100, each of which bends light received from the lens 115 toward a corresponding photodiode before the light reaches the photodiode. The focus setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus settings may be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus settings may be referred to as image capture settings and / or image processing settings.

[0050] Exposure control 125A of control mechanism 120 can obtain an exposure setting. In some cases, exposure control 125A stores the exposure setting in a memory register. Based on the exposure setting, exposure control 125A can control the size of the aperture (e.g., aperture size or F-stop), the duration the aperture is open (e.g., exposure time or shutter speed), the sensitivity of image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.

[0051] The zoom control 125C of the control mechanism 120 can obtain a zoom setting. In some examples, the zoom control 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control 125C can control the focal length of the lens assembly by actuating one or more motors or servos to move one or more of the lenses relative to one another. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly may include a parfocal zoom lens or a varifocal zoom lens. In some examples, the lens assembly may include a focusing lens (which may be the lens 115 in some cases) that first receives light from the scene 110, and then the light passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, an afocal zoom system may include two positive (e.g., converging, convex) lenses of equal or similar focal lengths (e.g., within a threshold difference), with a negative (e.g., diverging, concave) lens between them. In some cases, zoom control 125C moves one or more of the lenses in the afocal zoom system, such as the negative lens and one or both of the positive lenses.

[0052] Image sensor 130 includes one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a particular pixel in the image or frame produced by image sensor 130. In some cases, different photodiodes may be covered by different color filters, thus measuring light that matches the color of the filter covering the photodiode. For example, a Bayer color filter includes red, blue, and green color filters, and each pixel of the frame is generated based on red light data from at least one photodiode covered within the red color filter, blue light data from at least one photodiode covered within the blue color filter, and green light data from at least one photodiode covered within the green color filter. Other types of color filters may use yellow, magenta, and / or cyan (also called "emerald") color filters instead of or in addition to red, blue, and / or green color filters. Some image sensors may lack color filters entirely and instead use different photodiodes (possibly stacked vertically) throughout the pixel array. Different photodiodes across the pixel array can have different spectral sensitivity curves and therefore respond to different wavelengths of light. Monochrome image sensors may also lack color filters and therefore no color depth.

[0053] In some cases, image sensor 130 may alternatively or additionally include opaque and / or reflective masks that block light from reaching some photodiodes or portions of some photodiodes at some times and / or from some angles, which may be used for phase-detection autofocus (PDAF). Image sensor 130 may also include analog gain amplifiers to amplify analog signals output by the photodiodes and / or analog-to-digital converters (ADCs) and / or convert the analog signal output of the photodiodes (and / or amplified by the analog gain amplifiers) to digital signals. In some cases, instead of or in addition, some components or functions described with respect to one or more of control mechanisms 120 may be included in image sensor 130. The image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS), an n-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0054] Image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISPs 154), one or more host processors (including host processor 152), and / or one or more of any other types of processors 910 described with respect to computing system 900. Host processor 152 may be a digital signal processor (DSP) and / or other types of processor. In some implementations, image processor 150 is a single integrated circuit or chip (called a system-on-chip, or SoC) that includes host processor 152 and ISPs 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) ports 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth™, Global Positioning System (GPS), etc.), any combination thereof, and / or other components. The I / O ports 156 may include any suitable input / output ports or interfaces according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General-Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as MIPI CSI-2), a Physical (PHY) layer port or interface, an Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In one illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port, and the ISP 154 may communicate with the image sensor 130 using a MIPI port.

[0055] Host processor 152 can configure image sensor 130 (e.g., via an external control interface such as I2C, I3C, SPI, GPIO, and / or other interfaces) with the new parameter settings. In one illustrative example, host processor 152 can update the exposure settings used by image sensor 130 based on internal processing results of exposure control algorithms from past images or frames.

[0056] Host processor 152 may also dynamically configure parameter settings of pipelines or modules within ISP 154. For example, host processor 152 may configure a pipeline or module of ISP 154 to match the settings of one or more input frames from image sensor 130 so that the image or frame data is properly processed by ISP 154. The processing (or pipeline) blocks or modules of ISP 154 may include modules for lens (or sensor) noise correction, demosaicing, color space conversion, color correction, frame attribute enhancement and / or suppression, noise removal (e.g., using a noise removal filter), and sharpening (e.g., using a sharpening filter), among others. Based on the configured settings, the ISP 154 may perform one or more image processing tasks such as noise correction, demosaicing, color space conversion, frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance (AWB), merging frames to form an HDR frame or image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or any combination thereof.

[0057] The image processor 150 may store frames and / or processed frames in random access memory (RAM) 140 / 920, read-only memory (ROM) 145 / 925, cache 912, a memory unit (e.g., system memory 915), another storage device 930, or some combination thereof.

[0058] Various input / output (I / O) devices 160 may be connected to image processor 150. I / O devices 160 may include a display screen, a keyboard, a keypad, a touchscreen, a trackpad, a touch-sensitive surface, a printer, any other output device(s) 935, any other input device(s) 945, or some combination thereof. I / O 160 may include one or more ports, jacks, or other connectors that enable a wired connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. I / O 160 may include one or more wireless transceivers that enable a wireless connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. Peripheral devices may include any of the types of I / O devices 160 previously described, and may themselves be considered I / O devices 160 when coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector.

[0059] In some cases, frame capture and processing system 100 may be a single device. In some cases, frame capture and processing system 100 may be two or more separate devices including image capture device 105A (e.g., a camera) and image processing device 105B (e.g., a computing device coupled to a camera). In some implementations, image capture device 105A and image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors and / or wirelessly via one or more wireless transceivers. In some implementations, image capture device 105A and image processing device 105B may be separate from one another.

[0060] 1, a vertical dashed line divides frame capture and processing system 100 of FIG. 1 into two portions, representing image capture device 105A and image processing device 105B, respectively. Image capture device 105A includes lens 115, control mechanism 120, and image sensor 130. Image processing device 105B includes image processor 150 (including ISP 154 and host processor 152), RAM 140, ROM 145, and I / O 160. In some cases, some components shown in image capture device 105A, such as ISP 154 and / or host processor 152, may be included within image capture device 105A.

[0061] Frame capture and processing system 100 may include an electronic device such as a mobile or fixed telephone handset (e.g., a smartphone, a cellular telephone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, frame capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, image capture device 105A and image processing device 105B may be different devices. For example, image capture device 105A may include a camera device, and image processing device 105B may include a computing device, such as a mobile handset, a desktop computer, or other computing device.

[0062] Although frame capture and processing system 100 is shown as including several components, those skilled in the art will appreciate that frame capture and processing system 100 may include many more components than those shown in FIG. 1 . The components of frame capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of frame capture and processing system 100 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing frame capture and processing system 100.

[0063] FIG. 2 is a diagram illustrating various lux values ​​for different exemplary scenarios. While the example in FIG. 2 is shown using lux units, illumination may be measured in other units, such as high ISO (which represents the sensor's sensitivity to light). Generally, lux conditions may correspond to low illumination, standard illumination, and bright illumination, among others. The lux conditions illustrated in the diagram in FIG. 2 are relative terms used to distinguish lux values. For example, as used herein, a standard illumination range refers to a lux range that is relatively higher than a low illumination range, and an extra-low illumination range refers to a lux range that is relatively lower than a certain range. In some cases, the exemplary lux values ​​may be reassigned based on context. For example, the lux ranges may be reassigned by relative descriptors depending on frame processing and / or capture instruction system capabilities (e.g., sensor pixel size and / or sensitivity) and / or use case (e.g., based on scene conditions).

[0064] Referring to FIG. 2 , standard lighting conditions may correspond to lux values ​​of 50, 150, 300, 1,000, 10,000, and 30,000. For example, a lux value of 30,000 may occur in a scene including a sidewalk lit by direct sunlight, and a lux value of 1,000 may occur in a scene including a sidewalk during an overcast day. Low-light (LL) conditions may correspond to lux values ​​of 10 and 20. Ultra-low-light (ULL) conditions may correspond to lux values ​​of 0.1, 0.3, 0.6, 1, and 3. Although exemplary lux values ​​are shown in FIG. 2 to correspond to various lighting conditions, values ​​other than those shown in FIG. 2 may correspond to various lighting conditions. FIG. 3 is an image 300 (or frame) captured during an ultra-low lighting condition (e.g., a lux value of 3). As shown, image 300 is dark, the boat is depicted using dark pixels, and the bright portions of image 300 correspond to the light on the lamppost.

[0065] FIG. 4 is a block diagram illustrating an example of a frame processing and / or capture instruction system 400. One or more of the components of the frame processing and / or capture instruction system 400 of FIG. 4 may be similar to components of the frame capture and processing system 100 of FIG. 1 and may perform similar operations as such components. For example, the sensor 430 may be similar to the sensor 130 of the frame capture and processing system 100 and may perform similar operations as such sensor 130. As shown, a zero shutter lag (ZSL) buffer 432 may be used to store images or frames captured by the sensor 430. In some examples, the ZSL buffer 432 is a circular buffer. Generally, the ZSL buffer 432 may be used to store one or more frames recently captured by the sensor, which may compensate for lag time that may occur until the frame processing and / or capture instruction system 400 finishes encoding and storing a frame in response to a shutter (or capture) command being received (e.g., based on user input or automatically).

[0066] When the frame processing and / or capture instruction system 400 processes the shutter command, the frame processing and / or capture instruction system 400 can select one of the buffered frames and further process the selected frame for storage, display, transmission, etc. As shown, the ZSL frame is captured using a relatively short exposure and is referred to as a short-exposure frame 434 (referred to as a "Short Exp" frame or image in FIG. 4). The short-exposure frame 434 is output to a first MFxR engine 436 (e.g., an engine capable of performing MFNR and / or MFSR). The first MFxR engine 436 generates a blended frame 438 based on the short-exposure frame. The blended frame 438 is output to a multi-frame high dynamic range (MFHDR) engine 440. The MFHDR engine 440 can receive as input multiple frames or images of the same scene captured using different light sensitivity or exposure parameters (e.g., exposure time and / or gain), such as one or more frames captured using a relatively short exposure (e.g., 33 ms), one or more frames captured using a relatively medium exposure (e.g., 100 ms), and one or more frames captured using a relatively long exposure (e.g., 500 ms). The frames captured using the long exposure may be referred to as long-exposure frames or images, as described above. In some examples, in addition to or as an alternative to exposure duration, sensor gain may be adjusted for the short-exposure frames or images, the medium-exposure frames or images, and the long-exposure frames or images. The MFHDR engine can combine multiple frames or images into a single high dynamic range (HDR) frame. The MFHDR engine 440 can output the HDR frame to the post-IPE 442. In some examples, the MFHDR engine 440 can apply tone mapping to bring different portions of the frame to desired brightness levels.In some examples, each MFHDR input is a separate MFNR output (e.g., long-term and short-term input / output). The post-IPE 442 can perform additional image processing operations on the HDR frame from the MFHDR engine 440 to generate a final output frame or image. The additional image processing operations performed by the post-IPE 442 can include, for example, gamma, sharpening, color fine-touch, upscaling, and graining, among others.

[0067] The sensor 430 also outputs a positive shutter lag (PSL) frame or image during PSL capture 444. The PSL frame is captured using a relatively long exposure time (and / or corresponding gain) compared to the ZSL frame and is referred to as a long exposure frame 446 (referred to as a "Long Exp" frame or image in FIG. 4 ), as described above. AWB statistics 448 from the long exposure frame 446 are provided to a first AWB engine 450 (shown as "AWB1"). The first AWB engine 450 can generate a white balance (WB) scaler 451 and output the WB scaler 451 to the first MFxR engine 436 and / or a second MFxR engine 452 (e.g., an engine capable of performing MFNR and / or MFSR). In some examples, the WB scaler 451 can include three coefficients targeting red, green, and blue gain scales that can be applied to achieve a neutral gray for a particular observer. In one illustrative example, the WB scaler 451 may include a coefficient value of 1.9 for red (R), a coefficient value of 1.0 for green (G), and a coefficient value of 1.6 for blue (B). The long-exposure frame 446 is also output to a second MFxR engine 452. In some cases, the first MFxR engine 436 and the second MFxR engine 452 may be implemented by the same hardware using the same processing techniques, although the hardware may have different tuning settings when implementing the first MFxR engine compared to the second MFxR engine. As shown in FIG. 4 , the second MFxR engine 452 may output the long-exposure frame 446 as a preview frame 454 (e.g., displayed before a shutter command or capture command is received and / or while capture processing based on the shutter command is still being performed) on the frame processing and / or capture instruction system 400 or a display of a device including the frame processing and / or capture instruction system 400.In some examples, the second MFxR engine 452 can generate a blended frame 456 based on the long-exposure frame 446 (e.g., by fusing or blending the long-exposure frames). The blended frame 456 is output to the MFHDR engine 440 and also to a low-lighting (LL) engine 458, which can extract AWB statistics. The AWB statistics are output to a second AWB engine 460 (shown as “AWB2”). In some cases, the first AWB engine 450 and the second AWB engine 460 can be implemented by the same hardware using the same processing techniques, but the hardware can have different tuning settings when implementing the first AWB engine compared to the second AWB engine. The second AWB engine 460 can generate a WB scaler 461 and output the WB scaler 461 to the MFHDR engine 440. The MFHDR engine 440 outputs the combined frame (e.g., HDR frame) to the post-IPE 442, as described above.

[0068] As shown in FIG. 4 , the frame processing and / or capture instruction system can perform a low illumination (LL) decision 462 to determine the number of frames (or images) and exposure time (and / or corresponding gain) for each image / frame. Further details regarding determining the number of frames and exposure time (and / or gain) are described below with respect to FIGS. 12-17B . In some examples, the LL engine 458 can perform the LL decision 462. Based on the LL decision 462, an automatic exposure control (AEC) engine 464 can perform AEC to determine exposure parameters (e.g., exposure duration, gain, etc.) for the sensor 430. For example, the AEC engine 464 can send (to the sensor 430) an indication of the number of frames to capture and the exposure time (and / or sensor gain) for that number of frames based on the LL decision 462.

[0069] FIG. 5 is a timing diagram 500 illustrating an example of the timing of various processes executed by the frame processing and / or capture instruction system 400. For example, the timing diagram illustrates low lighting (LL) mode, when the shutter is activated (e.g., pressed or otherwise selected, corresponding to receiving a shutter or capture command), the inter-take time, and the total process time. The inter-take time refers to the duration between two consecutive user-initiated image captures (e.g., between activation of a shutter or capture option, such as selecting the shutter / capture button). As shown, a ZSL frame 502 (e.g., with short-term AEC) is captured before the shutter is pressed. In some cases, the AEC is a parent algorithm that controls exposure, gain, etc. The AEC also has a sensor interface. In some examples, the AEC has three metering modes: short-term metering, safety metering, and long-term metering. Short-term metering results in the brightest areas being preserved, safety metering provides a balanced exposure, and long-term metering prioritizes dark areas. The ZSL frame 502 is a preview image and may be used as a short-term exposure frame (e.g., to preserve the brightest parts) as described above. Therefore, the ZSL frame 502 may be captured according to short-term AEC photometry. When the shutter is pressed, a PSL frame 504 is captured, after which short-term multi-frame noise reduction (MFNR) 506 and long-term MFNR and preview 508 are applied. White balance (WB) refinement 510 is performed to refine the WB, and MFHDR and post-processing 512 may then be applied.

[0070] 6-11 illustrate an example implementation of the frame processing and / or capture instruction system 400 of FIG. 4 according to the timing diagram 500 of FIG. 5. FIG. 6 illustrates use of the frame processing and / or capture instruction system 400 during the LL mode timing of the timing diagram 500 shown in FIG. 5. As described above, the LL mode includes short-term AEC and ZSL (before the shutter is activated). In some cases, the LL mode (short-term AEC setting and ZSL capture) may be performed as the first step of the frame processing and / or capture instruction process. The AEC engine 464 may set the ZSL based on the LL decision 462 (e.g., by the LL engine 458). For example, the AEC engine 464 can determine a shutter priority (e.g., a setting that allows a specific shutter speed (corresponding to the exposure time) to be used, and then the AEC engine 464 calculates a gain to complement the shutter speed / exposure time), according to a user configuration or an automatic setting (e.g., within a valid ZSL range, such as between 1 / 7 and 1 / 15 seconds) to set the exposure (using short-term AEC) used to capture the short-exposure frame 434 (also called a ZSL frame or image). In some cases, the shutter speed can refer to the sensor exposure time or the effective sensor readout time (e.g., for devices that do not include a physical shutter). The AEC engine 464 can select a gain from an AEC “short-term” metric with a slight overexposure. The AEC engine 464 can determine a frame set based on the determined AEC setting. For example, a longer frame interval can be determined for a gain change (e.g., a sequence of 3 to 8 frames with the same gain). In some cases, the frame processing and / or capture instruction system 400 can determine (e.g., based on a dynamic range calculation) whether MFHDR is to be used to process the short-exposure frame 434 (e.g., whether MFHDR mode or non-MFHDR mode should be used by the MFHDR engine 440).In some examples, the frame processing and / or capture instruction system 400 can further calculate a dynamic range for the short-exposure frame 434. For example, the dynamic range may be determined from the AEC “long-term” to “short-term” ratio. In one illustrative example, to determine the desired mode (e.g., whether to use MFHDR or non-MFHDR) and some additional configurations, the frame processing and / or capture instruction system 400 can determine the “amount” of dynamic range in the scene, such as by calculating a ratio from the extreme photometry of the AEC (e.g., between short and long photometry). If the frame processing and / or capture instruction system 400 determines that the MFHDR mode will be used, the short-exposure frame 434 may be processed using adaptive dynamic range control (ADRC). The sensor 430 can capture the short-exposure frame 434 based on the exposure setting. The short-exposure frame 434 may then be stored in the ZSL buffer 432. In some cases, a previous raw frame (or image) with the same gain may be stored in the ZSL buffer 432. In some cases, a previous raw frame (or image) with a similar gain or similar sensitivity may be stored in the ZSL buffer 432.

[0071] FIG. 7 illustrates the use of the frame processing and / or capture instruction system 400 during the PSL capture timing of the timing diagram 500 shown in FIG. 5. As noted above and shown in FIG. 5, PSL capture may occur in response to a shutter command or a capture command being received. In some cases, PSL capture may be performed as a second step in the frame processing and / or capture instruction process. For PSL capture, the frame processing and / or capture instruction system 400 may measure scene motion or movement (e.g., local movement corresponding to movement of one or more objects in the scene for which frames are being captured). In some examples, the LL engine 458, or other components of the frame processing and / or capture instruction system 400, may use collected sensor measurements (e.g., from one or more inertial measurement units (IMUs), such as gyrometers or gyroscopes, accelerometers, any combination thereof, and / or other IMUs) and / or preview motion analysis statistics to determine global movement corresponding to movement of the device or a camera of the device. For example, the IMU provides a means for mechanically estimating camera motion, which may be used by the frame processing and / or capture command system 400 to determine how much shake is present when a frame (e.g., an image or picture) is taken. For example, the frame processing and / or capture command system 400 may analyze IMU samples (e.g., gyrometer or gyroscope samples, accelerometer samples, etc.) related to an optical image stabilizer (OIS). In some examples, the frame processing and / or capture command system 400 may cross-reference IMU measurements with a video analytics tripod detection mechanism to determine if a tripod was used.

[0072] In some examples, the frame processing and / or capture instruction system 400 can determine an amount of local motion associated with one or more preview frames (e.g., ZSL frames captured during the LL mode timing of the timing diagram 500 shown in FIG. 5 or shortly thereafter during a sensor delay period). To determine the amount of local motion, the frame processing and / or capture instruction system 400 can analyze a temporal filter instruction (TFI), sometimes referred to herein as a motion map. For example, a TFI (i.e., a motion map) may be included in each frame (e.g., in each of one or more preview frames) as metadata, such as first TFI metadata for a first frame, second TFI metadata for a second frame, etc. In some cases, the TFI includes an image (referred to as a TFI image or a motion map image). In some examples, the TFI image can have the same resolution (with the same number of pixels horizontally and vertically) as the frame for which the TFI image is associated (e.g., the frame for which the TFI image is included as metadata). In some examples, the TFI image may have a lower resolution (with fewer pixels horizontally and / or vertically) compared to the frame to which the TFI image is associated. For example, each pixel of the TFI image may include a value indicating the amount of motion for each corresponding pixel of the frame associated with the TFI image. In some cases, each pixel of the TFI may include a reliability value indicating the reliability associated with the value indicating the amount of motion. In some cases, the TFI may represent an image area that may not be temporally blended after global motion compensation during preview and may indicate local motion components in the scene.

[0073] In some examples, a semi-global match (SGM) image may be used in addition to or as an alternative to the TFI image. The SGM is a residual motion vector map (indicating local 2D and / or 3D motion) after global compensation has been performed. The SGM may be used as a local motion indicator similar to that described above with respect to the TFI image. For example, the SGM may be obtained (e.g., as an input) after the global alignment has been corrected (e.g., using OIS).

[0074] Using information from the TFI and / or SGM images, the frame processing and / or capture instruction system 400 can predict local motion during frame capture. As described in more detail herein, a combination of global motion (e.g., based on sensor measurements such as gyroscope measurements) and local motion (e.g., as indicated by the TFI), referred to as a final motion indication (FMI), can be determined.

[0075] The AEC engine 464 can set the long-term AEC to be used to capture the long-term exposure frame. For example, the LL engine 458 can provide the AEC engine 464 with a new combination of exposure and gain (e.g., a single exposure can reach 1 second on a tripod). The LL engine 458 can also determine the number of long-term frames. In some cases, as described herein, the LL engine 458 can determine the exposure and number of long-term frames based on local motion (e.g., indicated by TFI), global motion, or FMI. The sensor 430 can then capture a PSL image or frame (long-term exposure frame 446). In some examples, when capturing a PSL frame, an additional ZSL frame (short-term exposure frame) may be captured. The additional ZSL frame may be stored in the ZSL buffer 432. In some cases, the first AWB engine 450 can calculate an AWB from the first PSL frame captured by the sensor 430. The first AWB engine 450 may output the WB scaler 451 to the first MFxR engine 436 and / or to the second MFxR engine 452. After the first PSL frame is captured, the frame processing and / or capture instruction system 400 may begin the next portion of the capture process (e.g., the third step of the frame processing and / or capture instruction process, such as MFNR, as described with respect to FIG. 8 ).

[0076] FIG. 8 illustrates the use of the frame processing and / or capture instruction system 400 during short-term MFNR of the timing diagram 500 shown in FIG. 5 . For example, during PSL capture, the first MFxR engine 436 can perform MFNR on the short-exposure frame. In some cases, short-term MFNR may be performed as a third step of the frame processing and / or capture instruction process. In some examples, the first MFxR engine 436 can use a calculated WB (e.g., AWB from the first PSL frame) from the second step of the frame processing and / or capture instruction process. As described above, the frame processing and / or capture instruction system 400 can determine whether MFHDR mode is required to process the ZSL frame (e.g., the short-exposure frame 434). If the frame processing and / or capture instruction system 400 determines that MFHDR mode will be used, the first MFxR engine 436 can determine that MFNR should not be performed on the short-exposure frame.

[0077] FIG. 9 illustrates the use of the frame processing and / or capture instruction system 400 during the long-term MFNR and preview 508 portion of the timing diagram 500 shown in FIG. 5 . For example, the second MFxR engine can perform MFNR on the long-term exposure frame and output the frame for preview. In some cases, long-term MFNR and preview may be performed as a fourth step of the frame processing and / or capture instruction process. To process the long-term frame using MFNR, the second MFxR engine 452 can use the calculated WB (e.g., the AWB from the first PSL frame) from the second step of the frame processing and / or capture instruction process. In some examples, if the frame processing and / or capture instruction system 400 determines that a non-MFHDR mode will be used (when MFHDR is not required), the second MFxR engine 452 can use global tone mapping (GTM) in addition to local tone mapping (LTM). In some cases, the GTM applies a gain per pixel according to the pixel's luma value. In some cases, the LTM applies different gain values ​​according to the brightness of the region. The second MFxR engine 452 may continue to process each synthesized frame or image using a Position Processing Segment (PPS) (e.g., a location in the processing pipeline where local tone mapping, sharpening, color correction, upscaling, etc. are performed). In some cases, the synthesized frame may be the result of multiple fused frames being combined into a single image or frame (e.g., based on the application of MFNR and / or MFHDR processes). In some examples, in addition to IPE_TF_FULL_OUT (MFNR NPS out), the second MFxR engine 452 may continue to perform PPS on the IPE_DISP_OUT result. In some cases, each of the long exposure frames 446 may have a slightly stronger LTM (e.g., whitening effect).The second MFxR engine 452 can send each PPS result for display as a preview frame 454, as described above. In some cases, after all PSL frames (long exposure frames 446) have been captured by the frame processing and / or capture instruction system 400, the preview is returned to the ZSL buffer 432.

[0078] FIG. 10 illustrates the use of the frame processing and / or capture instruction system 400 during the WB refinement 510 portion of the timing diagram 500 shown in FIG. 5 . The WB refinement 510 can include refinement of the WB so that the long-exposure frame 446 produces better AWB statistics. In some cases, the WB refinement can be performed as a fifth step in the frame processing and / or capture instruction process. To produce better AWB statistics for the long-exposure frame 446, the LL engine 458 can perform an “inverse ISP” on the long-term MFNR result (e.g., the blended frame 456 from the MFxR engine 452) to generate improved AWB statistics. For example, to recalculate the WB coefficients in some cases, the frame must be an AWB-compatible image. In some cases, the inverse ISP function can include the reversal of the operation performed on the raw frame or image, resulting in a linear frame or image (e.g., a linear RGB frame or image). The resulting linear frame can be used to regenerate the statistics for the improved SNR frame. The inverse ISP can result in better statistics for all subsequent adjustments. Using the refined / improved AWB statistics for the long exposure frame, the second AWB engine 460 can calculate an improved WB scaler (e.g., as part of the WB scaler 461). In some examples, the LL engine 458 can calculate the improved WB scaler.

[0079] FIG. 11 illustrates the use of the frame processing and / or capture instruction system 400 during the MFHDR and post-processing 512 portion of the timing diagram 500 shown in FIG. 5 . In some cases, the MFHDR and post-processing 512 may be performed as a sixth step of the frame processing and / or capture instruction process. For example, the MFHDR engine 440 may use the improved WB scaler 461 from the second AWB engine 460 (e.g., determined during the fifth step of the frame processing and / or capture instruction process). In some examples, if the frame processing and / or capture instruction system 400 determines that a non-MFHDR mode will be used (MFHDR is not required), the frame processing and / or capture instruction system 400 may output only PPS frames to the post-IPE. In some examples, the post-IPE may be used to smooth out “chroma stain.” In some examples, the post-IPE includes a machine learning system (e.g., one or more neural network systems). For example, a machine learning based image signal processor may be used as a post-IPE for global frame or image refinement.

[0080] As mentioned above, in some implementations, a frame processing and / or capture instruction system (e.g., an LL engine) can determine scene motion (also referred to as intra-scene motion or local motion), such as based on the movement of one or more objects within a scene for which a frame or image (e.g., a preview or short-exposure frame or image) is being captured. In some cases, as mentioned above, the LL engine or other components of the frame processing and / or capture instruction system can evaluate motion using collected sensor measurements (e.g., from an inertial measurement unit (IMU), such as a gyroscope or gyrometer, an accelerometer, and / or other IMU) and / or preview CVP motion analysis statistics. In some cases, a motion recognition algorithm can be used to enhance the execution of the frame processing and / or capture instruction system and processes described herein. The motion recognition algorithm can optimize noise to motion blur for the time between shots. In some cases, the motion recognition algorithm can perform global motion analysis (e.g., based on camera movement) and / or local motion analysis (e.g., based on object movement within the scene) to determine an indication of motion. The local motion analysis can be based on a temporal filter instruction (TFI). In some cases, the TFI can be an image having a pixel value for each pixel indicating the amount of motion for each pixel (e.g., whether each pixel has motion or no motion and / or how much motion). In some cases, as described above, each pixel of the TFI can include a reliability value indicating the reliability associated with the value indicating the amount of motion. The TFI is sometimes referred to as a motion map. In some cases, the TFI can be provided as part of a native camera flow that produces a motion vector map, which can be used by a frame processing and / or capture instruction system (e.g., by an LL engine).In some cases, a TFI may include sparse motion vectors (e.g., undistorted and unstabilized or stabilized) indicating the amount of motion per pixel (e.g., in the horizontal and vertical directions), a dense motion map (e.g., undistorted and unstabilized) having motion vectors per pixel, and / or a distortion correction grid. In some examples, the local motion indication (for a given TFI) may be based on ghost detection, such as by averaging the amount of ghosts detected during the temporal filtering process. In some examples, the local motion indication (for a given TFI) may be based on residual dense motion map averaging and / or counting of significant motion vectors.

[0081] Based on the analysis of the global motion, the frame processing and / or capture command system can compensate for the global motion (e.g., reduced using image stabilization techniques, such as by using an optical image stabilizer (OIS)). In some cases, the frame processing and / or capture command system can analyze the local motion indicated by the TFI (in addition to or as an alternative to the global motion), such as to determine whether motion blur needs to be reduced (e.g., if the local motion exceeds a motion threshold). Based on the local and / or global motion analysis, the motion recognition algorithm can optimize one or more 3A settings, such as exposure parameters (e.g., exposure duration and / or gain), automatic white balance (AWB), automatic exposure control (AEC), and autofocus, and / or other parameters. The motion recognition algorithm can be used for low lighting conditions, very low lighting conditions, standard lighting conditions, and / or other lighting conditions. In some cases, an optional machine learning system can be used for texture and noise improvement, as described above.

[0082] FIG. 12 is an illustration of a graph 1200 plotting motion versus exposure time and number of frames. The x-axis of graph 1200 plots the amount of motion. Line 1202 represents exposure, and line 1204 represents number of frames. The values ​​in graph 1200 (or other values) can be used to determine a long exposure time and number of frames (e.g., for input into MFNR, MMF, etc.) based on the motion indicated by the TFI (local motion), global motion, or combined motion determined using global motion and local motion (based on the TFI). If the motion is determined to be small (e.g., a small amount of motion is detected, such as motion less than a motion threshold), the number of frames is reduced and the exposure (e.g., exposure time, aperture, etc.) is increased. When more motion (e.g., an amount of motion greater than a motion threshold), corresponding to moving from left to right in graph 1200, is detected, the exposure (e.g., exposure time, aperture, etc.) is reduced until a minimum exposure limit is reached, which can be used to achieve suitable motion blur results. In some cases, gain may be adjusted in addition to or instead of exposure (e.g., exposure time, aperture, etc.) based on the determined motion. In some cases, the minimum exposure limit may be equal to the long exposure time used for the preview image / frame (e.g., because the frame processing and / or capture command system may not expose frames or images shorter than the exposure used for the short exposure / preview / displayed frame). Additionally, when more motion is determined, the number of frames (corresponding to the increased number of frames captured) is increased to compensate for the frame brightness resulting from the shorter exposure (which increases the gain).

[0083] In some cases, the motion shown on the x-axis of graph 1200 may correspond to both local and global motion (e.g., a combination of local and global motion). For example, the frame processing and / or capture command system can calculate the global motion and the local motion separately and apply weights to the local and global motion (e.g., using a first weight for the local motion value and a second weight for the global motion value) to generate a final motion instruction. The frame processing and / or capture command system can use the final motion instruction to determine how much to reduce or increase the exposure (e.g., exposure time, aperture, etc.) and / or gain, and how much to reduce or increase the number of frames (e.g., for output to MFNR, MMF, etc.).

[0084] In some examples, the frame processing and / or capture command system can determine a global motion (referred to as a global motion indication or GMI) within the range [0,1]. In such examples, a value of 0 can indicate no motion for the pixel, and a value of 1 can indicate maximum motion for the pixel. The frame processing and / or capture command system can determine a local motion (referred to as a local motion indication or LMI) within the range [0,1], where a value of 0 indicates no motion and a value of 1 indicates maximum motion. In some cases, the LMI can be calculated by cropping the TFI image to some extent (e.g., to reduce the effect from global motion), averaging the cropped map, normalizing the value, and applying an exponent to reflect sensitivity. The LMI weight (referred to as LMI_weight) within the range [0,1] represents how sensitive the frame processing and / or capture command system is to the LMI. One exemplary LMI weight value is 0.4. A final motion instruction (FMI) may be determined based on local motion and global motion (in TFI). In one illustrative example, FMI may be determined as lin_blend(GMI, GMI*LMI, LMI_weight)^2, where lin_blend is a linear blending operation. In another illustrative example, FMI may be determined as lin_blend(GMI, GMI*LMI, LMI_weight), similar to the previous example, but without the nonlinear response (^2).

[0085] FIG. 13 illustrates an image 1302 (or frame) and a TFI image 1304. The motion shown in FIG. 13 is local motion (also called intra-scene motion). In image 1302, a person is waving. TFI image 1304 includes white pixels for portions of image 1302 that do not have motion and black pixels for portions of image 1302 that do have motion. The black pixels correspond to the portion of the user that is moving (the right hand) and some of the clouds in the background. The frame processing and / or capture command system can determine whether the motion indicated in TFI image 1304 (or the motion indicated by the FMI, which is based on the local motion and global motion of TFI image 1304) is greater than a motion threshold. An exemplary value for the motion threshold is 0.3, which indicates linear sensitivity to motion. For example, if the motion indicated by the TFI (or FMI) is 0.4, the motion is greater than the motion threshold of 0.3. If the frame processing and / or capture instruction system determines that the motion is less than the motion threshold, the exposure time and number of frames for capturing long exposure frames or images may remain unchanged. With reference to Figure 12, the frame processing and / or capture instruction system may determine a motion value of 0 when the motion is determined to be less than the motion threshold. With reference to Figure 12, the frame processing and / or capture instruction system may reduce the exposure (e.g., exposure time, aperture, etc.) and / or increase the number of frames (and thus increase the number of long exposure frames captured) when the motion is determined to be greater than the motion threshold. With reference to Figure 12, the frame processing and / or capture instruction system may determine a motion value of 0.4 based on a particular amount of motion when the motion is determined to be greater than the motion threshold.

[0086] FIG. 14A illustrates an example of a frame processing and / or capture instruction system 1400. One or more of the components of the frame processing and / or capture instruction system 1400 of FIG. 14 may be similar to components of the frame capture and processing system 100 of FIG. 1 and / or the frame processing and / or capture instruction system 400 of FIG. 4 and may perform similar operations as such components. As shown in FIG. 14A, inputs 1401 to a low lighting (LL) engine 1458 may include motion sensor data (e.g., from a gyrometer or gyroscope, an accelerometer, an IMU, and / or other sensors), zero shutter lag (ZSL) frames (preview / display frames that can be displayed in real time and go to the LL engine 1458), AEC decisions (e.g., including exposure settings) based on the ZSL frames, and display TFI statistics. In some cases, a per-frame histogram may also be provided to the LL engine 1458. It is not known how the user holds their device and position when the frames are captured. However, if the user holds the device in some manner at a first time point immediately (e.g., 0.5 seconds, 1 second, etc.) before selecting the shutter or capture option (e.g., pressing the shutter button or other option), it may be assumed that the user is holding the device in a manner similar to that when the user selects the capture option. Based on such an assumption, the LL engine 1458 may generate a new command and output the new command to the AEC engine 1464. The new command may include the number of frames, a long exposure value (corresponding to the exposure duration used to capture a long exposure frame), and a long-term DRC gain. The number of frames, the long exposure value, and the long-term DRC gain may be determined based on the values ​​shown in FIG. 12 or other similar values ​​based on motion. The AEC engine 1464 may perform AEC and output AEC information to the image sensor (e.g., sensor 430) to capture a long-exposure frame.

[0087] Long-term exposure frames may be stored in a long-term frame buffer 1410 (referred to as "buffer long-term" in FIG. 14A). The long-term frame buffer 1410 may include a single buffer or multiple long-term frame buffers. When long-term exposure frames (PSL frames) are not captured during the preview "real-time" pipeline, the long-term frame buffer 1410 may be considered an offline buffer. As described above with respect to FIG. 7, the long-term exposure images or frames stored in the long-term frame buffer 1410 are captured after a capture command is received. For example, assuming the LL engine 1458 instructs the AEC engine 1464 to apply a specific exposure duration, a specific gain, and a specific number of frames (e.g., 10 frames), when the capture command is received (e.g., based on a user selecting a capture option), the image sensor captures a specific number of long-exposure frames using the specific exposure and gain.

[0088] Similar to that described above and shown in FIG. 14A , a first frame 1411 (or image) from the long-term frame buffer 1410 can be output to the AWB engine 1408 for automatic white balance (AWB) calculation. In conventional systems, the frame used for AWB calculation is a frame from the preview (ZSL frame). However, the preview frame is poorer than the long-exposure frame in the long-term frame buffer 1410. For example, the preview frame is captured using a shorter exposure time compared to the long-exposure frame, and the long-exposure frame compensates for the brightness introduced by the increased exposure by reducing the gain. As a result, by using the long-exposure frame for AWB calculation, the SNR, statistics for AWB, and color in the output frame or image are better than when the preview frame is used by the AWB engine 1408 for AWB. In some cases, improvement can be enhanced (e.g., based on a larger SNR) when the long-exposure frame is used for AWB in low-light environments. Using the first long exposure frame 1411, the AWB engine 1408 performs AWB calculations and generates new AWB control parameters 1412. The AWB control parameters 1412 can then be used to apply automatic white balance. As shown in FIG. 14A, the AWB control parameters 1412 are output to the first multi-frame noise reduction (MFNR) engine 1402.

[0089] The long-exposure frames from the long-term frame buffer 1410 and the AWB control parameters 1412 are output to the first MFNR engine 1402. The first MFNR engine 1402 may be used to temporally blend or filter the long-exposure frames (to filter temporal noise from the long-exposure frames in the time domain) but may not perform spatial blending of the long-exposure frames to filter spatial noise from the long-exposure frames in the spatial domain. The first MFNR engine 1402 can perform temporal blending by spatially aligning a series of frames (or images) and averaging the values ​​of each pixel in the frames. The algorithm used by the first MFNR engine 1402 estimates relative motion between frames (because the frames were taken at different times) and aligns the frames so that pixels can be combined to improve the SNR. The second MFNR engine 1404 performs spatial blending (or filtering) and temporal blending (or filtering) of medium-exposure frames 1414 from one or more ZSL buffers 1413 used to store the medium-exposure frames 1414. The third MFNR engine 1406 performs spatial filtering and temporal filtering of short-exposure frames 1416 from one or more ZSL buffers 1415 used to store the medium-exposure frames 1414. The second MFNR engine 1404 and the third MFNR engine 1406 perform spatial blending to improve the frame by processing the pixels of each frame, for example, possibly by determining a measure (e.g., a statistical measure, pixel component value, average, or other measure) of each pixel with respect to neighboring pixels. In some examples, spatial blending can be used to perform edge-preserving noise reduction, which can be addressed using various algorithms, such as convolution kernels in the image or frame domain, processing in the frequency (or frequency-like) domain, wavelets, etc., among others.In some cases, the first MFNR engine 1402, the second MFNR engine 1404, and the third MFNR engine 1406 may be implemented by the same hardware using the same processing technique, but the hardware may have different tuning settings when implementing the first MFNR engine, the second MFNR engine, and the third MFNR engine. In some examples, the buffers used to store the medium exposure frames 1414 and the short exposure frames 1416 may be the same buffer or different buffers. In some cases, a storage mechanism other than a buffer may be used to store the short exposure frames, the medium exposure frames, and the long exposure frames (e.g., cache memory, RAM, etc.).

[0090] The first MFNR engine 1402 can also acquire or receive a TFI image (also called a TFI map) for the long-term exposure frame stored in the long-term frame buffer 1410. In some cases, as described above, the TFI image can indicate which pixels had motion and which pixels did not have motion (or the degree of motion for a pixel). The TFI image can also indicate whether temporal blending was applied to the pixel. For example, if temporal blending is applied, a ghosting effect may occur in the image, so if a pixel is indicated by the TFI image as having motion, temporal blending may not be applied. Thus, a pixel in the TFI image that indicates that the pixel in the captured frame has motion may also indicate that the pixel in the captured frame does not have temporal blending applied to it. Based on the motion indication, the first MFNR engine 1402 can integrate the TFI image and output an integrated TFI image 1418 (shown in FIG. 14A as an integrated DC4 TFI). For example, the goal is to reflect how many frames have been temporally blended for each pixel. The first stage may assume that each frame has the same noise variance and that there is no covariance between them. Variance processing and / or variance operations may be applied according to the TFI image. For each image or frame, the results may be stored in the combined TFI image 1418 (or TFI map). After all frames have been processed, the result of the processing is a pixel-by-pixel variance. For example, a variance map may be output (e.g., by the first MFNR engine 1402). The variance map may then be converted into a map indicating how many frames are blended per pixel. The term "DC4" in FIG. 14A for the combined TFI image 1418 indicates an image downscaled by a factor of 4 in each axis. For example, for an image with a total image size of 8000×4000, the DC4 size of the image is 2000×1000.The first MFNR engine 1402 may also output a blended single long-term exposure frame 1420 (shown in FIG. 14A as blended long-term 10b), as described above.

[0091] In some cases, the ML ISP 1422 can use the final ALGM map as input for each frame rather than the integrated TFI image 1418. For example, the ALGM is the same as the TFI integrated map before converting it into a "frame blend count" map as described herein. The ALGM can be generated by hardware for full resolution frames or images rather than downscaled (e.g., DC4) frames or images.

[0092] The blended long-exposure frame 1420 and the integrated TFI image 1418 are output to a machine learning-based image signal processor (ISP) 1422 (shown in FIG. 14A as an ML ISP node). An example of the ML ISP 1422 is shown in FIG. 15. In some examples, the ML ISP includes one or more neural network architectures, such as a convolutional neural network (CNN), as an illustrative example. FIG. 16 illustrates an example of a neural network for the ML ISP 1422. In one illustrative example, the neural network includes a UNet-like network with 3x3 convolution layers, parametric rectified linear unit (PReLU) activation, average pooling (AvgPool) for downsampling, and bilinear upsampling for upsampling. In some examples, training of the neural network may be performed based on selected sensor calibration data. In some cases, other training data may also be used.

[0093] The ML ISP 1422 performs spatial filtering on the long-exposure frame. For example, the ML ISP 1422 can perform spatial-domain edge-preserving noise filtering. The ML ISP 1422 can compensate for the amount of noise reduction using the blended input long-exposure frame 1420 and the integrated TFI image 1418, which can equalize the noise within the frame. In an illustrative example, for a given frame, if a portion of the frame has local motion above a motion threshold (as indicated by the integrated TFI image), that portion of the frame will not have temporal blending applied to it (otherwise it will have ghosting), as described above. As a result, the portion of the frame with local motion has more noise than other portions of the frame because temporal blending was not applied. Based on the increased noise in the portion of the frame with motion, the ML ISP 1422 can apply more spatial filtering to that portion of the frame. For another portion of the frame that does not have any motion (or motion below a motion threshold), the first MFNR engine 1402 can apply temporal blending that results in less noise for that portion of the frame. The ML ISP 1422 can perform less spatial filtering on parts of the frame that have little motion. The ML ISP 1422 can also process the frame to stabilize transitions between different portions of the frame to which different levels of spatial filtering are applied.

[0094] The ML ISP 1422 can provide better results than a conventional ISP with multiple filters or processing blocks (e.g., for noise removal, edge enhancement, color balancing, contrast, intensity adjustment, tone adjustment, sharpening, among others). Furthermore, tuning a conventional ISP can be cumbersome and time-consuming. Training the ML ISP 1422 may not take much time, based on the use of supervised or unsupervised learning techniques. An illustrative example of a deep neural network that may be used for the ML ISP 1422 is shown in FIG. 24. An illustrative example of a convolutional neural network (CNN) that may be used for the ML ISP 1422 is shown in FIG. 25. Examples of performing machine learning training are described below with respect to FIGS. 24 and 25.

[0095] The ML ISP 1422 outputs the filtered long-exposure frames 1424 to the multi-frame high dynamic range (MFHDR) engine 1440. The MFHDR engine 1440 applies MFHDR processing as previously described. As described above, the LL engine 1458 can determine whether MFHDR is used (MFHDR mode) or whether MFHDR is not used (non-MFHDR mode). In non-MFHDR mode, HDR functionality can be performed using a single frame. If the LL engine 1458 determines that MFHDR mode will be used, the input to the MFHDR engine 1440 uses two or more frames. Conventional systems perform MFHDR using alternating-frame MFHDR. Alternating-frame MFHDR uses an alternating sequence of frames, such as short-term, medium-term, long-term, short-term, medium-term, long-term, short-term, medium-term, long-term, etc. The sensor is configured to capture frames with different exposure durations (short-term, medium-term, long-term). However, latency is important when capturing frames (in low lighting or other conditions), and using additional short-term, medium-term, and long-term frames (e.g., 4, 10, 20, or other numbers of frames) increases latency. To address such issues, the LL engine 1458 uses preview frames (from the ZSL buffer) rather than using frames with alternating exposures. Preview frames are stored in the ZSL buffer before a shutter command is received, in which case they do not need to be captured after the shutter command is received and therefore do not add latency to the frame or image capture process. By using preview frames, the only frames captured during offline processing (after a shutter command is received) are long-exposure frames. The preview frames stored in the ZSL buffer have the same sensor configuration as the long-exposure frames in the long-term frame buffer.The frame settings for the preview and long exposure frames are therefore the same except for the exposure (the exposure for the preview frame is shorter than that for the long exposure frame).

[0096] As described above, the short-exposure frames and medium-exposure frames (from the ZSL buffer) are processed by the second MFNR engine 1404 and the third MFNR engine 1406, respectively, which may be conventional ISP processing blocks (they do not include ML ISP). When the MFHDR mode is used, the medium-exposure frames and short-exposure frames are processed by the second MFNR engine 1404 and the third MFNR engine 1406, respectively, and the processed medium-exposure frames and short-exposure frames are fused or combined inside the MFHDR engine 1440.

[0097] 14B illustrates another example of a frame processing and / or capture instruction system 1405. The frame processing and / or capture instruction system 1405 is similar to the frame processing and / or capture instruction system 1400 of FIG. 14A and includes like numerals indicating common components between the frame processing and / or capture instruction system 1400 and the frame processing and / or capture instruction system 1405. One difference between the frame processing and / or capture instruction system 1400 and the frame processing and / or capture instruction system 1405 is that the blended long exposure image 1447 is downscaled (to a smaller size) by a first downscale engine 1449 before being processed by the ML ISP 1422. The ALGM map 1448 is also downscaled (to a smaller size) by a second downscale engine 1451 before being processed by the ML ISP 1422. While an ALGM map 1448 is shown in Figure 14B, an integrated TFI image (e.g., integrated TFI image 1418 of Figure 14A) could be used instead. The ML ISP 1422 could process the downscaled and blended long-exposure image 1447 and the ALGM map 1448 (e.g., using spatial filtering) and output a filtered long-exposure frame to the upscale engine 1452, which could output the upscaled filtered long-exposure frame 1424. Another difference is that rather than using medium and short-exposure frames, the frame processing and / or capture command system 1405 of Figure 14B uses a short-exposure frame 1442 (e.g., from the PSL) using a single short-exposure 1441.

[0098] As mentioned above, FIG. 15 shows an example of an ML ISP from FIG. 14A and / or FIG. 14B. Based on the techniques described above, unique inputs are provided to the ML ISP. For example, the ML ISP can perform processing (e.g., spatial filtering or blending) based on motion indications from the TFI image and based on the amount of temporal processing that occurred in the previous stage (as indicated by the TFI image). As shown in FIG. 15, inputs to the ML ISP include the blended long-exposure frame, the integrated TFI image, user settings (shown as tuning configurations), the number of frames, gain, and white balance. Further details regarding exemplary neural networks (e.g., that may be used for ML ISP) are described below with reference to FIGS. 24 and 25.

[0099] 17A and 17B illustrate a frame processing and / or capture instruction system 1700 having an additional processing component 1710 used to refine the white balance processing of a frame. In some cases, the additional processing component 1710 may be part of and / or implemented by the AWB engine 1708. As described above, the AWB engine 1708 can generate AWB control parameters 1712 for the captured frame. The additional processing component 1710 may be used to further refine the AWB determined by the AWB engine 1708. For example, the frame processing and / or capture instruction system 1700 may capture 20 frames. 20 frames in the time domain can have a first-order effect on improving noise variance. For example, 20 frames may reduce noise variance by 10 to 20 times due to some noise profiles. The AWB engine 1708 can obtain values ​​in a 2D distribution and determine a linear correlation between the values. Reducing the noise variance effectively reduces scattering, resulting in more reliable statistics and better AWB decisions. The frame processing and / or capture instruction system 1700 can output frames resulting from the temporal blending performed by the first MFNR engine 1702 (which results in improved SNR due to noise reduction) back to the AWB engine 1708. The AWB engine 1708 can recalculate the AWB control parameters 1712 to obtain more accurate AWB parameters.

[0100] The additional processing component 1710 includes an inverse ISP for creating a frame or image of a specific format (a raw Bayer frame or image) that can be sent to an AWB hardware engine (shown in FIG. 17A as “Calc AWB control”). For example, the AWB hardware engine can be configured to process a frame or image having a specific format. The inverse ISP is described in more detail above. Using a frame having a specific format (e.g., a Bayer frame), the AWB hardware engine generates improved AWB control parameters (shown as WBC). The improved AWB control parameters are processed using a color correction matrix (CCM) engine (shown in FIG. 17A as “Calc CCM”). For example, the CCM engine can determine a difference or delta between the initial AWB control parameters and the improved AWB control parameters. The CCM engine can use the difference to generate a color correction matrix (CCM). The MFHDR engine 1740 can use the CCM to generate a final output frame or image.

[0101] The ML ISP 1722 of the frame processing and / or capture instruction system 1700 can use a combined TFI image 1718, similar to the ML ISP 1422 of FIGS. 14A and 14B, as shown in FIGS. 17A and 17B. In some cases, the ML ISP 1722 can use the final ALGM map as input for each frame rather than the TFI image. For example, as described above, the ALGM is the same as the TFI combined map before converting it into a “frame blend count” map, as described above. The ALGM can be generated by hardware for full-resolution frames or images rather than downscaled (e.g., DC4) frames or images. The first MFNR engine 1702 can tolerate a larger number of frames with shorter exposure times, leading to reduced motion blur and better texture and noise. The first MFNR engine 1702 can also use modified tuning, including disabling noise equalization. In some examples, the first MFNR engine 1702 can fuse multiple frames, such as 256 frames, to improve the SNR in the final blended frame 1720. Such an MFNR engine may sometimes be referred to as a massive multi-frame (MMF) engine. In some cases, the first MFNR engine 1702 may be followed by a stage that performs noise equalization. Referring to FIG. 17B, the MFNR-blended frame 1720 undergoes pseudo-inverse ISP, as described above, resulting in a linear raw Bayer frame (with RGB color components). The raw Bayer frame is the output for the AWB statistics regeneration and AWB algorithm, which results in an improved WB coefficient (WBC). The residual WB is calculated, converted to CCM, and sent to an MFHDR engine (including post-IPE) similar to that described above with respect to FIG. 17A.

[0102] FIG. 18 illustrates an example process for gradually displaying frames or images (providing an interactive preview). For example, as new frames are being buffered and / or processed, it may be beneficial to gradually display improved frames, which may allow a user to see how each frame contributes to quality improvement. A frame processing and / or capture instruction system (e.g., systems 400, 1400, 1405, 1700, or other systems) may use a preview frame (e.g., ZSL frame 1802) from the ZSL buffer and provide different inputs to a video motion compensated temporal filter (MCTF). For example, when an image capture option is selected and a capture or shutter command is received, starting from a given PSL frame (e.g., the second PSL frame), the frame processing and / or capture instruction system may change the temporal filtering (TF) blend mode, thereby switching between the previous and current frames (e.g., similar to MFNR). For example, as shown in FIG. 18 , a current frame 1804 and a previous frame 1806 are switched compared to a current frame 1808 and a previous frame 1810. For example, the current frame 1808 is output from the IFE, and the previous frame 1810 is output from the IPE. When a capture or shutter command is received, the current frame 1804 is output from the IPE, and the previous frame 1806 is output from the IFE. Switching between the previous frame and the current frame can be performed until the long frame capture is completed. “Previous” and “current” refer to the inputs to the temporal filtering process. Previous is denoted as “PREV” in FIG. 18 , and current is denoted as “CUR” in FIG. 18 . For example, a TF has three frame or image inputs: current, current_anr, and previous. Current is the current frame or image on top of which the previous image is blended and aligned.After all PSL frames have been collected, the frame processing and / or capture command system can switch back to MCTF blending. As a result, the preview display shows the target frame or image dynamically improving as new frames are acquired. As the image capture process progresses, the display becomes more visible and detail is restored.

[0103] One reason to switch to PSL frames (e.g., upon receiving a capture or shutter command) is due to the improved light sensitivity of PSL frames. As described above, light sensitivity (or exposure, image exposure, or image sensitivity) can be specified as a function of gain and exposure time or duration (e.g., light sensitivity = gain * exposure_time). For example, each incoming frame can be processed using an MCTF that can improve the signal-to-noise ratio (SNR), resulting in an interactive effect of preview frame improvement as frames are acquired. The preview results also provide an accurate "what you see is what you get" key performance indicator (KPI). Using such techniques, the preview display shows the target frame dynamically improving as new frames are acquired, without interruption to the preview pipeline (e.g., without the need to switch to a different mode).

[0104] 19 and 20 illustrate examples of raw temporal blending based on adding U and V channels to raw frames or image components. As described above, temporal filtering algorithms (e.g., applied by an IPE, such as a post-IPE or other component of a frame processing and / or capture command system) can be used to use consecutive frames and blend the frames in the time domain to improve the quality of the output frame or image, and thus improve the SNR of the frame. Existing algorithms for temporal filtering use frames in the YUV domain, which include pixels having a luminance (Y) component and a chrominance component (e.g., a chroma-blue (Cb) component and a chroma-red (Cr) component). Existing algorithms are also typically implemented in HW. Performing temporal filtering in software may result in greater latency and may result in inferior quality compared to hardware, due to the fact that hardware allows several paths and has greater processing power. Described herein are raw temporal filtering or blending systems and techniques that use existing YUV temporal blending hardware pipelines. Raw temporal filtering or blending systems and techniques operate by splitting a raw image or frame (having a color filter array (CFA) pattern) into individual color components (e.g., red (R), green (G), and blue (B) components) and treating each color component as a separate YUV frame or image by adding a U channel and a V channel to each color component. The U and V values ​​may contain redundant values ​​that are used to fill buffers from which the MFNR engine obtains frames for processing. Using raw frames allows blending to be performed at an earlier stage, avoiding various offset, clipping, and inaccurate destructive decision problems. Using raw frames also allows the ML ISP to process more operations.Many more operations than just noise reduction can be delegated with roughly the same cost (e.g. demosaicing, WB, tone mapping, among others).

[0105] As shown in FIG. 19 , a raw frame 1902 is provided as an input. An image sensor (e.g., image sensor 130 or 430) may be used to capture the raw frame 1902. The raw frame 1902 has a color filter array (CFA) pattern, such as a Bayer CFA pattern, that includes red, green, and blue color components. The same CFA pattern is repeated for the entire frame. A Bayer Processing Segment (BPS) 1904 receives the raw frame 1902. The BPS 1904 is a hardware block that performs various image processing operations, such as Phase Detection Pixel Correction (PDPC), Lens Shading Correction (LSC), DG, White Balance Correction (WBC), and Bin Correction (BinCorr), among others. In some cases, each of the operations (e.g., PDPC, LSC, WBC, etc.) may correspond to a filter in the BPS 1904 hardware. The output of the BPS 1904 is a Bayer 14 image or frame (represented using 14 bits). A digital signal processor (DSP) 1906 converts the frame from a 14-bit frame to a 10-bit frame (Bayer 10 frame). The DSP 1906 may perform the conversion using a gamma lookup table (LUT). In some cases, a processor other than a DSP may be used to convert the frame from 14 bits to 10 bits.

[0106] The frame processing and / or capture instruction system can divide or separate each color component of a frame into separate color components. Each color component is a frame of all pixels of that color component. For example, the red (R) component 1908 of a frame includes all of the red pixels from the raw frame and a resolution or dimension that is half the width and half the height of the frame (because the red component makes up half of the raw frame). The green (G) and blue (B) components have a resolution that is one-quarter the width and one-quarter the height of the raw frame. A plane 10 is a 10-bit single-channel frame or image (e.g., a grayscale frame). A plane 10 frame may be used because the system cannot distinguish between different channels based on the separation of the color channels.

[0107] The frame processing and / or capture instruction system adds the U and V channels to the R, G, and B color components to create a separate standard YUV frame or image. For example, the frame processing and / or capture instruction system can generate a YUV frame 1910 based on the R color channel by adding a U channel 1912 and a V channel 1914. The U and V values ​​added to different color components can include the same values ​​(e.g., 0 or 512 for U and 0 or 512 for V, where 512 is the center of the UV plane). The frame processing and / or capture instruction system adds the U and V channels so that the image is in the correct format for the MFNR algorithm. In some cases, once the U and V channels are added, the frame processing and / or capture instruction system converts the frame format to P010. The MFNR engine can perform a temporal blending process on the resulting YUV frame.

[0108] The exemplary configurations shown in Figures 19 and 20 are similar except for how the green channel is considered. In the example of Figure 19, each green channel Gr and Gb is processed as a separate YUV frame. In the example of Figure 20, the combined green channel 1914 is processed as a YUV frame. The combined green channel-based YUV frame can be upscaled (to a super-resolution frame) to accommodate more temporal blending. Figure 21 includes a frame 2102 resulting from raw temporal blending and a frame 2104 resulting from using a standard YUV frame. As shown by the comparison between frame 2102 and frame 2104, raw temporal filtering blending does not degrade the frame.

[0109] An advantage of using a color component-based YUV frame (e.g., YUV frame 1910) is that the YUV frame is smaller than a typical YUV frame and therefore more efficient to process. Another advantage is that if raw temporal blending is performed, the blended raw frame can be sent directly for AWB expansion by the AWB engine 1708 and / or additional processing component 1710 of the frame processing and / or capture instruction system 1700, in which case an inverse ISP is not required.

[0110] As mentioned above, the terms short, medium (or “mid”), safe, and long as used herein refer to relative characterizations between a first setting and a second setting. These terms do not necessarily correspond to defined ranges for particular settings. For example, a long exposure (e.g., a long exposure duration or a long exposure image or frame) simply refers to an exposure time that is longer than a second exposure (e.g., a short exposure or a medium exposure). In another example, a short exposure (e.g., a short exposure duration or a short exposure image or frame) refers to an exposure time that is shorter than a second exposure (e.g., a long exposure or a medium exposure). In yet another example, a mid exposure or a medium exposure (e.g., a medium exposure duration or a medium exposure image or frame) refers to an exposure time that is longer than a first exposure (e.g., a short exposure) and shorter than a second exposure (e.g., a long exposure or a medium exposure).

[0111] 22 is a flow diagram illustrating an example of a process 2200 for determining an exposure duration for a number of frames using the techniques described herein. At block 2202, the process 2200 includes obtaining a motion map for one or more frames. The one or more frames may be frames (called preview frames) obtained before a capture command to capture the number of frames is received, as shown in FIG. 14. In some examples, the preview frames may be from a ZSL buffer, as described above. In some cases, the preview frames may be short-exposure frames, as described above.

[0112] In some examples, a motion map may be obtained for each frame of one or more frames (e.g., a first motion map for a first frame, a second motion map for a second frame, etc.). For example, in some aspects, the motion map may be included as metadata in each frame. As described above, a motion map may also be referred to as a temporal filter indication (TFI). In some cases, the motion map includes an image (e.g., a TFI image as described above). In some examples, the TFI image may have the same resolution as one or more frames (and thus have the same number of pixels in the horizontal and vertical directions). In some examples, the TFI image may have a lower resolution (with fewer pixels in the horizontal and / or vertical directions) than one or more frames. For example, each pixel of the motion map image may include a value indicating the amount of motion for a corresponding pixel of a frame (e.g., a frame for which the motion map is included as metadata) from one or more frames associated with the motion map. In one illustrative example, the value for each pixel of the motion map (or TFI) image may be in the range [0,1], where a value of 0 indicates no motion for the pixel and a value of 1 indicates maximum motion for the pixel. Any other suitable range or value designation for the motion map (or TFI) image may be used.

[0113] At block 2204, process 2200 includes determining motion associated with one or more frames of the scene based on the motion map. The motion corresponds to movement of one or more objects in the scene relative to a camera used to capture the one or more frames. The motion may be referred to as local motion, as described herein. In some cases, motion may be determined for each pixel of the one or more frames by referencing values ​​of pixels in a motion map (e.g., a motion map image as described above) obtained (e.g., as metadata) for the frame.

[0114] In some cases, process 2200 may include determining global motion associated with a camera. For example, process 2200 may determine global motion based on one or more sensor measurements, such as measurements from a gyroscope or other inertial measurement unit (IMU) (e.g., an accelerometer, etc.) of a device used to perform process 2200. In some cases, global motion may be determined for each frame of one or more frames based on sensor measurements received during a time associated with each frame. For example, measurements from a gyroscope may include a vector of gyroscope samples along with timestamps collected during a particular frame. In one such example, process 2200 may determine global motion for a particular frame based on that vector.

[0115] In some examples, the process 2200 may include determining a final motion instruction based on the determined motion and a global motion. For example, the process 2200 may determine a final motion instruction based on a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion. In one illustrative example, the final motion instruction may be determined as lin_blend(GMI, GMI*LMI, LMI_weight)^2, where lin_blend is a linear blending operation. In another illustrative example, the FMI may be determined as lin_blend(GMI, GMI*LMI, LMI_weight).

[0116] At block 2206, process 2200 includes determining a number of frames and an exposure (e.g., exposure time or exposure duration) for capturing the number of frames based on the determined motion. In some examples, the determined exposure duration is based on the exposure duration (or exposure time) and a gain. As described above, process 2200 may include determining a global motion associated with the camera (e.g., based on one or more sensor measurements, such as gyroscope measurements). In such a case, at block 2206, process 2200 may include determining a number of frames and an exposure duration for capturing the number of frames based on the determined motion and the global motion. As further described above, process 2200 may include determining a final motion instruction based on the determined motion and the global motion. In such a case, at block 2206, process 2200 may include determining a number of frames and an exposure duration for capturing the number of frames based on the final motion instruction.

[0117] In some cases, the number of frames may include long exposure frames as described above. In some examples, as described above, the values ​​in graph 1200 of FIG. 12 (or other values) may be used to determine a long exposure time (or long exposure duration) and a number of frames based on the motion indicated by a motion map (or TFI image), based on global motion and / or based on a final motion indication. For example, in some examples, process 2200 may include determining, based on the determined motion and / or global motion, that an amount of motion in one or more frames is less than a motion threshold. For example, process 2200 may include determining, based on the final motion indication, that an amount of motion in one or more frames is less than a motion threshold. Based on the amount of motion in one or more frames being less than the motion threshold, process 2200 may include decreasing the number of frames for the number of frames and increasing the amount of exposure duration for the determined exposure duration. In some examples, process 2200 may include determining, based on the determined motion and / or global motion, that an amount of motion in one or more frames is greater than a motion threshold. For example, process 2200 may include determining, based on the final motion indication, that an amount of motion in one or more frames is greater than a motion threshold. Based on the amount of motion in one or more frames being greater than the motion threshold, process 2200 may include increasing the number of frames for that number of frames and decreasing the amount of exposure duration for the determined exposure duration.

[0118] At block 2208, process 2200 includes sending a request to capture the number of frames using the determined exposure duration. For example, a component of the frame processing and / or capture instruction system (e.g., low illumination engine 1458 or other component) can send a request to the MFNR and MMF, the MFHDR, the image sensor, the image signal processor, any combination thereof, and / or other components to capture the number of frames using the determined exposure duration.

[0119] In some aspects, the process 2200 includes performing temporal blending on the number of frames captured using the determined exposure duration to generate a temporally blended frame. In some cases, the process 2200 includes performing spatial processing on the temporally blended frames using a machine learning-based image signal processor (e.g., such as those shown in Figures 14A, 14B, 15, 16, 17A, 17B, 19, and / or 20). In some aspects, as described above, the machine learning-based image signal processor uses a motion map (e.g., TFI) as input to perform spatial blending on the temporally blended frames. For example, as shown in Figure 14A, the ML ISP node 1422 uses the DC4 TFI-integrated image 1418 as input to determine where motion exists in the blended long-exposure image 1420. The example of Figure 14B uses an ALGM map as input.

[0120] 23 is a flow diagram illustrating an example of a process 2300 for performing temporal blending on one or more frames using the techniques described herein. At block 2302, the process 2300 includes obtaining a raw frame. The raw frame includes a single color component for each pixel of the raw frame. In some aspects, the raw frame includes a color filter array (CFA) pattern, such as those shown in FIGS. 19 and 20.

[0121] At block 2304, the process 2300 includes dividing the raw frame into a first color component, a second color component, and a third color component. In some cases, the first color component includes a red color component, the second color component includes a green color component, and the third color component includes a blue color component. In some embodiments, the first color component includes all red pixels of the raw frame, the second color component includes all green pixels of the raw frame, and the third color component includes all blue pixels of the raw frame. For example, as shown in FIG. 19 , a raw image (from a plurality of raw images 1902) is divided into a red (R) component 1908, a green (G) component, and a blue (B) component. The R component 1908 of the raw image includes all of the red pixels from the raw image, in which case the R component 1908 has a resolution that is half the width and half the height of the raw image. The G component (which includes the Gr and Gb components) and the B component shown in Figure 19 each have a resolution that is one-quarter the width and one-quarter the height of the raw image. In the example shown in Figure 20, the G component (which combines the Gr and Gb components) has a resolution that is half the width and half the height of the raw image.

[0122] At block 2306, process 2300 includes generating a plurality of frames, at least in part, by adding at least a first chrominance value to a first color component, at least a second chrominance value to a second color component, and at least a third chrominance value to a third color component. For example, process 2300 may include generating a first frame, at least in part, by adding at least a first chrominance value to a first color component, generating a second frame, at least in part, by adding at least a second chrominance value to the second color component, and generating a third frame, at least in part, by adding at least a third chrominance value to the third color component. In some examples, the process 2300 may include generating a first frame, at least in part, by adding the first and second chrominance values ​​to a first color component, generating a second frame, at least in part, by adding the first and second chrominance values ​​to a second color component, and generating a third frame, at least in part, by adding the first and second chrominance values ​​to a third color component. In an illustrative example, referring to FIG. 19 , by adding a value for the U chrominance channel and a value for the V chrominance channel to the R component 1908, the U chrominance channel 1912 and the V chrominance channel 1914 are added to the R component 1908, resulting in an output frame that is processed by an MFNR engine for temporal filtering. In some aspects, the first chrominance value and the second chrominance value are the same value.

[0123] At block 2308, process 2300 includes performing temporal blending on the plurality of frames. For example, the MFNR engines shown in FIG. 19 and / or FIG. 20 may perform the temporal blending. In some aspects, to perform temporal blending on the plurality of frames, process 2300 may include temporally blending a first frame of the plurality of frames with one or more additional frames having a first color component, temporally blending a second frame of the plurality of frames with one or more additional frames having a second color component, and temporally blending a third frame of the plurality of frames with one or more additional frames having a third color component. For example, as shown in FIG. 19, a plurality of raw images 1902 are processed. A YUV image can be generated for each color component of each raw image (from the plurality of raw images 1902), resulting in multiple YUV images for each color component (e.g., multiple YUV images including an R color component from the raw image, multiple YUV images including a Gr color component from the raw image, multiple YUV images including a Gb color component from the raw image, and multiple YUV images including a B color component from the raw image). The multiple YUV images for each color component generated by the system of FIG. 19 can then be processed for temporal blending (e.g., by MFNR). For example, multiple YUV images including the R color component from the raw image can be temporally blended by MFNR, multiple YUV images including the Gr color component from the raw image can be temporally blended by MFNR, multiple YUV images including the Gb color component from the raw image can be temporally blended by MFNR, and multiple YUV images including the B color component from the raw image can be temporally blended by MFNR.

[0124] In some examples, the processes described herein (e.g., process 2200, process 2300, and / or other processes described herein) may be performed by a computing device or apparatus. In some examples, process 2200 and / or process 2300 may be performed by frame capture and processing system 100 of FIG. 1, frame processing and / or capture instruction system 400 of FIG. 4, frame processing and / or capture instruction system 1400 of FIG. 14A, frame processing and / or capture instruction system 1405 of FIG. 14B, frame processing and / or capture instruction system 1700 of FIG. 17A, the system of FIG. 19, and / or the system of FIG. 20. In another example, process 2200 and / or process 2300 may be performed by image processing device 105B of FIG. 1. In another example, process 2200 and / or process 2300 may be performed by a computing device or system having the architecture of computing system 2600 shown in FIG. 26. For example, a computing device having the architecture of computing system 2600 shown in FIG. 26 may include components of frame capture and processing system 100 of FIG. 1, frame processing and / or capture instruction system 400 of FIG. 4, frame processing and / or capture instruction system 1400 of FIG. 14A, frame processing and / or capture instruction system 1405 of FIG. 14B, frame processing and / or capture instruction system 1700 of FIG. 17A, the system of FIG. 19, and / or the system of FIG. 20, and may perform the operations of FIG. 22 and / or the operations of FIG. 23.

[0125] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, a vehicle or vehicular computing device, a robotic device, a television, and / or any other computing device having the resource capabilities to perform the processes described herein, including process 2200. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0126] Components of a computing device may be implemented in circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0127] Process 2200 and process 2300 are illustrated as logical flow diagrams, whose operations represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to perform the process.

[0128] Additionally, process 2200, process 2300, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0129] As described above, a machine learning-based image signal processor (e.g., the ML ISP of FIGS. 14A, 14B, 15, 16, and / or 17A) may be used in some cases. The ML ISP may include one or more neural networks. FIG. 24 is a block diagram illustrating an example of a neural network 2400, according to some examples. The neural network 2400 of FIG. 24 may be used to implement such an ML ISP. The neural network 2400 of FIG. 24 may be used to implement any of the operations described herein, such as those performed by any of the systems and techniques described above.

[0130] The input layer 2420 includes input data. In one illustrative example, the input layer 2420 can include data representing pixels of input images captured by one or more cameras. The images can be video frames. The neural network 2400 includes multiple hidden layers 2422a, 2422b, through 2422n. The hidden layers 2422a, 2422b, through 2422n include "n" hidden layers, where "n" is an integer greater than or equal to one. The number of hidden layers can be as many as required for a given application. The neural network 2400 further includes an output layer 2424 that provides output resulting from processing performed by the hidden layers 2422a, 2422b, through 2422n. In one illustrative example, the output layer 2424 can provide optical flow and / or weight maps for objects in the input video frames. In one illustrative example, the output layer 2424 can provide an encoded version of the input video frames.

[0131] Neural network 2400 is a multi-layered neural network of interconnected nodes. Each node can represent information. The information associated with a node is shared among different layers, with each layer retaining the information as it is processed. In some cases, neural network 2400 can include a feedforward network, in which there are no feedback connections where the output of the network is fed back to itself. In some cases, neural network 2400 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading at the input.

[0132] Information can be exchanged between nodes through node-to-node interconnections between various layers. Nodes in the input layer 2420 can activate a set of nodes in the first hidden layer 2422a. For example, as shown, each of the input nodes in the input layer 2420 is connected to each of the nodes in the first hidden layer 2422a. The nodes in the first hidden layer 2422a can transform each input node's information by applying an activation function to the input node information. The information derived from the transformation can then be passed to nodes in the next hidden layer 2422b, which can activate those nodes and perform their designated functions. Exemplary functions include convolution, upsampling, data transformation, and / or any other suitable functions. The output of the hidden layer 2422b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 2422n can activate one or more nodes in the output layer 2424, which provides the output. In some cases, a node in neural network 2400 (e.g., node 2426) is shown as having multiple output lines, but the node has a single output, and all lines shown as outputting from the node represent the same output value.

[0133] In some cases, each node or interconnections between nodes may have weights, which are sets of parameters derived from training the neural network 2400. Once the neural network 2400 is trained, it may be referred to as a trained neural network, which may be used to generate 2D optical flow, generate MS optical flow, generate weight maps, 2D warp frames based on the 2D optical flow, MS warp frames based on the MS optical flow, encode data, decode data, generate predicted frames, or any combination thereof. For example, the interconnections between nodes may represent information learned about the interconnected nodes. The interconnections may have tunable numerical weights that can be tuned (e.g., based on a training dataset), allowing the neural network 2400 to be adaptive to input and learn as more and more data is processed.

[0134] Neural network 2400 is pre-trained to process features from data in input layer 2420 using different hidden layers 2422a, 2422b, through 2422n to provide output through output layer 2424. In examples where neural network 2400 is used to identify objects in images, neural network 2400 may be trained using training data including both images and labels. For example, training images may be input into the network, with each training image having a label indicating the class of one or more objects in each image (essentially telling the network what the object is and what characteristics it has). In one illustrative example, the training images may include images of the number 2, in which case the label for the image may be [0 0 1 0 0 0 0 0 0 0].

[0135] In some cases, the neural network 2400 may adjust the node weights using a training process called backpropagation. Backpropagation may include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed during one training iteration. The process may be repeated for a number of iterations for each set of training images until the neural network 2400 is well trained enough that the layer weights are precisely tuned.

[0136] For the example of identifying objects in an image, the forward pass may include passing training images through neural network 2400. Before neural network 2400 is trained, the weights are first randomized. The image may include, for example, an array of numbers representing pixels of the image. Each number in the array may include a value from 0 to 255 representing the pixel intensity at that location in the array. In one example, the array may include a 28x28x3 array of numbers, with 28 rows and 28 columns of pixels and three color components (such as a red component, a green component, and a blue component, or a luma component and two chroma components).

[0137] For the first training iteration for the neural network 2400, due to the weights being randomly selected at initialization, the output may include values ​​that do not give preference to any particular class. For example, if the output is a vector with the probabilities that an object contains different classes, the probability values ​​for each of the different classes may be equal or at least very similar (e.g., for 10 possible classes, each class may have a probability value of 0.1). With the initial weights, the neural network 2400 cannot determine low-level features and therefore cannot make accurate decisions (e.g., of optical flow or weight mapping for a particular area of ​​a frame). A loss function may be used to analyze errors in the output. Any suitable loss function definition may be used. One example of a loss function includes the mean squared error (MSE). MSE is

[0138]

number

[0139] The loss is calculated by subtracting the predicted (output) solution from the actual solution and then squaring it. total may be set equal to the value of

[0140] The loss (or error) is large for the first training images because the actual values ​​are significantly different from the predicted outputs. The goal of training is to minimize the amount of loss so that the predicted outputs are the same as the training labels. The neural network 2400 can perform a backward pass by determining which inputs (weights) contributed most to the network's loss, and can adjust the weights so that the loss is smaller and eventually minimized.

[0141] To determine the weights that contributed most to the network's loss, the derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a particular layer) can be calculated. After the derivatives are calculated, a weight update can be performed by updating all of the weights of the filter. For example, the weights can be updated such that they change in the opposite direction of the gradient. The weight update can be

[0142]

number

[0143] where w denotes the weight and w i denotes the initial weights, and η denotes the learning rate. The learning rate can be set to any suitable value, with a higher learning rate inducing larger weight updates and a smaller value indicating smaller weight updates.

[0144] Neural network 2400 can include any suitable deep network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer with multiple hidden layers between the input and output layers. The hidden layers of a CNN include a series of convolutional layers, nonlinear layers, pooling layers (for downsampling), and fully connected layers. Neural network 2400 can also include any other deep network other than a CNN, such as an autoencoder, a deep belief net (DBN), or a recurrent neural network (RNN), among others.

[0145] 25 is a block diagram illustrating an example of a convolutional neural network (CNN) 2500, according to some examples. The input layer 2520 of the CNN 2500 includes data representing an image, such as an image captured by one of the one or more cameras 210. For example, the data may include an array of numbers representing pixels of the image, with each number in the array including a value from 0 to 255 representing the pixel intensity at that location in the array. Using the previous example from above, the array may include a 28×28×3 array of numbers, with 28 rows and 28 columns of pixels and three color components (e.g., a red component, a green component, and a blue component, or a luma component and two chroma components, etc.). The image is passed through a convolutional hidden layer 2522a, an optional nonlinear activation layer, a pooling hidden layer 2522b, and a fully connected hidden layer 2522c to obtain an output at the output layer 2524. 25, one skilled in the art will appreciate that multiple convolutional hidden layers, nonlinear layers, pooling hidden layers, and / or fully connected layers may be included in CNN 2500. As previously described, the output may generate a 2D optical flow, generate an MS optical flow, generate a weight map, 2D warp a frame based on the 2D optical flow, MS warp a frame based on the MS optical flow, encode data, decode data, generate a predicted frame, or a combination thereof.

[0146] The first layer of the CNN 2500 is a convolutional hidden layer 2522a. The convolutional hidden layer 2522a analyzes the image data of the input layer 2520. Each node of the convolutional hidden layer 2522a is connected to a region of nodes (pixels) in the input image, called the receptive field. The convolutional hidden layer 2522a can be thought of as one or more filters (each filter corresponds to a different activation map or feature map), and each convolutional iteration of a filter is a node or neuron in the convolutional hidden layer 2522a. For example, the region of the input image covered by a filter in each convolutional iteration is the receptive field of the filter. In one illustrative example, if the input image includes a 28x28 array and each filter (and corresponding receptive field) is a 5x5 array, there will be 24x24 nodes in the convolutional hidden layer 2522a. Each connection between a node and the receptive field for that node learns a weight and possibly an overall bias that each node learns to analyze its local, specific receptive field in the input image. Each node in hidden layer 2522a has the same weights and biases (called shared weights and shared biases). For example, a filter has an array of weights (numbers) and a depth the same as the input. The filter has a depth of 3 for the example video frame (according to the three color components of the input image). An illustrative example size of the filter array is 5 x 5 x 3, corresponding to the size of the node's receptive field.

[0147] The convolutional nature of the convolutional hidden layer 2522a results from the fact that each node in the convolutional layer is applied to its corresponding receptive field. For example, the filter in the convolutional hidden layer 2522a can start in the upper left corner of the input image array and convolve around the input image. As described above, each convolutional iteration of the filter can be considered a node or neuron in the convolutional hidden layer 2522a. In each convolutional iteration, the value of the filter is multiplied by a corresponding number of original pixel values ​​in the image (e.g., a 5x5 filter array is multiplied by a 5x5 array of input pixel values ​​in the upper left corner of the input image array). The multiplications from each convolutional iteration can be added together to obtain a total for that iteration or node. The process then continues at the next location in the input image according to the receptive field of the next node in the convolutional hidden layer 2522a. For example, the filter can move a certain step amount (called a stride) to the next receptive field. The stride can be set to 1 or another suitable amount. For example, if the stride is set to 1, the filter is moved one pixel to the right in each convolution iteration. Processing the filter at each unique location of the input volume produces a number representing the filter result for that location, resulting in an aggregate value being determined for each node in the convolutional hidden layer 2522a.

[0148] The mapping from the input layer to the convolutional hidden layer 2522a is called an activation map (or feature map). The activation map contains a per-node value representing the filter result at each location in the input volume. The activation map may include an array containing various aggregate values ​​resulting from each iteration of the filter on the input volume. For example, if a 5x5 filter is applied to each pixel of a 28x28 input image (with a stride of 1), the activation map will include a 24x24 array. The convolutional hidden layer 2522a may include several activation maps to identify multiple features in an image. The example shown in Figure 25 includes three activation maps. Using the three activation maps, the convolutional hidden layer 2522a can detect three different types of features, each detectable across the entire image.

[0149] In some examples, a nonlinear hidden layer may be applied after the convolutional hidden layer 2522a. The nonlinear layer may be used to introduce nonlinearity into a system that previously computed a linear operation. An illustrative example of a nonlinear layer is a rectified linear unit (ReLU) layer. The ReLU layer may apply a function f(x)=max(0,x) to all of the values ​​in the input volume, which changes all negative activations to 0. Thus, the ReLU can enhance the nonlinear characteristics of the CNN 2500 without affecting the receptive field of the convolutional hidden layer 2522a.

[0150] A pooling hidden layer 2522b may be applied after the convolutional hidden layer 2522a (and after the nonlinear hidden layer, if used). The pooling hidden layer 2522b is used to simplify the information in the output from the convolutional hidden layer 2522a. For example, the pooling hidden layer 2522b may take each activation map output from the convolutional hidden layer 2522a and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. Other forms of pooling functions, such as average pooling, L2-norm pooling, or other suitable pooling functions, may be used by the pooling hidden layer 2522a. A pooling function (e.g., a max pooling filter, an L2-norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 2522a. In the example shown in FIG. 25, three pooling filters are used for the three activation maps in the convolutional hidden layer 2522a.

[0151] In some examples, max pooling may be used by applying a max pooling filter (e.g., having a size of 2×2) with a stride (e.g., equal to the dimension of the filter, such as a stride of 2) to the activation map output from the convolutional hidden layer 2522a. The output from the max pooling filter contains the maximum number of all subregions around which the filter convolves. Using a 2×2 filter as an example, each unit in the pooling layer can summarize the region of 2×2 nodes (each node is a value in the activation map) in the previous layer. For example, four values ​​(nodes) in the activation map are analyzed by 2×2 max pooling at each iteration of the filter, and the maximum value from the four values ​​is output as the “max” value. If such a max pooling filter is applied to an activation filter from the convolutional hidden layer 2522a with dimensions of 24×24 nodes, the output from the pooling hidden layer 2522b is an array of 12×12 nodes.

[0152] In some examples, an L2 norm pooling filter may also be used, which involves calculating the square root of the sum of the squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (rather than calculating the maximum value as is done in max pooling) and using the calculated value as the output.

[0153] Intuitively, a pooling function (e.g., max pooling, L2 norm pooling, or other pooling functions) determines whether a given feature is found anywhere within a region of the image and discards the exact location information. This can be done without affecting the outcome of feature detection because, once the feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max pooling (as well as other pooling methods) offers the advantage of having far fewer features pooled, thus reducing the number of parameters required in later layers of CNN2500.

[0154] The final layer of connections in the network is a fully connected layer that connects every node from pooling hidden layer 2522b to every one of the output nodes in output layer 2524. Using the example above, the input layer includes 28x28 nodes that encode pixel intensities of the input image, convolutional hidden layer 2522a includes 3x24x24 hidden feature nodes based on applying 5x5 local receptive fields (for the filters) to three activation maps, and pooling hidden layer 2522b includes a layer of 3x12x12 hidden feature nodes based on applying max-pooling filters to 2x2 regions across each of the three feature maps. Extending this example, output layer 2524 could include 10 output nodes. In such an example, every node in the 3x12x12 pooling hidden layer 2522b is connected to every node in output layer 2524.

[0155] The fully connected layer 2522c can take the output of the previous pooling hidden layer 2522b (which should represent an activation map of high-level features) and determine the features that are most correlated to a particular class. For example, the fully connected layer 2522c can determine the high-level features that are most strongly correlated to a particular class and can include weights (nodes) for the high-level features. The products between the weights of the fully connected layer 2522c and the pooling hidden layer 2522b can be calculated to obtain probabilities for different classes. For example, if the CNN 2500 is being used to generate optical flow, there will be large values ​​in the activation map that represents the high-level feature of the movement of visual elements from one frame to another.

[0156] In some examples, output from output layer 2524 may include an M-dimensional vector (in the previous example, M=10), where M may include data corresponding to possible motion vector directions in the optical flow, possible motion vector amplitudes in the optical flow, possible weight values ​​in a weight map, etc. In one illustrative example, if a nine-dimensional output vector represents ten different possible values ​​as [0 0 0.05 0.8 0 0.15 0 0 0 0], the vector indicates that there is a third value with a 5% probability, a fourth value with an 80% probability, and a sixth value with a 15% probability. The probabilities for the possible values ​​may be considered as confidence or certainty levels for that value (e.g., for that motion vector direction, for that motion vector amplitude, for that weight value, etc.).

[0157] 26 is a diagram illustrating an example of a system for implementing some aspects of the present technology. In particular, FIG. 26 illustrates an example of a computing system 2600, which may be any computing device comprising, for example, an internal computing system, a remote computing system, a camera, or any component thereof, where components of the system communicate with each other using a connection 2605. The connection 2605 may be a physical connection using a bus, or a direct connection to a processor 2610, such as in a chipset architecture. The connection 2605 may also be a virtual connection, a network connection, or a logical connection.

[0158] In some embodiments, computing system 2600 is a distributed system in which the functionality described in this disclosure may be distributed among one or more data centers, peer networks, etc. In some embodiments, one or more of the system components described represent many components, each performing some or all of the functionality described for it. In some embodiments, the components may be physical or virtual devices.

[0159] The exemplary system 2600 includes at least one processing unit (CPU or processor) 2610, as well as connections 2605 coupling various system components to the processor 2610, including system memory 2615, such as read-only memory (ROM) 2620 and random access memory (RAM) 2625. The computing system 2600 may include a cache 2612 of high-speed memory directly connected to the processor 2610, in close proximity to the processor 2610, or integrated as part of the processor 2610.

[0160] Processor 2610 can include any general-purpose processor, as well as hardware or software services, such as services 2632, 2634, and 2636 stored in storage device 2630, configured to control processor 2610, and special-purpose processors where software instructions are embedded into the actual processor design. Processor 2610 may essentially be a completely self-contained computing system including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0161] To enable user interaction, computing system 2600 includes input device(s) 2645, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. Computing system 2600 may also include output device(s) 2635, which may be one or more of several output mechanisms. In some cases, a multi-modal system may allow a user to provide multiple types of input / output to communicate with computing system 2600. Computing system 2600 may include a communication interface 2640, which may generally govern and manage user input and system output.The communication interface may be an audio jack / plug, a microphone jack / plug, a Universal Serial Bus (USB) port / plug, an Apple® Lightning® port / plug, an Ethernet port / plug, an optical fiber port / plug, a proprietary wired port / plug, a Bluetooth® wireless signal transfer, a Bluetooth® Low Energy (BLE) wireless signal transfer, an IBEACON® wireless signal transfer, a radio frequency identification (RFID) wireless signal transfer, a near field communication (NFC) wireless signal transfer, a dedicated short range communication (DSRC) wireless signal transfer, an 802.11 The communication interface 2640 may perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including those utilizing Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or any combination thereof. The communication interface 2640 may also include one or more global navigation satellite system (GNSS) receivers or transceivers used to determine the location of the computing system 2600 based on reception of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite System (BDS), and the European Galileo GNSS.There is no constraint to operating on any particular hardware configuration, and therefore the basic features herein may be easily substituted for improved hardware or firmware configurations as they are developed.

[0162] The storage device 2630 may be a non-volatile and / or non-transitory and / or computer readable memory device, such as a hard disk, or a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, a flash memory, a memristor memory, any other solid state memory, a compact disc read only memory (CD-ROM) optical disk, a rewritable compact disc (CD) optical disk, a digital video disc (DVD) optical disk, a Blu-ray Disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, It may be another type of computer-readable medium capable of storing data that is accessible by a computer, such as an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or combinations thereof.

[0163] The storage device(s) 2630 may include software services, servers, services, etc. that cause the system to perform functions when code defining such software is executed by the processor(s) 2610. In some embodiments, hardware services that perform particular functions may include software components stored in computer-readable media in association with the necessary hardware components, such as the processor(s) 2610, connections 2605, output devices 2635, etc., to perform the functions.

[0164] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, storing, or transporting instructions and / or data. Computer-readable media may include non-transitory media on which data can be stored and that do not include carrier waves and / or transitory electronic signals propagating wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. A computer-readable medium may store code and / or machine-executable instructions, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted using any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0165] In some embodiments, computer-readable storage devices, media, and memories may include cable or wireless signals containing bitstreams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0166] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the present embodiments may be practiced without these specific details. For clarity of explanation, in some instances, the present technology may be presented as including individual functional blocks, including functional blocks comprising devices, device components, method steps or routines implemented in software, or combinations of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, or other components may be shown as components in block diagram form to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0167] Individual embodiments may be described above as a process or method that is depicted as a flow diagram, a flowchart, a data flow diagram, a structure diagram, or a block diagram. While the flow diagrams may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagrams. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.

[0168] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, special-purpose computer, or processing device to perform a certain function or group of functions, or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, etc.

[0169] Devices that implement processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., a computer program product) to perform the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in a peripheral device or add-in card. Such functionality may also be implemented on a circuit board, in different chips, or on different processes executing within a single device, as further examples.

[0170] The instructions, media for carrying such instructions, computing resources for executing the instructions, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.

[0171] In the foregoing description, aspects of the present application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Accordingly, while exemplary embodiments of the present application have been described in detail herein, it should be understood that the inventive concepts may be otherwise embodied and employed in various ways, and that the appended claims are intended to be construed to include such variations except insofar as limited by the prior art. Various features and aspects of the above-described applications may be used individually or jointly. Moreover, the embodiments may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present specification. Accordingly, the specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, methods have been described in a particular order. It should be appreciated that in alternative embodiments, methods may be performed in an order different from that described.

[0172] Those skilled in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of this description.

[0173] When a component is described as being "configured to" perform some operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof.

[0174] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or that is in communication, either directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).

[0175] Claim language or other language reciting "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language reciting "at least one of A and B" means A, B, or A and B. As another example, claim language reciting "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" can mean A, B, or A and B, and additionally can include items not listed in the set of A and B.

[0176] The various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0177] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium, such as a random access memory (RAM) such as a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a nonvolatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, etc. The techniques may additionally or alternatively be realized at least in part by a computer-readable communication medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.

[0178] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein, may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated software or hardware modules configured for encoding and decoding, or may be incorporated into a combined video encoder / decoder (codec).

[0179] Exemplary aspects of the present disclosure include the following.

[0180] Aspect 1. An apparatus for determining exposure durations for a number of frames. The apparatus includes a memory (e.g., implemented in circuitry) and one or more processors (e.g., one processor or multiple processors) coupled to the memory. The one or more processors are configured to: obtain a motion map for one or more frames; determine motion associated with the one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects in the scene relative to a camera used to capture the one or more frames; determine a number of frames and an exposure for capturing the number of frames based on the determined motion; and send a request to capture the number of frames using the determined exposure durations.

[0181] Embodiment 2. The apparatus of embodiment 1, wherein one or more frames are acquired before a capture command to capture the number of frames is received.

[0182] Embodiment 3. The apparatus of any one of embodiments 1 or 2, wherein the one or more processors are configured to perform temporal blending on the number of frames captured using the determined exposure duration to generate a temporally blended frame.

[0183] Embodiment 4. The apparatus of embodiment 3, wherein the one or more processors are configured to perform spatial processing on the temporally blended frames using a machine learning-based image signal processor.

[0184] Embodiment 5. The apparatus of embodiment 4, wherein the machine learning based image signal processor uses the motion map as input to perform spatial processing on the temporally blended frames.

[0185] Embodiment 6. The apparatus of any one of embodiments 1-5, wherein the determined exposure duration is based on gain.

[0186] Embodiment 7. The apparatus of any one of embodiments 1-6, wherein the motion map includes an image, and each pixel of the image includes a value indicative of at least one of an amount of motion for each pixel and a confidence value associated with the amount of motion.

[0187] Embodiment 8. The apparatus of any one of embodiments 1-7, wherein the one or more processors are configured to determine global motion associated with the camera based on one or more sensor measurements, and the number of frames and the exposure duration for capturing the number of frames are determined based on the determined motion and the global motion.

[0188] Embodiment 9. The apparatus of any one of embodiments 1 to 8, wherein, to determine a number of frames and an exposure duration for capturing the number of frames based on the determined motion and global motion, the one or more processors are configured to: determine a final motion instruction based on the determined motion and global motion; and determine, based on the final motion instruction, the number of frames and an exposure duration for capturing the number of frames.

[0189] Embodiment 10. The apparatus of embodiment 9, wherein to determine a final motion instruction based on the determined motion and the global motion, the one or more processors are configured to determine a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion.

[0190] Embodiment 11. The apparatus of any one of embodiments 9 or 10, wherein the one or more processors are configured to determine, based on the final motion indication, that the amount of motion in one or more frames is less than a motion threshold, and, based on the amount of motion in the one or more frames being less than the motion threshold, reduce the number of frames for that number of frames and increase the amount of exposure for the determined exposure duration.

[0191] Embodiment 12. The apparatus of any one of embodiments 9 or 10, wherein the one or more processors are configured to determine, based on the final motion indication, that an amount of motion in one or more frames is greater than a motion threshold, and, based on the amount of motion in the one or more frames being greater than the motion threshold, increase the number of frames for that number of frames and decrease the amount of exposure for the determined exposure duration.

[0192] Embodiment 13. The apparatus of any one of embodiments 1-12, further comprising at least one of a camera configured to capture at least one frame and a display configured to display the at least one frame.

[0193] Aspect 14. An apparatus for performing temporal blending on one or more frames. The apparatus includes a memory (e.g., implemented in circuitry) configured to store one or more frames, and one or more processors (e.g., one processor or multiple processors) coupled to the memory. The one or more processors are configured to: obtain a raw frame, the raw frame including a single color component for each pixel of the raw frame; divide the raw frame into a first color component, a second color component, and a third color component; generate a plurality of frames, at least in part, by adding at least a first chrominance value to the first color component, at least a second chrominance value to the second color component, and at least a third chrominance value to the third color component; and perform temporal blending on the plurality of frames.

[0194] Embodiment 15. The device of embodiment 14, wherein the raw frame includes a color filter array (CFA) pattern.

[0195] Embodiment 16. The device of any one of embodiments 14 or 15, wherein the first color component includes a red color component, the second color component includes a green color component, and the third color component includes a blue color component.

[0196] Embodiment 17. The device of any one of embodiments 14-16, wherein the first color component includes all red pixels of the raw frame, the second color component includes all green pixels of the raw frame, and the third color component includes all blue pixels of the raw frame.

[0197] Embodiment 18. The device of any one of embodiments 14-17, wherein, to generate the plurality of frames, the one or more processors are configured to generate a first frame, at least in part, by adding at least a first chrominance value to a first color component, generate a second frame, at least in part, by adding at least a second chrominance value to a second color component, and generate a third frame, at least in part, by adding at least a third chrominance value to a third color component.

[0198] Embodiment 19. The apparatus of embodiment 18, wherein to generate a first frame, the one or more processors are configured to apply the first chrominance value and the second chrominance value to a first color component, to generate a second frame, the one or more processors are configured to apply the first chrominance value and the second chrominance value to a second color component, and to generate a third frame, the one or more processors are configured to apply the first chrominance value and the second chrominance value to a third color component.

[0199] Embodiment 20. The apparatus of embodiment 19, wherein the first chrominance value and the second chrominance value are the same value.

[0200] Embodiment 21. The device of any one of embodiments 14-20, wherein, to perform temporal blending on the plurality of frames, the one or more processors are configured to temporally blend a first frame of the plurality of frames with one or more additional frames having a first color component, temporally blend a second frame of the plurality of frames with one or more additional frames having a second color component, and temporally blend a third frame of the plurality of frames with one or more additional frames having a third color component.

[0201] Embodiment 22. The device of any one of embodiments 14 to 21, wherein the device is a mobile device.

[0202] Embodiment 23. The apparatus of any one of embodiments 14-22, further comprising a camera configured to capture one or more frames.

[0203] Embodiment 24. The device of any one of embodiments 14-23, further comprising a display configured to display one or more frames.

[0204] Embodiment 25. An apparatus comprising the apparatus of any one of embodiments 1 to 13 and the apparatus of any one of embodiments 14 to 24.

[0205] Aspect 26. A method for determining an exposure duration for a number of frames. The method comprises obtaining a motion map for one or more frames, determining motion associated with one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects in the scene relative to a camera used to capture the one or more frames, determining a number of frames and an exposure for capturing the number of frames based on the determined motion, and sending a request to capture the number of frames using the determined exposure duration.

[0206] Embodiment 27. The method of embodiment 26, wherein one or more frames are acquired before a capture command to capture the number of frames is received.

[0207] Embodiment 28. The method of any one of embodiments 26 or 27, further comprising performing temporal blending on the number of frames captured using the determined exposure duration to generate a temporally blended frame.

[0208] Embodiment 29. The method of embodiment 28, further comprising performing spatial processing on the temporally blended frames using a machine learning-based image signal processor.

[0209] Embodiment 30. The method of embodiment 29, wherein the machine learning based image signal processor uses the motion map as an input to perform spatial processing on the temporally blended frames.

[0210] Embodiment 31. The method of any one of embodiments 26-30, wherein the determined exposure duration is based on gain.

[0211] Embodiment 32. The method of any one of embodiments 26-31, wherein the motion map includes an image, and each pixel of the image includes a value indicative of at least one of an amount of motion for each pixel and a confidence value associated with the amount of motion.

[0212] Embodiment 33. The method of any one of embodiments 26-32, further comprising determining global motion associated with the camera based on one or more sensor measurements, wherein the number of frames and the exposure duration for capturing the number of frames are determined based on the determined motion and the global motion.

[0213] Embodiment 34. The method of any one of embodiments 26 to 33, further comprising determining a final motion instruction based on the determined motion and global motion, wherein the number of frames and the exposure duration for capturing the number of frames are determined based on the final motion instruction.

[0214] Embodiment 35. The method of embodiment 34, wherein the final motion instruction is based on a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion.

[0215] Embodiment 36. The method of any one of embodiments 34 or 35, further comprising: determining, based on the final motion indication, that the amount of motion in one or more frames is less than a motion threshold; and, based on the amount of motion in the one or more frames being less than the motion threshold, reducing the number of frames for that number of frames and increasing the amount of exposure duration for the determined exposure duration.

[0216] Embodiment 37. The method of any one of embodiments 34 or 35, further comprising: determining, based on the final motion indication, that an amount of motion in one or more frames is greater than a motion threshold; and, based on the amount of motion in one or more frames being greater than the motion threshold, increasing the number of frames for that number of frames and decreasing the amount of exposure duration for the determined exposure duration.

[0217] Aspect 38. A method for performing temporal blending on one or more frames. The method comprises obtaining a raw frame, the raw frame including a single color component for each pixel of the raw frame, dividing the raw frame into a first color component, a second color component, and a third color component, generating a plurality of frames at least in part by adding at least a first chrominance value to the first color component, at least a second chrominance value to the second color component, and at least a third chrominance value to the third color component, and performing temporal blending on the plurality of frames.

[0218] Embodiment 39. The method of embodiment 38, wherein the raw frame includes a color filter array (CFA) pattern.

[0219] Embodiment 40. The method of any one of embodiments 38 or 39, wherein the first color component includes a red color component, the second color component includes a green color component, and the third color component includes a blue color component.

[0220] Embodiment 41. The method of any one of embodiments 38-40, wherein the first color component includes all red pixels of the raw frame, the second color component includes all green pixels of the raw frame, and the third color component includes all blue pixels of the raw frame.

[0221] Embodiment 42. The method of any one of embodiments 38-41, wherein generating the plurality of frames includes, at least in part, generating a first frame by adding at least a first chrominance value to a first color component, generating a second frame by adding at least a second chrominance value to a second color component, and generating a third frame by adding at least a third chrominance value to a third color component.

[0222] Embodiment 43. The method of embodiment 42, wherein generating the first frame includes adding the first chrominance value and the second chrominance value to a first color component, generating the second frame includes adding the first chrominance value and the second chrominance value to a second color component, and generating the third frame includes adding the first chrominance value and the second chrominance value to a third color component.

[0223] Embodiment 44. The method of embodiment 43, wherein the first chrominance value and the second chrominance value are the same value.

[0224] Embodiment 45. The method of any one of embodiments 38-44, wherein performing temporal blending on the plurality of frames includes temporally blending a first frame of the plurality of frames with one or more additional frames having a first color component, temporally blending a second frame of the plurality of frames with one or more additional frames having a second color component, and temporally blending a third frame of the plurality of frames with one or more additional frames having a third color component.

[0225] Embodiment 46. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations according to any of embodiments 26-37.

[0226] Embodiment 47. An apparatus comprising means for performing the operations according to any of embodiments 26-37.

[0227] Embodiment 48. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations according to any of embodiments 38-45.

[0228] Embodiment 49. An apparatus comprising means for performing the operations according to any of embodiments 38-45.

[0229] Embodiment 50. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations according to any of embodiments 26-35 and any of the operations of embodiments 36-43.

[0230] Embodiment 51. A method comprising the operation of any one of embodiments 26 to 37 and the operation of any one of embodiments 38 to 45.

[0231] Embodiment 52. An apparatus comprising means for performing the operations according to any of embodiments 26-37 and any of embodiments 38-45. [Explanation of symbols]

[0232] 100 Frame Capture and Processing System 105A Image Capture Device 105B Image Processing Device 110 scenes 115 Lens 120 Control Mechanism 125A Exposure control mechanism 125B Focus control mechanism 125C Zoom Control Mechanism 130 Image Sensor 140 Random Access Memory (RAM) 145 Read-Only Memory (ROM) 150 Image Processor 152 Host Processor 154 Image Signal Processor (ISP) 156 input / output (I / O) ports 160 Input / Output (I / O) Devices 400 Frame Processing and / or Capture Instruction System 430 Sensors 432 Zero Shutter Lag (ZSL) Buffer 434 short exposure frames 436 First MFxR engine 438 Blended Frames 440 Multi-frame High Dynamic Range (MFHDR) Engine 442 Post-IPE 444 PSL Capture 446 long exposure frames 448 AWB Statistics 450 First AWB engine 451 White Balance (WB) Scaler 452 Second MFxR engine 454 preview frames 456 Blended Frames 458 Low Light (LL) Engine 460 Second AWB Engine 461 WB Scaler 462 Low lighting (LL) determination 464 Automatic Exposure Control (AEC) Engine 502 ZSL frame 504 PSL Frame 506 Short-Term Multiframe Noise Reduction (MFNR) 508 Long-Term MFNR and Preview 510 White balance (WB) improvement 512 MFHDR and Post-Processing 1302 images 1304 TFI Images 1400 Frame Processing and / or Capture Instruction System 1401 Input 1402 First Multi-Frame Noise Reduction (MFNR) Engine 1404 Second MFNR Engine 1405 Frame Processing and / or Capture Instruction System 1406 Third MFNR Engine 1408 AWB engine 1410 Long Term Frame Buffer 1411 First frame, first long exposure frame 1412 AWB control parameters 1413 ZSL buffer 1414 Medium Exposure Frame 1415 ZSL buffer 1416 short exposure frames 1418 Integrated TFI Image 1420 Blended Long Exposure Frames 1422 Machine Learning-Based Image Signal Processor (ML ISP) 1424 filtered long exposure frames 1440 Multi-Frame High Dynamic Range (MFHDR) Engine 1441 Single Short Exposure 1442 short exposure frames 1447 Blended Long Exposure Images 1448 ALGM Map 1449 First Downscale Engine 1451 Second Downscale Engine 1452 Upscale Engine 1458 Low Light (LL) Engine 1464 AEC engine 1700 Frame Processing and / or Capture Instruction System 1702 First MFNR engine 1708 AWB engine 1710 Additional Processing Components 1712 AWB control parameters 1718 Integrated TFI Images 1720 Final Blended Frame 1722ML ISP 1740 MFHDR engine 1802 ZSL frame 1804 Current Frame Frames before 1806 1808 Current Frame Frames before 1810 1902 raw frames, raw images 1904 Bayer Processing Segment (BPS) 1906 Digital Signal Processor (DSP) 1908 Red (R) component 1910 YUV frames 1912 U chrominance channel 1914 V chrominance channels, combined green channel 2400 Neural Networks 2420 Input Layer 2422 Hidden Layer 2424 Output layer 2426 nodes 2500 Convolutional Neural Networks (CNN) 2520 input layer 2522a Convolutional Hidden Layer 2522b Pooling hidden layer 2522c Fully connected hidden layer, fully connected layer 2524 output layer 2600 Computing System 2605 Connection 2610 processor 2612 Cache 2615 system memory 2620 Read-Only Memory (ROM) 2625 Random Access Memory (RAM) 2630 Storage Device 2632 Service 2634 Service 2635 Output Device 2636 Service 2640 communication interface 2645 Input Devices

Claims

1. 1. An apparatus for processing one or more frames, comprising: Memory and one or more processors coupled to the memory, the one or more processors: obtaining a motion map for one or more frames; determining motion associated with the one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects within the scene; determining a global motion associated with a camera used to capture the one or more frames based on one or more sensor measurements; determining a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion; determining a number of frames and an exposure duration for capturing said number of frames based on said weighted combination of said determined motion and said global motion; sending a request to capture the number of frames using the determined exposure duration. Device.

2. The apparatus of claim 1 , wherein the one or more frames are obtained before a capture command to capture the number of frames is received.

3. the one or more processors: configured to perform temporal blending on the number of frames captured using the determined exposure duration to generate a temporally blended frame; performing the temporal blending includes spatially aligning the number of frames and averaging values ​​of each pixel in the number of frames; 10. The apparatus of claim 1.

4. the one or more processors: configured to perform spatial processing on the temporally blended frames using a machine learning based image signal processor.

4. The apparatus of claim 3.

5. 5. The apparatus of claim 4, wherein the machine learning based image signal processor uses the motion map as an input for performing the spatial processing on the temporally blended frames.

6. The apparatus of claim 1 , wherein the determined exposure duration is based on a gain.

7. The apparatus of claim 1 , wherein the motion map includes an image, and each pixel of the image includes a value indicative of at least one of an amount of motion for each pixel and a confidence value associated with the amount of motion.

8. the one or more processors: determining, based on the weighted combination of the determined motion and the global motion, that an amount of motion in the one or more frames is less than a motion threshold; and increasing an exposure amount for the determined exposure based on the amount of motion in the one or more frames being less than the motion threshold.

10. The apparatus of claim 1.

9. the one or more processors: determining that an amount of motion in the one or more frames is greater than a motion threshold based on the weighted combination of the determined motion and the global motion; increasing a number of frames for the number of frames and decreasing an exposure amount for the determined exposure based on the amount of motion in the one or more frames being greater than the motion threshold.

10. The apparatus of claim 1.

10. The apparatus of claim 1 , further comprising at least one of a camera configured to capture at least one frame and a display configured to display the at least one frame.

11. 1. A method of processing one or more frames, comprising: obtaining a motion map for one or more frames; determining motion associated with the one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects within the scene; determining a global motion associated with a camera used to capture the one or more frames based on one or more sensor measurements; determining a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion; determining a number of frames and an exposure duration for capturing said number of frames based on said weighted combination of said determined motion and said global motion; sending a request to capture the number of frames using the determined exposure duration; A method for providing the above.

12. The method of claim 11 , wherein the one or more frames are obtained before a capture command to capture the number of frames is received.

13. performing temporal blending on the number of frames captured using the determined exposure duration to generate a temporally blended frame; wherein the step of performing the temporal blending includes spatially aligning the number of frames and averaging values ​​of each pixel in the number of frames. The method of claim 11.

14. performing spatial processing on the temporally blended frames using a machine learning based image signal processor. The method of claim 13.

15. The method of claim 14 , wherein the machine learning based image signal processor uses the motion map as an input for performing the spatial processing on the temporally blended frames.

16. The method of claim 11 , wherein the determined exposure duration is based on a gain.

17. The method of claim 11 , wherein the motion map comprises an image, and each pixel of the image comprises a value indicative of at least one of an amount of motion per pixel and a confidence value associated with the amount of motion.

18. determining that an amount of motion in the one or more frames is less than a motion threshold based on the weighted combination of the determined motion and the global motion; decreasing the number of frames and increasing the amount of exposure duration for the determined exposure duration based on the amount of motion in the one or more frames being less than the motion threshold; The method of claim 11 further comprising:

19. determining that an amount of motion in the one or more frames is greater than a motion threshold based on the weighted combination of the determined motion and the global motion; increasing a number of frames for the number of frames and decreasing an amount of exposure duration for the determined exposure duration based on the amount of motion in the one or more frames being greater than the motion threshold; The method of claim 11 further comprising:

20. 1. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to: obtaining a motion map for one or more frames; determining motion associated with the one or more frames of a scene based on the motion map, the motion corresponding to movement of one or more objects within the scene; determining a global motion associated with a camera used to capture the one or more frames based on one or more sensor measurements; determining a weighted combination of the determined motion and the global motion using a first weight for the determined motion and a second weight for the global motion; determining a number of frames and an exposure duration for capturing said number of frames based on said weighted combination of said determined motion with said global motion; sending a request to capture the number of frames using the determined exposure duration. A computer-readable storage medium.

21. 21. The computer-readable storage medium of claim 20, wherein the one or more frames are obtained before a capture command to capture the number of frames is received.

22. When executed by the one or more processors, the one or more processors performing temporal blending on the number of frames captured using the determined exposure duration to generate a temporally blended frame; performing the temporal blending includes spatially aligning the number of frames and averaging values ​​of each pixel in the number of frames; 21. The computer-readable storage medium of claim 20.

23. When executed by the one or more processors, the one or more processors determining whether an amount of motion in the one or more frames is less than or greater than a motion threshold based on the weighted combination of the determined motion and the global motion; decreasing the number of frames and increasing the exposure amount for the determined exposure duration when the amount of motion in the one or more frames is less than the motion threshold; and further comprising instructions to increase a frame count for the number of frames and decrease an exposure amount for the determined exposure duration when the amount of motion in the one or more frames is greater than the motion threshold.

21. The computer-readable storage medium of claim 20.

Citation Information

Patent Citations

  • Imaging device

    JP2003259184A

  • Device, method and program for compositing image

    JP2010141486A

  • Motion vector detection apparatus, imaging apparatus, motion vector detection method, program, and storage medium

    JP2017195458A

  • Medical image processing apparatus, medical image processing method, and program

    JP2019216848A

  • Image capturing device and image capturing method

    US20120287310A1