Low light image enhancement using key frame and frame dependent neural networks
By generating enhanced keyframe and dependent frame images through a machine learning system, the problem of accuracy loss in image processing under low light conditions is solved, processing efficiency is improved and hardware requirements are reduced.
Patent Information
- Application Number
- CN202480028846.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-22
- Filing Date
- 2024-03-28
- Publication Date
- 2025-12-12
AI Technical Summary
Existing technologies suffer from accuracy loss and task delay when processing images under low-light conditions, and using additional hardware compensation increases cost and complexity.
A machine learning system is employed, including a first machine learning network that generates enhanced keyframe images and a second machine learning network that generates enhanced dependent frame images based on the hidden states of the keyframes. Image enhancement is performed through deep neural networks and convolutional neural networks.
Improving the accuracy and efficiency of image processing in low-light conditions reduces reliance on additional hardware, lowering costs and complexity.
Smart Images

Figure CN121127883A_ABST
Abstract
Description
[0001] TECHNICAL FIELD
[0002] This disclosure generally relates to image processing. For example, aspects of the disclosure relate to systems and techniques for performing low-light image enhancement using one or more machine learning models (e.g., neural networks). BACKGROUND
[0003] Many devices and systems allow for capturing a scene by generating images (or frames) and / or video data (including multiple frames) of the scene. For example, a camera or a device including a camera can capture a sequence of frames of a scene (e.g., a video of the scene). In some cases, the sequence of frames can be processed for performing one or more functions, can be output for display, can be output for processing and / or consumption by other devices, among other uses.
[0004] Artificial neural networks can be implemented using computer technology that is inspired by logical deductions performed by biological neural networks that make up animal brains. Deep neural networks, such as convolutional neural networks, are widely used for numerous applications, such as object detection, object classification, object tracking, big data analysis, etc. For example, a convolutional neural network is able to extract high-level features (such as face shape) from an input image and use these high-level features to output, for example, a probability that the input image includes a particular object.
[0005] SUMMARY
[0006] The following presents a simplified summary related to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose of presenting certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0007] Described herein are systems and techniques for performing low-light image (e.g., image, video frame, etc.) enhancement using machine learning systems (e.g., neural network systems or models) based on image key frames and image dependent frames. In some cases, a first neural network can be used to generate an enhanced image corresponding to a low-light key frame image. A second neural network can use a hidden state of the first neural network corresponding to the low-light key frame image to generate enhanced images corresponding to one or more low-light dependent frame images. The dependent frame images can be associated with the key frame image.
[0008] According to at least one illustrative example, there is provided an apparatus for processing image data (e.g., image data of a self-standing image or a video frame), comprising a memory (e.g., configured to store data, such as audio data, etc.) and one or more processors (e.g., implemented in circuitry) coupled to the memory. The one or more processors are configured to and capable of: classifying a first image as a key frame based on a difference between the first image and a previous image, wherein the first image and the previous image are included in a plurality of images; generating, using a first machine learning network, an enhanced key frame image corresponding to the first image and a hidden state output associated with the enhanced key frame image; classifying a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and generating, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, wherein the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced key frame image.
[0009] In another example, there is provided a method for processing image data (e.g., image data of a self-standing image or a video frame), the method comprising: classifying a first image as a key frame based on a difference between the first image and a previous image, wherein the first image and the previous image are included in a plurality of images; generating, using a first machine learning network, an enhanced key frame image corresponding to the first image and a hidden state output associated with the enhanced key frame image; classifying a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and generating, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, wherein the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced key frame image.
[0010] In another example, there is provided a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: classify a first image as a key frame based on a difference between the first image and a previous image, wherein the first image and the previous image are included in a plurality of images; generate, using a first machine learning network, an enhanced key frame image corresponding to the first image and a hidden state output associated with the enhanced key frame image; classify a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and generate, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, wherein the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced key frame image.
[0011] In another example, an apparatus for processing image data is provided. The apparatus includes means for classifying a first image as a keyframe based on a difference between the first image and a previous image, where the first image and the previous image are included in a plurality of images; means for generating, using a first machine learning network, an enhanced keyframe image corresponding to the first image and a hidden state output associated with the enhanced keyframe image; means for classifying a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and means for generating, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, where the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced keyframe image.
[0012] Aspects generally include methods, apparatus, systems, computer program products, non-transitory computer-readable media, user devices, user equipment, wireless communication devices, and / or processing systems as substantially described with reference to and as illustrated by the accompanying drawings and specification.
[0013] Some aspects include an apparatus having a processor configured to perform one or more operations of any of the methods summarized above. Further aspects include processing devices configured for use in an apparatus, the processing devices configured with processor-executable instructions for performing operations of any of the methods summarized above. Further aspects include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processor of an apparatus to perform operations of any of the methods summarized above. Further aspects include an apparatus having means for performing functions of any of the methods summarized above.
[0014] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows can be better understood. Additional features and advantages will be described hereinafter. The disclosed conception and specific examples can be readily utilized as bases for the designing of other structures for carrying out the same purposes of the disclosure. Such equivalent constructions not only follow from the scope of the claims but are intended to support that scope. The features of the concepts disclosed herein can be better understood with reference to the drawings following together with the description below. Each of the drawings is provided for the purpose of illustration and description and not as a definition of the limits of the claims. The foregoing summary, as well as other features and aspects of the concepts disclosed herein, will become better understood with regard to the following description, the accompanying drawings and the appended claims.
[0015] This Summary is neither intended nor should it be construed to identify any key or essential features, nor is it intended to be used in determining the scope of the subject matter. The subject matter should be understood from read ing the entire specification of the patent of which this Summary is a part. Various aspects are now described with reference to the drawings. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of one or more aspects. Accordingly, the novelty is not intended to be limited in scope by the novelty of any one description, but rather is intended to be construed as broadly as permitted by law. It should also be noted that the description is not intended to limit the novelty to the precise forms disclosed, and that clearly, many modifications, enhancements, changes, and alterations can be made to adapt the novelty to various uses, conditions, and circumstances. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings are included to provide a further understanding of the principles of the disclosure and are incorporated in and constitute a part of this specification, illustrate several aspects, and together with the description serve to explain various principles and operation of the aspects. In the drawings,
[0018] FIG. 1A An example implementation of a system-on-a-chip (SoC) is illustrated in accordance with some examples;
[0019] FIG. 1B is a block diagram illustrating an example architecture of an image capture and processing system in accordance with some examples;
[0020] FIG. 2A An example of a fully connected neural network is illustrated in accordance with some examples;
[0021] FIG. 2B An example of a locally connected neural network is illustrated in accordance with some examples;
[0022] FIG. 2C An example of a convolutional neural network is illustrated in accordance with some examples;
[0023] FIG. 3 is a block diagram illustrating another example DCN in accordance with some examples;
[0024] FIG. 4 is a diagram illustrating an example of a cascaded model pipeline that can be used to perform low-light image enhancement on image keyframes and image dependent frames in accordance with some examples;
[0025] FIG. 5 is a diagram illustrating a histogram corresponding to a low-light image and a histogram corresponding to a non-low-light image in accordance with some examples;
[0026] FIG. 6 is a diagram illustrating an example of a low-light detection machine learning architecture in accordance with some examples.
[0027] FIG. 7 is a diagram illustrating an example of a keyframe image enhancement machine learning architecture in accordance with some examples;
[0028] FIG. 8 is a diagram illustrating an example of a dependent frame image enhancement machine learning architecture according to some examples;
[0029] FIG. 9 is a diagram illustrating an example architecture of a low-light image enhancement machine learning model including a key frame image enhancement network and a dependent frame image enhancement network according to some examples;
[0030] FIG. 10 is a diagram illustrating an example architecture of a low-light image enhancement machine learning model according to some examples;
[0031] FIG. 11 is a flowchart illustrating an example process for generating an enhanced image from one or more images according to aspects of the present disclosure;
[0032] FIG. 12 is a block diagram illustrating an example of a deep learning network according to some examples;
[0033] FIG. 13 is a block diagram illustrating an example of a convolutional neural network according to some examples; and
[0034] FIG. 14 is a diagram illustrating an example system architecture for implementing certain aspects described herein.
[0035] DETAILED DESCRIPTION
[0036] Certain aspects and examples of the present disclosure are provided below. Some of these aspects and examples can be applied independently, and some of these aspects and examples can be applied in combination, as will be apparent to one of ordinary skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects and examples of the present disclosure. It is apparent, however, that various aspects and examples can be practiced in
[0037] The following description provides example aspects and examples only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the following description of the example aspects and examples will provide those skilled in the art with an enabling description for implementing aspects and examples of the disclosure. It is to be understood that various changes can be made in the function and arrangement of elements without departing from the scope of the application as set forth in the appended claims.
[0038] As previously mentioned, various devices and systems allow for capturing a scene by generating images (or frames) and / or video data (including multiple frames) of the scene. A camera is a device that uses an image sensor to receive light from a scene and capture an image, such as a still image or a video frame. The terms “image,” “image frame,” and “frame” are used interchangeably herein. A camera can include a processor, such as an image signal processor (ISP), that can receive and process one or more images. For example, a raw image frame captured by a camera sensor can be processed by an ISP to generate a final image. The processing by the ISP can be performed by applying multiple filters or processing blocks to the captured image, such as noise reduction or noise filtering, edge enhancement, color balance, contrast, intensity adjustment (such as darkening or brightening), tone adjustment, and so forth. Image processing blocks or modules can include lens / sensor noise correction, Bayer filter, demosaicing, color conversion, correction or enhancement / suppression of image properties, noise reduction filters, sharpening filters, and so forth.
[0039] In some cases, one or more machine learning networks can be used to implement various ISP operations (e.g., ISP processing blocks, including one or more of the processing blocks described above). Image processing machine networks can be included in an ISP and / or can be separate from an ISP. Machine learning systems (e.g., deep neural network systems or models) can be used to perform various tasks, such as (for example and without limitation) detection and / or recognition (e.g., scene or object detection and / or recognition, face detection and / or recognition, and so forth), depth estimation, pose estimation, image reconstruction, classification, three-dimensional (3D) modeling, dense regression tasks, data compression and / or decompression, and image processing tasks, among others. Moreover, machine learning models can be versatile and can achieve high quality results in a variety of tasks. In some cases, image processing machine learning networks can be trained and / or implemented based on use cases associated with input image data and / or output image data of the image processing machine learning networks.
[0040] Image data (e.g., images, image frames, frames, and so forth obtained using a camera) can be used for various purposes. In some examples, image data can be provided as input to decision algorithms associated with use cases such as monitoring, detection, and / or maneuvering, among others. For example, camera feeds (e.g., image data) can be provided as input to autonomous or semi-autonomous vehicle control systems. In some cases, image data can be provided in conjunction with various other sensor inputs associated with or corresponding to the image data (e.g., sensor inputs captured at the same or similar time as the images, in the same or similar location or environment as the images, and so forth). For example, image data can be used to perform tasks such as road boundary detection, sign detection, path detection, autonomous or semi-autonomous maneuvering, monitoring surveillance, and so forth.
[0041] Performance of a task or decision algorithm that utilizes image data input can be based on various characteristics and properties of the image data input. For example, one or more (or all) of the above tasks can suffer a decrease in performance or accuracy when the image data input is a low light image, and / or one or more (or all) of the above tasks can suffer a delay in prediction or other decision task when the image data input is a low light image. A low light image scenario can correspond to a decrease in availability of information represented in the resulting low light image. For example, it can be more difficult to perform a road boundary detection task using a low light image, in which case it can be challenging to distinguish road boundaries from other objects in the scene depicted by the low light image. It can additionally be challenging to perform a sign detection and path detection task using a low light image based on a decrease in visibility, contrast, and / or visual difference between the detection object of interest and other objects depicted in the low light image.
[0042] It can also be challenging to perform higher level tasks (such as autonomous or semi-autonomous maneuvering) based on low light images. For example, autonomous maneuvering of a vehicle can be based on a plurality of lower level detection and / or classification tasks that run directly on low light images (e.g., the autonomous maneuvering task can be based on lower level tasks such as road boundary detection, sign detection, path detection, road paint detection, etc.). In some examples, a degradation in output quality of each lower level task (e.g., based on the lower level task receiving a low light image as input) can impact the ability to perform the higher level task.
[0043] For example, existing techniques for implementing a camera (e.g., image processing) pipeline can result in a loss of accuracy during low light scenarios (e.g., when a low light image is provided as input to the camera pipeline). As mentioned above, a loss of accuracy in the camera or image processing pipeline can result in downstream delays and inconsistencies in higher level and / or decision tasks, which can be undesirable for real-time systems. In some cases, additional hardware can be used to compensate for the loss of accuracy associated with low light image processing. The additional hardware can be sensor hardware (e.g., infrared or other night vision camera sensors; sensing or mapping systems such as lidar, radar, sonar, time-of-flight (ToF) depth estimation, etc.) or auxiliary hardware (e.g., lighting systems that are activated during surrounding or ambient low light scenarios). Using additional hardware to compensate for the loss of accuracy associated with low light image scenarios can increase the cost of the camera or imaging device, can increase the complexity and decrease the robustness of the camera or imaging device, and / or can increase the power usage of the camera or imaging device.
[0044] Systems and techniques are needed that can be used to perform improved image processing (e.g., improved image processing for low light images) during low light image scenarios. Further, systems and techniques are needed that can be used to perform low light image processing without using additional hardware as described above. Still further, systems and techniques are needed that can be used to perform low light image processing in conjunction with existing cameras and image processing pipelines.
[0045] Systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to as “systems and techniques”) are described herein for using a machine learning system to perform low light image enhancement to generate enhanced image keyframes and enhanced image dependent frames. In some examples, the machine learning system can include a first machine learning network (e.g., a neural network) for generating an enhanced image corresponding to a low light image keyframe (e.g., a low light image identified as a keyframe). The machine learning system can additionally include a second machine learning network (e.g., a neural network) for generating an enhanced image corresponding to a low light dependent image or frame (e.g., a low light image identified as a non-keyframe and / or identified as dependent on a previously identified keyframe image).
[0046] In some aspects, the second machine learning network can use a hidden state of the first machine learning network (e.g., used to generate a low light enhanced image corresponding to a keyframe) to generate low light enhanced images corresponding to one or more dependent frames associated with the keyframe. In some examples, the first machine learning network can be a deep neural network (DNN). In some cases, the first machine learning network can be implemented using a quantized DNN model. The first machine learning network can perform keyframe enhancement to enhance a low light keyframe received as input. The first machine learning network can generate an enhanced output frame corresponding to the low light keyframe input as output. In some aspects, the first machine learning network can be referred to as a “keyframe network,” a “keyframe image enhancement network,” and / or a “keyframe enhancement network.”
[0047] The second machine learning network can be implemented as a recurrent neural network (RNN). In some aspects, the second machine learning network can generate an enhanced output frame corresponding to a low light dependent frame input. The second machine learning network can also be referred to as a “dependent frame network,” a “dependent frame image enhancement network,” and / or a “dependent frame enhancement network.” In one illustrative example, the dependent frame enhancement network can generate an enhanced output frame in a shorter inference time compared to the keyframe enhancement network. The dependent frame enhancement network can also generate an enhanced output frame with higher power efficiency compared to the keyframe enhancement network. In some examples, the dependent frame enhancement network can generate an enhanced output frame for one or more dependent frames, where the enhanced output frame is generated based on a hidden state associated with a particular keyframe corresponding to the one or more dependent frames.
[0048] In some aspects, the keyframe image enhancement model can generate an enhanced output frame corresponding to a low-light keyframe input based on processing luminance and chrominance components of the low-light keyframe input. For example, the low-light keyframe input can be split into luminance (Y) and chrominance (U, V) components. The luminance (e.g., Y) frame can be enhanced using a luminance enhancement machine learning network included in the keyframe image enhancement model. In some aspects, the luminance enhancement machine learning network can be implemented based on a UNet architecture. The U and V chrominance frames can be enhanced using a chrominance enhancement machine learning network included in the keyframe image enhancement model. In some aspects, the chrominance enhancement machine learning network can be implemented based on a residual Conv-Net architecture.
[0049] The luminance enhancement machine learning network can be used to preserve and / or enhance details represented in the low-light keyframe input. The chrominance enhancement machine learning network can be used to enhance color information represented in the low-light keyframe input. The chrominance enhancement output (e.g., generated by the chrominance enhancement network) can be fused with a hidden state of the luminance enhancement network. In some examples, the chrominance enhancement output can be concatenated with the hidden state of the luminance enhancement network and used to generate a final enhanced keyframe image output by the keyframe image enhancement model.
[0050] The hidden state of the keyframe image enhancement model can be output to the dependent frame image enhancement model. The hidden state of the keyframe image enhancement model can be different from the hidden state of the luminance enhancement network concatenated with the chrominance enhancement output. For example, the hidden state of the keyframe model can be obtained (and provided to the dependent frame model) after performing luminance-chrominance concatenation with an earlier (e.g., different) hidden state of the keyframe model.
[0051] A dependent frame image enhancement model can process a down-scaled version of a luminance (e.g., Y) frame of each dependent frame based on incorporating a hidden state output of a key frame image enhancement model to generate an enhanced output frame corresponding to a low-light dependent frame input. The key frame model hidden state can be a hidden state output generated while processing a particular key frame that is also associated with each dependent frame that is currently being processed by the dependent frame model. Based on receiving the key frame hidden state as input, the dependent frame image enhancement model can skip processing color components (e.g., U and V components of each dependent frame) of each dependent frame. In one illustrative example, the dependent frame image enhancement model can perform enhancement operations based on the down-scaled version of the luminance component of the dependent frame and can use corresponding key frame hidden state information (e.g., features) to generate accurate results with enhanced colors. In some aspects, performing enhancement operations based on the down-scaled version of the luminance component of the dependent frame can be associated with improved performance and power efficiency of the dependent frame image enhancement model. In some aspects, the dependent frame image enhancement model can down-scale the luminance (e.g., Y) frame of each dependent frame image by a factor of four (e.g., the down-scaled luminance frame has a height and width pixel size that is four times smaller than the height and width pixel size of the input luminance frame / input dependent frame image).
[0052] Various aspects of the disclosure will be described with reference to the drawings.
[0053] FIG. 1 illustrates an example implementation of a system on a chip (SOC) 100, which can include a central processing unit (CPU) 102 or a multi-core CPU configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), delays, frequency bin information, task information, and the like can be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, and / or can be distributed across multiple blocks. Instructions executed at the CPU 102 can be loaded from a program memory associated with the CPU 102 or can be loaded from the memory block 118.
[0054] The SOC 100 can also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which can include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that may, for example, detect and recognize gestures. In one implementation, the NPU is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 can also include a sensor processor 114, an image signal processor (ISP) 116, and / or a memory 120.
[0055] The SOC 100 can be based on an ARM instruction set. In an aspect of the disclosure, instructions loaded into the CPU 102 can include code to search a lookup table (LUT) for a stored product corresponding to a product of an input value and a filter weight. The instructions loaded into the CPU 102 can also include code to disable a multiplier during a multiplication operation of the product when a lookup table hit for the product is detected. Additionally, the instructions loaded into the CPU 102 can include code to store a computed product of the input value and the filter weight when a lookup table miss for the product is detected.
[0056] The SOC 100 and / or components thereof can be configured to perform image processing using machine learning techniques in accordance with aspects of the disclosure discussed herein. For example, the SOC 100 and / or components thereof can be configured to perform deep completion in accordance with aspects of the disclosure. In some cases, aspects of the disclosure can improve accuracy and efficiency of generating a dense depth map from image input and sparse depth input by using a graph-based neural network having a segmentation input and a depth input each associated with the same image.
[0057] The SOC 100 can be part of one or more computing devices. In some examples, the SOC 100 can be part of an electronic device (or multiple electronic devices), such as a camera system (e.g., a digital camera, an IP camera, a video camera, a security camera, etc.), a telephonic system (e.g., a smartphone, a cellular phone, a conferencing system, etc.), a desktop computer, an XR device (e.g., a head-mounted display, etc.), a smart wearable device (e.g., a smart watch, smart glasses, etc.), a laptop or notebook computer, a tablet computer, a set-top box, a television, a display device, a system on a chip (SOC), a digital media player, a gaming console, a video streaming device, a server, a drone, a computer in a car, an Internet of Things (IoT) device, or any other suitable electronic device.
[0058] In some implementations, CPU 102, GPU 104, DSP 106, NPU 108, connectivity block 110, multimedia processor 112, one or more sensors 114, ISP 116, memory block 118, and / or storage 120 can be part of the same computing device. For example, in some cases, CPU 102, GPU 104, DSP 106, NPU 108, connectivity block 110, multimedia processor 112, one or more sensors 114, ISP 116, memory block 118, and / or storage 120 can be integrated into a smartphone, a laptop computer, a tablet computer, a smart wearable device, a video game system, a server, and / or any other computing device. In other implementations, CPU 102, GPU 104, DSP 106, NPU 108, connectivity block 110, multimedia processor 112, one or more sensors 114, ISP 116, memory block 118, and / or storage 120 can be part of two or more separate computing devices.
[0059] FIG. 1B FIG. 1 is a block diagram illustrating an architecture of an image capture and processing system 100b. The image capture and processing system 100b includes various components for capturing and processing images of a scene (e.g., images of the scene 101). The image capture and processing system 100b can capture still images (or photos) and / or can capture video that includes multiple images (or video frames) in a particular sequence. The lens 115 of the system 100b faces the scene 101 and receives light from the scene 101. The lens 115 bends the light toward the image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 160 and is received by the image sensor 130.
[0060] The one or more control mechanisms 160 can control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. The one or more control mechanisms 160 can include multiple mechanisms and components; for example, the control mechanisms 160 can include one or more exposure control mechanisms 165 A, one or more focus control mechanisms 165B, and / or one or more zoom control mechanisms 165C. The one or more control mechanisms 160 can also include additional control mechanisms beyond the illustrated control mechanisms, such as control mechanisms that control analog gain, flash, HDR, depth of field, and / or other image capture attributes.
[0061] The focus control mechanism 165B of the control mechanism 160 can obtain a focus setting. In some examples, the focus control mechanism 165B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 165B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 165B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or servo system, thereby adjusting the focus. In some cases, additional lenses can be included in the system 100b, such as one or more micro lenses over each photodiode of the image sensor 130, each micro lens individually bending light received from the lens 115 toward a corresponding photodiode before the light reaches the photodiode. The focus setting can be determined via contrast-detect autofocus (CDAF), phase-detect autofocus (PDAF), or some combination thereof. The focus setting can be determined using the control mechanism 160, the image sensor 130, and / or the image processor 150. The focus setting can be referred to as an image capture setting and / or an image processing setting.
[0062] The exposure control mechanism 165A of the control mechanism 160 can obtain an exposure setting. In some cases, the exposure control mechanism 165A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 165A can control the size of the aperture (e.g., aperture size or f / stop), the time duration that the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting can be referred to as an image capture setting and / or an image processing setting.
[0063] The zoom control 165C of the control mechanism 160 can obtain a zoom setting. In some examples, the zoom control 165C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control 165C can control a focal length of a lens element assembly (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control 165C can control the focal length of the lens assembly by actuating one or more motors or servo systems to move one or more lenses relative to one another. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a varifocal zoom lens. In some examples, the lens assembly can include a focusing lens (which in some cases can be the lens 115) that first receives light from the scene 101 and then the light passes through a parfocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before the light reaches the image sensor 130. In some cases, the parfocal zoom system can include two positive (e.g., converging, convex) lenses with equal or similar (e.g., within a threshold difference) focal lengths and a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control 165C moves one or more lenses in the parfocal zoom system, such as the negative lens, and one or both of the positive lenses.
[0064] The image sensor 130 includes one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures an amount of light that ultimately corresponds to a particular pixel in an image produced by the image sensor 130. In some cases, different photodiodes can be covered by different color filters, and thus can measure light that matches a color of the filter covering the photodiode. For example, a Bayer color filter includes red filters, blue filters, and green filters, where each pixel of an image is generated based on red light data from at least one photodiode covered in a red filter, blue light data from at least one photodiode covered in a blue filter, and green light data from at least one photodiode covered in a green filter. Other types of color filters can use yellow, magenta, and / or cyan (also referred to as “emerald green”) color filters instead of or in addition to red, blue, and / or green color filters. Some image sensors can be completely devoid of color filters, and instead can use different photodiodes (in some cases vertically stacked) throughout the pixel array. The different photodiodes throughout the pixel array can have different spectral sensitivity curves, and thus respond to different wavelengths of light. Monochrome image sensors can also lack color filters, and thus lack color depth.
[0065] In some cases, image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching a specific photodiode or portions thereof at a specific time and / or from a specific angle, which may be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output from the photodiode and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiode (and / or amplified by the analog gain amplifier) into a digital signal. In some cases, specific components or functions discussed with respect to one or more of the control mechanisms 160 may alternatively or additionally be included in image sensor 130. Image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide-semiconductor (CMOS), an N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0066] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or any other type of processor 900 discussed with respect to the computing system 900. The host processor 152 may be a digital signal processor (DSP) and / or other types of processors. In some implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) including the host processor 152 and ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). TM The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an Integrated Circuit 2-to-Chip (I2C) interface, an Integrated Circuit 3-to-Chip (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General Purpose Input / Output (GPIO) interface, a Mobile Industrial Processor Interface (MIPI) (such as a MIPI CSI-2 physical (PHY) layer port or interface, an Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports). In an illustrative example, the host processor 152 may use the I2C port to communicate with the image sensor 130, and the ISP 154 may use the MIPI port to communicate with the image sensor 130.
[0067] The image processor 150 can perform several tasks, such as demosaicing, color space conversion, image down-sampling, pixel interpolation, auto exposure (AE) control, auto gain control (AGC), CDAF, PDAF, auto white balance, merging images to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. The image processor 150 can store images and / or processed images in random access memory (RAM) 140 / 1425, read only memory (ROM) 145 / 1420, cache 1412, a memory unit (e.g., system memory 1415), another storage device 1430, or some combination thereof.
[0068] Various input / output (I / O) devices 170 can be connected to the image processor 150. The I / O devices 170 can include a display screen, a keyboard, a keypad, a touchscreen, a touchpad, a touch-sensitive surface, a printer, any other output device 1435, any other input device 1445, or some combination thereof. In some cases, subtitles can be entered into the image processing device 105B through a physical keyboard or keypad of the I / O devices 170, or through a virtual keyboard or keypad of a touchscreen of the I / O devices 170. I / O 156 can include one or more ports, jacks, or other connectors that enable wired connections between system 100b and one or more peripheral devices from which system 100b can receive data and / or to which system 100b can transmit data. I / O 156 can include one or more wireless transceivers that enable wireless connections between system 100b and one or more peripheral devices from which system 100b can receive data and / or to which system 100b can transmit data. Peripheral devices can include any of the types of I / O devices 170 discussed previously, and once they are coupled to the ports, jacks, wireless transceivers, or other wired and / or wireless connectors, they themselves can be considered I / O devices 170.
[0069] In some cases, image capture and processing system 100b can be a single device. In some cases, image capture and processing system 100b can be two or more separate devices, including: an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, image capture device 105A and image processing device 105B can be coupled together, e.g., via one or more wires, cables, or other electrical connectors and / or wirelessly via one or more wireless transceivers. In some implementations, image capture device 105A and image processing device 105B can be disconnected from one another.
[0070] As shown in FIG. 1, image capture and processing system 100b includes a lens 115, an image sensor 130, an image signal processor (ISP) 154, a host processor 152, a random access memory (RAM) 140, a read-only memory (ROM) 145, an input / output (I / O) 156, and a control mechanism 160. FIG. 1B As shown in FIG. 1, image capture and processing system 100b includes a lens 115, an image sensor 130, an image signal processor (ISP) 154, a host processor 152, a random access memory (RAM) 140, a read-only memory (ROM) 145, an input / output (I / O) 156, and a control mechanism 160. FIG. 1B As shown in FIG. 1, a vertical dashed line divides image capture and processing system 100b into two parts representing image capture device 105A and image processing device 105B, respectively. Image capture device 105A includes lens 115, control mechanism 160, and image sensor 130. Image processing device 105B includes image processor 150 (including ISP 154 and host processor 152), RAM 140, ROM 145, and I / O 156. In some cases, certain components illustrated in image capture device 105A, such as ISP 154 and / or host processor 152, can be included in image capture device 105A.
[0071] Image capture and processing system 100b can include an electronic device, such as a mobile or stationary telephone handset (e.g., a smartphone, a cellular telephone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, image capture and processing system 100b can include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, image capture device 105A and image processing device 105B can be different devices. For example, image capture device 105A can include a camera device and image processing device 105B can include a computing device, such as a mobile handset, a desktop computer, or other computing device.
[0072] Although image capture and processing system 100b is shown as including certain components, one of ordinary skill in the art will appreciate that image capture and processing system 100b can include more, fewer, or different components than those shown in FIG. 1. FIG. 1BThe components of image capture and processing system 100b can include software, hardware, or one or a combination of software and hardware. For example, in some implementations, the components of image capture and processing system 100b can include and / or can be implemented using electronic circuitry or other electronic hardware (which can include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits)) and / or can include and / or can be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware can include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device that implements image capture and processing system 100b.
[0073] The host processor 152 can configure new parameter settings for the image sensor 130 (e.g., via an external control interface such as I2C, I3C, SPI, GPIO, and / or other interfaces). In one illustrative example, the host processor 152 can update the exposure settings used by the image sensor 130 based on internal processing results from an exposure control algorithm of past images. The host processor 152 can also dynamically configure parameter settings of the internal pipeline or modules of the ISP 154 to match the settings of one or more input images from the image sensor 130 so that the image data is correctly processed by the ISP 154. The processing (or pipeline) blocks or modules of the ISP 154 can include modules for lens (or sensor) noise correction, demosaicing, color conversion, correction or enhancement / suppression of image properties, denoising filters, sharpening filters, etc. Each module of the ISP 154 can include a large number of tunable parameter settings. In addition, the modules can be interdependent because different modules can affect similar aspects of the image. For example, denoising and texture correction or enhancement can both affect high frequency aspects of the image. As a result, the ISP uses a large number of parameters to generate a final image from a captured raw image.
[0074] Machine learning (ML) can be considered a subset of artificial intelligence (AI). ML systems can include algorithms and statistical models that computer systems can use to perform various tasks by relying on patterns and inferences without using explicit instructions. One example of a ML system is a neural network (also referred to as an artificial neural network), which can include a population of interconnected artificial neurons (e.g., neuron models). Neural networks can be used for various applications and / or devices, such as image and / or video encoding, image analysis and / or computer vision applications, Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, service robots, and so on.
[0075] Individual nodes in a neural network can mimic biological neurons by taking in input data and performing a simple operation on the data. The results of the simple operation performed on the input data are selectively passed to other neurons. Weight values are associated with each vector and node in the network, and these values constrain how the input data relates to the output data. For example, the input data for each node can be multiplied by a corresponding weight value, and the products can be summed. The sum of these products can be adjusted by an optional bias, and an activation function can be applied to the result, resulting in an output signal or output activation (sometimes referred to as an activation map or feature map) for the node. The weight values can be initially determined by an iterative flow of training data through the network (e.g., the weight values are established during a training phase in which the network learns how to identify a particular class through typical input data characteristics for the class).
[0076] There are different types of neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), multi-layer perceptron (MLP) neural networks, transformer neural networks, etc. For example, a convolutional neural network (CNN) is a type of feed-forward artificial neural network. A convolutional neural network can include a collection of artificial neurons, each with a receptive field (e.g., a spatially local region of the input space), and they collectively tile the input space. RNNs work by saving the output of a layer and feeding that output back to the input to help predict the results of the layer. A GAN is a form of generative neural network that can learn patterns in input data such that the neural network model can generate new synthetic outputs that can reasonably come from the original dataset. A GAN can include two neural networks that operate together, including a generative neural network that generates synthetic outputs and a discriminative neural network that evaluates the authenticity of the outputs. In an MLP neural network, data can be fed into an input layer, and one or more hidden layers provide several levels of abstraction to the data. Predictions can then be made to an output layer based on the abstracted data.
[0077] Deep learning (DL) is an example of a machine learning technique, and can be considered a subset of ML. Many DL methods are based on neural networks, such as RNNs or CNNs, and utilize multiple layers. The use of multiple layers in a deep neural network can permit progressively higher-level features to be extracted from a given raw data input. For example, the output of a first layer of artificial neurons becomes the input to a second layer of artificial neurons, the output of the second layer of artificial neurons becomes the input to a third layer of artificial neurons, and so on. Layers that are between the input and the output of the entire deep neural network are often referred to as hidden layers. The hidden layers learn (e.g., are trained) to transform intermediate inputs from a previous layer into a slightly more abstract and composite representation that can be provided to a subsequent layer, until an ultimate or desired representation is obtained as the final output of the deep neural network.
[0078] As mentioned above, a neural network is an example of a machine learning system and can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes of the input layer, processing is performed by hidden nodes of the one or more hidden layers, and output is produced through output nodes of the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network can include a feature map or activation map, which can include artificial neurons (or nodes). A feature map can include filters, kernels, etc. A node can include one or more weights used to indicate the importance of the node or nodes of a respective layer. In some cases, a deep learning network can have a series of many hidden layers, with early layers used to determine simple and low-level characteristics of the input, and later layers build a hierarchy of more complex and abstract characteristics.
[0079] Deep learning architectures can learn a hierarchy of features. For example, if visual data is presented to a first layer, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if auditory data is presented to a first layer, the first layer can learn to recognize spectral power in certain frequencies. A second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as recognizing simple shapes for visual data or sound combinations for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Still higher layers can learn to recognize common visual objects or spoken phrases. Deep learning architectures can perform particularly well when applied to problems that have a natural hierarchical structure. For example, classification of motor vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher layers to recognize cars, trucks, and airplanes.
[0080] Neural networks can be designed with various connectivity patterns. In feedforward networks, information is passed from lower to higher layers, with each neuron in a given layer communicating to neurons in higher layers. As described above, a hierarchical representation can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can be helpful in recognizing patterns that span more than one chunk of input data delivered to the neural network in sequence. Connections from a neuron in a given layer to a neuron in a lower layer are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when recognition of high-level concepts can aid in discriminating particular low-level features of the input.
[0081] Connections between layers of a neural network can be fully connected or locally connected. FIG. 2AAn example of a fully connected neural network 202 is illustrated. In a fully connected neural network 202, a neuron in a first hidden layer can convey its output to every neuron in a second hidden layer, so that every neuron in the second layer will receive input from every neuron in the first layer. FIG. 2B An example of a locally connected neural network 204 is illustrated. In a locally connected neural network 204, a neuron in a first hidden layer can be connected to a limited number of neurons in a second hidden layer. More generally, the locally connected layers of a locally connected neural network 204 can be configured so that every neuron in a layer will have the same or a similar connectivity pattern, but its connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern can result in spatially distinct receptive fields in higher layers, since higher layer neurons in a given region can receive input that is tuned through training to the nature of a restricted portion of the total input to the network.
[0082] One example of a locally connected neural network is a convolutional neural network. FIG. 2C An example of a convolutional neural network 206 is illustrated. A convolutional neural network 206 can be configured so that the connection strengths associated with input to each neuron in a second layer are shared (e.g., 208). Convolutional neural networks can be well suited for problems in which the spatial location of input is meaningful. Convolutional neural networks 206 can be used to perform one or more aspects of image processing in accordance with aspects of the present disclosure. Reference is made to FIG. 12 An example block diagram of a deep learning network is described in greater depth with reference to FIG. 13 An example block diagram of a convolutional neural network is described in greater depth with reference to
[0083] A deep convolutional network (DCN) is a network of convolutional networks configured with additional pooling and normalization layers. DCNs can achieve high performance on many tasks. DCNs can be trained using supervised learning, in which both the input and the output target are known for many canonical examples and are used to modify the weights of the network by using a gradient descent method. DCNs can be feedforward networks. In addition, as described above, the connections from neurons in a first layer of a DCN to a group of neurons in a next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of a DCN can be used to perform fast processing. The computational burden of a DCN can be much less than, for example, a similarly sized neural network that includes recurrent or feedback connections.
[0084] The processing at each layer of a convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be considered three-dimensional, having two spatial dimensions along the image's axes and a third dimension capturing color information. The output of the convolutional connections can be considered as forming a feature map in subsequent layers, where each element in the feature map receives input from a range of neurons in the previous layer and from each of those multiple channels. The values in the feature map can be further processed non-linearly (such as correction, max(0, x)). Values from neighboring neurons can be further pooled (which corresponds to downsampling) and provide additional local invariance and dimensionality reduction.
[0085] FIG. 3 This is a block diagram illustrating an example of a deep convolutional network 350. A deep convolutional network 350 can include multiple layers of different types based on connectivity and weight sharing. For example... FIG. 3 As shown, the deep convolutional network 350 includes convolutional blocks 354A and 354B. Each of the convolutional blocks 354A and 354B may be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.
[0086] Convolutional layer 356 may include one or more convolutional filters that can be applied to input data 352 to generate feature maps. Although only two convolutional blocks 354A and 354B are shown, this disclosure is not limited thereto, and any number of convolutional blocks (e.g., blocks 354A and 354B) may be included in the deep convolutional network 350 according to design preferences. Normalization layer 358 may normalize the output of the convolutional filters. For example, normalization layer 358 may provide whitening or lateral suppression. Max pooling layer 360 may provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.
[0087] For example, parallel filter banks of deep convolutional networks can be loaded onto the CPU 102 or GPU 104 of image processing system 100 and / or image capture and processing system 100b to achieve high performance and low power consumption. In some examples, as depicted in Figure 1, parallel filter banks can be loaded onto the DSP 106 or ISP 116 of image processing system 100, or onto... FIG. 1B The image processing device 105B has an image processor 150 on it and / or can be loaded on it. FIG. 1B The image capture device 105A can access the deep convolutional network 350, which can exist on the image capture device 105A. FIG. 1A Image processing system 100 and / or FIG. 1B Other processing blocks on the image capture and processing system 100b.
[0088] The deep convolutional network 350 can include one or more fully connected layers, such as layer 362A (labeled “FC1”) and layer 362B (labeled “FC2”). The deep convolutional network 350 can include a logistic regression (LR) layer 364. Between each layer 356, 358, 360, 362, 364 of the deep convolutional network 350 are weights (not shown) to be updated. The output of each layer (e.g., 356, 358, 360, 362, 364) can be used as input to a subsequent layer (e.g., 356, 358, 360, 362, 364) in the deep convolutional network 350 to learn hierarchical feature representations from the input data 352 (e.g., images, audio, video, sensor data, and / or other input data) supplied at the first convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 can be a set of probabilities. For example, in some classification tasks, each probability is a probability that the input data includes a feature from a set of features. In some examples, each probability is a probability that the input data belongs to a particular class (e.g., a classification) in one or more classes.
[0089] As previously mentioned, the systems and techniques described herein can be used to perform low-light image enhancement that uses a first machine learning network to generate an enhanced output image corresponding to a keyframe image and a second machine learning network to generate an enhanced output image corresponding to a dependent frame image. The keyframe image and the dependent frame image can be included in a plurality of input images. The plurality of input images can be associated with a sequential order, such as a time-based sequential order. In some cases, the plurality of input images can be a time series of video frames (e.g., images). A keyframe can be associated with one or more dependent frames that are located sequentially after the keyframe. As used herein, a “keyframe” can refer to an image (e.g., in the plurality of input images) that is identified as a keyframe, and a “dependent frame” can refer to an image (e.g., in the plurality of images) that is identified as a dependent frame and / or that corresponds to a keyframe.
[0090] FIG. 4 FIG. 400 is a diagram 400 illustrating an example of a cascaded model pipeline 410 that can be used to perform low-light image enhancement on a keyframe image and a dependent frame image, according to some examples. The cascaded model pipeline 410 can be implemented as a cascaded low-light image enhancement pipeline. For example, the cascaded model pipeline 410 can include a low-light detection engine 410, a frame type identification engine 430, a keyframe enhancement network 440, and a dependent frame enhancement network 460.
[0091] The low-light detection engine 410 can use a low-light detection machine learning network (e.g., such as described with respect to FIG. 3) to determine whether the input image 402 is a keyframe image or a dependent frame image. The low-light detection engine 410 can output a frame type 412 that indicates whether the input image 402 is a keyframe image or a dependent frame image. The frame type 412 can be a binary value (e.g., 0 or 1) that indicates whether the input image 402 is a keyframe image or a dependent frame image. In some examples, the frame type 412 can be a value that indicates a probability that the input image 402 is a keyframe image or a dependent frame image. FIG. 6The low-light detection machine learning model 410 described herein is used to implement this. The low-light detection engine 410 can receive one or more images (e.g., one or more images or frames at regular intervals) as input and generate an output indicating whether each corresponding input image (e.g., a batch of images) is a low-light image or a non-low-light image.
[0092] In some aspects, the low-light detection engine 410 (e.g., and / or FIG. 6 The low-light detection machine learning model 600 can be implemented on a CPU. For example, the low-light detection engine 410 can be implemented on the CPU of an image processing device, such as... FIG. 1A CPU102, FIG. 1B The host processor 152, FIG. 14 Processors such as the 1410.
[0093] In some respects, the frame type identification engine 430, the keyframe enhancement network 440, and / or the dependent frame enhancement network 460 can be implemented on the NPU 420. For example, the NPU 420 can be included in the same image processing device as a CPU for implementing the low-light detection engine 410. In some cases, the NPU 420 can be combined with... FIG. 1A The NPU 108 is the same as or similar to the NPU 420. In some examples, the NPU 420 may be included. FIG. 1B The image processing device 105B may include one or more of the image processor 150, host processor 152, and / or ISP 154. In another example, the NPU 420 may be included in or implemented by the image processing device 105B, the NPU 420 being integrated with... FIG. 1B The components of the image processing device 105B depicted are separate. In some cases, the NPU 420 may be included in or implemented by the image capture device 105A, and the NPU 420 in FIG. 1B The image capture device 105A described herein is either within or separate from the components thereof.
[0094] In an illustrative example, the low-light detection engine 410 can be compared with... FIG. 6 The low-light detection machine learning network 600 described herein is the same as or similar to a machine learning network. In some aspects, low-light detection can be performed based on a subset 620 of images selected from a plurality of images 610. The plurality of images 610 can be a time series or other sequential collection of images. For example, FIG. 6 The rightmost depiction of "frame n" can be compared to FIG. 6 The leftmost frame "1" is drawn at a later time point (or a later point in the sequence).
[0095] In some examples, the image subset 620 can be obtained from the plurality of images 610 using a predetermined interval. For example, the image subset 620 can be obtained using an interval of 3 frames, where the image subset 620 includes every third or fourth frame of the plurality of images 610.
[0096] In some cases, the plurality of images 610 can be YUV images (also referred to as YUV frames or YUV image frames). For example, the plurality of images 610 can be image data in YUV 420 format. In YUV 420 image format, the luminance (e.g., Y) samples of an image are separated from the chrominance (e.g., U and V) samples of the same image. The Y samples can indicate the grayscale (e.g., luminance) information for each pixel of the image. The U and V chrominance samples can indicate the color difference information for the pixels of the image (e.g., the U samples can indicate the blue luminance and the V samples can indicate the red luminance). The YUV 420 image format can also be referred to as YUV 4:2:0. The luminance information (e.g., Y samples) can have the same resolution (e.g., size or number of pixels) as a full YUV image. The chrominance information can be subsampled relative to the full resolution. For example, in YUV 420 image format, both the U and V chrominance samples can contain four times fewer pixels than the Y luminance samples (e.g., the vertical resolution of the U and V samples is 1 / 2 of the Y luminance samples and the horizontal resolution is 1 / 2 of the Y luminance samples).
[0097] In one illustrative example, low light detection can be performed based on the Y luminance samples (e.g., luminance information) of the images. In some aspects, the image subset 620 can include only the luminance frames (e.g., Y samples) of the plurality of YUV images 610. In some cases, the image subset 620 can include a luminance frame corresponding to each YUV image of the plurality of YUV images. In another example, the image subset 620 can include a luminance frame corresponding to YUV images selected from the plurality of YUV images at a predetermined interval (e.g., every third or fourth luminance frame (e.g., Y sample) of the plurality of YUV frames 610).
[0098] The low light detection model 600 can perform low light detection based on histogram information corresponding to the luminance samples 620. For example, a histogram engine 630 can be used to generate one or more histograms corresponding to each of the luminance samples 620. The histograms can be generated with a constant number of bins, such as 64 bins. In some aspects, each histogram can be generated based on distributing the pixels (e.g., of each luminance sample 620) into the 64 bins using a BitShift hash. The histograms of the luminance information can indicate information of image features such as contrast, luminance, intensity distribution, etc.
[0099] FIG. 5This is an illustration of example histogram 550 corresponding to a low-light image and histogram 510 corresponding to a non-low-light image, based on some examples. In some examples, the low-light image may be associated with a histogram similar to example low-light histogram 550, wherein the histogram distribution (e.g., based on the low pixel intensity / brightness values associated with and present in the low-light image) is shifted to the left. FIG. 5 The non-low-light image depicted can be associated with a histogram such as a non-low-light histogram 510, wherein the histogram distribution includes a larger number of higher pixel intensity / brightness values (e.g., the histogram bars on the right side of each histogram).
[0100] FIG. 6 The low-light detection model 600 may include a matrix multiplication engine 640 (e.g., "MatMul") with learned weights that amplify low-intensity pixel bars in the histogram generated by the histogram engine 630 and suppress normal and high-intensity pixel bars in the histogram generated by the histogram engine 630. The low-light detection model 600 may perform low-light detection based on binary classification using the matrix multiplication output 640 and a sigmoid function 650. Binary classification may be implemented using the sigmoid function 650, which outputs a strong 1 for low-light images and a 0 for standard (e.g., non-low-light images). The binary classification of the low-light detection model 600 may be based on a learned percentile threshold above the outputs of the matrix multiplication output 640 and / or the sigmoid function 650.
[0101] In some respects, FIG. 4 The low-light detection engine 410 of the cascaded model pipeline 410 depicted in the diagram can be implemented based on the lux index, which can be determined by the ISP and / or image capture device associated with the input image being processed by the cascaded model pipeline 410. For example, the low-light detection engine 410 can be based on the ISP (e.g., FIG. 1B The image received by and / or from the image capture device (e.g., ISP 154) is a type of image capture device. FIG. 1B The image capture device 105A receives lux index or other low-light identification information to output a value '1', which indicates a low-light image. In some cases, lux index information and / or other low-light identification information can be received from the ISP node 490 to bypass the low-light detection engine 410 of the cascaded model pipeline 410.
[0102] For example, such as FIG. 4As depicted herein, the cascaded model pipeline 410 may be associated with an ISP node 490. The ISP node 490 may provide one or more images (e.g., from multiple images) as input to the cascaded model pipeline 410. Additionally, the ISP node 490 may receive output enhanced images generated by the cascaded model pipeline 410. In an illustrative example, the ISP node 490 may be a processing block or other node that can be used to implement the cascaded model pipeline 410 in conjunction with various image processing pipelines, ISPs, etc. For example, the ISP node 490 may be included in... FIG. 1A The DSP 106 is used to implement a cascaded model pipeline 410 in one or more image processing pipelines associated with the DSP 106. In another example, an ISP node 490 may be included. FIG. 1B In the ISP 154 and / or image processor 150, and for use in conjunction with FIG. 1B A cascaded model pipeline 410 is implemented in one or more image processing pipelines associated with the image processing device 105B. In some examples, the ISP node 490 can provide images in YUV 4:2:0 format as input to the cascaded model pipeline 410 and / or can receive enhanced images in YUV 4:2:0 format as output from the cascaded model pipeline 410. In some aspects, the output from the cascaded model pipeline 410 can be an enhanced image generated using a keyframe enhancement network 440 or an enhanced image generated using a dependent frame enhancement network 460. The enhanced image can have the same format and resolution as the input image obtained from the ISP node 490. An enhanced dependent frame image (e.g., generated by the dependent frame image enhancement network 460) can have the same format and resolution as an enhanced keyframe image (e.g., generated by the keyframe image enhancement network 440).
[0103] When a low-light image is determined or otherwise identified by the low-light detection engine 410 and / or based on lux index information obtained from the ISP node 490, the frame type identification engine 430 can be used to identify the low-light image as a keyframe 432 or a dependent frame 434. In some aspects, if the input image has previously been identified as a low-light image by the low-light detection engine 410, the frame type identification engine 430 may only receive the input image. In some examples, the frame type identification engine 430 may receive each of a plurality of images as input, wherein each corresponding image provided to the frame type identification engine 430 is associated with a corresponding indicator of low-light identification or non-low-light identification determined by the low-light detection engine 410 for the corresponding image.
[0104] In an explanatory example, frame type identification engine 430 can classify an image as a keyframe 432 or a dependent frame 434 based on comparing the current image with one or more previous images (e.g., where the current image and one or more previous images are included in the same set of images). For example, frame type identification engine 430 can classify the current image frame F as a keyframe 432 or a dependent frame 434. t (For example, associated with time t) with the previous image frame F t-1 (For example, in relation to time t-1) for comparison.
[0105] In some aspects, the frame type identification engine 430 may include one or more convolutional layers trained to extract (e.g., generate) image features. Pre-trained residual blocks (e.g., associated with one or more convolutional layers) can generate features corresponding to the current image frame F. t The first feature set and the corresponding previous image frame F t-1 The second set of features. The generated features can be provided from the output of one or more convolutional layers to the input of one or more max-pooling layers, which are also included in the frame type identification engine. The max-pooling layers can output the first pooled feature set (e.g., corresponding to the current image frame F). t The generated features) and the second pooling feature set (e.g., corresponding to the previous image frame F) t-1 (Generated features).
[0106] The frame type identification engine 430 can use the corresponding first pooling feature set and second pooling feature set to perform F. t and F t-1 Motion detection between frames. In an illustrative example, the frame type identification engine 430 can be based on F... t Pooling features and F t-1 The Euclidean distance between pooling features is used to determine the current frame F. t And previous frame F t-1 The amount of motion between them. In some respects, the Euclidean distance between two pooling feature sets can be determined based on subtracting the first pooling feature set from the second pooling feature set, or vice versa. An indication of the current frame F at each pixel location in both frames can be generated. t Features and previous frame F t-1 The residual frames of the Euclidean distance (e.g., difference) between features (e.g., two frames may each have the same pixel size, and F t A specific pixel position in F t-1 (It has a corresponding specific position).
[0107] Indicates the current frame F t Features and previous frame F t-1Residuals of the Euclidean distances between the features can be provided to one or more fully connected (FC) layers of the frame type identification engine 430. The one or more fully connected layers can be used to classify the motion represented in the Euclidean distance residuals as a "small motion" class or a "large motion" class. In one illustrative example, the one or more fully connected layers generate an output that includes a single number with a value between 0 and 1. In some aspects, the value between 0 and 1 indicates a probability that there is large motion between the current image frame F t and the previous image frame F t-1 . In some aspects, the frame type identification engine 430 can be trained using a training dataset of video decoded frames obtained from a plurality of video files (e.g., MP4 files). I-frames of the video files can be used as key frames 432 during training. Dependent frames are frames based on motion, and P-frames and B-frames of the video files can be used as dependent frames 434 during training.
[0108] In one illustrative example, the value output by the fully connected layers can be referred to as a "perceived similarity index" S, where a value of "1" indicates a highest similarity (e.g., no motion between F t and F t-1 ) and a value of "0" indicates no similarity. In some aspects, based on comparing the perceived similarity index S to one or more thresholds, the one or more thresholds can be used to classify a low light image as a key frame 432 or a dependent frame 434.
[0109] For example, a key frame 432 can be a low light image frame that is dissimilar to a previous image frame (e.g., the frame type identification engine 430 determines a perceived similarity index S that is less than a threshold). A dependent frame 434 can be a low light image frame that is similar to a previous image (e.g., the frame type identification engine 430 determines a perceived similarity index that is greater than a threshold). In some aspects, a key frame 432 can follow a dependent frame 434 (e.g., based on a comparison to a previous dependent frame F t-1 , the current frame F t is identified as a key frame). In another example, a key frame can follow another key frame (e.g., both the current frame F t and the previous frame F t-1 are key frames). A dependent frame can follow a key frame (e.g., based on a comparison to a previous key frame F t-1 , the F t is identified as a dependent frame). A dependent frame can also follow another dependent frame (e.g., both F t and F t-1 are dependent frames).
[0110] Based on the identification performed by the frame type identification engine 430, each low-light image is identified as a key frame 432 or a dependent frame 434. In one illustrative example, the key frames 432 can be processed (e.g., augmented) using a key frame image augmentation machine learning network 440, and the dependent frames 434 can be processed (e.g., augmented) using a dependent frame image augmentation machine learning network 460. The key frame image augmentation network 440 can be the same as or similar to the key frame image augmentation network 700 of FIG. 7 and / or the DNN 930 of FIG. 9 . The dependent frame image augmentation network 460 can be the same as or similar to the dependent frame image augmentation network 800 of FIG. 8 and / or the RNN 940 of FIG. 9 .
[0111] FIG. 7 is a diagram illustrating an example of a key frame image augmentation machine learning network 700 in accordance with some examples. In one illustrative example, the key frame image augmentation machine learning network 700 can be the same as or similar to the key frame augmentation network 440 of FIG. 4 and / or the DNN 930 of FIG. 9 .
[0112] The key frame image augmentation machine learning network 700 (e.g., also referred to as a “key frame augmentation network” or a “key frame network”) can receive a low-light YUV image as input. As described above, the YUV image can be identified as a low-light image, for example, using the low-light detection engine 410 of FIG. 4 , the low-light detection model 600 of FIG. 6 , the lux index value from the ISP node 490 of FIG. 4 . In one illustrative example, the system and techniques can perform one or more pre-processing operations to split the low-light YUV image into its respective Y (e.g., luminance), U (e.g., blue-difference chroma), and V (e.g., red-difference chroma) components.
[0113] For example, the low-light YUV image processed by the key frame augmentation network 700 can be the same as or similar to the low-light YUV image classified as a key frame 432 by the frame type identification engine 430 of FIG. 4 . The key frame 432 can be pre-processed and split into a Y-plane luminance component 702, a U-plane chrominance component 704, and a V-plane chrominance component 706.
[0114] In some aspects, the keyframe image enhancement network 700 can include a luminance enhancement subnetwork and a chrominance enhancement subnetwork. The luminance enhancement machine learning network can be used to preserve and / or enhance details represented in the low-light keyframe input. For example, the luminance enhancement subnetwork can be used to process and enhance the Y-plane luminance component 702 of the keyframe image. In some aspects, the Y-plane luminance component 702 can be provided as an input to a luminance enhancement subnetwork based on a UNet architecture, which can be implemented to help preserve image details during the enhancement image processing operation. The chrominance enhancement subnetwork 720 can also be referred to as a “ChromaNet.” In some aspects, the ChromaNet 720 can be implemented based on a residual Conv-Net architecture. The ChromaNet 720 can be used to enhance color information represented in the low-light keyframe input. For example, the ChromaNet 720 can be used to enhance and restore daylight color tones of the color information. The chrominance enhancement output (e.g., generated by the chrominance enhancement network 720) can be fused with the hidden state of the luminance enhancement network. In some examples, the chrominance enhancement output can be concatenated with the hidden state of the luminance enhancement network and used to generate the final enhanced keyframe image 750 output by the keyframe image enhancement network 700.
[0115] The luminance enhancement subnetwork (and the chrominance enhancement subnetwork 720) can include a plurality of machine learning layers. For example, the Y-plane luminance component 702 can be provided as an input to a pair of convolutional layers 762. The convolutional layers 762 can include one or more convolutional layers, one or more batch normalization (BN) layers, and one or more rectified linear unit (ReLu) layers. The luminance enhancement subnetwork can include a plurality of concatenation operations or concatenation layers 764, which can be provided between pairs of the remaining machine learning layers of the luminance enhancement subnetwork and / or the entire keyframe image enhancement machine learning network 700.
[0116] One or more max-pooling layers 766 can be provided to perform pooling between the output of one machine layer and the input of another machine learning layer. For example, one or more max-pooling layers 766 can be provided between at least a portion of the convolutional + BN + ReLu layers 762.
[0117] One or more up-convolutional layers 768 can be provided to implement the keyframe image enhancement network 700. For example, an up-convolutional layer 768 can be provided at the output of the ChromaNet 720 and used to fuse or concatenate the output of the ChromaNet 720 with the internal hidden state of the luminance enhancement performed by the keyframe enhancement network 700.
[0118] After the fused concatenation of the ChromaNet 720 output and the internal hidden states of the luma enhancement subnetwork, the keyframe enhancement network 700 may perform one or more additional processing steps. In some aspects, the keyframe enhancement network 700 may include one or more convolutional layers 769. For example, the convolutional layer 769 may be used to generate the hidden state output 730 of the keyframe image enhancement network 700. The hidden state output 730 may be different from the internal hidden states fused with the ChromaNet 720 output (and generated thereafter). As will be described in more detail below, the hidden state output 730 of the keyframe enhancement network 700 may be provided as a dependent frame enhancement network (e.g., FIG. 8 The input of the dependent frame enhancement network (800) is used to generate FIG. 8 One or more enhancement-dependent frames 860 (e.g., for) FIG. 8 One or more enhancements depend on frame 860 (which contributes).
[0119] The enhanced keyframe image output 750 can be generated by the keyframe enhancement network 700 based on a fused concatenated hidden state 730 representing the hidden state within the luminance enhancement subnetwork and the output of the chroma enhancement subnetwork 720. In some aspects, the keyframe image enhancement network 700 can be trained based on a perceptual quality index. For example, the perceptual quality index (PQI) can be used as a loss function for model pre-training associated with training the keyframe image enhancement network 700. In an illustrative example, the luminance enhancement subnetwork and ChromaNet 720 can be trained separately. Subsequently, the separately trained luminance enhancement subnetwork and ChromaNet 720 can be combined and trained end-to-end. After PQI-based pre-training, GAN-based fine-tuning can be performed to train the keyframe image enhancement network 700 (e.g., GAIN-based fine-tuning can be performed after end-to-end training of the combined luminance enhancement subnetwork and ChromaNet 720).
[0120] FIG. 8 This is an illustration illustrating an example of a frame-dependent image augmentation machine learning network 800. In an illustrative example, the frame-dependent image augmentation machine learning network 800 can be compared with... FIG. 4 Dependent frame image enhancement network 460 and / or FIG. 9 The RNN 940 is the same as or similar to it.
[0121] A frame-dependent image enhancement machine learning network 800 (e.g., also referred to as a "frame-dependent enhancement network" or "frame-dependent network") can receive the Y-plane luminance component 802 of a low-light YUV image as input. As described above, for example using... FIG. 4 Low-light detection engine 410 FIG. 6low light detection model 600, from FIG. 4 the ISP node 490, etc., can identify the YUV image as a low light image. In one illustrative example, the low light YUV image processed by the dependent frame enhancement network 800 can be the same or similar to the low light YUV image classified as a dependent frame 434 by the frame type identification engine 430. FIG. 4
[0122] In some aspects, the Y plane luminance component 802 can be obtained based on splitting the low light YUV image (identified as a dependent frame) into its respective Y plane luminance component, U plane chrominance component, and V plane chrominance in the same or similar manner as described above with respect to splitting the YUV key frame image into Y, U, V components 702, 704, 706, respectively.
[0123] In one illustrative example, the dependent frame image enhancement network 800 receives the Y plane luminance component 802 of the low light dependent frame image as input, but does not receive the U plane or V plane chrominance components of the low light dependent frame image. In some aspects, the dependent frame enhancement network 800 can receive the Y plane luminance component 802 and the previous frame hidden state 835, and can use the previous frame hidden state 835 to recover or otherwise generate accurate color information (e.g., chrominance information) corresponding to the Y plane luminance component 802.
[0124] In some aspects, the dependent frame enhancement network 800 can receive the Y plane luminance component 802, the previous frame hidden state 835, and a histogram distribution 803 corresponding to the Y plane luminance component 802 as input.
[0125] In the example where the previous frame F t-1 is a key frame, the previous frame hidden state 835 can be the same as the key frame hidden state 730 generated by the key frame enhancement network 700. FIG. 7 When the previous frame F t-1 is a dependent frame, the previous frame hidden state 835 can be an updated hidden state 838 generated by the dependent frame enhancement network 800 in processing a previous dependent frame F t-1 as will be described in greater depth below.
[0126] The Y plane luminance component 802 can have a pixel size of H x W (e.g., the same as the pixel size of the underlying YUV image identified as a dependent frame). A downscaling engine 805 can be used to generate a downscaled version of the Y plane luminance component 802. For example, the Y plane luminance component 802 can be downsampled by a factor of four, where the downsampled Y plane luminance component 810 has a pixel size of H / 4 x W / 4. In some aspects, the downsampled Y plane luminance component 810 can have 16 times fewer pixels than the original or full resolution Y plane luminance component 802.
[0127] The reduction in the number of pixels associated with the downscaling engine 805 (and the processing of only the luminance component of the dependent frame image) can reduce the computational cost and complexity associated with using the dependent frame enhancement network 800 to generate the enhanced dependent frame image 860. Additionally, the reduction in the number of pixels and the processing of only the luminance component of the dependent frame image can reduce the inference time associated with using the dependent frame enhancement network 800 to generate the enhanced dependent frame image 860. For example, the dependent frame enhancement network 800 can skip the color components (e.g., the chrominance components U and V) of the dependent frame image to improve computational efficiency, and the previous frame hidden state features 835 can be used to generate accurate results with enhanced colors. Additionally, performing convolutions and / or convolution operations on full-sized images (e.g., the original resolution Y plane luminance component 802) can be computationally expensive or costly operations. Based on the downsizing of the input Y plane luminance component 802 to ¼ size (e.g., the down-scaled Y plane luminance component 810), the dependent frame network 800 can perform the convolutions and / or convolution operations with lower computational cost.
[0128] In some aspects, the previous frame hidden state 835 (and the updated hidden state 838 generated for the current dependent frame) can have the same pixel size as the down-scaled Y plane luminance component 810. For example, the downscaling engine 805 can perform the downsizing to adjust the size of the full resolution Y plane luminance component 802 to match the pixel size of the previous frame hidden state 835. The previous frame hidden state 835 can have the same pixel size as the pixel size of the U and V chrominance components of the image. In the example where the image is a YUV 4:2:0 format image, the U and V chrominance components have ½ horizontal resolution and ½ vertical resolution of the Y plane luminance component (and the YUV image itself, which has the same resolution as the Y plane luminance component). In such an example, the downscaling engine 805 can downscale the full resolution dependent frame Y plane component 802 by a factor of two, such that the down-scaled Y plane component 810 has the same size as the previous frame hidden state 835 (e.g., H / 2 x W / 2).
[0129] The down-scaled luminance component 810 can be provided as input to the feature extractor 840, which generates a plurality of features corresponding to the down-scaled luminance component 810 of the dependent frame image. The one or more correlation layers 842 can receive the plurality of features from the feature extractor 840 and the previous frame hidden state features 835 as input. The correlation layers 842 can perform a functional convolution operation with the previous frame hidden state features 835 as the convolution kernel. The functional convolution operation compares the Y plane extracted features (e.g., the features generated by the feature extractor 840) to the previous frame hidden state features 835 at each spatial location. The output of the one or more correlation layers 842 can be a correlation output tensor.
[0130] The relevance output tensor generated by the relevance layer 842 can be provided as input to a cross-attention unit 844. The cross-attention unit 844 can include a convolution + sigmoid activation that takes as input the relevance output tensor from the relevance layer 842 and generates as output an attention map. In some aspects, the cross-attention unit 844 can generate an attention map that indicates a mapping between the Y-plane luminance component of the dependent frame image and the U-plane and V-plane chrominance components of the previous frame (e.g., based on the previous frame hidden state features 835).
[0131] The attention map from the cross-attention unit 844 can be provided as input to a feature alignment unit 846. The feature alignment unit 846 can additionally receive the downsampled Y-plane luminance component 810 as input. The feature alignment unit 846 applies the attention map (e.g., from the cross-attention unit 844) to features of the downsampled Y-plane luminance component 810 while preserving temporal consistency. In some aspects, the feature alignment unit 846 can apply the attention map to features generated by the feature extractor 840 for the downsampled Y-plane luminance component 810.
[0132] The warping engine 850 can include a plurality of convolutional layers (e.g., Conv2D) and one or more ReLu layers. The warping engine 850 can receive as input the output of the feature alignment unit 846 and can be used to compensate for relative motion between the previous frame F t-1 (e.g., associated with the previous frame hidden state 835) and the current frame F t (e.g., a low-light dependent frame image processed by the dependent frame enhancement network 800). For example, the warping engine 850 can be used to scale and update the previous frame hidden state 835 to correspond to the current frame (e.g., to correspond to the Y-frame luminance component of the current dependent frame). In some aspects, the warped hidden state is the same as the updated hidden state 838 (e.g., the updated hidden state 838 can be generated by warping the previous frame hidden state 835 based on the current dependent frame Y-plane luminance component 810 using the warping engine 850).
[0133] The downsampled Y-plane luminance component 810 of the current low-light dependent frame can be warped with the chrominance features of the updated hidden state 838 and used to generate a downsampled enhanced image output 852. The downsampled enhanced image output 852 can have the same pixel dimensions as the downsampled Y-plane luminance component 810, the previous frame hidden state 835, and the current frame updated hidden state 838.
[0134] In one illustrative example, the dependency frame enhancement network 800 can include an upscaling engine 855. The upscaling engine 855 can generate an enhanced image output 860 corresponding to the original resolution dependency frame (e.g., the original resolution Y-plane luminance component corresponding to the low light dependency frame). For example, the upscaling engine 855 can use the same upscaling factor as the downscaling factor applied by the downscaling engine 805. The enhanced image output 860 can have the same pixel dimensions (e.g., H x W) as the original resolution Y-plane luminance component 802 (and thus, the same pixel dimensions as the original resolution low light dependency frame processed using the dependency frame enhancement network 800). The histogram distribution 863 corresponding to the enhanced image output 860 has a wider distribution (e.g., ranging to all pixel intensity regions) than the histogram distribution 803 corresponding to the input image 802 (e.g., which is a low light image and has a distribution concentrated to the left, ranging only to the low intensity pixel region).
[0135] FIG. 9 is a diagram illustrating an example of a low light image enhancement machine learning network 900 in accordance with some examples. The low light image enhancement network 900 includes a DNN 930, which can be the same as or similar to the key frame image enhancement network 700 of FIG. 7 . The low light image enhancement network 900 further includes an RNN 940, which can be the same as or similar to the dependency frame image enhancement network 800 of FIG. 8 .
[0136] A low light input camera feed 902 can be provided to the low light image enhancement network 900. The low light input camera feed 902 can be associated with a plurality of images. Some (or all) of the plurality of images can be low light images. The plurality of images can be provided in a YUV image format (e.g., YUV 4:2:0), among various other image formats.
[0137] The low light input camera feed 902 (e.g., the plurality of images) can be provided to a frame extractor 904. The frame extractor 904 can be used to perform low light determinations, in which each input frame of the plurality of images provided by the low light input camera feed 902 is identified as a low light frame or a non-low light frame. In one illustrative example, the frame extractor 904 can identify the extracted frames Fl, F2, F3,..., Fn based on a corresponding plurality of histograms 906 associated with the extracted frames. The plurality of histograms 906 can be generated by the frame extractor 904 based on the plurality of images provided by the low light input camera feed 902. n perform low light identification.
[0138] In one illustrative example, the frame extractor 904 can be the same as or similar to the low light detection engine 410 of FIG. 4 and / or the low light detection model 600 of FIG. 6 . The plurality of histograms 906 can be similar to the example histograms of FIG. 5 , and can use the same or similar histogram generation techniques as the low light detection engine 410 of FIG. 6The histogram generator 630 is used to generate the histogram. In some examples, the frame extractor 904 can be used to generate a selected subset 620 of luminance frames from multiple images 610, such as... FIG. 6 As depicted in the text.
[0139] In some respects, frame extractor 904 can couple input images in pairs and compare the histograms of each frame in the pair. For example, frame extractor 904 can extract the histogram corresponding to the current frame F. t Histogram 906 and corresponding to the previous frame F t-1 The histograms 906 are compared. Based on the generated histogram pairs, the frame extractor 904 can determine whether the current frame is a low-light frame or a non-low-light frame, as previously described above.
[0140] Low-light frames (e.g., identified by frame extractor 904) can proceed to frame type identifier 910, which determines the current frame F. t And previous frame F t-1 The similarity S between them. In an explanatory example, FIG. 9 The frame type identification engine 910 can be used with FIG. 4 The frame type identification engine 430 is the same as or similar to the frame type identification engine 910. The frame type identification engine 910 can compare the similarity score S between the current frame and the previous frame with a threshold and determine the current frame F. t Is it a keyframe or a dependent frame (e.g., as previously referenced)? FIG. 4 (As described by the frame type identification engine 430).
[0141] If the similarity score S is less than the threshold, the frame type identification engine 910 will assign the current frame F to the frame type identification engine. t Identified as keyframe 920 (e.g., in FIG. 9 (As shown in the underlying YUV image 921 associated with the identified keyframe 920). If the similarity score S is not less than a threshold (e.g., greater than a threshold; greater than or equal to a threshold, etc.), the frame type identification engine 910 will identify the current frame F. t It is identified as dependent frame 950.
[0142] The YUV image 921, identified as keyframe 920, is provided as input to the keyframe DNN 930, which generates an enhanced image 961 and the corresponding hidden state 935 as output. The keyframe DNN 930 can be used with... FIG. 7 The keyframe image enhancement network 700 is the same as or similar to the one used. The enhanced keyframe image output 961 can be compared with... FIG. 7 The enhanced image output 750 is the same as or similar to the keyframe hidden state 935. FIG. 7the same or similar to the keyframe hidden state 730. In some aspects, the low-light keyframe 921 is enhanced based on the quantized DNN model used to implement the DNN 930. The keyframe hidden state 935 can include enhanced luminance and chrominance features that can be used to perform (e.g., using the dependent frame RNN 940) enhancement of low-light dependent frames.
[0143] One or more dependent frames can be associated with the same keyframe. For example, FIG. 9 A first image 921 is depicted that is identified as a keyframe 920. Subsequently, a next image (e.g., the next image in a sequence of images obtained from the low-light input camera feed 902) is identified as a dependent frame 950, and a down-scaled version of the luminance component 952a is generated and provided to the dependent frame RNN 940. The down-scaled version of the luminance component 952a can be the same or similar to the down-scaled luminance component 810 described above with respect to FIG. 8 The dependent frame RNN 940 can be the same or similar to the dependent frame image enhancement network 800 of FIG. 8
[0144] The dependent frame RNN 940 uses the down-scaled dependent frame luminance component 952a and the previous frame hidden state 935 (e.g., in this example, the previous frame hidden state 935 is the keyframe hidden state associated with generating the enhanced keyframe image output 961) as inputs to generate an enhanced image output of the low-light dependent frame. The output of the dependent frame RNN 940 has the same down-sampled resolution as the down-scaled dependent frame luminance component 952a, and can be provided to the upscaling engine 948. FIG. 9 The upscaling engine 948 of FIG. 8 may be the same or similar to the upscaling engine 855 of The output of the upscaling engine 948 is a dependent frame enhanced image output 963a, which corresponds to the input dependent frame 950 / down-scaled dependent frame luminance component 952a.
[0145] The updated hidden state 945a of the dependent frame RNN 940 can be provided as input for potentially use in processing a next frame in the plurality of low-light frames obtained from the low-light input camera feed 902. For example, if the next frame is also a dependent frame 950, the image enhancement for the next dependent frame 952b can be performed using the updated hidden state determined for the current frame dependent frame 952a (e.g., as the previous frame hidden state).
[0146] The above process can be repeated for each additional dependent frame 950 identified with respect to the same keyframe 920 (e.g., each of the dependent frames images 952a, 952b, 952c, 952d are similar to and associated with the keyframe image 921). The same dependent frame RNN 940 can be used to generate a respective enhanced image output for each of the plurality of dependent frame images (e.g., 952a-d).
[0147] In some examples, the dependent frame RNN 940 uses the key frame hidden state 935 as FIG. 8 the previous frame hidden state 835 depicted in FIG. 8, and uses the updated hidden state 945a of the RNN as FIG. 8 the updated hidden state 838 of the RNN to generate an enhanced image output for a first dependent frame 952a.
[0148] The dependent frame RNN 940 can use the updated hidden state 945a of the first dependent frame 952a as FIG. 8 the previous frame hidden state 835 depicted in FIG. 8 to generate an enhanced image output for a second dependent frame 952b. The updated hidden state 945b of the second dependent frame 952b can be used as the previous frame hidden state 835 to generate an enhanced image output for a third dependent frame 952c. The updated hidden state 945c of the third dependent frame 952c can be used as the previous frame hidden state 835 to generate an enhanced image output for a fourth dependent frame 952d.
[0149] When a new key frame is identified after one or more previous dependent frames, a new hidden state is computed. For example, the dependent frame RNN 940 processes the dependent frames 952a-d based on iteratively or successively updating an initial hidden state provided as the key frame hidden state 935 corresponding to the key frame image 921. When a later low-light image 921b is identified as a new (e.g., second) key frame 920b, the key frame DNN 930 generates a brand new hidden state 935b based on processing Y, U, and V components of the new key frame image 921b (e.g., as described with reference to the key frame image enhancement network 700 of FIG. 7). For example, the new key frame hidden state 935b can be generated without reference to (e.g., independent of) the previous key frame hidden state 935 or any intermediate updated dependent frame hidden states 945a-d. FIG. 7
[0150] In one illustrative example, the dependent frame image enhancement described herein can be performed in a shorter inference time than the key frame image enhancement described herein. For example, the key frame enhancement DNN 930 can be associated with an inference time of about 19 milliseconds (ms) for generating an enhanced output image 961 corresponding to a low-light key frame image 921. The dependent frame enhancement RNN 940 can be associated with an inference time of about 10 milliseconds (ms) for generating an enhanced output image 963a corresponding to a low-light dependent frame image 952a.
[0151] FIG. 10 is a diagram illustrating an example architecture of a low-light image enhancement machine learning model 1000 in accordance with some examples. The keyframe image 1020 can be associated with or include a luminance component (Y plane) 1020a, a blue color difference chroma component (Cb) 1020b, and a red color difference chroma component (Cr) 1020c. In some aspects, the keyframe image 1020 can be split to generate the Y component 1020a, the Cb component 1020b, and the Cr component 1020c.
[0152] The Y plane luminance component 1020a can be provided as input to a luminance enhancement engine 1030, and the two chroma components 1020b, 1020c can be provided as input to a chroma enhancement engine 1040. In some aspects, the luminance enhancement engine 1030 can be the same as or similar to the luminance enhancement subnetwork included in the keyframe image enhancement machine learning network 700 of FIG. 7, described above. FIG. 7 In some cases, the chroma enhancement engine 1040 can be the same as or similar to the ChromaNet 720 described above with respect to the keyframe image enhancement machine learning network 700 of FIG. 7. FIG. 7
[0153] The luminance enhancement output (e.g., associated with the luminance enhancement 1030) can be combined with the chroma enhancement output (e.g., associated with the chroma enhancement 1040) in a hidden activation state 1050. The hidden activation state 1050 can include the final activation layer from one or more convolutions associated with implementing the luminance enhancement 1030 and the chroma enhancement 1040. For example, the hidden activation state 1050 can include the final activation layer from the keyframe image enhancement network 700 of FIG. 7, described above. FIG. 7 In some aspects, the hidden activation state 1050 can be the same as or similar to the luminance enhancement internal or hidden state that is output fused with the ChromaNet 720 as described above with respect to the keyframe enhancement network 700 of FIG. 7. For example, the hidden activation state 1050 can include the luminance enhancement internal state and the ChromaNet 720 output concatenated. FIG. 7 FIG. 7
[0154] Image enhancement can be performed on one or more (e.g., multiple) dependency images 1050. As previously described, each of the dependency images 1050 can be down-sized. For example, a 4x hardware down-sampled image 1052 can be generated for each of the dependency images 1050. Various other down-sampling or down-sizing factors other than 4x can also be utilized. The down-sampled images 1052 can be generated based on down-sampling only the luminance (e.g., Y plane) component of each of the dependency images 1050. The down-sampled dependency image frame luminance 1052 can be warped using the hidden activation state 1050 using a warping engine 1070. The warping engine 1070 can be the same as or similar to the warping engine 850 described above with respect to the dependency frame image enhancement machine learning network 800. The warping engine 1070 can generate an updated hidden state 1052 that is the same as or similar to the updated hidden state 838 of FIG. 8 and / or the updated hidden state 945a of FIG. 8 and / or the updated hidden state 945a of FIG. 9 .
[0155] FIG. 11 is a flowchart illustrating an example process 1100 for processing image data. The process 1100 can be performed by a computing device (or apparatus), or a component of a computing device (e.g., a chip set, a processor such as a neural processing unit (NPU), a digital signal processor (DSP), etc.) utilizing or implementing one or more of the neural networks and / or machine learning models described herein.
[0156] At block 1102, the process 1100 can include classifying a first image as a key frame based on a difference between the first image and a previous image, where the first image and the previous image are included in a plurality of images. For example, the plurality of images can be the same as or similar to the plurality of images 610 of FIG. 6 . The plurality of images can be still images (e.g., self-standing images) and / or can be video frames (e.g., frames of video data). In some cases, the first image can be the same as or similar to one or more of the key frame images 921, 921b of FIG. 9 . In some cases, the first image can be the same as or similar to one or more of the key frame images 1020 of FIG. 10 . When the first image is the key frame image 921b of FIG. 10 , the previous image can be the image 952d of FIG. 9 .
[0157] In some examples, the first image is a low light image. For example, a respective luminance associated with the first image can be less than a low light threshold. In some examples, each respective image of the plurality of images can be classified as a low light image or a non-low light image based on a lux index associated with each respective image. The lux index can be obtained from an ISP or ISP node, such as the ISP node 490 of FIG. 4 .
[0158] In some examples, a frame identification neural network can be used to classify each low-light image of the plurality of images as a key frame or a dependent frame. For example, the frame identification neural network can be the same as or similar to the frame type identification network 430 of FIG. 4 .
[0159] In some cases, a low-light detection engine can be used to classify each respective image of the plurality of images as a low-light image or a non-low-light image (and a frame identification neural network can be used to classify each low-light image of the plurality of images as a key frame or a dependent frame). For example, the low-light detection engine can be the same as or similar to the low-light detection engine 410 of FIG. 4 and / or the low-light detection model 600 of FIG. 6 In some cases, the low-light detection engine can be used to classify each respective image based on generating a first luminance histogram for the respective image and a second luminance histogram for a previous image, where the previous image and the respective image are consecutive images of the plurality of images. For example, the luminance histograms can be the same as or similar to the histograms 630 generated by the low-light detection model 600 of FIG. 6 . The binary classification for the respective image can be determined based on the first luminance histogram and the second luminance histogram. For example, the binary classification can classify the respective image as a low-light image or a non-low-light image, and the sigmoid 650 of FIG. 6 may be used.
[0160] At block 1104, the process 1100 can include generating, using a first machine learning network, an enhanced key frame image corresponding to the first image and a hidden state output associated with the enhanced key frame image. For example, the first machine learning network can be a deep neural network (DNN). The first machine learning network can be the same as or similar to the key frame enhancement network 440 of FIG. 4 , the key frame enhancement network 700 of FIG. 7 , and / or the DNN 930 of FIG. 9 In some cases, the first machine learning network includes a luminance enhancement subnetwork and a chrominance enhancement subnetwork. For example, the luminance enhancement subnetwork can be the same as or similar to the luminance enhancement subnetwork of the key frame enhancement network 700 of FIG. 7 , and the chrominance enhancement subnetwork can be the same as or similar to the ChromaNet 720 of FIG. 7 .
[0161] The enhanced key frame image can be the same as or similar to the enhanced image output 750 of FIG. 7 and / or the enhanced image output 961, 961b of FIG. 9 . The hidden state output can be the same as or similar to the hidden state 730 of FIG. 7 , the hidden state 730 of FIG. 9 .Hidden states 935, 935b, and / or FIG. 10 The hidden state 1052 is the same as or similar to the hidden state 1052. In some cases, the hidden state output associated with the enhancement keyframe includes enhanced luminance features associated with the enhancement keyframe image, enhanced luminance features generated using the luminance enhancement sub-network, and enhanced chroma features associated with the enhancement keyframe image, enhanced chroma features generated using the chroma enhancement sub-network. In some aspects, the luminance enhancement sub-network can be similar to... FIG. 10 The brightness enhancement network 1030 is the same as or similar to it, and the brightness enhancement features can be the same as... FIG. 10 The enhanced brightness features of 1020a are the same or similar. The chroma enhancement subnetwork can be similar to... FIG. 10 The color correction chromaticity enhancement network 1040 is the same as or similar to it, and the enhanced chromaticity features can be the same as... FIG. 10 The enhanced chromaticity features of 1020b and 1020c are the same or similar.
[0162] In some examples, to generate the hidden state output associated with the enhanced keyframe image, the chroma enhancement subnetwork (e.g., FIG. 7 ChromaNet 720 can be used to generate enhanced chroma features associated with the chroma information of the first image. The chroma information can be... FIG. 7 The U component 704 and V component 706 are the same as or similar to, and / or may be related to FIG. 10 The chromaticity information 1020b and 1020c is the same or similar. Enhanced chromaticity features can be combined with a luminance enhancement subnetwork (e.g., FIG. 7 Brightness enhancement subnetwork and / or FIG. 10 The internal hidden state of the brightness enhancement network (1030) is fused. The internal hidden state can be different from the hidden state output. For example, the internal state can be different from the hidden state output. FIG. 7 The internal hidden states at 768 of the fusion cascade between ChromaNet 720 and the luminance enhancement subnetwork are the same or similar.
[0163] In box 1106, process 1100 may include classifying the second image into a dependent frame based on the similarity between the second image and the first image among the plurality of images. For example, the second image may be compared with... FIG. 9 Dependent frame images 952a, 952b, 952c, 952d, 921b, and / or FIG. 10 The second image is identical or similar to one or more of the dependent frame images 1050. In some cases, the second image may be based on... FIG. 9 The similarity calculation 910 determines that the dependent frame images 950 are the same or similar.
[0164] In some examples, the first and second images are low-light images. For instance, the corresponding brightness associated with the first image may be less than a low-light threshold, and the corresponding brightness associated with the second image may also be less than a low-light threshold. In some examples, each corresponding image among multiple images can be classified as a low-light image or a non-low-light image based on the lux index associated with each corresponding image. The lux index can be obtained from the ISP or ISP node, such as... FIG. 4 ISP node 490. In some examples, the same frame labeling neural network can be used to classify each low-light image in the multiple images as a keyframe or a dependent frame. For example, the frame labeling neural network can be used with... FIG. 4 The frame type identifiers are the same as or similar to those in network 430.
[0165] In some scenarios, a low-light detection engine can be used to classify each corresponding image among multiple images as a low-light image or a non-low-light image (and a frame identification neural network can be used to classify each low-light image among multiple images as a keyframe or a dependent frame). For example, a low-light detection engine can be combined with... FIG. 4 Low light detection engine 410 and / or FIG. 6 The low-light detection model 600 is the same as or similar to the previous low-light detection model. In some cases, the low-light detection engine can be used to classify each corresponding image based on a first brightness histogram of the generated corresponding image and a second brightness histogram of the previous image, where the previous image and the corresponding image are coherent images among the multiple images. For example, the brightness histogram can be compared with that of the previous image and the corresponding image. FIG. 6 The histograms 630 generated by the low-light detection model 600 are the same or similar. A binary classification of the corresponding image can be determined based on the first and second brightness histograms. For example, binary classification can classify the corresponding image as a low-light image or a non-low-light image, and can be used... FIG. 6 sigmoid 650.
[0166] In some examples, the first image can be classified as a keyframe based on first motion information between the first image and the previous image. For example, the first image can be classified as a keyframe 920 based on first motion information between the first image and the previous image, as described by... FIG. 9 The similarity calculation is determined by 910. The second image can be classified as a dependent frame based on the second motion information between the second image and the first image. For example, when the second image is... FIG. 9 When relying on frames 950 and 952a, it can be based on the second image and the first keyframe image (e.g., FIG. 9 The second motion information between frames 920 and 921 classifies the second image into dependent frame 950. The second motion information can be obtained from... FIG. 9 The similarity is calculated to determine 910. In some cases, FIG. 4The frame type identification engine 430 can be used to classify the first image as a key frame based on the first motion information and to classify the second image as a dependent frame based on the second motion information.
[0167] In some examples, the first motion information includes a first perceptual similarity index value determined using a frame identification neural network (e.g., such as the frame type identification network 430) that is based on the first perceptual similarity index value being less than a threshold (e.g., such as the perceptual similarity index value S and FIG. 4 the difference between the threshold associated with the similarity computation 910 depicted in FIG. 9B) to indicate a difference between the first image and the previous image. The second motion information can include a second perceptual similarity index value determined using a frame identification neural network (e.g., such as the frame type identification network 430) that is based on the second perceptual similarity index value being greater than the threshold to indicate a similarity between the second image and the first image. FIG. 9 FIG. 4 In some examples, the first motion information includes a first perceptual similarity index value determined using a frame identification neural network (e.g., such as the frame type identification network 430) that is based on the first perceptual similarity index value being less than a threshold (e.g., such as the perceptual similarity index value S and
[0168] At block 1108, the process 1100 can include generating, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, where the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced key frame image. The first machine learning network can be different than the second machine learning network. For example, the second machine learning network can be a recurrent neural network (RNN). In some cases, the second machine learning network can be the same as or similar to the dependent frame enhancement network 460 of FIG. 4, FIG. 4 the dependent frame enhancement network 800 of FIG. 8, and / or FIG. 8 the RNN 940 of FIG. 9. In some cases, a resolution of the luminance information provided as input to the second machine learning network is less than or equal to half of a resolution of the luminance information provided as input to the first machine learning network. For example, the luminance information provided as input to the second machine learning network can be the same as or similar to the downsampled luminance information 810 of FIG. 8, which can be downsampled by a factor of 4 from the original input image 802. The luminance information provided as input to the first machine learning network can be the same as or similar to the luminance information 702 of FIG. 7, which can have the same resolution as the input image of the first machine learning network. In another example, the luminance information provided as input to the second machine learning network is the same as or similar to the 4x downsampled luminance information 952a-d of FIG. 9, which is one quarter of the resolution of the luminance information 921 provided as input to the first machine learning. In another example, the luminance information provided as input to the second machine learning network is the same as or similar to the 2x downsampled luminance information 952a-d of FIG. 9, which is one half of the resolution of the luminance information 921 provided as input to the first machine learning. FIG. 9 FIG. 8 FIG. 7 FIG. 9 FIG. 10 the 4x downsampled luma information 1052 of the first image 1050 is the same or similar to the luma information 1020 provided as input to the first machine learning network, which is one quarter of the resolution of the luma information 1020. In some aspects, the resolution of the first image is the same as the resolution of the second image, and each image of the plurality of images has the resolution.
[0169] In some cases, the second machine learning network can be used to generate an updated hidden state output associated with an enhancement dependent frame image. For example, the updated hidden state output associated with the enhancement dependent frame image can be the same or similar to the updated hidden state output 838 of the first image 1050. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 860 of the first image 1050, and / or the enhancement dependent frame image 963a of the second image 960. FIG. 8 In some cases, the second machine learning network can be used to generate an updated hidden state output associated with an enhancement dependent frame image. For example, the updated hidden state output associated with the enhancement dependent frame image can be the same or similar to the updated hidden state output 838 of the first image 1050. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 860 of the first image 1050, and / or the enhancement dependent frame image 963a of the second image 960. FIG. 8 In some cases, the second machine learning network can be used to generate an updated hidden state output associated with an enhancement dependent frame image. For example, the updated hidden state output associated with the enhancement dependent frame image can be the same or similar to the updated hidden state output 838 of the first image 1050. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 860 of the first image 1050, and / or the enhancement dependent frame image 963a of the second image 960. FIG. 9 In some cases, the second machine learning network can be used to generate an updated hidden state output associated with an enhancement dependent frame image. For example, the updated hidden state output associated with the enhancement dependent frame image can be the same or similar to the updated hidden state output 838 of the first image 1050. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 860 of the first image 1050, and / or the enhancement dependent frame image 963a of the second image 960. FIG. 8 In some cases, the second machine learning network can be used to generate an updated hidden state output associated with an enhancement dependent frame image. For example, the updated hidden state output associated with the enhancement dependent frame image can be the same or similar to the updated hidden state output 838 of the first image 1050. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 860 of the first image 1050, and / or the enhancement dependent frame image 963a of the second image 960. FIG. 9 In some cases, the second machine learning network can be used to generate an updated hidden state output associated with an enhancement dependent frame image. For example, the updated hidden state output associated with the enhancement dependent frame image can be the same or similar to the updated hidden state output 838 of the first image 1050. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 860 of the first image 1050, and / or the enhancement dependent frame image 963a of the second image 960.
[0170] In some examples, a third image of the plurality of images can be classified as a dependent frame based on a similarity between the third image and the second image. For example, the similarity can be determined based on the similarity computation 910 of the first image 1050, and based on the similarity between the second image and the third image being greater than a threshold, the third image can be classified as a dependent frame. In some cases, the second machine learning network can generate an enhancement dependent frame image corresponding to the third image based on the third image and the updated hidden state output. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 963a of the second image 960, and / or the enhancement dependent frame image 860 of the first image 1050. FIG. 9 In some examples, a third image of the plurality of images can be classified as a dependent frame based on a similarity between the third image and the second image. For example, the similarity can be determined based on the similarity computation 910 of the first image 1050, and based on the similarity between the second image and the third image being greater than a threshold, the third image can be classified as a dependent frame. In some cases, the second machine learning network can generate an enhancement dependent frame image corresponding to the third image based on the third image and the updated hidden state output. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 963a of the second image 960, and / or the enhancement dependent frame image 860 of the first image 1050. FIG. 9 In some examples, a third image of the plurality of images can be classified as a dependent frame based on a similarity between the third image and the second image. For example, the similarity can be determined based on the similarity computation 910 of the first image 1050, and based on the similarity between the second image and the third image being greater than a threshold, the third image can be classified as a dependent frame. In some cases, the second machine learning network can generate an enhancement dependent frame image corresponding to the third image based on the third image and the updated hidden state output. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 963a of the second image 960, and / or the enhancement dependent frame image 860 of the first image 1050. FIG. 8 In some examples, a third image of the plurality of images can be classified as a dependent frame based on a similarity between the third image and the second image. For example, the similarity can be determined based on the similarity computation 910 of the first image 1050, and based on the similarity between the second image and the third image being greater than a threshold, the third image can be classified as a dependent frame. In some cases, the second machine learning network can generate an enhancement dependent frame image corresponding to the third image based on the third image and the updated hidden state output. The enhancement dependent frame image can be the same or similar to the enhancement dependent frame image 963a of the second image 960, and / or the enhancement dependent frame image 860 of the first image 1050.
[0171] In some cases, the first machine learning network can be provided with respective luma information, respective first chroma information, and respective second chroma information corresponding to a key frame image. For example, the key frame enhancement network 700 of the first image 1050 can be provided with the luma information 702, the first chroma information 704, and the second chroma information 706. FIG. 7 In some cases, the first machine learning network can be provided with respective luma information, respective first chroma information, and respective second chroma information corresponding to a key frame image. For example, the key frame enhancement network 700 of the first image 1050 can be provided with the luma information 702, the first chroma information 704, and the second chroma information 706. FIG. 8The dependent frame enhancement network 800 may be provided with luminance information 802. To generate an enhanced dependent frame image, the luminance information corresponding to the dependent frame image can be reduced. For example, to generate... FIG. 8 The enhanced dependent frame image 860 can be used as follows FIG. 8 The depicted reduction engine 805 reduces the brightness information 802. The reduced brightness information can be compared with... FIG. 8 The reduced brightness information 810 is the same as or similar to the reduced brightness information. For example, the reduced brightness information can be reduced by a factor of 4 by the reduction engine 805. The reduced brightness information can have the same resolution as the hidden state output associated with the enhanced keyframe image. For example, FIG. 8 The reduced brightness information at a resolution of 810 can be compared with FIG. 8 The previous frame hidden state 835 has the same resolution (e.g., it can itself be the same as...). FIG. 7 (The hidden state is the same as 730).
[0172] In some cases, the convolutional layers of a second machine learning network can be used to generate correlation output tensors that indicate the correlation between the reduced brightness information and the hidden state output associated with the enhanced keyframe image. The convolutional layers can be used to implement... FIG. 8 The correlations of the 842 convolutional layers are the same or similar. The correlation output tensor, which indicates correlation information, can be generated as... FIG. 8 The output of the correlation engine 842.
[0173] The cross-attention layer of the second machine learning network can be used to determine the attention map based on the relevance output tensor. For example, FIG. 8 The cross-attention layer 844 can be used to determine the attention map based on the correlation output tensor. Features of reduced brightness information can be determined (e.g., such as using...). FIG. 8 The feature alignment between the features generated by the feature extractor 840 and the hidden state output is based on an attention map. In some cases, a second machine learning network can be used to generate an updated hidden state output associated with an enhancement-dependent frame image by warping the feature alignment with the hidden state output associated with the enhancement-dependent frame image. For example, the updated hidden state output associated with the enhancement-dependent frame image can be... FIG. 8 The updated hidden state output 838 is the same as or similar to it.
[0174] As mentioned above, the processes described herein (e.g., process 1100 and / or any other process described herein) can be utilized or implemented by a computing device or apparatus using machine learning and / or neural network models (e.g., including one or more of the following: FIG. 4 Cascaded model pipeline 410 FIG. 6 Low-light detection model 600 FIG. 7Keyframe image enhancement model 700 FIG. 8 Dependent frame image enhancement model 800 FIG. 9 Low-light image enhancement model 900, and / or FIG. 10 The low-light image enhancement model 1000 is used to perform this.
[0175] In one example, process 1100 can be... FIG. 1A The process is executed by electronic device 100. In another example, process 1100 can be performed by... FIG. 1B The image capture and processing system 100b performs this process. In another example, process 1100 can be performed by an image capture and processing system 100b. FIG. 14 The computing system 1400 shown is a computing device architecture used to execute computing systems. For example, it has... FIG. 14 The computing device architecture of the computing system 1400 shown can realize the computing device. FIG. 11 The operation and / or this article about FIG. 4 to 11 The components and / or operations described in any of them.
[0176] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, XR devices (e.g., VR headsets, AR headsets, AR glasses, etc.), wearable devices (e.g., network-connected watches or smartwatches, or other wearable devices), server computers, vehicles (e.g., autonomous vehicles) or vehicles with computing devices, robotic devices, laptop computers, smart TVs, cameras, drones or unmanned aerial vehicles (UAVs), and / or any other computing device with the resources to perform the processes described herein (including process 1100 and / or any other processes described herein). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. A network interface can be configured to transmit and / or receive Internet Protocol (IP)-based data or other types of data.
[0177] The components of computing device can be implemented with circuitry. For example, the components can include and / or can use electronic circuitry or other electronic hardware (which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits)) to implement, and / or can include and / or can use computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0178] The process 1100 is illustrated as a logic flow graph, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, the operations represent computer-executable instructions stored, for example, on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The described operations do not necessarily have to be performed in the order they are described. Furthermore, various described operations can be performed concurrently, or in parallel, to save time.
[0179] Additionally, the process 1100 and / or other processes described herein can be performed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processing units, by hardware, or combinations thereof. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, such as in the form of a computer program comprising instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0180] FIG. 12is an illustrative example of a deep learning neural network 1200 that can be used by one or more of the machine learning and / or neural network models, systems, and / or architectures described herein. The input layer 1220 includes input data. In one illustrative example, the input layer 1220 can include data representing pixels of an input video frame. The neural network 1200 includes a plurality of hidden layers 1222a, 1222b, through 1222n. The hidden layers 1222a, 1222b, through 1222n include an “n” number of hidden layers, where “n” is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for a given application. The neural network 1200 further includes an output layer 1224 that provides an output resulting from processing performed by the hidden layers 1222a, 1222b, through 1222n. In one illustrative example, the output layer 1224 can provide a classification for an object in an input video frame. The classification can include a class that identifies an object type (e.g., a person, a dog, a cat, or other object).
[0181] The neural network 1200 is a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with these nodes is shared between different layers, and each layer retains information as it is processed. In some cases, the neural network 1200 can include a feedforward network, in which there are no feedback connections in which the output of the network is fed back to itself. In some cases, the neural network 1200 can include a recurrent neural network, which can have loops that allow information to be carried across nodes as input is read in.
[0182] Information can be exchanged between nodes through node-to-node interconnections between layers. Nodes of the input layer 1220 can activate a set of nodes in the first hidden layer 1222a. For example, as shown, each input node of the input layer 1220 is connected to every node of the first hidden layer 1222a. Nodes of the hidden layers 1222a, 1222b, through 1222n can transform information by applying an activation function to the information from each input node. Information derived from the transformation can then be passed to and can activate nodes of the next hidden layer 1122b, which can perform their own designated functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of the hidden layer 1222b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 1222n can activate one or more nodes of the output layer 1224, where the output is provided. In some cases, while nodes in the neural network 1200 (e.g., node 1226) are shown as having multiple output lines, the nodes have a single output and all lines are shown as outputting from the node representing the same output value.
[0183] In some cases, each node or interconnections between nodes can have a weight that is a set of parameters derived from training of the neural network 1200. Once the neural network 1200 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, the interconnections between nodes can represent pieces of information learned about the nodes of the interconnections. The interconnections can have adjustable numerical weights that can be tuned (e.g., based on a training data set), allowing the neural network 1200 to adapt to inputs and be able to learn as more and more data is processed.
[0184] The neural network 1200 is pre-trained to process features from data in the input layer 1120 using different hidden layers 1222a, 1222b, through 1222n, in order to provide an output through the output layer 1224. In an example where the neural network 1200 is used to identify objects in an image, the neural network 1200 can be trained using training data that includes both images and labels. For example, training images can be input into the network, where each training image has a label indicating the class of one or more objects in each image (essentially, indicating to the network what the objects are and what features they have). In one illustrative example, the training image can include an image of a number 2, in which case the label for the image can be [0 01 0 0 0 0 0 0 0].
[0185] In some cases, the neural network 1200 can adjust the weights of the nodes using a training process called backpropagation. Backpropagation can include forward pass, loss function, backward pass, and weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. For each set of training images, the process can repeat for a certain number of iterations until the neural network 1200 is trained well enough such that the weights of the layers are accurately tuned.
[0186] For the example of identifying objects in an image, the forward pass can include passing a training image through the neural network 1200. The weights are initially randomized before the neural network 1200 is trained. The image can include, for example, an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 that describes the intensity of the pixel at that location in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chrominance components, etc.).
[0187] For the first training iteration of the neural network 1200, the output will likely include values that do not give preference to any particular class (as the weights were randomly chosen at initialization). For example, if the output is a vector with probabilities of the object including different classes, the probability values for each of the different classes can be equal or at least very similar (e.g., for ten possible classes, each class can have a probability value of 0.1). With the initial weights, the neural network 1200 is unable to determine low-level features and, as such, cannot accurately determine what the classification of the object can be. A loss function can be used to analyze the error in the output. Any suitable loss function definition can be used. One example of a loss function includes mean squared error (MSE). MSE is defined as which computes one-half the sum of the squared differences between the true value output (e.g., the actual answer) minus the predicted output (e.g., the predicted answer). The loss can be set equal to the value of (E 总共 ).
[0188] For the first training image, the loss (or error) will be high as the actual value will be significantly different from the predicted output. The goal of the training is to minimize the amount of loss so that the predicted output is the same as the training label. The neural network 1200 can perform backpropagation by determining which inputs (weights) contribute the most to the loss of the network and can adjust the weights so that the loss is reduced and eventually minimized.
[0189] The derivative of the loss with respect to the weights (denoted as dL / dW, where W is a weight of a particular layer) can be computed to determine the weights that contribute the most to the loss of the network. After the derivative is computed, the weight update can be performed by updating the weights of all the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted as where w denotes the weight, w i denotes the initial weight, and η denotes the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes a large weight update, and a lower value indicates a smaller weight update.
[0190] The neural network 1200 can include any suitable deep network. One example includes a convolutional neural network (CNN) that includes an input layer and an output layer with a plurality of hidden layers between the input layer and the output layer. An example of a CNN is described below with respect to FIG. 13 The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for down-sampling), and fully connected layers. The neural network 1200 can include any other deep network other than a CNN, such as an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), etc.
[0191] FIG. 13is a illustrative example of a convolutional neural network 1300 (CNN 1300). The input layer 1320 of the CNN 1300 includes data representing an image. For example, the data can include an array of numbers representing the pixels of an image, where each number in the array includes a value from 0 to 255 that describes the intensity of the pixel at that location in the array. Using the previous example from above, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chrominance components, etc.). The image can be passed through a convolutional hidden layer 1322a, an optional non-linear activation layer, a pooling hidden layer 1322b, and a fully connected hidden layer 1322c to obtain an output at the output layer 1324. While only one of each hidden layer is shown in FIG. 13
[0192] The first layer of the CNN 1300 is the convolutional hidden layer 1322a. The convolutional hidden layer 1322a analyzes the image data of the input layer 1320. Each node of the convolutional hidden layer 1322a is connected to a region of nodes (pixels) of the input image, referred to as a receptive field. The convolutional hidden layer 1322a can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolution iteration of a filter is a node or neuron of the convolutional hidden layer 1322a. For example, the region of the input image that a filter covers at each convolution iteration will be the receptive field of that filter. In one illustrative example, if the input image includes a 28 x 28 array, and each filter (and corresponding receptive field) is a 5 x 5 array, there will be 24 x 24 nodes in the convolutional hidden layer 1322a. Each connection between a node and the receptive field of that node learns a weight, and in some cases, a global bias, such that each node learns to analyze its particular local receptive field in the input image. Each node of the hidden layer 1322a will have the same weights and biases (referred to as shared weights and shared biases). For example, a filter has an array of weights (numbers) and the same depth as the input. For the video frame example, the filter will have a depth of 3 (according to the three color components of the input image). An illustrative example size of the filter array is 5 x 5 x 3, which corresponds to the size of the receptive field of a node.
[0193] The convolutional nature of the convolutional hidden layer 1322a is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 1322a can start at the top left corner of the input image array and can convolve around the input image. As mentioned above, each convolution iteration of the filter can be considered a node or neuron of the convolutional hidden layer 1322a. At each convolution iteration, the values of the filter are multiplied by the corresponding number of raw pixel values of the image (e.g., a 5x5 filter array is multiplied by a 5x5 array of input pixel values at the top left corner of the input image array). The multiplications from each convolution iteration can be added together to obtain a sum for that iteration or node. Next, the process continues at the next location in the input image according to the receptive field of the next node in the convolutional hidden layer 1322a.
[0194] For example, the filter can move one step amount to the next receptive field. The step amount can be set to one or other suitable amount. For example, if the step amount is set to one, the filter will move one pixel to the right at each convolution iteration. Processing the filter at each unique location of the input amount results in a number representing the filter result for that location, resulting in a sum value being determined for each node of the convolutional hidden layer 1322a.
[0195] The mapping from the input layer to the convolutional hidden layer 1322a is referred to as an activation map (or feature map). The activation map includes the value of each node representing the filter result at each location of the input amount. The activation map can include an array including various sum values resulting from each iteration of the filter on the input amount. For example, if a 5x5 filter is applied to each pixel of a 28x28 input image (with a step amount of one), the activation map will include a 24x24 array. The convolutional hidden layer 1322a can include several activation maps in order to identify multiple features in the image. FIG. 13 The example shown in FIG. 13B includes three activation maps. By using three activation maps, the convolutional hidden layer 1322a can detect three different types of features, where each feature is detectable across the entire image.
[0196] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 1322a. A non-linear layer can be used to introduce non-linearity into a system that has been computing linear operations. One illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. A ReLU layer can apply the function f(x) = max(0, x) to all values in the input amount, which changes all negative activations to 0. Thus, a ReLU can increase the non-linear properties of the CNN 1300 without affecting the receptive fields of the convolutional hidden layer 1322a.
[0197] A pooling hidden layer 1322b can be applied after the convolution hidden layer 1322a (and after the non-linear hidden layer, if used). The pooling hidden layer 1322b is used to simplify the information in the output from the convolution hidden layer 1322a. For example, the pooling hidden layer 1322b can take each activation map output from the convolution hidden layer 1322a and use a pooling function to generate a compressed activation map (or feature map). Max pooling is one example of a function performed by the pooling hidden layer. Other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions, can be used by the pooling hidden layer 1322a. A pooling function (e.g., a max pooling filter, an L2 norm filter, or other suitable pooling filter) is applied to each activation map included in the convolution hidden layer 1322a. In the example shown in FIG. 13B, three pooling filters are used for the three activation maps in the convolution hidden layer 1322a. FIG. 14
[0198] In some examples, max pooling can be used by applying a max pooling filter (e.g., having a size of 2x2) with a stride amount (e.g., equal to the dimension of the filter, such as a stride amount of 2) to the activation maps output from the convolution hidden layer 1322a. The output from the max pool filter includes the maximum number in each sub-region of the filter convolution. Using a 2x2 filter as an example, each cell in the pooling layer can summarize a region of 2x2 nodes in the previous layer (where each node is a value in the activation map). For example, four values (nodes) in the activation map would be analyzed by the 2x2 max pooling filter in each iteration of the filter, with the maximum of the four values being output as the "maximum" value. If such a max pooling filter is applied to the activation filter from the convolution hidden layer 1322a having a 24x24 node size, the output from the pooling hidden layer 1322b would be an array of 12x12 nodes.
[0199] In some examples, an L2 norm pooling filter can also be used. The L2 norm pooling filter includes calculating the square root of the sum of the squares of the values in a 2x2 region (or other suitable region) of the activation map (rather than calculating the maximum value as in max pooling), and using the calculated value as the output.
[0200] Intuitively, the pooling function (e.g., max pooling, L2 norm pooling, or other pooling function) determines whether a given feature is found anywhere in a region of the image, and discards the exact location information. This can be done without affecting the feature detection results, as once a feature is found, the exact location of the feature is not as important as its approximate location relative to other features. The benefit of max pooling (and other pooling methods) is that the pooled features are much fewer, thus reducing the number of parameters needed in the subsequent layers of the CNN 1300.
[0201] The final layer of connections in the network is a fully connected layer that connects each node in the pooling hidden layer 1322b to each output node in the output layer 1324. Using the example above, the input layer includes 28 x 28 nodes that encode the pixel intensities of an input image; the convolutional hidden layer 1322a includes 3 x 24 x 24 hidden feature nodes that are based on applying a 5 x 5 local receptive field (for a filter) to three activation maps; and the pooling hidden layer 1322b includes a 3 x 12 x 12 hidden feature node layer that is based on applying a max-pooling filter to a 2 x 2 region across each of the three feature maps. Extending this example, the output layer 1324 can include ten output nodes. In such an example, each node of the 3 x 12 x 12 pooling hidden layer 1322b is connected to each node of the output layer 1324.
[0202] The fully connected layer 1322c can take the output of the previous pooling layer 1322b (which should represent an activation map of high-level features) and determine the features that are most relevant to a particular class. For example, the fully connected layer 1322c layer can determine the high-level features that are most strongly correlated to a particular class and can include weights (nodes) of the high-level features. The product between the weights of the fully connected layer 1322c and the pooling hidden layer 1322b can be computed to obtain probabilities for different classes. For example, if the CNN 1300 is used to predict that an object in a video frame is a person, then higher values in the activation map will occur for high-level features that represent a person (e.g., two legs present, a face present at the top of the object, two eyes present at the top left and top right of the face, a nose present in the middle of the face, a mouth present at the bottom of the face, and / or other features common to a person).
[0203] In some examples, the output from the output layer 1324 can include an M-dimensional vector (in the present example, M = 10), where M can include the number of classes that the program must choose from in classifying an object in an image. Other example outputs can also be provided. Each number in the N-dimensional vector can represent a probability that the object belongs to a particular class. In one illustrative example, if a 10-dimensional output vector representing ten different classes of objects is [0 0 0.05 0.8 0 0.15 0 0 0 0], then this vector indicates that there is a 5% probability that the image is of a third class of object (e.g., a dog), an 80% probability that the image is of a fourth class of object (e.g., a person), and a 15% probability that the image is of a sixth class of object (e.g., a kangaroo). The probability for a class can be thought of as a level of confidence that the object is part of that class.
[0204] FIG. 14 is a diagram illustrating an example of a system for implementing certain aspects of the present disclosure. In particular, An example of a computing system 1400 has been described. This computing system 1400 can be, for example, any computing device, camera system, or any component thereof constituting a computing system, wherein the components of the system communicate with each other using connection 1405. Connection 1405 can be a physical connection using a bus, or a direct connection to processor 1410 (such as in a chipset architecture). Connection 1405 can also be a virtual connection, a networking connection, or a logical connection.
[0205] In some examples, computing system 1400 is a distributed system, wherein the functionality described herein can be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent a plurality of such components, each performing some or all of the functionality described for that component. In some examples, the components can be physical or virtual devices.
[0206] Example system 1400 includes at least one processing unit (CPU or processor) 1410 and connectivity 1405 that couples various system components, including system memories 1415 such as read-only memory (ROM) 1420 and random access memory (RAM) 1425, to processor 1410. Computing system 1400 may include a cache 1412 of high-speed memory that is directly connected to, adjacent to, or integrated into processor 1410.
[0207] Processor 1410 may include any general-purpose processor and hardware or software services, such as services 1432, 1434, and 1436 stored in storage device 1430 and configured to control processor 1410, as well as dedicated processors in which software instructions are incorporated into the actual processor design. Processor 1410 may be a substantially self-contained computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0208] To enable user interaction, the computing system 1400 includes input devices 1445 that can represent any number of input mechanisms, such as a microphone for voice, a touchscreen for gesture or graphic input, a keyboard, a mouse, motion input, voice input, etc. The computing system 1400 may also include output devices 1435, which can be one or more of several output mechanisms. In some instances, a multimodal system allows users to provide multiple types of input / output to communicate with the computing system 1400. The computing system 1400 may include a communication interface 1440, which generally manages and controls user input and system output.
[0209] The communication interface can perform or facilitate receiving and / or transmitting wired or wireless communications using wired and / or wireless transceivers, including those utilizing an audio jack / plug, a microphone jack / plug, a Universal Serial Bus (USB) port / plug, an Apple® Lightning® port / plug, an Ethernet port / plug, a fiber optic port / plug, a dedicated wired port / plug, Bluetooth® wireless signal transmission, Bluetooth® Low Energy (BLE) wireless signal transmission, iBEACON® wireless signal transmission, Radio Frequency Identification (RFID) wireless signal transmission, Near Field Communication (NFC) wireless signal transmission, Dedicated Short-Range Communications (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, Wireless Local Area Network (WLAN) signal transmission, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof.
[0210] The communication interface 1440 can also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine a location of the computing system 1400 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States based Global Positioning System (GPS), the Russian based Global Navigation Satellite System (GLONASS), the Chinese based BeiDou Navigation Satellite System (BDS), and the European based Galileo GNSS. There is no limitation on operating on any particular hardware arrangement, and thus the underlying features herein can be readily substituted for improved hardware or firmware arrangements as they are developed.
[0211] Storage device 1430 may be a non-volatile and / or non-transient and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as magnetic tape, flash memory cards, solid-state storage devices, digital versatile discs, cartridges, floppy disks, hard disks, magnetic tapes, magnetic stripes, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, CD-ROM discs, rewritable CD discs, digital video discs (DVD discs), Blu-ray discs (BDD discs), holographic discs, another optical medium, secure digital storage (SD) cards, micro-secure digital storage (microSD) cards, memory. Stick® cards, smart card chips, EMV chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, other integrated circuit (IC) chips / cards, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase-change memory (PCM), spin-transfer torque RAM (STT-RAM), other memory chips or cartridges, and / or combinations thereof.
[0212] The storage device 1430 can include software services, servers, services, and the like that, when code defining such software is executed by the processor 1410, cause the system to perform a function. In some examples, a hardware service that performs a particular function can include the software components stored in a computer-readable medium that are necessary to perform the function coupled with the necessary hardware components, such as a processor 1410, the connection 1405, the output device 1435, and the like, to perform the function. The term computer-readable medium includes, but is not limited to, portable or fixed storage devices, optical storage devices, and various other mediums capable of storing, containing or carrying instruction(s) and / or data. A computer-readable medium can include a non-transitory medium in which data can be stored and which does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wires, cables, or other communication mediums. Examples of a non-transitory medium can include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium can have stored thereon code and / or machine-executable instructions that can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0213] In some examples, computer-readable storage devices, media, and memories can include cables or wireless signals containing bitstreams, etc. However, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se as media in this context.
[0214] In the above description, specific details are provided to provide a thorough understanding of the aspects and examples provided herein. However, one skilled in the relevant art will recognize that the aspects and examples can be practiced without some or all of the specific details, some of which are provided in the interest of conciseness and clarity. In some instances, the techniques of the application can be presented in terms of functional blocks that can include various functional means, components, steps, or routines in a software or firmware implementation. Additional components can be used, in addition to those shown and / or described herein. For example, circuitry, systems, networks, processes, and other components can be presented in block diagram form as components to avoid obscuring the aspects and examples of the present disclosure. In other instances, well-known structures have not been shown or described in order to avoid unnecessarily obscuring the aspects and examples being presented.
[0215] Individual aspects and examples above may be described as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations may be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process may correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0216] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise available from a computer-readable medium. These instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. Parts of the computer resources used are accessible via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, or source code. Examples of computer-readable media that can be used to store instructions, information used during the methods according to the described examples, and / or information created include hard disks or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, etc.
[0217] Devices implementing the various processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include: laptop devices, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mount devices, self-standing devices, etc. The functionality described herein may also be implemented using peripheral devices or plug-in cards. As a further example, such functionality may also be implemented on a circuit board within different chips or different processes executed on a single device.
[0218] Instructions, media for conveying these instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0219] In the descriptions above and in the claims, aspects of the applications are described with reference to particular examples. Those skilled in the art will appreciate that the applications are not limited to these specific examples and that there are other ways of implementing the applications. Hence, the specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a basis for the claims and as a representative way by which the claims can be practiced. The methods of the applications can be practiced in other ways, and that the appended claims are not construed as being limited to the specific examples disclosed herein.
[0220] Those of ordinary skill in the art will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein can be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of the specification.
[0221] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing the components to perform the operation, by programming the components to perform the operation, or any combination thereof. For example, a component can be a processor configured to perform operations by use of software instructions stored in memory.
[0222] The phrase “coupled to” means that any component is directly or indirectly connected to another component, and / or that any component is in communication with another component (e.g., connected to the other component by a wired or wireless connection, and / or other suitable communication interface).
[0223] The claim language of “at least one of” and / or “one or more of” a set, as recited in the disclosure, indicates that one member from the set or multiple members of the set (in any combination) satisfy the claims. For example, the claim language of “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, the claim language of “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language of “at least one of” and / or “one or more of” a set does not limit the set to the items listed in the set. For example, the claim language of “at least one of A and B” or “at least one of A or B” can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.
[0224] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0225] The techniques described herein can be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of various devices such as a general purpose computer, a wireless communication device handsets, or an integrated circuit device having a multitude of applications, including wireless communication devices handsets and other devices. Any features described as modules or components can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium can form part of a computer program product, which can include packaging materials. The computer-readable medium can comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, can be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0226] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein can refer to any of the above structures or any other structure suitable for implementation of the techniques described herein.
[0227] Illustrative aspects of the present disclosure include:
[0228] Aspect 1. An apparatus for processing image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: classify a first image as a keyframe based on a difference between the first image and a previous image, wherein the first image and the previous image are included in a plurality of images; generate, using a first machine learning network, an enhanced keyframe image corresponding to the first image and a hidden state output associated with the enhanced keyframe image; classify a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and generate, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, wherein the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced keyframe image.
[0229] Aspect 2. The apparatus of Aspect 1, wherein the at least one processor is further configured to: generate, using the second machine learning network, an updated hidden state output associated with the enhanced dependent frame image.
[0230] Aspect 3. The apparatus of Aspect 2, wherein, to generate the updated hidden state output, the at least one processor is configured to warp chrominance information of the second image with the hidden state output associated with the enhanced keyframe image.
[0231] Aspect 4. The apparatus of any of Aspects 2-3, wherein the at least one processor is further configured to: classify a third image of the plurality of images as a dependent frame based on a similarity between the third image and the second image; and generate, using the second machine learning network, an enhanced dependent frame image corresponding to the third image, wherein the enhanced dependent frame image corresponding to the third image is based on the third image and the updated hidden state output.
[0232] Aspect 5. The apparatus of any one of aspects 1 through 4, wherein the first image and the second image are low-light images, and wherein respective luminances associated with the first image and the second image are less than a low-light threshold.
[0233] Aspect 6. The apparatus of any one of aspects 1 through 5, wherein the at least one processor is further configured to: classify each respective image of the plurality of images as a low-light image or a non-low-light image using a low-light detection engine; and classify each low-light image of the plurality of images as a key frame or a dependent frame using a frame identification neural network.
[0234] Aspect 7. The apparatus of aspect 6, wherein, to classify each respective image using the low-light detection engine, the at least one processor is configured to: generate a first luminance histogram of the respective image and a second luminance histogram of a previous image, wherein the previous image and the respective image are consecutive images of the plurality of images; and determine a binary classification of the respective image, the binary classification based on the first luminance histogram and the second luminance histogram.
[0235] Aspect 8. The apparatus of any one of aspects 1 through 7, wherein the first machine learning network is different than the second machine learning network.
[0236] Aspect 9. The apparatus of any one of aspects 1 through 8, wherein the first machine learning network is a deep neural network (DNN), and wherein the first machine learning network comprises: a luminance enhancement subnetwork; and a chrominance enhancement subnetwork.
[0237] Aspect 10. The apparatus of aspect 9, wherein the hidden state output associated with the enhanced key frame image comprises: an enhanced luminance feature associated with the enhanced key frame image, the enhanced luminance feature generated using the luminance enhancement subnetwork; and an enhanced chrominance feature associated with the enhanced key frame image, the enhanced chrominance feature generated using the chrominance enhancement subnetwork.
[0238] Aspect 11. The apparatus of any one of aspects 9 through 10, wherein, to generate the hidden state output associated with the enhanced key frame image, the at least one processor is configured to: generate an enhanced chrominance feature associated with chrominance information of the first image using the chrominance enhancement subnetwork; and fuse the enhanced chrominance feature with an internal hidden state of the luminance enhancement subnetwork, wherein the internal hidden state is different than the hidden state output.
[0239] Aspect 12. The apparatus of any one of aspects 1 through 11, wherein the second machine learning network is a recurrent neural network (RNN); and a resolution of luminance information provided as input to the second machine learning network is less than or equal to half a resolution of luminance information provided as input to the first machine learning network.
[0240] Aspect 13. The apparatus of aspect 12, wherein: the resolution of the first image is the same as the resolution of the second image; and each image of the plurality of images has the resolution.
[0241] Aspect 14. The apparatus of any of aspects 1-13, wherein the at least one processor is configured to: provide, to the first machine-learned network, respective luma information, respective first chroma information, and respective second chroma information corresponding to the keyframe image; and provide, to the second machine-learned network, luma information corresponding to the dependent frame image.
[0242] Aspect 15. The apparatus of aspect 14, wherein, to generate the enhanced dependent frame image, the at least one processor is configured to: downscale the luma information corresponding to the dependent frame image, wherein the down scaled luma information has a resolution that is the same as a resolution associated with the hidden state output of the enhanced keyframe image.
[0243] Aspect 16. The apparatus of aspect 15, wherein the at least one processor is further configured to: generate, using a convolutional layer of the second machine-learned network, a correlation output tensor indicative of a correlation between the down scaled luma information and the hidden state output associated with the enhanced keyframe image; determine, using a cross-attention layer of the second machine-learned network, an attention map based on the correlation output tensor; and determine a feature alignment between features of the down scaled luma information and the hidden state output, the feature alignment based on the attention map.
[0244] Aspect 17. The apparatus of aspect 16, wherein the at least one processor is further configured to: generate, using the second machine-learned network, an updated hidden state output associated with the enhanced dependent frame image based on warping the feature alignment with the hidden state output associated with the enhanced keyframe image.
[0245] Aspect 18. The apparatus of any of aspects 1-17, wherein the at least one processor is configured to: classify the first image as a keyframe based on first motion information between the first image and a previous image; and classify the second image as a dependent frame based on second motion information between the second image and the first image.
[0246] Aspect 19. The apparatus of aspect 18, wherein: the first motion information includes a first perceptual similarity index value determined using the frame identification neural network, wherein the first motion information indicates a difference between the first image and the previous image based on the first perceptual similarity index value being less than a threshold value; and the second motion information includes a second perceptual similarity index value determined using the frame identification neural network, wherein the second motion information indicates a similarity between the second image and the first image based on the second perceptual similarity index value being greater than the threshold value.
[0247] Aspect 20. The apparatus of any one of Aspects 1 to 19, wherein the at least one processor is further configured to: classify each respective image of the plurality of images as a low-light image or a non-low-light image based on a lux index associated with the respective image; and classify each low-light image of the plurality of images as a keyframe or a dependent frame using a frame identification neural network.
[0248] Aspect 21. A method for processing image data, the method comprising: classifying a first image as a keyframe based on a difference between the first image and a previous image, wherein the first image and the previous image are included in a plurality of images; generating, using a first machine learning network, an enhanced keyframe image corresponding to the first image and a hidden state output associated with the enhanced keyframe image; classifying a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and generating, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, wherein the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced keyframe image.
[0249] Aspect 22. The method of Aspect 21, further comprising: generating, using the second machine learning network, an updated hidden state output associated with the enhanced dependent frame image.
[0250] Aspect 23. The method of Aspect 22, wherein generating the updated hidden state output comprises: warping chrominance information of the second image with the hidden state output associated with the enhanced keyframe image.
[0251] Aspect 24. The method of any one of Aspects 22 to 23, further comprising: classifying a third image of the plurality of images as a dependent frame based on a similarity between the third image and the second image; and generating, using the second machine learning network, an enhanced dependent frame image corresponding to the third image, wherein the enhanced dependent frame image corresponding to the third image is based on the third image and the updated hidden state output.
[0252] Aspect 25. The method of any one of Aspects 21 to 24, wherein the first image and the second image are low-light images, and wherein respective luminances associated with the first image and the second image are less than a low-light threshold.
[0253] Aspect 26. The method of any one of Aspects 21 to 25, further comprising: classifying each respective image of the plurality of images as a low-light image or a non-low-light image using a low-light detection engine; and classifying each low-light image of the plurality of images as a keyframe or a dependent frame using a frame identification neural network.
[0254] Aspect 27. The method of aspect 26, wherein classifying each respective image using the low-light detection engine comprises: generating a first luminance histogram of the respective image and a second luminance histogram of a previous image, wherein the previous image and the respective image are consecutive images in the plurality of images; and determining a binary classification of the respective image, the binary classification based on the first luminance histogram and the second luminance histogram.
[0255] Aspect 28. The method of any of aspects 21 to 27, wherein the first machine learning network is different from the second machine learning network.
[0256] Aspect 29. The method of any of aspects 21 to 28, wherein the first machine learning network is a deep neural network (DNN), and wherein the first machine learning network comprises: a luminance enhancer network; and a chrominance enhancer network.
[0257] Aspect 30. The method of aspect 29, wherein the hidden state output associated with the enhanced keyframe image comprises: an enhanced luminance feature associated with the enhanced keyframe image, the enhanced luminance feature generated using the luminance enhancer network; and an enhanced chrominance feature associated with the enhanced keyframe image, the enhanced chrominance feature generated using the chrominance enhancer network.
[0258] Aspect 31. The method of any of aspects 29 to 30, wherein generating the hidden state output associated with the enhanced keyframe image comprises: generating, using the chrominance enhancer network, an enhanced chrominance feature associated with chrominance information of the first image; and fusing the enhanced chrominance feature with an internal hidden state of the luminance enhancer network, wherein the internal hidden state is different from the hidden state output.
[0259] Aspect 32. The method of any of aspects 21 to 31, wherein the second machine learning network is a recurrent neural network (RNN); and a resolution of the luminance information provided as input to the second machine learning network is less than or equal to half of a resolution of the luminance information provided as input to the first machine learning network.
[0260] Aspect 33. The method of aspect 32, wherein: the resolution of the first image is the same as a resolution of the second image; and each image in the plurality of images has the resolution.
[0261] Aspect 34. The method of any of aspects 21 to 33, further comprising: providing the respective luminance information, the respective first chrominance information, and the respective second chrominance information corresponding to the keyframe image to the first machine learning network; and providing the luminance information corresponding to the dependent frame image to the second machine learning network.
[0262] Aspect 35. The method of aspect 34, wherein generating the augmented dependent frame image comprises: downscaling luminance information corresponding to the dependent frame image, wherein the downscaled luminance information has a same resolution as a resolution associated with the hidden state output of the augmented key frame image.
[0263] Aspect 36. The method of aspect 35, further comprising: generating, using a convolutional layer of the second machine learning network, a correlation output tensor indicative of a correlation information between the downscaled luminance information and the hidden state output associated with the augmented key frame image; determining, using a cross-attention layer of the second machine learning network, an attention map based on the correlation output tensor; and determining a feature alignment between features of the downscaled luminance information and the hidden state output, the feature alignment based on the attention map.
[0264] Aspect 37. The method of aspect 36, further comprising: generating, using the second machine learning network, an updated hidden state output associated with the augmented dependent frame image based on warping the feature alignment with the hidden state output associated with the augmented key frame image.
[0265] Aspect 38. The method of any one of aspects 21 to 37, further comprising: classifying the first image as a key frame based on first motion information between the first image and a previous image; and classifying the second image as a dependent frame based on second motion information between the second image and the first image.
[0266] Aspect 39. The method of aspect 38, wherein: the first motion information comprises a first perceptual similarity index value determined using the frame identification neural network, wherein the first motion information indicates a difference between the first image and the previous image based on the first perceptual similarity index value being less than a threshold value; and the second motion information comprises a second perceptual similarity index value determined using the frame identification neural network, wherein the second motion information indicates a similarity between the second image and the first image based on the second perceptual similarity index value being greater than the threshold value.
[0267] Aspect 40. The method of any one of aspects 21 to 39, further comprising: classifying each respective image of the plurality of images as a low-light image or a non-low-light image based on a lux index associated with the respective image; and classifying each low-light image of the plurality of images as a key frame or a dependent frame using the frame identification neural network.
[0268] Aspect 41: A computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations in accordance with any one of aspects 1 to 20.
[0269] Aspect 42: A computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations of any of aspects 21 through 40.
[0270] Aspect 43: A device for processing image data comprising one or more means for performing operations of any of aspects 1 through 20.
[0271] Aspect 44: A device for processing image data comprising one or more means for performing operations of any of aspects 21 through 40.
Claims
1. An apparatus for processing image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor configured to: classify a first image as a keyframe based on a difference between the first image and a previous image, wherein the first image and the previous image are included in a plurality of images; generate, using a first machine learning network, an enhanced keyframe image corresponding to the first image and a hidden state output associated with the enhanced keyframe image; classify a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and generate, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, wherein the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced keyframe image.
2. The apparatus of claim 1, wherein the at least one processor is further configured to: generate, using the second machine learning network, an updated hidden state output associated with the enhanced dependent frame image. To generate the updated hidden state output, the at least one processor is configured to:
3. The apparatus of claim 2, wherein, warp chrominance information of the second image with the hidden state output associated with the enhanced keyframe image.
4. The apparatus of claim 2, wherein the at least one processor is further configured to: classify a third image of the plurality of images as a dependent frame based on a similarity between the third image and the second image; and generate, using the second machine learning network, an enhanced dependent frame image corresponding to the third image, wherein the enhanced dependent frame image corresponding to the third image is based on the third image and the updated hidden state output.
5. The apparatus of claim 1, wherein the first image and the second image are low light images, and wherein respective luminances associated with the first image and the second image are less than a low light threshold.
6. The apparatus of claim 1, wherein the at least one processor is further configured to: classify each respective image of the plurality of images as a low light image or a non-low light image using a low light detection engine; and classify each low light image of the plurality of images as a keyframe or a dependent frame using a frame identification neural network. To classify each respective image using the low light detection engine, the at least one processor is configured to:
7. The apparatus of claim 6, wherein, generate a first luminance histogram of the respective image and a second luminance histogram of a previous image, wherein the previous image and the respective image are consecutive images of the plurality of images; and determine a binary classification of the respective image, the binary classification based on the first luminance histogram and the second luminance histogram.
8. The apparatus of claim 1, wherein the first machine learning network is different than the second machine learning network.
9. The apparatus of claim 1, wherein the first machine learning network is a deep neural network (DNN), and wherein the first machine learning network comprises: a brightness enhancer network; and a chrominance enhancer network.
10. The apparatus of claim 9, wherein the hidden state output associated with the enhanced key frame image comprises: enhanced brightness features associated with the enhanced key frame image, the enhanced brightness features generated using the brightness enhancer network; and enhanced chrominance features associated with the enhanced key frame image, the enhanced chrominance features generated using the chrominance enhancer network.
11. The apparatus of claim 9, wherein, To generate the hidden state output associated with the enhanced key frame image, the at least one processor is configured to: generate, using the chrominance enhancer network, enhanced chrominance features associated with chrominance information of the first image; and fuse the enhanced chrominance features with an internal hidden state of the brightness enhancer network, wherein the internal hidden state is different from the hidden state output.
12. The apparatus of claim 1, wherein: the second machine learning network is a recurrent neural network (RNN); and a resolution of brightness information provided as input to the second machine learning network is less than or equal to half of a resolution of brightness information provided as input to the first machine learning network.
13. The apparatus of claim 12, wherein: the resolution of the first image is the same as a resolution of the second image; and each image of the plurality of images has the resolution.
14. The apparatus of claim 1, wherein the at least one processor is configured to: provide, to the first machine learning network, respective brightness information, respective first chrominance information, and respective second chrominance information corresponding to key frame images; and provide, to the second machine learning network, brightness information corresponding to dependent frame images.
15. The apparatus of claim 14, wherein, To generate the enhanced dependent frame image, the at least one processor is configured to: downscale the brightness information corresponding to the dependent frame image, wherein the down-scaled brightness information has a resolution that is the same as a resolution of the hidden state output associated with the enhanced key frame image.
16. The apparatus of claim 15, wherein the at least one processor is further configured to: generate, using a convolutional layer of the second machine learning network, a correlation output tensor indicative of correlation information between the down-scaled brightness information and the hidden state output associated with the enhanced key frame image; determine, using a cross-attention layer of the second machine learning network, an attention map based on the correlation output tensor; and determine a feature alignment between features of the down-scaled brightness information and the hidden state output, the feature alignment based on the attention map.
17. The apparatus of claim 16, wherein the at least one processor is further configured to: generate, using the second machine learning network, an updated hidden state output associated with the enhanced dependent frame image based on warping the feature alignment with the hidden state output associated with the enhanced key frame image.
18. The apparatus of claim 1, wherein the at least one processor is configured to: classify the first image as a key frame based on first motion information between the first image and the previous image; and classify the second image as a dependent frame based on second motion information between the second image and the first image.
19. The apparatus of claim 18, wherein: the first motion information includes a first perceptual similarity index value determined using a frame identification neural network, wherein the first motion information indicates a difference between the first image and the previous image based on the first perceptual similarity index value being less than a threshold value; and the second motion information includes a second perceptual similarity index value determined using the frame identification neural network, wherein the second motion information indicates a similarity between the second image and the first image based on the second perceptual similarity index value being greater than the threshold value.
20. The apparatus of claim 1, wherein the at least one processor is further configured to: classify each respective image of the plurality of images as a low-light image or a non-low-light image based on a lux index associated with the respective image; and classify each low-light image of the plurality of images as a key frame or a dependent frame using a frame identification neural network.
21. A method for processing image data, the method comprising: classifying a first image as a key frame based on a difference between the first image and a previous image, wherein the first image and the previous image are included in a plurality of images; generating, using a first machine learning network, an enhanced key frame image corresponding to the first image and a hidden state output associated with the enhanced key frame image; classifying a second image of the plurality of images as a dependent frame based on a similarity between the second image and the first image; and generating, using a second machine learning network, an enhanced dependent frame image corresponding to the second image, wherein the enhanced dependent frame image is based on the second image and the hidden state output associated with the enhanced key frame image.
22. The method of claim 21, further comprising: generating, using the second machine learning network, an updated hidden state output associated with the enhanced dependent frame image.
23. The method of claim 22, wherein generating the updated hidden state output comprises: warping chrominance information of the second image with the hidden state output associated with the enhanced key frame image.
24. The method of claim 22, further comprising: classifying a third image of the plurality of images as a dependent frame based on a similarity between the third image and the second image; and generating, using the second machine learning network, an enhanced dependent frame image corresponding to the third image, wherein the enhanced dependent frame image corresponding to the third image is based on the third image and the updated hidden state output.
25. The method of claim 21, wherein the first image and the second image are low-light images, and wherein respective luminances associated with the first image and the second image are less than a low-light threshold.
26. The method of claim 21, further comprising: classifying each respective image of the plurality of images as a low-light image or a non-low-light image using a low-light detection engine; and classifying each low-light image of the plurality of images as a keyframe or a dependent frame using a frame identification neural network.
27. The method of claim 26, wherein classifying each respective image using the low-light detection engine comprises: generating a first luminance histogram of the respective image and a second luminance histogram of a previous image, wherein the previous image and the respective image are consecutive images of the plurality of images; and determining a binary classification of the respective image, the binary classification based on the first luminance histogram and the second luminance histogram.
28. The method of claim 1, wherein the first machine learning network is a deep neural network (DNN), and wherein the first machine learning network comprises: a luminance enhancer network; and a chrominance enhancer network.
29. The method of claim 28, wherein the hidden state output associated with the enhanced keyframe image comprises: an enhanced luminance feature associated with the enhanced keyframe image, the enhanced luminance feature generated using the luminance enhancer network; and an enhanced chrominance feature associated with the enhanced keyframe image, the enhanced chrominance feature generated using the chrominance enhancer network.
30. The method of claim 29, wherein generating the hidden state output associated with the enhanced keyframe image comprises: generating an enhanced chrominance feature associated with chrominance information of the first image using the chrominance enhancer network; and fusing the enhanced chrominance feature with an internal hidden state of the luminance enhancer network, wherein the internal hidden state is different from the hidden state output.