Method and device for processing one or more frames and non-transitory computer readable storage medium

By obtaining frames of different setting fields in the image capture system and transforming the initial frame based on the reference frame, the problem of high power consumption and difficulty in capturing unexpected events when capturing images or videos is solved, and video capture with lower power consumption and higher frame quality is achieved.

CN120111366APending Publication Date: 2025-06-06QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510288083.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-09-15
Filing Date
2022-08-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing AON cameras have high overall power consumption when capturing images or videos, resulting in a shorter battery life of mobile devices and it is difficult to capture clear frames of unexpected events when users cannot start video recording in advance.

Method used

By obtaining frames associated with different settings domains from the image capture system, including initial frames captured before capturing input, reference frames close to capture input time, and frames captured thereafter, transformed frames associated with the second settings domain are generated based on the reference frame.

Benefits of technology

Effectively reduces the power consumption of AON cameras, extends the battery life of mobile devices, and improves the ability to capture clear frames when unexpected events occur.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111366A_ABST
    Figure CN120111366A_ABST
Patent Text Reader

Abstract

Systems and techniques for processing one or more frames are provided. For example, a process may include obtaining a first plurality of frames associated with a first setup domain from an image capture system, where the first plurality of frames are captured prior to obtaining a capture input. The process may include obtaining a reference frame associated with a second set domain from the image capture system, where the reference frame is captured proximate to obtaining the capture input. The process may include obtaining a second plurality of frames associated with a second setting domain from the image capture system, where the second plurality of frames is captured after the reference frame. The process may include transforming, based on the reference frame, at least a portion of the first plurality of frames to generate a transformed plurality of frames associated with the second setting domain.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application with a filing date of August 10, 2022, entitled “A method, device and non-transitory computer-readable storage medium for processing one or more frames” and application number 202280061210.4. Technical Field

[0002] The present disclosure relates generally to the capture of images and / or video, and more particularly to systems and techniques for performing film shutter lag capture. Background Art

[0003] Many devices and systems allow for capturing a scene by generating images (or frames) and / or video data (including multiple frames) of the scene. For example, a camera or device including a camera can capture a sequence of frames of a scene (e.g., a video of the scene). In some cases, the sequence of frames can be processed to perform one or more functions, can be output for display, can be output for processing and / or consumption by other devices, and other uses.

[0004] Devices (e.g., mobile devices) and systems increasingly utilize dedicated ultra-low power camera hardware for "always on" (AON) camera use cases, where the camera can remain on to continuously record while maintaining a lower power usage footprint. AON cameras can capture images or videos of unexpected events, where a user may not be able to initiate a video recording before the event occurs. However, the overall power consumption of an AON camera setting to capture images or videos can still significantly reduce the battery life of the mobile device, which typically has a limited battery life AON. In some cases, an AON camera setting can utilize low-power camera hardware to reduce power consumption. Summary of the invention

[0005] In some examples, systems and techniques for providing film shutter lag video and / or image capture are described. According to at least one illustrative example, a method for processing one or more frames is provided. The method includes: obtaining a first plurality of frames associated with a first setup domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtaining at least one reference frame associated with a second setup domain from the image capture system, wherein the at least one reference frame is captured proximate to obtaining the capture input; obtaining a second plurality of frames associated with the second setup domain from the image capture system, wherein the second plurality of frames are captured after the at least one reference frame; based on the at least one reference frame, transforming at least a portion of the first plurality of frames to generate a transformed plurality of frames associated with the second setup domain.

[0006] In another example, an apparatus for processing one or more frames is provided, comprising: at least one memory (e.g., configured to store data, such as virtual content data, one or more images, etc.); and at least one processor (e.g., implemented in a circuit system) coupled to the at least one memory. The one or more processors are configured to execute instructions to and are capable of: obtaining a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtaining at least one reference frame associated with a second setting domain from the image capture system, wherein the at least one reference frame is captured close to obtaining the capture input; obtaining a second plurality of frames associated with the second setting domain from the image capture system, wherein the second plurality of frames are captured after the at least one reference frame; based on the at least one reference frame, transforming at least a portion of the first plurality of frames to generate a transformed plurality of frames associated with the second setting domain.

[0007] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided, which, when executed by one or more processors, causes the one or more processors to: obtain a first plurality of frames associated with a first setup domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtain at least one reference frame associated with a second setup domain from the image capture system, wherein the at least one reference frame is captured close to obtaining the capture input; obtain a second plurality of frames associated with the second setup domain from the image capture system, wherein the second plurality of frames are captured after the at least one reference frame; and transform at least a portion of the first plurality of frames based on the at least one reference frame to generate a transformed plurality of frames associated with the second setup domain.

[0008] In another example, an apparatus for processing one or more frames is provided. The apparatus includes: a unit for obtaining a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; a unit for obtaining at least one reference frame associated with a second setting domain from the image capture system, wherein the at least one reference frame is captured close to obtaining the capture input; a unit for obtaining a second plurality of frames associated with the second setting domain from the image capture system, wherein the second plurality of frames are captured after the at least one reference frame; and a unit for transforming at least a portion of the first plurality of frames based on the at least one reference frame to generate a transformed plurality of frames associated with the second setting domain.

[0009] In some aspects, the first setting domain includes a first resolution and the second setting domain includes a second resolution. In some cases, to transform at least a portion of the first plurality of frames, the methods, apparatus, and computer-readable media described above may include upscaling at least a portion of the first plurality of frames from the first resolution to a second resolution to generate the transformed plurality of frames, wherein the transformed plurality of frames has the second resolution.

[0010] In some aspects, the methods, devices, and computer-readable media described above may also include obtaining an additional reference frame having the second resolution from the image capture system, wherein the additional reference frame is captured before obtaining the capture input, and generating the enlarged multiple frames having the second resolution is based on at least the portion of the first multiple frames, the at least one reference frame, and the additional reference frame, and wherein the at least one reference frame provides a reference for enlarging at least a first portion of at least the portion of the first multiple frames, and the additional reference frame provides a reference for enlarging at least a second portion of at least the portion of the first multiple frames.

[0011] In some aspects, the methods, devices, and computer-readable media described above may further include combining the transformed plurality of frames and the second plurality of frames to generate a video associated with the second settings domain.

[0012] In some aspects, the methods, devices, and computer-readable media described above may also include: obtaining motion information associated with the first plurality of frames, wherein generating the transformed plurality of frames associated with the second setting domain is based on at least the portion of the first plurality of frames, the at least one reference frame, and the motion information.

[0013] In some aspects, the methods, devices, and computer-readable media described above may also include determining a translation direction based on motion information associated with the first plurality of frames; and applying the translation direction to the transformed plurality of frames.

[0014] In some aspects, the first setting domain includes a first frame rate and the second setting domain includes a second frame rate. In some cases, in some cases, to transform at least a portion of the first plurality of frames, the methods, apparatus, and computer-readable media described above may include: frame rate converting at least the portion of the first plurality of frames from the first frame rate to the second frame rate.

[0015] In some aspects, a first subset of the first plurality of frames is captured at the first frame rate, and a second subset of the first plurality of frames is captured at a third frame rate that is different from the first frame rate. In some cases, the third frame rate is equal to or different from the second frame rate. In some aspects, a change between the first frame rate and the third frame rate is based at least in part on motion information associated with at least one of the first subset of the first plurality of frames and the second subset of the first plurality of frames.

[0016] In some aspects, the first setting domain includes a first resolution and a first frame rate, and the second setting domain includes a second resolution and a second frame rate. In some cases, in some cases, to transform at least a portion of the first plurality of frames, the above methods, apparatus, and computer-readable media may include: upscaling at least a portion of the first plurality of frames from the first resolution to the second resolution, and frame rate converting at least the portion of the first plurality of frames from the first frame rate to the second frame rate.

[0017] In some aspects, the methods, devices, and computer-readable media described above also include: obtaining an additional reference frame associated with the second setting domain from the image capture system, wherein the additional reference frame is captured before obtaining the capture input, wherein generating the transformed multiple frames associated with the second setting domain is based on at least the portion of the first multiple frames, at least one reference frame and the additional reference frame, and wherein the at least one reference frame provides a reference for transforming at least a first subset of at least the portion of the first multiple frames, and the additional reference frame provides a reference for transforming at least a second subset of at least the portion of the first multiple frames.

[0018] In some aspects, the methods, devices, and computer-readable media described above also include: obtaining a second reference frame associated with the second setting domain from the image capture system, wherein the second reference frame is captured close to obtaining the capture input; based on the first reference frame, transforming at least the portion of the first plurality of frames to generate the transformed plurality of frames associated with the second setting domain; and based on the second reference frame, transforming at least another portion of the first plurality of frames to generate a second transformed plurality of frames associated with the second setting domain.

[0019] In some aspects, the methods, devices, and computer-readable media described above may also include: obtaining a motion estimate associated with the first plurality of frames; obtaining a third reference frame associated with the second setup domain from the image capture system, wherein the third reference frame is captured before obtaining the capture input; and based on the third reference frame, transforming a third portion of the first plurality of frames to generate a third transformed plurality of frames associated with the second setup domain; wherein the amount of time between the first reference frame and the third reference frame is based on the motion estimate associated with the first plurality of frames.

[0020] In some aspects, the first setting domain includes at least one of a first resolution, a first frame rate, a first color depth, a first noise reduction technique, a first edge enhancement technique, a first image stabilization technique, and a first color correction technique, and the second setting domain includes at least one of a second resolution, a second frame rate, a second color depth, a second noise reduction technique, a second edge enhancement technique, a second image stabilization technique, and a second color correction technique.

[0021] In some aspects, the methods, devices, and computer-readable media described above may also include generating the transformed multiple frames using a trainable neural network, wherein the neural network is trained using a training data set including image pairs, each pair of images including a first image associated with the first setting domain and a second image associated with the second setting domain.

[0022] In some aspects, capturing the at least one reference frame proximate to obtaining the capture input includes: capturing a first available frame associated with the second setup domain after receiving the capture input; capturing a second available frame associated with the second setup domain after receiving the capture input; capturing a third available frame associated with the second setup domain after receiving the capture input; or capturing a fourth available frame associated with the second setup domain after receiving the capture input.

[0023] In some aspects, capturing the at least one reference frame in proximity to obtaining the capture input includes capturing a frame associated with the second setup domain within 10 milliseconds (ms), within 100 ms, within 500 ms, or within 1000 ms after receiving the capture input.

[0024] According to at least one other example, a method for processing one or more frames is provided. The method includes: obtaining a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtaining a reference frame associated with a second setting domain from the image capture system, wherein the reference frame is captured close to obtaining the capture input; obtaining a selection of one or more selected frames associated with the first plurality of frames; and transforming the one or more selected frames to generate one or more transformed frames associated with the second setting domain based on the reference frame.

[0025] In another example, an apparatus for processing one or more frames is provided, comprising: at least one memory (e.g., configured to store data, such as virtual content data, one or more images, etc.); and at least one processor (e.g., implemented in a circuit system) coupled to the at least one memory. The one or more processors are further configured to execute instructions and may: obtain a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtain a reference frame associated with a second setting domain from the image capture system, wherein the reference frame is captured close to obtaining the capture input; obtain a selection of one or more selected frames associated with the first plurality of frames; based on the reference frame, transform the one or more selected frames to generate one or more transformed frames associated with the second setting domain.

[0026] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided, the instructions, when executed by one or more processors, cause the one or more processors to: obtain a first plurality of frames associated with a first setup domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtain a reference frame associated with a second setup domain from the image capture system, wherein the reference frame is captured close to obtaining the capture input; obtain a selection of one or more selected frames associated with the first plurality of frames; and transform the one or more selected frames based on the reference frame to generate one or more transformed frames associated with the second setup domain.

[0027] In another example, an apparatus for processing one or more frames is provided. The apparatus includes: means for obtaining a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; means for obtaining a reference frame associated with a second setting domain from the image capture system, wherein the reference frame is captured close to obtaining the capture input; means for obtaining a selection of one or more selected frames associated with the first plurality of frames; and means for transforming the one or more selected frames to generate one or more transformed frames associated with the second setting domain based on the reference frame.

[0028] In some aspects, selection of the one or more selected frames is based on a selection from a user interface.

[0029] In some aspects, the user interface includes a thumbnail gallery, a slider, or a frame-by-frame review.

[0030] In some aspects, the methods, devices, and computer-readable media described above may also include: determining one or more suggested frames from the first plurality of frames based on one or more of the following: determining whether the amount of motion or the amount of change in motion exceeds a threshold; determining which frames contain content of interest based on the presence of one or more human faces; and determining which frames contain content similar to a set of labeled image sets.

[0031] In some aspects, one or more of the devices described above are or are part of a vehicle (e.g., a computing device of a vehicle), a mobile device (e.g., a mobile phone or so-called "smart phone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some aspects, a device includes a camera or multiple cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device may include one or more sensors, which may be used to determine the position and / or posture of the device, the state of the device, and / or for other purposes.

[0032] This summary is neither intended to identify key or essential features of the claimed subject matter nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0033] The foregoing and other features and embodiments will become more fully apparent after reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Illustrative embodiments of the present application are described in detail below with reference to the following drawings:

[0035] Figure 1 is a block diagram illustrating the architecture of an image capture and processing system according to some examples;

[0036] Figure 2 is a schematic diagram illustrating the architecture of an example extended reality (XR) system according to some examples;

[0037] Figure 3 is a block diagram illustrating an example of an image processing system according to some examples;

[0038] Figure 4A is a schematic diagram illustrating an example Always-On (AON) camera use case according to some examples;

[0039] Figure 4B is a diagram illustrating an example film shutter hysteresis use case according to some examples;

[0040] Figure 5A is a block diagram illustrating an example film shutter hysteresis system according to some examples.

[0041] Figure 5B is a block diagram illustrating another example film shutter hysteresis system according to some examples;

[0042] Figure 6 is a flow chart illustrating an example of a process for processing one or more frames according to some examples;

[0043] Figure 7 is a flow chart illustrating an example of a process for processing one or more frames according to some examples;

[0044] Fig. 8A is a flow chart illustrating an example of a process for performing film shutter lag capture according to some examples;

[0045] Figure 8B is shown according to some examples Fig. 8A A schematic diagram of an example of relative power consumption levels during a film shutter lag capture process shown in;

[0046] Fig. 9 is a block diagram illustrating another example film shutter hysteresis system according to some examples;

[0047] Fig.10is a flow chart illustrating another example of a process for processing one or more frames according to some examples;

[0048] Fig.11 is a block diagram illustrating another example film shutter hysteresis system according to some examples;

[0049] Fig.12 is a flow chart illustrating an example of a process for processing one or more frames according to some examples;

[0050] Fig.13 is a block diagram illustrating another example film shutter hysteresis system according to some examples;

[0051] Fig.14 is a flow chart illustrating another flow chart showing an example of a film shutter lag frame capture sequence according to some examples;

[0052] Fig.15 is a diagram illustrating an example of relative power consumption levels during a film shutter lag frame capture sequence according to some examples;

[0053] Fig.16 is a flow chart illustrating an example of a process for processing one or more frames according to some examples;

[0054] Fig.17 is a block diagram illustrating an example of a deep learning network according to some examples;

[0055] Fig.18 is a block diagram illustrating an example of a conventional neural network according to some examples;

[0056] Fig.19 is a schematic diagram illustrating an example of a computing system for implementing certain aspects described herein. DETAILED DESCRIPTION

[0057] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some of them can be applied in combination, as will be apparent to those skilled in the art. In the following description, for ease of explanation, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it is apparent that each embodiment can be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.

[0058] The following description only provides exemplary embodiments, and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of the exemplary embodiments will provide a feasible description for implementing the exemplary embodiments for those skilled in the art. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0059] Systems and techniques for providing film shutter lag video and / or image capture systems are described herein. In some examples, an image capture system can implement a lower power or "always on" (AON) camera that operates continuously or periodically to automatically detect certain objects in an environment. For example, an image capture system that can capture video using an AON camera can be useful where an AON camera is always pointed at a target of interest. For example, one or more cameras included in a head-mounted device (e.g., a virtual reality (VR) or augmented reality (AR) head-mounted display (HMD), AR glasses, etc.) can always point to where the user is looking, based on the movement of the user's head. Although examples are described herein with reference to AON cameras, such aspects can be applied to any camera or image sensor operating in a low-power mode.

[0060] In some cases, AON cameras can utilize low-power camera hardware to reduce power consumption. In some cases, AON cameras can operate with low-power settings and / or perform different or fewer image processing steps for reducing power consumption. The amount of power consumed by AON cameras may depend on the domain (also referred to as setting domain in this article) of image and / or video frame capture. In some cases, the domain of the AON camera that captures images and / or video frames may include configuration parameters (e.g., resolution, frame rate, color depth, etc.) and / or other parameters of the AON camera. In some cases, the domain may also include processing steps (e.g., noise reduction, edge enhancement, image stabilization, color correction, etc.) performed by the camera pipeline of the AON camera to the captured image and / or video frame. In some cases, after the user initiates video and / or image capture, non-AON cameras can capture still images or start capturing video frames. Compared with AON or other low-power cameras, non-AON cameras can utilize higher-power camera hardware and / or can operate with higher power settings and / or different or more image processing steps.

[0061] It should be understood that, although specific examples of the present disclosure are discussed in terms of AON cameras (or AON camera sensors) and main cameras (or main camera sensors), the systems and techniques described herein can be applied to many different camera or sensor configurations without departing from the scope of the present disclosure. In an illustrative example, a single camera or sensor can be configured to operate in different operating modes (e.g., low-power AON mode and non-AON mode). The "AON operation" mentioned herein can be understood to include capturing images and / or frames with one or more AON cameras and / or operating one or more cameras in AON mode. In the present disclosure, frames captured during AON operation (e.g., video frames or images) are sometimes referred to as low-power frames (e.g., low-power video frames). Similarly, references to "standard operation" can be understood to include capturing needles using one or more non-AON cameras and / or one or more cameras operating in non-AON mode. In addition, non-AON cameras or sensors and / or cameras or sensors operating in non-AON mode can sometimes be referred to as one or more "main cameras" or "main camera sensors." In this disclosure, frames captured during standard camera operation (e.g., video frames) are sometimes referred to as high-power frames (e.g., high-power video frames). Systems and techniques will be described herein as performing with respect to video frames. However, it should be understood that the systems and techniques can operate using any sequence of images or frames, such as continuously captured still images.

[0062] A user may initiate image and / or video capture by, for example, pressing a capture or record button (also referred to as a shutter button in some cases), performing a gesture, and / or providing any other method of capture input. The time at which the image capture system receives the capture input may be defined as time t=0. Video frames captured during AON operation (e.g., at time t<0) may be stored in a memory (e.g., a buffer, a circular buffer, a video memory, etc.) that maintains the captured video frames for a period of time (e.g., 30 seconds, 1 minute, 5 minutes, etc.). Frames captured during AON operation may be associated with a first setup domain (hereinafter also referred to as the first domain). In one illustrative example, the first setup domain may include a wide video graphics array (WVGA) resolution (e.g., 800 pixels x 480 pixels) and a frame rate of 30 frames per second (fps).

[0063] After the image capture system receives the capture input, the image capture system may begin capturing video frames in standard operation. In some cases, the video frames captured during standard operation may be associated with a second setting domain (hereinafter also referred to as the second domain) that is different from the first setting domain. In an illustrative example, the second setting domain may include an ultra-high definition (UHD) resolution (e.g., 3840 pixels x 2160 pixels) and a frame rate of 30fps. The frames captured during standard operation may include capturing a key frame (also referred to herein as a reference frame) associated with the second setting domain (e.g., with UHD resolution) at or near t=0. As an illustrative example, capturing a key frame close to the capture input may include capturing a first available frame associated with the second setting domain after receiving the capture input. In some cases, capturing a key frame close to the capture input may include capturing a second frame, a third frame, a fourth frame, or a fifth frame or any other suitable available frame associated with the second setting domain after receiving the capture input. In some cases, capturing a key frame approaching a capture input may include capturing a frame associated with the second setting domain within 10 milliseconds (ms), within 100ms, within 500ms, within 1000ms, or within any other suitable time window after receiving the capture input. If the capture input occurs in response to an unexpected event (e.g., a goal in a child scoring soccer, a pet performing a skill, etc.), the capture input may occur after the user is interested in capturing an event (e.g., time t<0). In some cases, the image processing system may have captured an event during AON operation. In some cases, the video frame captured during AON operation associated with the first domain may have a significantly different appearance (e.g., different resolutions, frame rates, color depths, sharpness, etc.) from the video frame associated with the second domain and the video frame captured during standard operation. In some cases, using key frames as a guide, the video frame captured during AON operation associated with the first domain may be transformed to the second setting domain. In an illustrative example, transforming a video frame from the first setting domain to the second setting domain may include resolution scaling, frame rate scaling, changing color depth, and / or other transformations described herein. In some cases, a composite (or stitched) video may be formed from the transformed video frames associated with the second set domain and the video frames associated with the second set domain that were captured after receiving the capture input.

[0064] The process of transforming a video frame from a first domain to a second domain using a key frame may be referred to as a domain transform or a guided domain transform. In some cases, a deep learning neural network (e.g., a domain transform model) may be trained to perform a guided domain transform at least in part by converting a first portion of a video frame associated with a second domain (e.g., from a video included in a training data set) to a first domain, providing a key frame associated with the second domain from the original video, and transforming the first portion of the video frame back to the second domain using the key frame as a guide. The resulting transformed video frame may be directly compared to the original video frame, and a loss function may be used to determine the amount of error between the transformed video frame and the original frame. The parameters (e.g., weights, biases, etc.) of the deep learning network may be adjusted (or tuned) based on the error. Such a training process may be referred to as supervised learning using back propagation, which may be performed until the tuned parameters provide the desired results. In some cases, a deep generative neural network model (e.g., a generative adversarial network (GAN)) may be used to train the domain transform model.

[0065] In some cases, AON operation can include capturing video frame data at a low resolution (e.g., VGA resolution of 640 pixels x 480 pixels or WVGA resolution). In some cases, standard operation can include capturing video data at a higher resolution than AON operation (e.g., 720p, 1080i, 1080p, 4K, 8K, etc.). Capturing and processing video data at a lower resolution during AON operation can consume less power than capturing and processing video data at a higher resolution during standard operation. For example, the amount of power required to read the video frame data captured by the image capture system can increase with the increase of the number of pixels read out. In addition, the amount of power required for the image post-processing of the captured video frame data can also increase with the increase of the number of pixels in the video frame. In some cases, low-resolution video frames can be enlarged to higher resolution. The examples of low-resolution capture during AON operation and higher-resolution capture during standard operation described above provide an illustrative example of different domains that can be used to capture video frames in a film shutter hysteresis system.

[0066] The process of enlarging using high-resolution key frames can be referred to as a guided super-resolution process. The guided super-resolution process is an illustrative example of a guided domain transformation. In some cases, a deep learning neural network (e.g., a guided super-resolution model) can be trained to perform a guided super-resolution process in a training process similar to the process of training the domain transformation model described above. For example, the guided super-resolution model can be trained at least in part by reducing a portion of a video frame from a high-resolution video (e.g., a video included in a training data set), providing a high-resolution key frame from the high-resolution video, and using the high-resolution key frame as a guide to enlarge the reduced portion of the video frame from the high-resolution video back to the original resolution. The enlarged video frame can be directly compared with the original high-resolution video frame, and the loss function can be used to determine the amount of error between the enlarged video frame and the original high-resolution video frame of the video. The parameters (e.g., weights, biases, etc.) of the deep learning network can be adjusted (or tuned) based on the error. Such a training process can be referred to as supervised learning using back propagation, which can be performed until the tuned parameters provide the desired results. In some cases, a GAN can be used to train a guided super-resolution model.

[0067] In some cases, an image capture system operating in AON mode can capture video with a lower frame rate than the video captured during high-resolution mode. By reducing the frame rate and / or resolution during AON mode, the image capture system can operate with lower power consumption during AON mode when compared to high-resolution mode. In some cases, additional power can be saved in AON mode by utilizing an adaptive frame rate. In some implementations, the image capture system can determine the amount of motion of the image capture system (e.g., the motion of the head-mounted device) using inertial motion estimation from an inertial sensor. In some implementations, the image capture system can perform optical motion estimation (e.g., by analyzing captured video frames) to determine the amount of motion of the image capture system. If the image capture system still or only has a small amount of motion, the frame rate for capturing video frames can be set to a lower frame rate setting (e.g., to 15fps). On the other hand, if the image capture system and / or the scene being captured has a large amount of motion, the frame rate for capturing video frames during AON operation can be set to a high frame rate setting (e.g., 30fps, 60fps). In some cases, the maximum frame rate during AON mode can be equal to the frame rate setting of the high-resolution mode. In some cases, in addition to storing video frames in memory during AON mode, information indicating the frame rate at which the video frames were captured may also be stored.

[0068] In some implementations, in addition to capturing key frames near the capture input at the moment of receiving the capture input (e.g., at t=0) or capturing key frames near the capture input, the image capture system can periodically capture additional high-resolution key frames during AON operation. As an illustrative example, capturing key frames near the capture input may include capturing the first available high-resolution frame after receiving the capture input. In some cases, capturing key frames near the capture input may include capturing the second, third, fourth, or fifth high-resolution available frame after receiving the capture input. In some cases, capturing key frames near the capture input may include capturing high-resolution frames within 10ms, within 100ms, within 500ms, within 1000ms, or within any other suitable time window after receiving the capture input. For example, key frames may be captured every half second, every second, every five seconds, every ten seconds, or any other suitable duration. The additional high-resolution key frames captured during AON operation may be stored in a memory (e.g., a buffer) together with the low-resolution video frames. Each captured high-resolution key frame may be used as a guide to the video frames captured during AON operation that are close to the key frame in time. In some cases, additional high-resolution keyframes can improve guided super-resolution processes in which the image capture system moves over time and / or objects in the scene move over time. By periodically capturing high-resolution keyframes, the scene captured in each keyframe is more likely to capture at least a portion of the same scene captured during AON operation within a particular time period (e.g., within half a second, one second, five seconds, or ten seconds). In some cases, the capture rate of high-resolution keyframes can be determined based on the amount of motion of the image capture system. In some cases, the amount of motion can be determined based on readings from an inertial motion sensor and / or based on performing optical motion detection on video frames captured during AON operation.

[0069] In some cases, inertial sensor data may also be stored in memory along with video frames captured during AON mode and utilized in the guided super-resolution process. For example, if data from an inertial sensor captured during AON mode indicates that the image capture system is moving to the left, guided super-resolution may utilize information about the motion in the scene to ensure, for example, that the video it outputs is appropriately panned to the left to match the measured motion.

[0070] Although specific example domain transformations (e.g., resolution upscaling, frame rate adjustment, and color depth adjustment) are described herein as being used to perform film shutter lag capture, the systems and techniques described herein can be used to perform film shutter lag capture using other domain transformations without departing from the scope of the present disclosure. For example, a film shutter lag system can capture monochrome video frames during AON operation and color video frames during standard operation. In such an example, the domain transformation can include colorizing the monochrome frames.

[0071] Various aspects of the technology described herein are discussed below with respect to the accompanying figures. Figure 1 1 is a block diagram illustrating the architecture of image capture and processing system 100. Image capture and processing system 100 includes various components for capturing and processing an image of a scene (e.g., an image of scene 110). Image capture and processing system 100 can capture independent images (or photographs) and / or can capture a video including multiple images (or video frames) in a specific sequence. Lens 115 of system 100 faces scene 110 and receives light from scene 110. Lens 115 bends the light toward image sensor 130. Light received by lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by image sensor 130.

[0072] The one or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. The one or more control mechanisms 120 may include a plurality of mechanisms and components; for example, the control mechanisms 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. The one or more control mechanisms 120 may also include additional control mechanisms beyond those shown, such as control mechanisms to control analog gain, flash, HDR, depth of field, and / or other image capture attributes.

[0073] A focus control mechanism 125B in the control mechanism 120 may obtain a focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a storage register. Based on the focus setting, the focus control mechanism 125B may adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B may move the lens 115 closer to the image sensor 130 or further away from the image sensor 130 by driving a motor or servo device (or other lens mechanism), thereby adjusting the focus. In some cases, additional lenses may be included in the system 100, such as one or more microlenses above each photodiode of the image sensor 130, each of which bends the light received from the lens 115 toward the corresponding photodiode before it reaches the photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid focus (HAF), or some combination thereof. The focus setting may be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus setting may be referred to as an image capture setting and / or an image processing setting.

[0074] Exposure control mechanism 125A of control mechanism 120 may obtain an exposure setting. In some cases, exposure control mechanism 125A stores the exposure setting in a storage register. Based on this exposure setting, exposure control mechanism 125A may control the size of the aperture (e.g., aperture size or f / stop), the duration that the aperture is open (e.g., exposure time or shutter speed), the sensitivity of image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.

[0075] The zoom control mechanism 125C of the control mechanism 120 can obtain the zoom setting. In some examples, the zoom control mechanism 125C stores the zoom setting in a storage register. Based on the zoom setting, the zoom control mechanism 125C can control the focal length of the lens element assembly (lens assembly), which includes the lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servo mechanisms (or other lens mechanisms) to move one or more lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a zoom zoom lens. In some examples, the lens assembly can include a focusing lens (which can be a lens 115 in some cases) that first receives light from the scene 110, wherein the light then passes through an afocal zoom system between the focusing lens (e.g., lens 115) and the image sensor 130 before the light reaches the image sensor 130. In some cases, the afocal zoom system can include two positive (e.g., converging, convex) lenses having equal or similar focal lengths (e.g., within a threshold difference of each other) with a negative (e.g., diverging, concave) lens between them. In some cases, zoom control mechanism 125C moves one or more of the lenses in the afocal zoom system, such as a negative lens and one or both of the positive lenses.

[0076] The image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by the image sensor 130. In some cases, different photodiodes may be covered by different color filters, and thus may measure light that matches the color of the filter covering the photodiode. For example, a Bayer color filter includes a red filter, a blue filter, and a green filter, wherein each pixel of the image is generated based on red light data from at least one photodiode covered in the red filter, blue light data from at least one photodiode covered in the blue filter, and green light data from at least one photodiode covered in the green filter. Other types of color filters may use yellow, magenta, and / or cyan (also referred to as "emerald") color filters instead of red, blue, and / or green color filters, or use yellow, magenta, and / or cyan (also referred to as "emerald") color filters in addition to red, blue, and / or green color filters. Some image sensors (e.g., image sensor 130) may lack color filters entirely, and may instead use different photodiodes (stacked vertically in some cases) throughout the pixel array. The different photodiodes throughout the pixel array may have different spectral sensitivity curves, responding to different wavelengths of light. Monochrome image sensors may also lack color filters, and therefore lack color depth.

[0077] In some cases, the image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which may be used for phase detection autofocus (PDAF). The image sensor 130 may also include an analog gain amplifier for amplifying an analog signal output by the photodiode and / or an analog-to-digital converter (ADC) for converting the analog signal output of the photodiode (and / or amplified by the analog gain amplifier) ​​into a digital signal. In some cases, certain components or functions discussed with respect to one or more of the control mechanisms 120 may alternatively or additionally be included in the image sensor 130. The image sensor 130 may be a charge coupled device (CCD) sensor, an electron multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal oxide semiconductor (CMOS), an N-type metal oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0078] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more processors of any other type of processor 1910 discussed with respect to the computing system 1900. The host processor 152 may be a digital signal processor (DSP) and / or other types of processors. In some implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system on a chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), a memory, a connection component (e.g., Bluetooth™, a global positioning system (GPS), etc.), any combination thereof, and / or other components. The I / O ports 156 may include any suitable input / output ports or interfaces according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof), and / or other input / output ports. In one illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port, and the ISP 154 may communicate with the image sensor 130 using a MIPI port.

[0079] The image processor 150 may perform a number of tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. The image processor 150 may store image frames and / or processed images in a random access memory (RAM) 140 / 1925, a read-only memory (ROM) 145 / 1920, a cache, a memory unit, another storage device, or some combination thereof.

[0080] Various input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touch pad, a touch-sensitive surface, a printer, any other output device 1935, any other input device 1945, or some combination thereof. In some cases, subtitles may be input into the image processing device 105B via a physical keyboard or keypad of the I / O device 160 or via a virtual keyboard or keypad of a touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that implement a wired connection between the system 100 and one or more peripheral devices, through which the system 100 may receive data from one or more peripheral devices and / or transmit data to one or more peripheral devices. The I / O 160 may include one or more wireless transceivers that implement a wireless connection between the system 100 and one or more peripheral devices, through which the system 100 may receive data from one or more peripheral devices and / or transmit data to one or more peripheral devices. The peripheral devices may include any of the types of I / O devices 160 discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they themselves may be considered I / O devices 160.

[0081] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some implementations, the image capture device 105A and the image processing device 105B may be disconnected from each other.

[0082] like Figure 1 As shown, the vertical dashed line will Figure 1 The image capture and processing system 100 is divided into two parts, which are respectively represented by the image capture device 105A and the image processing device 105B. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, some of the components shown in the image capture device 105A (e.g., the ISP 154 and / or the host processor 152) can be included in the image capture device 105A.

[0083] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communications (such as cellular network communications, 802.11 wi-fi communications, wireless local area network (WLAN) communications, or some combination thereof). In some implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.

[0084] Although the image capture and processing system 100 is shown as including certain components, one of ordinary skill will appreciate that the image capture and processing system 100 may include more than Figure 1 . The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, the components of the image capture and processing system 100 may include electronic circuits or other electronic hardware and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device that implements the image capture and processing system 100.

[0085] In some examples, Figure 2 The extended reality (XR) system 200 may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.

[0086] Figure 22 is a schematic diagram illustrating the architecture of an XR system 200 according to some aspects of the present disclosure. The XR system 200 can run (or execute) XR applications and implement XR operations. In some examples, as part of the XR experience, the XR system 200 can perform tracking and positioning, mapping of the environment in the physical world (e.g., a scene), and positioning and rendering of virtual content on a display 209 (e.g., a screen, a visible plane / area, and / or other display). For example, the XR system 200 can generate a map of the environment in the physical world (e.g., a three-dimensional (3D) map), track the posture (e.g., position and positioning) of the XR system 200 relative to the environment (e.g., relative to a 3D map of the environment), position and / or anchor virtual content in a specific location on the map of the environment, and render the virtual content on the display 209 so that the virtual content appears to be at a location in the environment corresponding to a specific location on the map of the scene where the virtual content is positioned and / or anchored. Display 209 may include glass, screens, lenses, projectors, and / or other means that allow a user to see the real world environment and also allow XR content to be overlaid, superimposed, blended, or otherwise displayed thereon.

[0087] In this illustrative example, XR system 200 includes one or more image sensors 202, accelerometer 204, gyroscope 206, storage device 207, computing component 210, XR engine 220, interface layout and input management engine 222, image processing engine 224, and rendering engine 226. It should be noted that Figure 2 The components 202-226 shown in FIG. 1 are non-limiting examples provided for purposes of illustration and explanation, and other examples may include, for example, Figure 2 For example, in some cases, XR system 200 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radars, light detection and ranging (LIDAR) sensors, radio unit detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 2 Although various components of XR system 200 (e.g., accelerometer 204) may be referred to herein in the singular, it should be understood that XR system 200 may include multiple of any component discussed herein (e.g., multiple accelerometers 204).

[0088] XR system 200 includes or communicates (wired or wirelessly) with input device 208. Input device 208 may include any suitable input device, such as a touch screen, a pen or other pointer device, a keyboard, mouse buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, a video game controller, a steering wheel, a joystick, a set of buttons, a trackball, a remote control, any other input device 1945 discussed herein, or any combination thereof. In some cases, image sensor 202 may capture images that may be processed for interpretation of gesture commands.

[0089] In some embodiments, one or more image sensors 202, accelerometers 204, gyroscopes 206, storage 207, multimedia component 203, computing component 210, XR engine 220, interface layout and input management engine 222, image processing engine 224, and rendering engine 226 may be part of the same computing device. For example, in some cases, one or more image sensors 202, accelerometers 204, gyroscopes 206, storage 207, computing component 210, XR engine 220, interface layout and input management engine 222, image processing engine 224, and rendering engine 226 may be integrated into a head mounted display (HMD), extended reality glasses, a smart phone, a laptop computer, a tablet computer, a gaming system, and / or any other computing device. However, in some implementations, one or more of image sensor 202, accelerometer 204, gyroscope 206, storage 207, computing component 210, XR engine 220, interface layout and input management engine 222, image processing engine 224, and rendering engine 226 may be part of two or more separate computing devices. For example, in some cases, some of components 202-226 may be part of or implemented by one computing device, while the remaining components may be part of or implemented by one or more other computing devices.

[0090] Storage 207 may be any storage device for storing data. In addition, storage 207 may store data from any component of XR system 200. For example, storage 207 may store data from image sensor 202 (e.g., image or video data), data from accelerometer 204 (e.g., measurements), data from gyroscope 206 (e.g., measurements), data from computing component 210 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, facial recognition data, occlusion data, etc.), data from XR engine 220, data from interface layout and input management engine 222, data from image processing engine 224, and / or data from rendering engine 226 (e.g., output frames). In some examples, storage 207 may include a buffer for storing frames for processing by computing component 210.

[0091] One or more computing components 210 may include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, an image signal processor (ISP) 218, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). The computing component 210 may perform various operations, such as image enhancement, computer vision, graphics rendering, extended reality operations (e.g., tracking, positioning, pose estimation, mapping, content anchoring, content rendering, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, facial recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, the computing component 210 may implement (e.g., control, operate, etc.) an XR engine 220, an interface layout and input management engine 222, an image processing engine 224, and a rendering engine 226. In other examples, the computing component 210 may also implement one or more other processing engines.

[0092] Image sensor 202 may include any image and / or video sensor or capture device. In some examples, image sensor 202 may be part of a multi-camera assembly (such as a dual-camera assembly). Image sensor 202 may capture image and / or video content (e.g., raw image and / or video data), which may then be processed by compute component 210, XR engine 220, interface layout and input management engine 222, image processing engine 224, and / or rendering engine 226, as described herein. In some examples, image sensor 202 may include image capture and processing system 100, image capture device 105A, image processing device 105B, or a combination thereof.

[0093] In some examples, image sensor 202 may capture image data and may generate an image (also referred to as a frame) based on the image data and / or may provide the image data or frame to XR engine 220, interface layout and input management engine 222, image processing engine 224, and / or rendering engine 226 for processing. An image or frame may include a video frame of a video sequence or a still image. An image or frame may include an array of pixels representing a scene. For example, an image may be: a red-green-blue (RGB) image having red, green, and blue components per pixel; a luminance, chrominance red, chrominance blue (YCbCr) image having a luminance component and two chrominance (color) components (chrominance red and chrominance blue) per pixel; or any other suitable type of color or monochrome image.

[0094] In some cases, the image sensor 202 (and / or other cameras of the XR system 200) may also be configured to capture depth information. For example, in some implementations, the image sensor 202 (and / or other cameras) may include an RGB depth (RGB-D) camera. In some cases, the XR system 200 may include one or more depth sensors (not shown) that are separate from the image sensor 202 (and / or other cameras) and may capture depth information. For example, such a depth sensor may obtain depth information independently of the image sensor 202. In some examples, the depth sensor may be physically mounted in the same general location as the image sensor 202, but may operate at a different frequency or frame rate than the image sensor 202. In some examples, the depth sensor may take the form of a light source that may project a structured or textured light pattern onto one or more objects in the scene, the structured or textured light pattern may include one or more narrow light bands. Depth information may then be obtained by exploiting the geometric distortion of the projected pattern caused by the surface shape of the object. In one example, depth information may be obtained from a stereo sensor (e.g., a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera)).

[0095] The XR system 200 may also include other sensors in one or more of its sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. One or more sensors may provide velocity, orientation, and / or other position-related information to the computing component 210. For example, the accelerometer 204 may detect the acceleration of the XR system 200 and may generate acceleration measurements based on the detected acceleration. In some cases, the accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, front / back) that may be used to determine the position or attitude of the XR system 200. The gyroscope 206 may detect and measure the orientation and angular velocity of the XR system 200. For example, the gyroscope 206 may be used to measure the pitch, roll, and yaw of the XR system 200. In some cases, the gyroscope 206 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, image sensor 202 and / or XR engine 220 may use measurements obtained by accelerometer 204 (e.g., one or more translation vectors) and / or measurements obtained by gyroscope 206 (e.g., one or more rotation vectors) to calculate the pose of XR system 200. As previously described, in other examples, XR system 200 may also include other sensors, such as an inertial measurement unit (IMU), a magnetometer, gaze and / or eye tracking sensors, machine vision sensors, smart scene sensors, voice recognition sensors, collision sensors, impact sensors, position sensors, tilt sensors, and the like.

[0096] As described above, in some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure specific forces, angular velocities, and / or directions of XR system 200. In some examples, the one or more sensors may output measured information associated with the capture of images captured by image sensor 202 (and / or other cameras of XR system 200) and / or depth information obtained using one or more depth sensors of XR system 200.

[0097] XR engine 220 may use outputs of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) to determine the pose (also referred to as head pose) of XR system 200 and / or the pose of image sensor 202 (or other camera of XR system 200). In some cases, the pose of XR system 200 and the pose of image sensor 202 (or other camera) may be the same. The pose of image sensor 202 refers to the position and orientation of image sensor 202 relative to a reference frame (e.g., about an object detected by image sensor 202). In some implementations, the camera pose may be determined for 6 degrees of freedom (6DoF), which refers to three translational components (e.g., which may be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame (e.g., an image plane)) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some implementations, the camera pose can be determined for 3 degrees of freedom (3DoF), which refers to three angular components (eg, roll, pitch, and yaw).

[0098] In some cases, a device tracker (not shown) may track the pose (e.g., 6DoF pose) of the XR system 200 using measurements from one or more sensors and image data from the image sensor 202. For example, the device tracker may fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and motion of the XR system 200 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the XR system 200, the device tracker may generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates to the 3D map of the scene. The 3D map updates may include, for example, but not limited to, new or updated features and / or features or landmark points associated with the scene and / or the 3D map of the scene, positioning updates for identifying or updating the position of the XR system 200 within the scene and the 3D map of the scene, and the like. The 3D map may provide a digital representation of the scene in the real / physical world. In some examples, the 3D map may anchor location-based objects and / or content to coordinates and / or objects in the real world. The XR system 200 may use a mapped scene (e.g., a scene in the physical world represented by a 3D map and / or associated with a 3D map) to merge the physical and virtual worlds and / or merge virtual content or objects with the physical environment.

[0099] In some aspects, computing component 210 can use a visual tracking solution to determine and / or track the pose of image sensor 202 and / or the XR system 200 as a whole based on images captured by image sensor 202 (and / or other cameras of XR system 200). For example, in some examples, computing component 210 can perform tracking using computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques. For example, computing component 210 can perform SLAM or can communicate (wired or wirelessly) with a SLAM system (not shown). SLAM refers to a class of technologies that create a map of an environment (e.g., a map of an environment modeled by XR system 200) while simultaneously tracking the pose of a camera (e.g., image sensor 202) and / or XR system 200 relative to the map. The map can be referred to as a SLAM map and can be three-dimensional (3D). SLAM techniques may be performed using color or grayscale image data captured by image sensor 202 (and / or other cameras of XR system 200) and may be used to generate estimates of 6DoF pose measurements for image sensor 202 and / or XR system 200. Such SLAM techniques configured to perform 6DoF tracking may be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) may be used to estimate, correct, and / or otherwise adjust the estimated pose.

[0100] In some cases, 6DoF SLAM (e.g., 6DoF tracking) can associate features observed from certain input images from the image sensor 202 (and / or other cameras) to a SLAM map. For example, 6DoF SLAM can use feature point associations from the input image to determine the pose (position and orientation) of the image sensor 202 and / or the XR system 200 for the input image. 6DoF map building can also be performed to update the SLAM map. In some cases, a SLAM map maintained using 6DoF SLAM can contain 3D feature points triangulated from two or more images. For example, a keyframe can be selected from an input image or video stream to represent the observed scene. For each keyframe, a corresponding 6DoF camera pose associated with the keyframe can be determined. The pose of the image sensor 202 and / or the XR system 200 can be determined by projecting features from a 3D SLAM map into an image or video frame, and updating the camera pose based on a verified 2D-3D correspondence.

[0101] In one illustrative example, the computing component 210 can extract feature points from certain input images (e.g., each input image, a subset of the input images, etc.) or from each key frame. Feature points (also referred to as registration points) as used herein are unique or identifiable parts of an image, such as a portion of a hand, an edge of a table, and other examples. Features extracted from captured images can represent different feature points along three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and each feature point can have an associated feature location. Feature points in a key frame match (are identical to or correspond to) or fail to match feature points of a previously captured input image or key frame. Feature detection can be used to detect feature points. Feature detection can include image processing operations for examining one or more pixels of an image to determine whether a feature exists at a particular pixel. Feature detection can be used to process the entire captured image or certain portions of an image. For each image or key frame, once a feature has been detected, a local image block around the feature can be extracted. Features may be extracted using any suitable technique, such as Scale-Invariant Feature Transform (SIFT) (which locates features and generates their descriptions), Learned Invariant Feature Transform (LIFT), Speeded Up Robust Features (SURF), Gradient Location Orientation Histogram (GLOH), Oriented Fast and Rotated Blocks (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), KAZE, Accelerated KAZE (AKAZE), Normalized Cross Correlation (NCC), descriptor matching, another suitable technique, or a combination thereof.

[0102] In some cases, the XR system 200 may also track the user's hands and / or fingers to allow the user to interact with and / or control virtual content in the virtual environment. For example, the XR system 200 may track the gestures and / or movements of the user's hands and / or fingertips to recognize or interpret the user's interactions with the virtual environment. User interactions may include, for example, but are not limited to, moving virtual content items, resizing virtual content items, selecting input interface elements in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), providing input through a virtual user interface, and the like.

[0103] Figure 3 An example block diagram of image processing system 300 is shown. In some cases, image processing system 300 may include or may be included in image capture and processing system 100, image capture device 105A, image processing device 105B, XR system 200, portions thereof, or any combination thereof. Figure 3In the illustrative example of , image processing system 300 includes AON camera processing subsystem 302 , main camera processing subsystem 304 , graphics processing subsystem 306 , video processing subsystem 308 , central processing unit (CPU) 310 , DRAM subsystem 312 , and SRAM 320 .

[0104] In some implementations, the AON camera processing subsystem 302 may receive input from the AON camera sensor 316, and the main camera processing subsystem 304 may receive input from the main camera sensor 318. The AON camera sensor 316 and the main camera sensor 318 may include any image and / or video sensor or capture device. In some cases, the AON camera sensor 316 and the main camera sensor 318 may be part of a multi-camera assembly (such as a dual camera assembly). In some examples, the AON camera sensor 316 and the main camera sensor 318 may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof. In some implementations, the AON camera processing subsystem 302 of the image processing system 300 may communicate with the AON camera sensor 316 to send and / or receive operating parameters to / from the AON camera sensor 316. Similarly, in some implementations, the main camera processing subsystem 304 of the image processing system 300 can communicate with the main camera sensor 318 to send operating parameters to the main camera sensor 318 and / or receive operating parameters from the main camera sensor 318. The DRAM subsystem 312 of the image processing system 300 can communicate with the DRAM 314 via the data bus 315. For example, the DRAM subsystem 312 can send video frames to the DRAM 314 and / or retrieve video frames from the DRAM 314. In some implementations, the image processing system 300 can include a local SRAM 320.

[0105] In some cases, the AON camera sensor 316 may include optimizations for reducing power consumption. In some cases, the AON camera processing subsystem 302 may be configured to store data (e.g., video frame data) in an SRAM 320 located within the image processing system 300. In some cases, storing data in the SRAM 320 can save power by reducing the power required to drive data and address lines compared to driving signals through the data bus 315 to communicate with the DRAM 314. In some implementations, the island voltage rail can be used to power the AON camera sensor 316 and the AON camera processing subsystem 302. In some cases, using the island rail can save power by preventing inactive components of the image processing system 300 from drawing power. In some examples, the AON camera sensor 316 can be clocked with a low-power clock source (such as one or more ring oscillators). In some implementations, the images and / or videos captured by the AON camera sensor 316 can be associated with different domains from the images and / or videos captured by the main camera sensor 318. As described above, the domains may include, but are not limited to, characteristics or parameters of frames captured by the camera sensor, such as resolution, color depth, and / or frame rate. In one illustrative example, the AON camera sensor 316 may capture images at a lower resolution than the main camera sensor 318. In some cases, capturing lower resolution frames with the AON camera sensor 316 may save power by reducing the amount of data (e.g., pixel data) that needs to be read out from the AON camera sensor 316. In some implementations, the AON camera processing subsystem 302 may perform similar processing steps as the main camera processing subsystem 304 on fewer pixels, thereby generating fewer calculations and thereby reducing power consumption.

[0106] In some cases, a domain may include a collection of image processing steps (e.g., noise reduction, edge enhancement, image stabilization, color correction) performed on an image or video frame captured by a camera sensor. In some implementations, different image processing steps may be utilized to process video frames captured by the AON camera sensor 316 and the main camera sensor 318. For example, the AON camera processing subsystem 302 may perform fewer and / or different processing steps than the main camera processing subsystem 304. In some cases, utilizing the AON camera processing subsystem 302 to perform fewer and / or different processing operations may save power during AON operation.

[0107] In another example, the AON camera sensor 316 may capture monochrome video frames, while the main camera sensor 318 may capture red, green, blue (RGB) color video frames. In some cases, reading out and processing monochrome video frames may consume less power than reading out and processing RGB color video frames. In some cases, the video frames captured by the AON camera sensor 316 and the main camera sensor 318 may be based on data captured from different parts of the light spectrum, such as visible, ultraviolet (UV), near infrared (NIR), short wave infrared (SWIR), other parts of the light spectrum, or any combination thereof.

[0108] exist Figure 3 In the illustration of FIG. 3 , the AON camera sensor 316 and the main camera sensor 318 can each provide frames to different camera processing subsystems. However, in some cases, a single camera processing subsystem can process frames from the AON camera sensor 316 and the main camera sensor 318 without departing from the scope of the present disclosure. In addition, although the AON camera sensor 316 and the main camera sensor 318 are shown as two different sensors, one or more camera sensors that can operate in two or more different modes (e.g., AON mode, medium power mode, high power mode, etc.) can also be used without departing from the scope of the present disclosure.

[0109] Figure 4A and Figure 4B An example AON camera system implementation according to some examples of the present disclosure is shown. Figure 4A In one example shown in FIG. , an AON camera sensor (e.g., Figure 3 The AON camera sensor 316 shown in FIG. 4 may be used to determine whether an authorized user is interacting with a device (e.g., a mobile phone) in the AON face unlock implementation 402. Figure 4A In another example shown in , an AON camera sensor can be used to provide a vision-based context sensing implementation 404. Figure 4AIn the illustrated vision-based context sensing implementation 404, audio data, AON camera data, and / or data from one or more additional sensors may be used to determine whether a user is participating in a meeting and to place the device into silent mode. In another illustrative example, a vision-based context sensing implementation may include using a combination of inertial sensor data, audio sensor data, AON camera data, and / or data from one or more additional sensors to determine whether a user is speaking to the device (e.g., giving a voice command) or speaking to another person or device in the room. Another illustrative example of a vision-based context sensing implementation may be a combination of inertial sensor data, audio sensor data, AON camera data, and / or data from one or more additional sensor cameras to prevent device use by a driver of a car. For example, a driver may be prevented from using a mobile phone while a passenger may be allowed to use a mobile phone. In Figure 4A In another example shown in , an AON camera sensor may be used as part of an AON gesture detection implementation 406. For example, an AON camera sensor may be included in an XR system (e.g., an HMD), a mobile device, or any other device in the AON gesture detection implementation 406. In such an implementation, the AON camera may capture images and perform gesture detection to allow a user to interact with the device without making physical contact with the device.

[0110] Figure 4B An example film shutter lag video capture sequence 408 is shown. Figure 4B In the illustrative example of FIG. 4 , a user wearing an HMD or AR glasses is depicted. In the example shown, the user can observe 410 a scene. In some cases, while the user is observing 410 the scene, an event 412 can occur. In some cases, event 412 can occur during an AON operation included in the HMD (e.g., the above Figure 3 In the example shown, the user can provide a capture input 414 to the HMD to initiate video capture. In some cases, after the user initiates video capture, the HMD can begin capturing video frames in standard operation (e.g., using the above Figure 3). As will be described in more detail with respect to the following figures, frames captured during AON operation can be combined with frames captured during standard operation into a combined video that begins before the input is captured. In some cases, frames captured during AON operation can be transformed to provide a consistent appearance of the combined video frames before and after the input is captured. As described above, references to AON operation herein can be understood to include capturing images and / or frames using one or more AON cameras and / or operating one or more cameras in AON mode. Similarly, standard operation can be understood to include capturing needles using one or more non-AON cameras and / or one or more cameras operating in non-AON mode.

[0111] Figure 5A An example block diagram of a film shutter hysteresis system 500 according to some examples of the present disclosure is shown. As shown, components of the film shutter hysteresis system 500 may include one or more camera sensors 502, a camera pipeline 504, a domain switch 506, a frame buffer 512, a domain transformation model 514, and a combiner 518. In some cases, the one or more camera sensors 502 may include an AON camera sensor (e.g., Figure 3 AON camera sensor 316 shown) and the main camera sensor (e.g., Figure 3 10 ). In some examples, the AON camera sensor may capture video frames 508 associated with the first domain, and the main camera sensor may capture video frames 510 associated with the second domain. In some cases, the one or more camera sensors 502 may include one or more camera sensors that can operate in two or more modes (e.g., an AON mode, a medium power mode, and a high power mode). In some implementations, the one or more camera sensors 502 may capture video frames 508 associated with the first domain during AON operation, and capture captured video frames 510 associated with the second domain during standard operation. In the example shown, the time at which the film shutter hysteresis system 500 receives the capture input 505 may be marked as time t=0. Video frames captured before the capture input 505 occur during a range of time t<0, and video frames captured after the capture input 505 occur during a range of time t>0.

[0112] In some cases, frames captured by one or more camera sensors 502 may be processed by camera pipeline 504. For example, camera pipeline 504 may perform noise reduction, edge enhancement, image stabilization, color correction, and / or other image processing operations on raw video frames provided by one or more camera sensors 502.

[0113] In some examples, the domain switch 506 can be communicatively coupled to the one or more camera sensors 502 and can control which camera sensor and / or camera sensor mode captures the video frame. In some cases, the domain switch 506 can receive the video frame associated with the first domain before receiving the capture input 505. For example, the video frame associated with the first domain can be received during AON operation. In some cases, the domain switch 506 can route the video frame 508 associated with the first domain to the frame buffer 512. In some implementations, the domain switch 506 can enable the addition of additional frames to the frame buffer 512 during AON operation and disable the addition of additional frames to the frame buffer 512 during standard operation. In some examples, the frame buffer 512 can receive the video frame 508 associated with the first domain directly from the one or more camera sensors 502 and / or the camera pipeline 504.

[0114] In some cases, the film shutter hysteresis system 500 may receive a capture input 505 (e.g., from a user pressing a button, performing a gesture, etc.). Based on the capture input 505, the film shutter hysteresis system 500 may begin capturing video frames associated with the second domain. In some cases, the first video frame (or subsequent video frames) associated with the second domain captured after receiving the capture input 505 may be stored or marked as a key frame 511. In some cases, the key frame 511 may be captured at approximately the same time as the capture input 505 is received. After receiving the capture input 505, the one or more camera sensors 502 and the camera pipeline 504 may output video frames associated with the second domain until an end capture input is received. In some cases, upon receiving the capture input 505, the domain switch 506, the one or more camera sensors 502, and / or the camera pipeline 504 may stop providing new frames to the frame buffer 512. In this case, the video frames captured in the frame buffer 512 may span a period of time B based on the buffer length of the frame buffer 512. In some cases, the video frames stored in frame buffer 512 may span a period of time between t=−B and the time t≈0 at which the last video frame associated with the first domain was captured prior to capturing input 505 .

[0115] In some cases, the domain transformation model 514 can be implemented as a deep learning neural network. In some implementations, the domain transformation model 514 can be trained to transform video frames stored in the frame buffer 512 associated with the first domain into transformed video frames associated with the second domain. In some implementations, the domain transformation model 514 can be trained to use the key frames 511 associated with the second domain as a guide for transforming video frames associated with the first domain stored in the frame buffer 512.

[0116] In an illustrative example, the training data set may include an original video and a training video. In some cases, all frames of the original video may be associated with a second domain. In some cases, a portion of each training video in the training video (e.g., the first 30 frames, the first 100 frames, the first 300 frames, the first 900 frames, or any other number of frames) may be associated with a first domain to simulate data stored in a frame buffer 512 during AON operation. In some cases, a portion of a video frame associated with a first domain may be generated from the original video by transforming a portion of the original video frame from the second domain to the first domain (e.g., by reducing the resolution, converting from color to monochrome, reducing the frame rate, simulating different steps in a camera pipeline, etc.). The original video and the training video may include a key frame associated with the second domain (e.g., key frame 511), which may be used by a domain transformation model 514 to transform a portion of the training video associated with the first domain to the second domain.

[0117] In some cases, the resulting transformed video frame generated by the domain transform model 514 can be directly compared to the original video frame, and a loss function can be used to determine the amount of error between the transformed video frame and the original video frame. The parameters (e.g., weights, biases, etc.) of the deep learning network can be adjusted (or tuned) based on the error. Such a training process can be referred to as supervised learning using back-propagation, which can be performed until the tuned parameters provide the desired results.

[0118] In some examples, the domain transformation model 514 may be trained using a deep generative neural network model (e.g., a generative adversarial network (GAN)). A GAN is a generative neural network that can learn patterns in input data so that the neural network model can generate new synthetic outputs that are reasonably likely to come from the original data set. A GAN may include two neural networks operating together. One of the neural networks (referred to as a generative neural network or generator, denoted as G(z)) generates a synthetic output, while the other neural network (referred to as a discriminative neural network or discriminator, denoted as D(X)) evaluates the synthetic output for authenticity (whether the synthetic output comes from the original data set, such as a training data set, or is generated by the generator). The generator G(z) may correspond to the domain transformation model 514. The generator is trained to try and deceive the discriminator to determine that the synthetic video frame (or group of video frames) generated by the generator is a real video frame (or group of video frames) from a training data set (e.g., a first set of training video data). The training process continues, and the generator becomes better at generating synthetic video frames that look like real video frames. The discriminator continues to find defects in the synthetic video frames, and the generator figures out what the discriminator is looking at to determine the defects in the image. Once the network is trained, the generator is able to produce realistic video frames that the discriminator cannot distinguish from real video frames.

[0119] One example of a neural network that can be used for training of the domain transformation model 514 is a conditional GAN. In a conditional GAN, the generator is learning a conditional distribution. A conditional GAN ​​can condition the generator and discriminator neural networks with some vector y, in this case, the vector y is input to both the generator network and the discriminator network. Based on the vector y, the generator and the discriminator become G(z,y) and D(X,y), respectively. The generator G(z,y) models the distribution of data given z and y, in this case, the data X is generated as X~G(X|z,y). The discriminator D(X,y) tries to find a distribution for XandX G The discriminator D(X,y) and the generator G(z,y) are thus jointly conditioned into two variables: z or X and y.

[0120] In one illustrative example, a conditional GAN ​​can generate a video frame (or multiple video frames) conditioned on a label (e.g., represented as a vector y), wherein the label indicates the category of an object in the video frame (or multiple video frames), and the goal of the GAN is to transform a video frame (or multiple video frames) associated with a first domain into a video frame (or multiple video frames) associated with a second domain including the class of the object. There is a competitive aspect between the generator G and the discriminator D with respect to maximizing the parameters of the discriminator D. The discriminator D will try to distinguish between real video frames (from a first set of training video data) and pseudo video frames (generated by the generator G based on the second set of training video data) and possible, and the generator G should minimize the ability of the discriminator D to identify pseudo images. The parameters of the generator G (e.g., the weights of the nodes of the neural network and in some cases other parameters, such as bias) can be adjusted during the training process so that the generator G will output a video frame that is indistinguishable from a real video frame associated with the second domain. A loss function can be used to analyze the errors in the generator G and the discriminator D. In one illustrative example, a binary cross entropy loss function can be used. In some cases, other loss functions can be used.

[0121] During inference (after the domain transformation model 514 has been trained), the parameters of the domain transformation model 514 (e.g., the generator G trained by the GAN) can be fixed, and the domain transformation model 514 can use the keyframes 511 as a guide to transform the video frames associated with the first domain (e.g., by upscaling the resolution, colorizing, and / or performing any other transformation) to generate transformed video frames 516 associated with the second domain.

[0122] In some cases, the combiner 518 can convert the transformed video frames 516 and the captured video frames 510 associated with the second domain for time t>=0 into a combined video 520 with the video frames associated with the second domain. In one illustrative example, the combiner 518 can perform a concatenation of the transformed video frames 516 and the captured video frames 510 associated with the second domain to generate the combined video 520.

[0123] Figure 5B A block diagram of another example film shutter hysteresis system 550 is shown. Figure 5A , the film shutter hysteresis system 500 shown in FIG. 5 , the film shutter hysteresis system 550 may include one or more camera sensors 502, a frame buffer 512, a domain transform model 514, and a combiner 518. Figure 5B In the example shown, Figure 5A The domain switch 506 shown in has been removed. In addition, instead of including Figure 5A , the film shutter hysteresis system 550 may include a first camera pipeline 522 and a second camera pipeline 524 communicatively coupled to the one or more camera sensors 502 .

[0124] As shown, the first camera pipeline 522 can receive video frames from one or more camera sensors 502. In some cases, the first camera pipeline 522 can receive video frames from one or more camera sensors 502 during AON operation of the one or more camera sensors 502. The first camera pipeline 522 can perform noise reduction, edge enhancement, image stabilization, color correction, and / or other image processing operations on the video frames from the one or more camera sensors 502, and the output of the first camera pipeline 522 can be routed to the frame buffer 512. The video frames output from the first camera pipeline 522 can be associated with the first domain. In some cases, the first domain can include specific processing steps performed by the first camera pipeline 522. The output of the first camera pipeline 522 can include the video frames 508 associated with the first domain.

[0125] The second camera pipeline 524 may also receive video frames from the one or more camera sensors 502. In some cases, the second camera pipeline 524 may receive video frames from the one or more camera sensors 502 during standard operation of the one or more camera sensors 502. The second camera pipeline 524 may perform one or more image processing operations on the received video frames. In some cases, the image processing operations performed by the second camera pipeline 524 may perform different processing steps and / or a different number of processing steps compared to the first camera pipeline 522. In some cases, the second domain may include specific processing steps performed by the second camera pipeline 524. The output of the second camera pipeline 524 may include a key frame 511 associated with the second domain and a captured video frame 510 associated with the second domain. In some cases, when the capture input 505 is received (e.g., at time t=0), the one or more camera sensors 502 may pause outputting video frames to the first camera pipeline 522 and begin outputting video frames to the second camera pipeline 524.

[0126] As mentioned above about Figure 5A As described, the domain transformation model 514 can be a deep learning neural network trained to transform video frames from a first domain to a second domain. In the case of the film shutter lag system 550, generating a portion of each of the training videos associated with the first domain can include simulating the differences in processing steps between the first camera pipeline 522 and the second camera pipeline 524. In some cases, the differences in processing steps between the first camera pipeline 522 and the second camera pipeline 524 can be simulated in addition to replicating any other differences between the first domain and the second domain (e.g., resolution, color depth, frame rate, etc.). Using the above training set, the domain transformation model 514 can be trained in a supervised training process and / or trained using a GAN, as described in the above examples, as well as any other suitable training techniques.

[0127] During inference (after the domain transformation model 514 has been trained), the parameters of the domain transformation model 514 (e.g., the generator G trained by the GAN) can be fixed, and the domain transformation model 514 can use the keyframes 511 as a guide to transform video frames associated with the first domain (e.g., by upscaling the resolution, colorizing, emulating the image processing steps of a camera pipeline, and / or performing any other transformations) to generate transformed video frames 516 associated with the second domain.

[0128] Figure 6 6 is a flow chart illustrating an example of a process 600 for processing one or more frames. At block 602, process 600 includes obtaining a frame from an image capture system (e.g., Figure 1 The image capture and processing system 100 shown in Figure 2 In some cases, the first setting domain includes a first resolution. In some cases, the first setting domain includes a first frame rate.

[0129] At block 604, process 600 includes obtaining at least one reference frame (eg, a key frame) associated with a second settings domain from an image capture system. The at least one reference frame is captured proximate to when the capture input was obtained.

[0130] At block 606, process 600 includes obtaining, from the image capture system, a second plurality of frames associated with the second settings domain. The second plurality of frames are captured after at least one reference frame. In some cases, the second settings domain includes a second resolution. In some cases, the second settings domain includes a second frame rate.

[0131] At block 608, process 600 includes generating a reference frame based on the at least one reference frame (e.g., using Figure 5A and Figure 5B 514) to transform at least a portion of the first plurality of frames to generate the transformed plurality of frames associated with the second set domain. In some cases, transforming at least a portion of the first plurality of frames includes upscaling at least a portion of the first plurality of frames from the first resolution (e.g., using Fig. 9 In some cases, transforming at least a portion of the first plurality of frames includes converting at least a portion of the first plurality of frames from a first frame rate (e.g., using Fig.11 Although the following examples include discussion of a plurality of frames, the examples may be associated with a portion of a plurality of frames (eg, at least a portion of a first plurality of frames).

[0132] In some cases, process 600 includes combining the transformed plurality of frames and the second plurality of frames to generate (eg, using combiner 1118 ) a video associated with the second settings domain.

[0133] In some cases, process 600 includes (e.g., Fig.13 1326 and / or the optical motion estimator 1324 shown in ), wherein generating the transformed plurality of frames associated with the second setting domain is based on the first plurality of frames, the at least one reference frame, and the motion information. In some cases, process 600 includes determining a translation direction based on the motion information associated with the first plurality of frames, and applying the translation direction to the transformed plurality of frames.

[0134] In some cases, process 600 includes capturing a first subset of a first plurality of frames at a first frame rate, and capturing a second subset of the first plurality of frames at a second frame rate different from the first frame rate. In some implementations, the change between the first frame rate and the second frame rate is based at least in part on motion information associated with the first subset of the first plurality of frames, the second subset of the first plurality of frames, or both.

[0135] In some cases, process 600 includes obtaining an additional reference frame associated with the second set domain from the image capture system. In some implementations, the additional reference frame is captured prior to obtaining the capture input. In some cases, generating the transformed plurality of frames associated with the second domain is based on the first plurality of frames, the at least one reference frame, and the additional reference frame. In some examples, the at least one reference frame provides a reference for transforming at least a first portion of the first plurality of frames, and the additional reference frame provides a reference for magnifying at least a second portion of the first plurality of frames.

[0136] While the examples of this disclosure describe various techniques for film shutter hysteresis associated with video frames, the techniques of this disclosure can also be used to provide similar film shutter hysteresis for still images (also referred to herein as frames). For example, a user may have missed an interesting event, and after the event occurs, the user may initiate a capture input for still image capture (e.g., Figure 5A and Figure 5B ). In such a case, such as Figure 5A The film shutter hysteresis system 500 shown in Figure 5B The film shutter hysteresis system 550 shown in FIG. 5 may receive a selection of one or more selected frames stored in the frame buffer 512 to transform from the first domain to the second domain. In some cases, the film shutter hysteresis (or, for example, a system including a film shutter hysteresis system such as the XR system 200) may present an interface for selecting the one or more selected frames. For example, a user may be provided with an interface for stepping through and viewing the video frames stored in the frame buffer 512 frame by frame to select one or more selected frames for transformation. Other illustrative examples of interfaces for selecting one or more selected frames may include a slider for advancing through the frames stored in the frame buffer 512 and a thumbnail gallery of frames stored in the frame buffer 512.

[0137] In some cases, the film shutter hysteresis system may suggest one or more frame transformations. In one illustrative example, the film shutter hysteresis system may determine whether there is a significant change or motion in a particular frame compared to other recent frames that exceeds a threshold amount of change or motion, and suggest a frame (or frames) that exceeds the threshold. For example, the film shutter hysteresis system may use a change detection algorithm to determine whether there is a significant change or motion to analyze the change or motion of multiple detected and / or tracked features. In another illustrative example, a deep learning neural network may be trained to determine which frames contain relevant and / or interesting content based on training with a dataset of labeled image data. In some cases, the deep learning neural network may be trained on a personalized dataset of images that have been taken by a particular user. In another illustrative example, the film shutter hysteresis system may determine a change in the number of human faces as a basis for suggesting one or more frame transformations. For example, a face detection neural network may be trained on a dataset of facial and non-facial images to determine the number of human faces in the frames stored in the frame buffer 512. Once the film shutter hysteresis system receives a selection of one or more selected frames, the domain transformation model 514 may transform the one or more selected frames to generate one or more transformed frames associated with the second domain.

[0138] Figure 7 700 is a flow chart illustrating an example of a process 700 for processing one or more frames. At block 702, process 700 includes obtaining a frame from an image capture system (e.g., Figure 1 The image capture and processing system 100 shown in Figure 2 The image sensor 202 shown in FIG. 1 , and / or one or more camera sensors 502 ) obtains a first plurality of frames associated with a first setting domain, wherein the first plurality of frames are captured before obtaining a capture input.

[0139] At block 704 , process 700 includes obtaining a reference frame (eg, a key frame) associated with a second setting domain from an image capture system, wherein the reference frame is captured proximate to when the capture input was obtained.

[0140] At block 706, process 700 includes obtaining a selection of one or more selected frames associated with the first plurality of frames. For example, a user may be provided with a menu for stepping through and viewing frames stored in a frame buffer (e.g., Figure 5A and Figure 5B An interface for selecting one or more selected frames for transformation.

[0141] At block 708, process 700 includes transforming (eg, using Figure 5A and Figure 5BThe one or more selected frames are used to generate one or more transformed frames associated with the second set domain.

[0142] Fig. 8A 8 is a flow chart illustrating an example process 800 for performing film shutter hysteresis capture. At block 802, process 800 may disable a camera (e.g., Figure 3 Capture of video frames by the AON camera sensor 316 and / or the main camera sensor 318) shown in FIG.

[0143] At block 804, process 800 may determine whether AON logging is enabled. If AON logging is disabled, process 800 may return to block 802. If AON logging is enabled, process 800 may proceed to block 806.

[0144] At block 806, process 800 may capture frames associated with the first domain during AON operation. In one illustrative example, the first domain may include capturing monochrome frames from a NIR sensor. In some cases, the frames associated with the first domain may be stored in a video buffer (e.g., Figure 3 The SRAM 320, DRAM 314 and / or Fig.19 1925). For example, the video buffer may include a circular buffer that stores video frames for a particular buffer length (e.g., 1 second, 5 seconds, 30 seconds, 1 minute, 5 minutes, or any other amount of time). In some cases, the video buffer may accumulate video frames until the amount of video stored in the video buffer spans the buffer length. Once the video buffer is filled, each new frame captured at block 806 may replace the oldest frame held in the video buffer. As a result, the video buffer may include the most recent video frame captured at block 806 for a time span based on the buffer length.

[0145] At block 808, process 800 may determine whether a capture input has been received. If process 800 determines that a capture input has not been received, process 800 may return to block 806 and continue AON operation. If process 800 determines that a capture input has been received, process 800 may continue to block 810.

[0146] At block 810, process 800 may capture video frames associated with the second domain during standard operation. In one illustrative example, the second domain may include capturing RGB frames from a visible light sensor. In some cases, at block 810, the contents of the video buffer may remain fixed during standard operation. In some cases, the video frames associated with the second domain captured at block 810 may be stored in a separate memory or portion of memory from the video buffer.

[0147] At block 812, process 800 may determine whether an end recording input has been received. If an end recording input has been received, process 800 may return to block 804. If an end recording input has not been received, process 800 may return to block 810 and continue capturing frames during standard operation.

[0148] Figure 8B Shows Fig. 8A 800 during different stages of the process 800. The curve 850 is not shown to scale and is provided for illustrative purposes. In the example shown, the height of the bar 822 indicates the image processing system (e.g., above Figure 3 806 ). In the example shown, bar 822 may include the power consumed by storing video frames captured during AON operation in a video buffer (e.g., a circular buffer). Similarly, in the example shown, the height of bar 824 may include the power consumed by one or more camera sensors (e.g., the above) during AON operation. Figure 3 In the example shown, the video buffer (e.g., Figure 3 The buffer length 826 of the DRAM 314 and / or SRAM 320 shown in FIG. 8 is shown by an arrow that ends at the time the input 827 is captured and extends back in time by an amount based on the buffer length B of the video buffer.

[0149] After receiving the capture input 827, the process 800 can capture a video frame associated with the second domain during standard operation (e.g., at block 810). The height of bar 828 can represent the relative power consumed by the image processing system during standard operation, and shows an increase in power consumption relative to the power consumed during buffering of video frames associated with the first domain. The increased power consumption of the image processing system shown by bar 828 can be the result of, for example, processing video frames with a greater number of pixels, more color information, a higher frame rate, more image processing steps, a power difference associated with any other difference between the first domain and the second domain, or any combination thereof. Similarly, bar 830 can represent the power of one or more camera sensors during standard operation. The increased power consumption of one or more camera sensors shown by bar 830 can result from capturing and transmitting data for a greater number of pixels, with more color information, at a higher frame rate, a power difference associated with any other difference between the first domain and the second domain, or any combination thereof.

[0150] In some cases, AON operation may continue for minutes, hours, or days without receiving a capture input by process 800. In such cases, reducing the power associated with capturing frames during AON operation compared to standard operation may significantly increase the available battery life of the film shutter hysteresis system. In some cases, an AON camera system that is always capturing and processing high-power video frames during both standard and AON operation may consume available power (e.g., from a battery) more quickly by comparison.

[0151] Video frame capture during standard operation can continue until process 800 receives (e.g., at block 812) an end recording input. Figure 5A The film shutter hysteresis system 500 shown in Figure 5B As described in the film shutter hysteresis system 550 shown in FIG. 8 , video frames associated with a first domain captured prior to capturing input 827 and stored in a video buffer may be transformed based on the key frames to generate transformed video frames associated with a second domain. In some cases, the transformed video frames associated with the second domain may be combined (e.g., at combiner 518) with the video frames associated with the second domain to form a combined video. Brackets 832 show the video frames associated with the first domain captured prior to capturing input 827 and stored in a video buffer based on the key frames. Figure 8B Total duration of the example combined video for the example shown.

[0152] Fig. 9 An example film shutter hysteresis system 900 according to an example of the present disclosure is shown. As shown, the film shutter hysteresis system 900 may include one or more camera sensors 902, a camera pipeline 904, a resolution switch 906, a frame buffer 912, an amplification model 914, and a combiner 918. Fig. 9 In the example of , the frame 908 captured before the film shutter hysteresis system 900 receives the capture input 905 can have a first resolution, and the frame 910 captured after receiving the capture input 905 can have a second resolution different from the first resolution. In one illustrative example, the first resolution can be lower than the second resolution. Figure 5A and Figure 5B , low resolution video frame 908 may be an illustrative example of a frame associated with a first domain (e.g., video frame 508), and high resolution video frame 910 may be an illustrative example of a video frame associated with a second domain (e.g., captured video frame 510). Resolution switch 906 may be similar to and perform the same operations as Figure 5A 906. For example, the resolution switch 906 can be communicatively coupled to one or more camera sensors 902, and can control which (which) camera sensors and / or camera sensor modes are used to capture video frames during AON operation and standard operation. In some cases, the resolution switch 906 can receive video frames from one or more camera sensors 902 processed by the camera pipeline 904 at a first resolution before receiving the capture input 905. For example, low-resolution video frames can be received from AON camera sensors and / or from camera sensors operating in AON mode. In some cases, the resolution switch 906 can route low-resolution video frames 908 to the frame buffer 912. In some implementations, the resolution switch 906 can enable or disable the frame buffer 912. In some cases, the frame buffer 912 can receive low-resolution video frames 908 directly from one or more camera sensors 902 and / or the camera pipeline 904.

[0153] In some implementations, the frame buffer 912 can be similar to and perform the same Figure 5A 912 may have a buffer length B (e.g., one second, five seconds, ten seconds, thirty seconds, or any other selected buffer length). In some cases, after capture input 905 is received by film shutter hysteresis system 900, the video frames stored in frame buffer 912 may span a period of time between t=-B and the time when the last low-resolution video frame was captured before capture input 905 (e.g., t≈0).

[0154] In some cases, the upscaling model 914 can be implemented as a deep learning neural network. In some cases, the upscaling model 914 can be trained to upscale the low-resolution video frames 908 stored in the frame buffer 912 to high resolution. In some implementations, the upscaling model 914 can be trained to use the high-resolution keyframes 911 as a guide for upscaling the low-resolution video frames 908 stored in the frame buffer 912.

[0155] In some cases, you can use something similar to Figure 5A The upscaling model 914 may be trained by the process described for training the domain transform model 514 described above. In the case of the film shutter lag system 900, the training data set may include the original video and the training video. In some cases, all frames of the original video may be high resolution frames. In some cases, a portion of each of the original videos (e.g., the first 30 frames, the first 100 frames, the first 300 frames, the first 900 frames, or any other number of frames) may be scaled down to a low resolution to simulate data stored in the frame buffer 512 during AON operation. The original video and the training video may include high resolution key frames (e.g., key frames 511) that may be used by the domain transform model 514 to upscale the low resolution portion of the training video to a high resolution.

[0156] During inference (after the upscaling model 914 has been trained), the parameters of the upscaling model 914 (e.g., the generator G trained using a GAN) can be fixed, and the upscaling model 914 can upscale the low-resolution frame (e.g., the frame stored in the frame buffer 912) using the high-resolution key frame 911 as a guide to generate an upscaling video frame 916.

[0157] In some cases, combiner 918 may combine up-scaled video frame 916 and high-resolution video frame 910 into high-resolution combined video 920. In one illustrative example, combiner 918 may perform concatenation of up-scaled video frame 916 and high-resolution video frame 910 to generate combined video 920.

[0158] Fig.10 1 is a flow chart illustrating an example of a process 1000 for processing one or more video frames. At block 1002, process 1000 includes capturing an image from an image capture system (e.g., Figure 1 The image capture and processing system 100 shown in Figure 2 The image sensor 202 shown in FIG. 1 , and / or one or more camera sensors 502 ) obtains a first plurality of frames having a first resolution, wherein the first plurality of frames are captured before obtaining a capture input.

[0159] At block 1004 , process 1000 includes obtaining a reference frame (eg, a key frame) having a second resolution from an image capture system, wherein the reference frame is captured proximate to where the capture input was obtained.

[0160] At block 1006 , process 1000 includes obtaining, from the image capture system, a second plurality of frames having a second resolution, wherein the second plurality of frames are captured after the reference frame.

[0161] At block 1008, process 1000 includes upscaling the first plurality of frames from a first resolution based on the reference frame (e.g., using Figure 5A and Figure 5B The domain transformation model 514 shown in FIG. 5 and / or Fig. 9 The upscaling model 914 shown in FIG. 9 is further illustrated in FIG. 1 ) to a second resolution to generate a plurality of upscaled frames having the second resolution.

[0162] Fig.11 is a schematic diagram illustrating another example film shutter hysteresis system 1100 according to some examples. As shown, the film shutter hysteresis system 1100 may include one or more camera sensors 1102, a camera pipeline 1104, a domain switch 1106, a frame buffer 1112, a frame rate converter 1122, an amplification model 1114, and a combiner 1118. Fig.11 One or more components of the film shutter hysteresis system 1100 may be similar to Fig. 9 The same numbered components of the film shutter hysteresis system 900 and perform similar operations. For example, the one or more camera sensors 1102 can be similar to and perform similar operations as the one or more camera sensors 902. As another example, the camera pipeline 1104 can be similar to and perform similar operations as the camera pipeline 904. Fig.11 In the example of , frames 1108 captured before receiving capture input 1105 can have a first resolution and a first frame rate. Video frames 1110 and key frames 1111 captured after receiving capture input 1105 can have a second resolution different from the first resolution and a second frame rate different from the second frame rate. In one illustrative example, the second resolution can be higher than the first resolution and the second frame rate can be higher than the first frame rate.

[0163] To generate upscaled and frame rate adjusted video frames 1116 having a second resolution and a second frame rate, frame rate converter 1122 may adjust the frame rate of frame 1108 from the first frame rate to the second frame rate. In one illustrative example, frame rate converter 1122 may interpolate data from adjacent frame pairs stored in frame buffer 1112 to generate additional frames and thereby generate frame rate adjusted frames 1125. In some cases, upscaling model 1114 guided by key frame 1111 may upscale the resolution of frame rate adjusted frame 1125 to the second resolution. Upscaling model 1114 may be similar to Fig. 9 The enlarged model 914 and the execution Fig. 9 The enlarged model 1114 can also be trained in a similar manner to the enlarged model 914, such as Fig. 9In some implementations, the resolution upscaling performed by upscaling model 1114 and the frame rate adjustment performed by frame rate converter 1122 may be performed in reverse order without departing from the scope of the present disclosure. In some cases, upscaling model 1114, frame rate converter 1122, or a combination thereof may be collectively considered a domain transformation model (e.g., corresponding to Figure 5A and Figure 5B An illustrative example of a domain transformation model 514) is shown in FIG.

[0164] In some cases, combiner 1118 may combine upscaled and frame rate adjusted video frames 1116 and video frames 1110 having the second resolution and the second frame rate into a combined video 1120 having the second resolution and the second frame rate. In one illustrative example, combiner 1118 may perform concatenation of upscaled and frame rate adjusted video frames 1116 and video frames 1110 at the second resolution and the second frame rate to generate combined video 1120.

[0165] Fig.12 1 is a flow chart illustrating an example of a process 1200 for processing one or more frames. At block 1202, process 1200 includes obtaining a frame from an image capture system (e.g., Figure 1 The image capture and processing system 100 shown in Figure 2 The image sensor 202 shown in, and / or one or more camera sensors 502) obtains a first plurality of frames having a first resolution and a first frame rate, wherein the first plurality of frames are captured before obtaining a capture input.

[0166] At block 1204 , process 1200 includes obtaining a reference frame (eg, a key frame) having a second resolution from an image capture system, wherein the reference frame is captured proximate to where the capture input was obtained.

[0167] At block 1206 , process 1200 includes obtaining, from the image capture system, a second plurality of frames having a second resolution and a second frame rate, wherein the second plurality of frames are captured after the reference frame.

[0168] At block 1208 , process 1200 includes upscaling the first plurality of frames from the first resolution to the second resolution based on the reference frame and frame rate adjusting the first plurality of frames from the first frame rate to the second frame rate to generate a transformed plurality of frames having the second resolution and the second frame rate.

[0169] Fig.131 is a schematic diagram illustrating another example film shutter hysteresis system 1300 according to an example of the present disclosure. As shown, the film shutter hysteresis system 1300 may include one or more camera sensors 1302, a camera pipeline 1304, a domain switch 1306, a frame buffer 1312, an amplification model 1314, a combiner 1318, and a frame rate converter 1322. Fig.13 One or more components of the film shutter hysteresis system 1300 may be similar to Fig.11 The camera pipeline 1304 may be similar to and perform similar operations as the film shutter hysteresis system 1100 of the embodiment of the present invention. For example, the one or more camera sensors 1302 may be similar to and perform similar operations as the one or more camera sensors 1102. As another example, the camera pipeline 1304 may be similar to and perform similar operations as the camera pipeline 1104.

[0170] As shown, film shutter hysteresis system 1300 may also include frame rate controller 1323, optical motion estimator 1324, and inertial motion estimator 1326. In film shutter hysteresis system 1300, frames 1308 captured prior to receiving capture input 1305 may be captured at a variable frame rate. In some cases, the variable frame rate of frames 1308 may be determined by frame rate controller 1323. Frame rate controller 1323 may receive input from one or more motion estimators. Fig.13 In the illustrative example shown, frame rate controller 1323 may receive input from inertial motion estimator 1326 and optical motion estimator 1324. In some cases, frame rate controller 1323 may receive input from fewer (e.g., one), more (e.g., three or more), and / or different types of motion estimators without departing from the scope of the present disclosure. In some cases, sensor data 1328 may be received as input from one or more sensors (e.g., Figure 2 1305 , the accelerometer 204 and / or the gyroscope 206, an IMU, and / or any other motion sensor as shown) are provided to the inertial motion estimator 1326. In some cases, the inertial motion estimator 1326 can determine an estimated amount of motion of the one or more camera sensors 1302 based on the sensor data 1328. In some cases, the optical motion estimator 1324 can determine an amount of motion of the one or more camera sensors 1302 and / or the scene captured by the one or more camera sensors 1302 based on the frames 1308 captured prior to receiving the capture input 1305. In some cases, the optical motion estimator can estimate the amount of motion of the one or more camera sensors 1302 and / or the scene using one or more optical motion estimation techniques. In one illustrative example, the optical motion estimator 1324 can detect features in two or more of the frames 1308 (e.g., as described with respect to FIG. 1305 ). Figure 2), and determining an amount of motion of one or more features between two or more frames in frame 1308.

[0171] In some cases, based on the motion estimates received from the inertial motion estimator 1326 and / or the optical motion estimator 1324, the frame rate controller 1323 can change the frame rate of the frames 1308 captured by the one or more camera sensors 1302 before receiving the capture input 1305. For example, if the frame rate controller 1323 detects a relatively small amount of motion (or no motion) of the one or more camera sensors 1302 and / or objects in the scene, the frame rate controller 1323 can reduce the frame rate of the frames 1308. In one illustrative example, if the one or more camera sensors 1302 are stationary and capturing a static scene, the frame rate controller 1323 can reduce the frame rate. In some cases, the frame rate controller 1323 can reduce the frame rate of the frames 1308 to one-eighth, one-quarter, one-half, or any other fraction of the second frame rate. In some cases, by reducing the frame rate of the frames 1308 before receiving the capture input 1305 (e.g., frames captured during AON mode), power consumption can be reduced.

[0172] On the other hand, the frame rate controller 1323 may increase the frame rate if it detects a lot of motion from the inertial motion estimator 1326 and / or the optical motion estimator 1324. In one illustrative example, the frame rate controller 1323 may increase the frame rate if one or more camera sensors 1302 are in motion and are viewing a sporting event (e.g., a scene containing a lot of motion).

[0173] In some cases, frame rate controller 1323 may increase the frame rate of frames 1308 captured prior to receiving capture input 1305 to be as high as the frame rate used to capture frames 1310 after receiving capture input 1305. In some cases, frame rate controller 1323 may output a variable frame rate associated with each of frames 1308 captured prior to receiving capture input 1305 in frame buffer 1312.

[0174] In some cases, by increasing the frame rate, power consumption for capturing frames 1308 prior to receiving capture input 1305 may be increased. Although power consumption may be increased by increasing the frame rate of frames 1308 during periods of increased motion, the quality of frames obtained after frame rate conversion during AON operation and the magnification of frames captured during AON operation may be improved.

[0175] In some cases, frame rate converter 1322 can utilize frame rate information stored in frame buffer 1312 to properly perform frame rate conversion. For example, if frame rate controller 1323 reduces the frame rate of frame 1308, frame rate converter 1322 can generate additional frames (e.g., using interpolation techniques) to match the frame rate of frame 1310 captured after receiving capture input 1305.

[0176] In some cases, frame rate converter 1322 may adjust the frame rate of frame 1308 from the variable frame rate to a second frame rate. In one illustrative example, frame rate converter 1322 may interpolate data from adjacent frame pairs stored in frame buffer 1112 to generate additional frames and thereby generate frame rate adjusted frame 1325.

[0177] In an illustrative example, frames 1308 captured prior to receiving capture input 1305 may have a variable frame rate and a first resolution. Additionally, frames 1310 and key frames 1311 captured after receiving capture input 1305 may have a second resolution and a second frame rate. In such an example, upscaling model 1314 guided by key frames 1311 may upscale the resolution of frame rate adjusted frames 1325 to the second resolution to generate frame rate converted and upscaled frames 1316. Upscaling model 1314 may be similar to Fig. 9 The enlarged model 914 and the execution Fig. 9 The enlarged model 1314 can also be trained in a similar manner to the enlarged model 914, such as Fig. 9 In some implementations, the resolution upscaling performed by upscaling model 1314 and the frame rate adjustment performed by frame rate converter 1322 may be performed in reverse order without departing from the scope of the present disclosure. In some cases, upscaling model 1314, frame rate converter 1322, or a combination thereof may be considered a domain transformation model (e.g., corresponding to Figure 5A and Figure 5B An illustrative example of a domain transformation model 514) is shown in FIG.

[0178] In some cases, inertial motion estimation data (e.g., from inertial motion estimator 1326) can be provided to magnification model 1314. In some cases, the inertial motion estimation data can be stored in an inertial motion circular buffer (not shown). In some cases, the inertial estimation data can be associated with each video frame stored in frame buffer 1312. In some cases, the magnification model can be trained to utilize the inertial motion estimation data along with key frames 1311 as a guide for magnifying frame rate adjusted frames 1325 to a second resolution. For example, if the inertial motion estimation data indicates that one or more camera sensors 1302 moved to the left immediately prior to receiving capture input 1305, magnification model 1314 can include a translational motion to the left in the transformed and magnified frame 1316. In some cases, magnification model 1314 can be trained using training data that includes training data related to magnification model 1314 for the video frame 1325. Fig. 9 The simulated inertial motion estimation data is used in a process similar to the training process described for training the augmented model 914. It should be understood that the use of the inertial motion estimation data as a guide can be used with any of the examples described herein. For example, the inertial motion data can be used as Figure 5A and Figure 5B A guide to the domain transformation model 514 is shown.

[0179] Fig.14 1 is a flow chart illustrating an example of a process 1400 for processing one or more frames. At block 1402, the process 1400 includes capturing video frames associated with a first domain during AON operation. In some cases, the video frames captured during AON operation may correspond to Figure 5A Video frame 508 in.

[0180] At block 1404, process 1400 may determine whether a capture input (e.g., a user presses a record button, performs a gesture, etc.) is received. If no capture input is received, process 1400 may proceed to block 1406. At block 1406, process 1400 may determine whether a key frame associated with a second domain different from the first domain needs to be captured. For example, process 1400 may determine the amount of time since the most recent key frame was acquired. In some cases, the period between consecutively captured key frames may have a fixed value. In some cases, the period between consecutively captured key frames may be variable. In an illustrative example, the period between consecutively captured key frames may be determined based on motion estimates from one or more motion estimators (e.g., inertial motion estimator 1326, optical motion estimator 1324, any other motion estimator, or a combination thereof).

[0181] If a capture input has been received at block 1404, process 1400 may proceed to block 1410. At block 1410, process 1400 may receive a capture input at approximately the time of the capture input (e.g., Figure 5A ) captures a key frame associated with the second domain (e.g., Figure 5A 511 as shown in FIG.

[0182] At block 1412, process 1400 may capture a video frame associated with a second domain (e.g., Figure 5A ) associated with the second domain shown in FIG.

[0183] At block 1414, process 1400 may determine whether an end recording input is received. If an end recording input is not received, process 1400 may continue to capture video frames associated with the second domain during standard operation at block 1412. If an end recording input is received, process 1400 may return to block 1402 if AON operation is enabled.

[0184] In some cases, after receiving the end recording input at block 1414, process 1400 may include utilizing a domain transformation model (e.g., Figure 5A1408 ) to transform video frames associated with the first domain captured at block 1402 to a second domain. In some cases, each of the key frames captured at block 1408 may be used as a guide for transforming a portion of the video frames associated with the first domain. In some cases, each of the key frames captured at block 1408 may be used as a guide for transforming frames from a first plurality of frames that are temporally local to each of the key frames captured at block 1408. In one implementation, a particular key frame captured at block 1408 may be used as a guide to transform a subset of video frames associated with the first domain captured prior to the particular key frame captured at block 1402, and also captured after the key frame immediately prior to the particular key frame. In one illustrative example, the particular key frame used to transform each video frame associated with the first domain captured at block 1402 may be a key frame captured closest in time to each respective video frame. In another illustrative example, selecting which of the key frames to use as guides for a particular subset of video frames associated with the first domain captured at block 1402 can be based at least in part on performing feature detection in the key frames and one or more of the video frames associated with the first domain and comparing the detected features to determine the best key frame (or key frames) to use as guides. For example, feature matching techniques can be used to determine which key frame selected between two key frames closest in time is closest in content to each corresponding video frame transformed at block 1414. In such an example, the key frame closest in content to the corresponding video frame can be used as a guide.

[0185] Fig.15 Shows Fig.14 1400 during different stages of the process 1400. The curve 1550 is not shown to scale and is provided for illustration purposes. In the example shown, the height of the bar 1522 indicates the image processing system (e.g., above Figure 3 1400) while process 1400 is capturing video frames associated with a first domain during AON operation (e.g., at block 1402). In the example shown, bar 1522 may include power consumption of storing video frames associated with the first domain in a video buffer (e.g., a circular buffer). Similarly, in the example shown, the height of bar 1524 may indicate the power consumption of one or more camera sensors in capturing frames associated with the first domain (e.g., above Figure 3 The power consumed by the AON camera sensor 316 and / or the main camera sensor 318 shown in FIG. Fig.15In the example shown, after receiving the capture input 1527 (e.g., at block 1408), the process 1400 may capture a video frame associated with the second domain (e.g., at block 1410). In the example shown, the video buffer (e.g., Figure 3 The buffer length 1526 of the DRAM 314, SRAM 320) shown in FIG. 1 is shown by an arrow that ends at the time the input 1527 is captured and extends backward in time by an amount based on the buffer length B of the video buffer.

[0186] After receiving the capture input 1527, the process 1400 may capture a video frame associated with the second domain (e.g., at block 1410). Bar 1528 may represent the power consumed by the image processing system during the capture of the video frame associated with the second domain, and show the increase in power consumption relative to the power consumed during the buffering of the video frame associated with the first domain. The increased power consumption of the image processing system may be the result of, for example, processing a video frame having a greater number of pixels, more color information, a higher frame rate, more and / or different image processing steps, a power difference associated with any other difference between the first domain and the second domain, or any combination thereof. Bar 1529 may represent the power consumed by the image processing system during the capture of a key frame associated with the second domain before receiving the capture input 1527 (e.g., at block 1406). As shown, capturing a key frame associated with the second domain may consume a considerable amount of power after receiving the capture input 1527 to capture the video frame associated with the second domain.

[0187] Bar 1530 may represent power of one or more camera sensors while capturing video frames associated with the second domain. Increased power consumption of the one or more camera sensors may result from capturing and transmitting data for a greater number of pixels, with more color information, at a higher frame rate, power differences associated with any other differences between the first domain and the second domain, or any combination thereof. Bar 1531 may represent power consumption by the one or more camera sensors during capture of a key frame associated with the second domain prior to receiving capture input 1527 (e.g., at block 1406).

[0188] Prior to receiving capture input 1527 , a period 1534 between consecutive key frames can be determined based at least in part on an estimated motion of the one or more camera sensors and / or a scene captured by the one or more camera sensors.

[0189] The capture of video frames associated with the second domain may continue until process 1400 receives an end recording input (e.g., at block 1412). In some cases, portions of video frames associated with the first domain captured prior to capture input 1527 stored in the video buffer may be converted by a domain transformation model (e.g., Figure 5A and Figure 5B 514) from a first domain to a second domain to generate a transformed video frame associated with the second domain. The transformed video frame can be combined with a video frame associated with the second domain captured after the capture input 1527 (e.g., by Figure 5A and Figure 5B The square bracket 1532 shows the combiner 518 shown in FIG. Fig.15 Total duration of the example combined video for the example shown.

[0190] Fig.16 1 is a flow chart illustrating an example of a process 1600 for processing one or more frames. At block 1602, process 1600 includes obtaining a frame from an image capture system (e.g., Figure 1 The image capture and processing system 100 shown in Figure 2 The image sensor 202 shown in FIG. 1 , and / or one or more camera sensors 502 ) obtains a first plurality of frames associated with a first setting domain, wherein the first plurality of frames are captured before obtaining a capture input.

[0191] At block 1604 , process 1600 includes obtaining, from an image capture system, a first reference frame (eg, a key frame) associated with a second settings domain, wherein the reference frame is captured prior to obtaining the capture input.

[0192] At block 1606 , process 1600 includes obtaining, from the image capture system, a second reference frame associated with a second setting domain, wherein the second reference frame is captured proximate to where the capture input was obtained.

[0193] At block 1608 , process 1600 includes obtaining, from the image capture system, a second plurality of frames associated with a second setting domain, wherein the second plurality of frames are captured after the second reference frame.

[0194] At block 1610 , process 1600 includes transforming at least a portion of the first plurality of frames based on the first reference frame to generate a first transformed plurality of frames associated with a second set domain.

[0195] At block 1612 , process 1600 includes transforming at least another portion of the first plurality of frames based on a second reference frame to generate a second transformed plurality of frames associated with a second set domain.

[0196] In some cases, process 1600 includes obtaining a motion estimate associated with the first plurality of frames and obtaining a third reference frame associated with the second setting domain from the image capture system, wherein the third reference frame was captured prior to obtaining the capture input. In some cases, an amount of time between the first reference frame and the third reference frame is based on the motion estimate associated with the first plurality of frames.

[0197] In some examples, the processes described herein (e.g., processes 600, 700, 800, 1000, 1200, 1400, 1600, and / or other processes described herein) can be performed by a computing device or apparatus. In one example, one or more processes can be performed by Figure 3 In another example, one or more processes may be performed by the image processing system 300. Fig.19 The computing system 1900 shown executes. For example, Fig.19 The computing device of computing system 1900 shown in FIG. 1 may include components of film shutter lag system 500, film shutter lag system 550, film shutter lag system 900, film shutter lag system 1100, film shutter lag system 1300, or any combination thereof, and may implement Figure 6 The process of 600 Figure 7 The process 700 Fig. 8A The process of 800 Fig.10 The process of 1000 Fig.12 The process 1200 Fig.14 The process of 1400 Fig.16 Operations of process 1600 and / or other processes described herein.

[0198] The computing device may include any suitable device, such as a vehicle or a computing device of the vehicle (e.g., a driver monitoring system (DMS) of the vehicle), a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a connected watch or smart watch, or other wearable device), a server computer, a robotic device, a television, and / or any other computing device with resource capabilities to perform the processes described herein, including the processes 600, 700, 800, 1000, 1200, 1400, 1600, and / or other processes described herein. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or (multiple) other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or (multiple) other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP) based data or other types of data.

[0199] The components of the computing device may be implemented in circuits. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or the components may include and / or may be implemented using computer software, firmware, or a combination thereof for performing the various operations described herein.

[0200] Processes 600, 700, 800, 1000, 1200, 1400, and 1600 are illustrated as logical flow diagrams, the operations of which represent a series of operations that can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer executable instructions stored on one or more computer-readable storage media that perform the described operations when executed by one or more processors. Typically, computer executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations may be combined in any order and / or in parallel to implement the process.

[0201] In addition, processes 600, 700, 800, 1000, 1200, 1400, and 1600 and / or any other processes described herein may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors by hardware or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes multiple instructions that may be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0202] As described above, various aspects of the present disclosure may utilize a machine learning model or system. Fig.17 is an illustrative example of a deep learning neural network 1700 that can be used to implement machine learning-based feature extraction and / or activity recognition (or classification). In one illustrative example, in training a guided domain transformation model (e.g., as described above) using a GAN as described above, Figure 5A ) and / or guide super-resolution models (e.g., as described with respect to Fig. 9 The discriminator network may use feature extraction and / or activity recognition during the process described above. Input layer 1720 includes input data. In an illustrative example, input layer 1720 may include data representing pixels of an input video frame. Neural network 1700 includes multiple hidden layers 1722a, 1722b to 1722n. Hidden layers 1722a, 1722b to 1722n include "n" hidden layers, where "n" is an integer greater than or equal to 1. The number of hidden layers may include as many layers as required for a given application. Neural network 1700 also includes an output layer 1721, which provides an output generated by the processing performed by hidden layers 1722a, 1722b to 1722n. In an illustrative example, output layer 1721 may provide a classification of objects in an input video frame. Classification may include identifying a category of activity type (e.g., looking up, looking down, closing eyes, yawning, etc.).

[0203] Neural network 1700 is a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with a node is shared between different layers, and each layer retains the information as it is processed. In some cases, neural network 1700 may include a feedforward network in which there is no feedback connection where the output of the network is fed back to itself. In some cases, neural network 1700 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read in.

[0204] Information can be exchanged between nodes by node-to-node interconnection between each layer.The nodes of input layer 1720 can activate the node set in the first hidden layer 1722a.For example, as shown in the figure, each input node of input layer 1720 is connected to each node of the first hidden layer 1722a.The nodes of the first hidden layer 1722a can transform the information of each input node by applying an activation function to the input node information.Then the information derived from this transformation can be transferred to the node of the next hidden layer 1722b and the node of the next hidden layer 922b can be activated, and the node of the next hidden layer 922b can perform its own specified function.Example functions include convolution, upsampling, data conversion and / or any other suitable function.Then, the output of hidden layer 1722b can activate the node of the next hidden layer, and so on.The output of last hidden layer 1722n can activate one or more nodes of output layer 1721, and output is provided at the one or more nodes. In some cases, although a node in neural network 1700 (e.g., node 1726) is shown as having multiple output lines, the node has a single output, and all lines shown as output from the node represent the same output value.

[0205] In some cases, each node or interconnection between nodes can have a weight, which is a set of parameters derived from the training of neural network 1700. Once neural network 1700 is trained, it can be referred to as a trained neural network, which can be used to classify one or more activities. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnections can have adjustable digital weights that can be adjusted (e.g., based on a training data set), allowing neural network 1700 to adapt to the input and be able to learn as more and more data is processed.

[0206] The neural network 1700 is pre-trained to process features from the data in the input layer 1720 using different hidden layers 1722a, 1722b to 1722n to provide an output through the output layer 1721. In an example where the neural network 1700 is used to identify an activity performed by a driver in a frame, the neural network 1700 can be trained using training data including both frames and labels, as described above. For example, training frames can be input into the network, where each training frame has a label indicating features in the frame (for a feature extraction machine learning system) or a label indicating the class of the activity in each frame. In one example using object classification for illustrative purposes, the training frames may include an image of the number 2, in which case the labels for the image may be [0 0 1 0 0 0 0 0 0 0].

[0207] In some cases, the neural network 1700 can use a training process called back propagation to adjust the weights of the nodes. As described above, the back propagation process can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, the loss function, the backward pass, and the parameter update are performed for one training iteration. For each training image set, the process can be repeated a certain number of iterations until the neural network 1700 is trained well enough so that the weights of each layer are accurately adjusted.

[0208] For the example of identifying an object in a frame, a forward pass may include passing a training frame through the neural network 1700. Prior to training the neural network 1700, the weights are initially randomized. As an illustrative example, a frame may include an array of numbers representing pixels of an image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that position in the array. In one example, the array may include a 28×28×3 digital array having 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chrominance components, etc.).

[0209] As described above, for the first training iteration of the neural network 1700, the output will likely include values ​​that do not favor any particular class due to the weights randomly selected at initialization. For example, if the output is a vector with probabilities that the object includes different classes, the probability values ​​for each different class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 1700 cannot determine low-level features and, therefore, cannot accurately determine what the classification of the object may be. A loss function may be used to analyze the errors in the output. Any suitable loss function definition may be used, such as a cross entropy loss. Another example of a loss function includes the mean squared error (MSE), defined as The loss can be set equal to E tital The value of .

[0210] For the first training images, the loss (or error) will be high because the actual values ​​will be very different from the predicted outputs. The goal of training is to minimize the amount of loss so that the predicted outputs are the same as the training labels. The neural network 1700 can perform a backward pass by determining which inputs (weights) contribute the most to the loss of the network, and the weights can be adjusted so that the loss decreases and is ultimately minimized. The derivative of the loss with respect to the weights (expressed as dL / dW, where W is the weight at a particular layer) can be calculated to determine the weights that contribute the most to the loss of the network. After calculating the derivatives, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be expressed as Where w represents the weight, wi represents the initial weight, and η represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes a larger weight update and a lower value indicates a smaller weight update.

[0211] The neural network 1700 may include any suitable deep network. An example includes a convolutional neural network (CNN), which includes an input layer and an output layer with multiple hidden layers between the input layer and the output layer. The hidden layers of the CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. The neural network 1700 may include any other deep network other than a CNN, such as an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), etc.

[0212] Fig.18 is an illustrative example of a convolutional neural network (CNN) 1800. The input layer 1820 of the CNN 1800 includes data representing an image or frame. For example, the data may include an array of numbers representing pixels of the image, where each number in the array includes a value from 0 to 255 describing the intensity of the pixel at that position in the array. Using the previous example from above, the array may include a 28×28×3 digital array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chrominance components, etc.). The image may be passed through a convolutional hidden layer 1822a, an optional non-linear activation layer, a pooling hidden layer 1822b, and a fully connected hidden layer 1822c to obtain an output at an output layer 1824. Although Fig.18 Only one hidden layer is shown in the various hidden layers, but a person of ordinary skill will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in CNN 1800. As previously described, the output may indicate a single category of an object, or may include a probability of a category that best describes an object in the image.

[0213] The first layer of CNN 1800 is a convolutional hidden layer 1822. Convolutional hidden layer 1822a analyzes the image data of input layer 1820. Each node of convolutional hidden layer 1822a is connected to an area of ​​nodes (pixels) of the input image called a receptive field. Convolutional hidden layer 1822a can be considered as one or more filters (each filter corresponds to a different activation or feature map), where each convolution iteration of the filter is a node or neuron of convolutional hidden layer 1822a. For example, the area of ​​the input image covered by the filter at each convolution iteration will be the receptive field of the filter. In an illustrative example, if the input image includes a 28×28 array and each filter (and corresponding receptive field) is a 5×5 array, there will be 24×24 nodes in convolutional hidden layer 1822a. Each connection between a node and the receptive field of that node learns a weight, and in some cases learns an overall bias, so that each node learns to analyze its specific local receptive field in the input image. Each node of hidden layer 1822a will have the same weights and biases (referred to as shared weights and shared biases). For example, the filter has an array of weights (numbers) and the same depth as the input. For the video frame example, the filter will have a depth of 3 (according to the three color components of the input image). An illustrative example size of the filter array is 5×5×3, corresponding to the size of the node's receptive field.

[0214] The convolutional nature of the convolutional hidden layer 1822a is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filter of the convolutional hidden layer 1822a can start at the upper left corner of the input image array and can be convolved around the input image. As described above, each convolution iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 1822a. At each convolution iteration, the value of the filter is multiplied by the original pixel value of the corresponding number of the image (for example, a 5×5 filter array multiplied by a 5×5 array of input pixel values ​​at the upper left corner of the input image array). The multiplications from each convolution iteration can be added together to obtain the sum of the iteration or node. Next, the process continues at the next position in the input image according to the receptive field of the next node in the convolutional hidden layer 1822a. For example, the filter can move a step amount (called a stride) to the next receptive field. The stride can be set to 1 or other suitable amount. For example, if the stride is set to 1, the filter will move 1 pixel to the right at each convolution iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, resulting in a summed value being determined for each node of the convolutional hidden layer 1822a.

[0215] The mapping from the input layer to the convolutional hidden layer 1822a is called an activation map (or feature map). The activation map includes a value for each node that represents the filter result at each location of the input volume. The activation map may include an array that includes various summed values ​​produced by each iteration of the filter over the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map will include a 24×24 array. The convolutional hidden layer 1822a may include several activation maps in order to identify multiple features in an image. Fig.18 The example shown in includes three activation maps. Using these three activation maps, the convolutional hidden layer 1822a can detect three different kinds of features, each of which is detectable over the entire image.

[0216] In some examples, a nonlinear hidden layer can be applied after the convolutional hidden layer 1822a. A nonlinear layer can be used to introduce nonlinearity into a system that already computes linear operations. An illustrative example of a nonlinear layer is a rectified linear unit (ReLU) layer. The ReLU layer can apply the function f(x) = max(0, x) to all values ​​in the input volume, which changes all negative activations to 0. Thus, the ReLU can increase the nonlinear properties of the CNN 1800 without affecting the receptive field of the convolutional hidden layer 1822a.

[0217] The pooling hidden layer 1822b may be applied after the convolutional hidden layer 1822a (and after the non-linear hidden layer when used). The pooling hidden layer 1822b is used to simplify the information in the output from the convolutional hidden layer 1822a. For example, the pooling hidden layer 1822b may take each activation map output from the convolutional hidden layer 1822a and generate a condensed activation map (or feature map) using a pooling function. Maximum pooling is an example of a function performed by a pooling hidden layer. The pooling hidden layer 1822a uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. A pooling function (e.g., a maximum pooling filter, an L2 norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 1822a. In Fig.18 In the example shown, three pooling filters are used for the three activation maps in the convolutional hidden layer 1822a.

[0218] In some examples, max pooling can be used by applying a max pooling filter (e.g., having a size of 2×2) with a stride (e.g., equal to the dimension of the filter, such as a stride of 2) to the activation map output from the convolutional hidden layer 1822a. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes in the previous layer (each node is a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by a 2×2 max pooling filter at each iteration of the filter, where the maximum of the four values ​​is output as the "maximum" value. If such a max pooling filter is applied to an activation filter with a dimension of 24×24 nodes from the convolutional hidden layer 1822a, the output from the pooling hidden layer 1822b will be an array of 12×12 nodes.

[0219] In some examples, an L2 norm (L-norm) pooling filter may also be used. The L2 norm pooling filter involves calculating the square root of the sum of the squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (rather than calculating the maximum value as done in max pooling), and using the calculated value as output.

[0220] Intuitively, a pooling function (e.g., max pooling, L2 norm pooling, or other pooling functions) determines whether a given feature is found anywhere in a region of an image, and discards the exact location information. This can be done without affecting the results of feature detection, because once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max pooling (as well as other pooling methods) provides the following benefits: far fewer features are pooled, thereby reducing the number of parameters required in subsequent layers of CNN 1800.

[0221] The last layer of connections in the network is a fully connected layer that connects each node from the pooling hidden layer 1822b to each output node in the output layer 1824. Using the above example, the input layer includes 28x 28 nodes that encode the pixel intensities of the input image, the convolutional hidden layer 1822a includes 3×24×24 hidden feature nodes based on applying a 5x5 local receptive field (for the filter) to the three activation maps, and the pooling hidden layer 1822b includes a layer of 3×12×12 hidden feature nodes based on applying a max pooling filter to a 2×2 region across each of the three feature maps. Extending this example, the output layer 1824 may include ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 1822b is connected to each node of the output layer 1824.

[0222] The fully connected layer 1822c can obtain the output of the previous pooled hidden layer 1822b (which should represent an activation map of high-level features) and determine the features most relevant to a particular category. For example, the fully connected layer 1822c layer can determine the high-level features that are most strongly associated with a particular category, and can include weights (nodes) for the high-level features. The product between the weights of the fully connected layer 1822c and the pooled hidden layer 1822b can be calculated to obtain probabilities for different categories. For example, if CNN 1800 is used to predict that an object in a video frame is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., there are two legs, a face at the top of the object, two eyes at the upper left and upper right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0223] In some examples, the output from the output layer 1824 may include an M-dimensional vector (in the previous example, M=10). M indicates the number of classes that the CNN 1800 must choose from when classifying an object in an image. Other example outputs may also be provided. Each number in the M-dimensional vector may represent the probability that an object belongs to a certain category. In an illustrative example, if the 10-dimensional output vector represents that objects of ten different categories are [0 0 0.05 0.8 0 0.15 0 0 0 0], the vector indicates that the probability that the image is an object of the third category (e.g., dog) is 5%, the probability that the image is an object of the fourth category (e.g., person) is 80%, and the probability that the image is an object of the sixth category (e.g., kangaroo) is 15%. The probability for a category can be considered as the confidence level that the object is part of that category.

[0224] Fig.19 is a schematic diagram showing an example of a system for implementing certain aspects of the present technology. Specifically, Fig.19 An example of a computing system 1900 is shown, which may be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1905. Connection 1905 may be a physical connection using a bus, or a direct connection into processor 1910 (e.g., in a chipset architecture). Connection 1905 may also be a virtual connection, a networked connection, or a logical connection.

[0225] In some embodiments, computing system 1900 is a distributed system, where the functions described in the present disclosure may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent a number of such components, each of which performs some or all of the functions described for that component. In some embodiments, a component may be a physical device or a virtual device.

[0226] The example computing system 1900 includes at least one processing unit (CPU or processor) 1910 and connections 1905 coupling various system components including system memory 1915, such as read-only memory (ROM) 1920 and random access memory (RAM) 1925, to the processor 1910. The computing system 1900 may include a cache 1912 of high-speed memory directly connected to the processor 1910, close to the processor 1510, or integrated as part of the processor 1510.

[0227] Processor 1910 may include any general purpose processor and hardware or software services, such as services 1932, 1934, and 1936 stored in storage device 1930, configured to control processor 1910 as well as special purpose processors in which software instructions are incorporated into the actual processor design. Processor 1910 may essentially be a completely self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, and the like.

[0228] To enable user interaction, the computing system 1900 includes an input device 1945, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keypad, a mouse, motion input, voice, and the like. The computing system 1900 can also include an output device 1935, which can be one or more of a plurality of output mechanisms. In some instances, a multimodal system can enable a user to provide multiple types of input / output to communicate with the computing system 1900. The computing system 1900 can include a communication interface 1940, which can generally control and manage user input and system output. The communication interface can use a wired and / or wireless transceiver to perform or facilitate the reception and / or transmission of wired or wireless communications, including the use of an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Proprietary wired ports / plugs, Wireless signal transmission, Low power (BLE) wireless signal transmission, Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or a wired and / or wireless transceiver of some combination thereof. The communication interface 1940 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 1900 based on receiving one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States-based Global Positioning System (GPS), the Russian-based Global Navigation Satellite System (GLONASS), the Chinese-based BeiDou Navigation Satellite System (BDS), and the European-based Galileo GNSS. There are no restrictions on the operation of any particular hardware arrangement, so the basic features here can be easily replaced with improved hardware or firmware arrangements as they are developed.

[0229] The storage device 1930 may be a non-volatile and / or non-transitory and / or computer-readable memory device, and may be a hard disk, or other type of computer-readable medium that can store data accessible by a computer, such as a magnetic tape cartridge, a flash memory card, a solid-state storage device, a digital versatile disk, a cassette, a floppy disk, a floppy disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) disc, a rewritable compact disc (CD) disc, a digital video disc (DVD) disc, a Blu-ray disc (BDD) disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin-transfer torque RAM (STT-RAM), another memory chip or box, and / or a combination thereof.

[0230] Storage device 1930 may include software services, servers, services, etc., which cause the system to perform functions when the code defining such software is executed by processor 1910. In some embodiments, hardware services that perform specific functions may include software components stored in a computer-readable medium that interface with the necessary hardware components (such as processor 1910, connection 1905, output device 1935, etc.) to perform the functions.

[0231] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying (multiple) instructions and / or data. Computer-readable media may include non-transitory media that can store data, but does not include carrier waves and / or temporary electronic signals transmitted wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to: disks or tapes, optical storage media (e.g., compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or storage devices. Computer-readable media may store code and / or machine-executable instructions that may represent any combination of procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be delivered, forwarded, or sent using any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0232] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0233] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be appreciated by those of ordinary skill in the art that these embodiments can be implemented without these specific details. For clarity of explanation, in some cases, the technology herein can be presented as a separate functional block including the following functional blocks, which include equipment, device components, steps or routines in the method embodied in software, or a combination of hardware and software. Additional components other than those shown in the accompanying drawings and / or described herein can be used. For example, circuits, systems, networks, processes, and other components can be shown as components in block diagram form, so as not to obscure these embodiments in unnecessary details. In other cases, known circuits, processes, algorithms, structures, and techniques can be shown as not having unnecessary details, so as to avoid obscuring these embodiments.

[0234] Each embodiment can be described as a process or method above, and this process or method is described as flow chart, flow diagram, data flow chart, structure diagram or block diagram.Although the operation is described as a sequential process using flow chart, many operations can be parallel or carried out simultaneously.Additionally, the order of these operations can be rearranged.When these operations end, the processing process also ends, but it can have other steps not included in the accompanying drawings.Process can correspond to method, function, process, subroutine, subroutine etc.When process corresponds to function, its termination can correspond to function returning to calling function or main function.

[0235] The process and method according to the above-mentioned example can be implemented using computer executable instructions stored in a computer readable medium or otherwise obtainable from a computer readable medium. For example, such instructions may include instructions and data that cause or configure a general-purpose computer, a special-purpose computer, or a processing device to perform a specific function or a functional group. The computer resource part used can be accessed through a network. The computer executable instructions can be, for example, binary, intermediate format instructions (such as assembly language, firmware, source code, etc.). Examples of computer readable media that can be used to store instructions, information used, and / or information created during the method according to the described examples include disks or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, etc.

[0236] The equipment implementing the process and method according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language or any combination thereof, and may adopt any of a variety of form factors. When implemented in software, firmware, middleware or microcode, the program code or code segment (e.g., computer program product) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. One (or more) processors may perform the necessary tasks. Typical examples of form factors include personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc., of laptop computers, smart phones, mobile phones, tablet devices or other small form factors. The functions described herein may also be embodied in peripheral devices or add-on cards. By further example, such functions may also be implemented on circuit boards between different processes performed in different chips or in a single device.

[0237] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are exemplary means for providing the functions described in this disclosure.

[0238] In the foregoing description, various aspects of the present application are described with reference to the specific embodiments of the present application, but those of ordinary skill in the art will recognize that the present application is not limited thereto. Therefore, although the illustrative embodiments of the present application have been described in detail herein, it should be understood that these inventive concepts can be embodied and adopted differently in other ways, and the appended claims are intended to be interpreted as including such variations, except as limited by the prior art. The various features and aspects of the above-mentioned applications can be used individually or in combination. In addition, without departing from the broader spirit and scope of protection of this specification, embodiments can be used in any number of environments and applications outside of those described herein. Therefore, the description and the accompanying drawings should be considered illustrative rather than restrictive. For illustration purposes, the method is described in a specific order. It should be appreciated that in an alternative embodiment, the method can be performed in an order different from the order described.

[0239] It should be understood by those of ordinary skill in the art that the less than ("<") and greater than (">") symbols or terms used in this document may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of this specification.

[0240] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0241] The phrase "coupled to" refers to any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other appropriate communication interface).

[0242] Claim language or other language that refers to "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that recites "at least one of A and B" or "at least one of A or B" refers to A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" or "at least one of A, B, or C" refers to A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set of items listed in the set. For example, claim language that recites "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0243] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. The technician may implement the described functions in different ways for each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0244] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handheld devices, or integrated circuit devices with multiple uses, including applications in wireless communication device handheld devices and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or individually as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium, which includes a program code, which includes instructions for executing one or more of the methods described above when executed. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a FLASH memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the techniques may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0245] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated logic circuits or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor; however, in an alternative manner, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the techniques described herein.

[0246] Illustrative aspects of the disclosure include:

[0247] Aspect 1: A method for processing one or more frames, comprising: obtaining a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtaining at least one reference frame associated with a second setting domain from the image capture system, wherein the at least one reference frame is captured close to obtaining the capture input; obtaining a second plurality of frames associated with the second setting domain from the image capture system, wherein the second plurality of frames are captured after the at least one reference frame; and based on the at least one reference frame, transforming at least a portion of the first plurality of frames to generate a transformed plurality of frames associated with the second setting domain.

[0248] Aspect 2: A method according to Aspect 1, wherein: the first setting domain includes a first resolution; the second setting domain includes a second resolution; and transforming at least the portion of the first plurality of frames includes upscaling at least the portion of the first plurality of frames from the first resolution to the second resolution to generate the transformed plurality of frames, wherein the transformed plurality of frames have the second resolution.

[0249] Aspect 3: The method according to any one of Aspects 1 to 2 further includes: obtaining an additional reference frame having the second resolution from the image capture system, wherein the additional reference frame is captured before obtaining the capture input, wherein generating the enlarged multiple frames having the second resolution is based on at least the portion of the first multiple frames, the at least one reference frame and the additional reference frame, and wherein the at least one reference frame provides a reference for enlarging at least a first portion of at least the portion of the first multiple frames, and the additional reference frame provides a reference for enlarging at least a second portion of at least the portion of the first multiple frames.

[0250] Aspect 4: The method according to any one of aspects 1 to 3, further comprising: combining the transformed plurality of frames and the second plurality of frames to generate a video associated with the second setting domain.

[0251] Aspect 5: The method according to any one of Aspects 1 to 4 further includes: obtaining motion information associated with the first plurality of frames, wherein generating the transformed plurality of frames associated with the second setting domain is based on at least the portion of the first plurality of frames, the at least one reference frame and the motion information.

[0252] Aspect 6: The method according to any one of aspects 1 to 5, further comprising: determining a translation direction based on the motion information associated with the first plurality of frames; and applying the translation direction to the transformed plurality of frames.

[0253] Aspect 7: A method according to any one of Aspects 1 to 6, wherein: the first setting domain includes a first frame rate; the second setting domain includes a second frame rate; and transforming at least the portion of the first plurality of frames includes converting at least the portion of the first plurality of frames from the first frame rate to the second frame rate.

[0254] Aspect 8: A method according to any one of Aspects 1 to 7, wherein a first subset of the first plurality of frames is captured at the first frame rate and a second subset of the first plurality of frames is captured at a third frame rate different from the first frame rate, wherein the third frame rate is equal to or not equal to the second frame rate, and wherein the change between the first frame rate and the third frame rate is at least partially based on motion information associated with at least one of the first subset of the first plurality of frames and the second subset of the first plurality of frames.

[0255] Aspect 9: A method according to any one of Aspects 1 to 8, wherein: the first setting domain includes a first resolution and a first frame rate; the second setting domain includes a second resolution and a second frame rate; and transforming at least the portion of the first plurality of frames includes enlarging at least the portion of the first plurality of frames from the first resolution to the second resolution, and converting at least the portion of the first plurality of frames from the first frame rate to the second frame rate.

[0256] Aspect 10: The method according to any one of Aspects 1 to 9 further includes: obtaining an additional reference frame associated with the second setting domain from the image capture system, wherein the additional reference frame is captured before obtaining the capture input, wherein the transformed multiple frames associated with the second setting domain are generated based on at least the portion of the first multiple frames, at least one reference frame and the additional reference frame, and wherein the at least one reference frame provides a reference for transforming at least a first subset of at least the portion of the first multiple frames, and the additional reference frame provides a reference for transforming at least a second subset of at least the portion of the first multiple frames.

[0257] Aspect 11: A method according to any one of Aspects 1 to 10, wherein the at least one reference frame includes a first reference frame, and the method further includes: obtaining a second reference frame associated with the second setting domain from the image capture system, wherein the second reference frame is captured close to obtaining the capture input; based on the first reference frame, transforming at least the portion of the first plurality of frames to generate the transformed plurality of frames associated with the second setting domain; and based on the second reference frame, transforming at least another portion of the first plurality of frames to generate a second transformed plurality of frames associated with the second setting domain.

[0258] Aspect 12: The method according to any one of Aspects 1 to 11 further includes: obtaining a motion estimate associated with the first plurality of frames; obtaining a third reference frame associated with the second setting domain from the image capture system, wherein the third reference frame is captured before obtaining the capture input; and based on the third reference frame, transforming a third portion of the first plurality of frames to generate a third transformed plurality of frames associated with the second setting domain; wherein the amount of time between the first reference frame and the third reference frame is based on the motion estimate associated with the first plurality of frames.

[0259] Aspect 13: A method according to any one of Aspects 1 to 12, wherein the first setting domain includes at least one of a first resolution, a first frame rate, a first color depth, a first noise reduction technology, a first edge enhancement technology, a first image stabilization technology, and a first color correction technology, and the second setting domain includes at least one of a second resolution, a second frame rate, a second color depth, a second noise reduction technology, a second edge enhancement technology, a second image stabilization technology, and a second color correction technology.

[0260] Aspect 14: The method according to any one of Aspects 1 to 13 further includes: generating the transformed multiple frames using a trainable neural network, wherein the neural network is trained using a training data set including image pairs, each pair of images including a first image associated with the first setting domain and a second image associated with the second setting domain.

[0261] Aspect 15: A method according to any one of Aspects 1 to 14, wherein capturing the at least one reference frame close to obtaining the capture input includes: capturing a first available frame associated with the second setting domain after receiving the capture input; capturing a second available frame associated with the second setting domain after receiving the capture input; capturing a third available frame associated with the second setting domain after receiving the capture input; or capturing a fourth available frame associated with the second setting domain after receiving the capture input.

[0262] Aspect 16: A method according to any one of Aspects 1 to 15, wherein capturing the at least one reference frame near obtaining the capture input includes: capturing a frame associated with the second setting domain within 10 milliseconds (ms), within 100 ms, within 500 ms, or within 1000 ms after receiving the capture input.

[0263] Aspect 17: A device for processing one or more frames, comprising: at least one memory; and one or more processors, coupled to the at least one memory and configured to: obtain a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtain at least one reference frame associated with a second setting domain from the image capture system, wherein the at least one reference frame is captured close to obtaining the capture input; obtain a second plurality of frames associated with the second setting domain from the image capture system, wherein the second plurality of frames are captured after the at least one reference frame; and based on the at least one reference frame, transform at least a portion of the first plurality of frames to generate a transformed plurality of frames associated with the second setting domain.

[0264] Aspect 18: An apparatus according to Aspect 17, wherein: the first setting domain includes a first resolution; the second setting domain includes a second resolution; and in order to transform at least the portion of the first plurality of frames, the one or more processors are configured to upscale at least the portion of the first plurality of frames from the first resolution to the second resolution to generate the transformed plurality of frames, wherein the transformed plurality of frames have the second resolution.

[0265] Aspect 19: An apparatus according to any one of Aspects 17 to 18, wherein the one or more processors are configured to: obtain an additional reference frame having the second resolution from the image capture system, wherein the additional reference frame is captured before obtaining the capture input, wherein generating the enlarged multiple frames having the second resolution is based on at least the portion of the first multiple frames, the at least one reference frame and the additional reference frame, and wherein the at least one reference frame provides a reference for enlarging at least a first portion of at least the portion of the first multiple frames, and the additional reference frame provides a reference for enlarging at least a second portion of at least the portion of the first multiple frames.

[0266] Aspect 20: The apparatus according to any one of aspects 17 to 19, wherein the one or more processors are configured to: combine the transformed plurality of frames and the second plurality of frames to generate a video associated with the second setting domain.

[0267] Aspect 21: An apparatus according to any one of Aspects 17 to 20, wherein the one or more processors are further configured to: obtain motion information associated with the first plurality of frames, wherein generating the transformed plurality of frames associated with the second setting domain is based on at least the portion of the first plurality of frames, the at least one reference frame and the motion information.

[0268] Aspect 22: An apparatus according to any one of Aspects 17 to 21, wherein the one or more processors are configured to: determine a translation direction based on the motion information associated with the first plurality of frames; and apply the translation direction to the transformed plurality of frames.

[0269] Aspect 23: An apparatus according to any one of Aspects 17 to 22, wherein: the first setting domain includes a first frame rate; the second setting domain includes a second frame rate; and in order to transform at least the portion of the first plurality of frames, the one or more processors are configured to: convert at least the portion of the first plurality of frames from the first frame rate to the second frame rate.

[0270] Aspect 24: An apparatus according to any one of Aspects 17 to 23, wherein a first subset of the first plurality of frames is captured at the first frame rate and a second subset of the first plurality of frames is captured at a third frame rate different from the first frame rate, wherein the third frame rate is equal to or not equal to the second frame rate, and wherein a change between the first frame rate and the third frame rate is at least partially based on motion information associated with at least one of the first subset of the first plurality of frames and the second subset of the first plurality of frames.

[0271] Aspect 25: An apparatus according to any one of Aspects 17 to 24, wherein: the first setting domain includes a first resolution and a first frame rate; the second setting domain includes a second resolution and a second frame rate; and in order to transform at least the portion of the first plurality of frames, the one or more processors are configured to: enlarge at least the portion of the first plurality of frames from the first resolution to the second resolution, and convert at least the portion of the first plurality of frames from the first frame rate to the second frame rate.

[0272] Aspect 26: According to an apparatus described in any one of Aspects 17 to 25, the one or more processors are configured to: obtain an additional reference frame associated with the second setting domain from the image capture system, wherein the additional reference frame is captured before obtaining the capture input, and the generation of the transformed multiple frames associated with the second setting domain is based on at least the portion of the first multiple frames, at least one reference frame and the additional reference frame, and wherein the at least one reference frame provides a reference for transforming at least a first subset of at least the portion of the first multiple frames, and the additional reference frame provides a reference for transforming at least a second subset of at least the portion of the first multiple frames.

[0273] Aspect 27: An apparatus according to any one of Aspects 17 to 26, wherein the one or more processors are configured to: obtain a second reference frame associated with the second setting domain from the image capture system, wherein the second reference frame is captured close to obtaining the capture input; based on the first reference frame, transform at least the portion of the first plurality of frames to generate the transformed plurality of frames associated with the second setting domain; and based on the second reference frame, transform at least another portion of the first plurality of frames to generate a second transformed plurality of frames associated with the second setting domain.

[0274] Aspect 28: According to an apparatus described in any one of Aspects 17 to 27, the one or more processors are configured to: obtain a motion estimate associated with the first plurality of frames; and obtain a third reference frame associated with the second setting domain from the image capture system, wherein the third reference frame is captured before obtaining the capture input; and based on the third reference frame, transform a third portion of the first plurality of frames to generate a third transformed plurality of frames associated with the second setting domain; wherein the amount of time between the first reference frame and the third reference frame is based on the motion estimate associated with the first plurality of frames.

[0275] Aspect 29: An apparatus according to any one of Aspects 17 to 28, wherein the first setting domain includes at least one of a first resolution, a first frame rate, a first color depth, a first noise reduction technology, a first edge enhancement technology, a first image stabilization technology, and a first color correction technology; and the second setting domain includes at least one of a second resolution, a second frame rate, a second color depth, a second noise reduction technology, a second edge enhancement technology, a second image stabilization technology, and a second color correction technology.

[0276] Aspect 30: According to the apparatus described in any one of Aspects 17 to 29, the one or more processors are configured to: generate the transformed multiple frames using a trainable neural network, wherein the neural network is trained using a training data set including image pairs, each pair of images including a first image associated with the first setting domain and a second image associated with the second setting domain.

[0277] Aspect 31: An apparatus according to any one of Aspects 17 to 30, wherein, in order to capture the at least one reference frame close to obtaining the capture input, the one or more processors are configured to: capture a first available frame associated with the second setting domain after receiving the capture input; capture a second available frame associated with the second setting domain after receiving the capture input; capture a third available frame associated with the second setting domain after receiving the capture input; or capture a fourth available frame associated with the second setting domain after receiving the capture input.

[0278] Aspect 32: An apparatus according to any one of Aspects 17 to 31, wherein, in order to capture the at least one reference frame near obtaining the capture input, the one or more processors are configured to: capture a frame associated with the second setting domain within 10 milliseconds (ms), within 100 ms, within 500 ms, or within 1000 ms after receiving the capture input.

[0279] Aspect 33: Aspect 31: A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to perform any of the operations of Aspect 1 to Aspect 32.

[0280] Aspect 34: An apparatus comprising means for performing any of the operations described in aspects 1 to 32.

[0281] Aspect 35: A method for processing one or more frames, comprising: obtaining a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtaining a reference frame associated with a second setting domain from the image capture system, wherein the reference frame is captured close to obtaining the capture input; obtaining a selection of one or more selected frames associated with the first plurality of frames; and based on the reference frame, transforming the one or more selected frames to generate one or more transformed frames associated with the second setting domain.

[0282] Aspect 36: The method of aspect 35, wherein the selection of the one or more selected frames is based on a selection from a user interface.

[0283] Aspect 37: A method according to any one of aspects 35 to 36, wherein the user interface comprises a thumbnail gallery, a slider, or a frame-by-frame review.

[0284] Aspect 38: A device for processing one or more frames, comprising: at least one memory; and one or more processors coupled to the at least one memory and configured to: obtain a first plurality of frames associated with a first setting domain from an image capture system, wherein the first plurality of frames are captured before obtaining a capture input; obtain a reference frame associated with a second setting domain from the image capture system, wherein the reference frame is captured close to obtaining the capture input; obtain a selection of one or more selected frames associated with the first plurality of frames; and based on the reference frame, transform the one or more selected frames to generate one or more transformed frames associated with the second setting domain.

[0285] Aspect 39: The apparatus of aspect 38, wherein the selection of the one or more selected frames is based on a selection from a user interface.

[0286] Clause 40: The apparatus of any one of Clauses 38 to 39, wherein the user interface comprises a thumbnail gallery, a slider, or a frame-by-frame review.

[0287] Aspect 41: Aspect 31: A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to perform any of the operations of Aspect 35 to Aspect 40.

[0288] Aspect 42: An apparatus comprising means for performing any of the operations described in aspects 35 to 40.

[0289] Aspect 43: A method comprising the operations according to any of Aspects 1 to 32 and any of Aspects 35 to 40.

[0290] Aspect 44: An apparatus for processing one or more frames. The apparatus comprises at least one memory (e.g., implemented in a circuit) configured to store the one or more frames and one or more processors (e.g., one processor or multiple processors) coupled to the at least one memory. The one or more processors are configured to perform the operations according to any one of aspects 1 to 32 and any one of aspects 35 to 40.

[0291] Aspect 45: A computer-readable storage medium storing instructions, wherein when the instructions are executed by one or more processors, the one or more processors are caused to perform the operations according to any of Aspects 26 to 35 and any of Aspects 36 to 43.

[0292] Aspect 46: An apparatus comprising means for performing the operations according to any of Aspects 26 to 37 and any of Aspects 38 to 45.

Claims

1. A method for processing one or more frames, include: obtaining a first plurality of frames from an image capture system, wherein the first plurality of frames is captured prior to obtaining a capture input; Responsive to the capture input, obtaining at least one additional frame; determining a change in at least one frame relative to one or more frames captured in temporal proximity to the at least one frame; displaying at least one suggested frame in an interface for selecting at least one selected frame based on determining the change to the at least one frame, each selected frame being associated with one or more frames of the first plurality of frames; receiving, via the interface, a selection of the at least one selected frame; and Based on the selection, an image is generated based on the at least one selected frame and the at least one additional frame.

2. The method according to claim 1, in, The at least one additional frame is obtained from an additional image capture system different from the image capture system.

3. The method according to claim 1, in, The interface includes one or more of a slider for advancing through frames or a gallery corresponding to the first plurality of frames.

4. The method according to claim 1, in, The first plurality of frames are stored in a frame buffer.

5. The method according to claim 1, further comprising: include: determining at least one additional selected frame associated with one or more frames in the first plurality of frames, wherein determining the at least one additional selected frame comprises one or more of: determining a change of the at least one additional selected frame relative to one or more frames captured in temporal proximity to the at least one additional selected frame; determining that an amount of motion in the at least one additional selected frame exceeds a motion threshold; determining that the at least one additional frame contains relevant content; determining that the at least one additional frame contains a human face; or The at least one additional selected frame is processed by a neural network.

6. The method according to claim 1, further comprising: include: determining that an amount of motion in the at least one frame exceeds a motion threshold; as well as Based on determining that the amount of motion in the at least one frame exceeds the motion threshold, the at least one suggested frame is displayed.

7. The method according to claim 1, in, The at least one proposal frame is generated using a deep learning neural network.

8. The method according to claim 1, in: The at least one selected frame is associated with a first setting domain; The at least one additional frame is associated with a second setup field; as well as Generating the image includes transforming the at least one selected frame from the first setting domain to the second setting domain.

9. The method according to claim 8, in: The first setting domain includes a first resolution; The second setting field includes a second resolution; as well as Converting the at least one selected frame from the first setting domain to the second setting domain includes upscaling the at least one selected frame from the first resolution to the second resolution.

10. The method according to claim 8, in: The first setting field includes a first frame rate; The second setting field includes a second frame rate; as well as Transforming the at least one selected frame from the first setting domain to the second setting domain includes frame rate converting the at least one selected frame from the first frame rate to the second frame rate.

11. The method according to claim 8, in: The first setting domain includes a first resolution and a first frame rate; The second setting field includes a second resolution and a second frame rate; as well as Converting the at least one selected frame from the first setting domain to the second setting domain includes upscaling the at least one selected frame from the first resolution to the second resolution and frame rate converting the at least one selected frame from the first frame rate to the second frame rate.

12. The method according to claim 1, in, Displaying the interface for selecting the at least one selected frame includes displaying one or more frames of the first plurality of frames.

13. The method according to claim 1, in, The at least one selected frame comprises a frame selected from the first plurality of frames.

14. The method according to claim 1, in, The at least one additional frame is a key frame.

15. An apparatus for processing one or more frames, include: Memory; as well as one or more processors coupled to the memory and configured to: obtaining a first plurality of frames from an image capture system, wherein the first plurality of frames is captured prior to obtaining a capture input; Responsive to the capture input, obtaining at least one additional frame; determining a change in at least one frame relative to one or more frames captured in temporal proximity to the at least one frame; displaying at least one suggested frame in an interface for selecting at least one selected frame based on determining the change to the at least one frame, each selected frame being associated with one or more frames of the first plurality of frames; receiving, via the interface, a selection of the at least one selected frame; and Based on the selection, an image is generated based on the at least one selected frame and the at least one additional frame.

16. The device according to claim 15, in, The at least one additional frame is obtained from an additional image capture system different from the image capture system.

17. The device according to claim 15, in, The interface includes one or more of a slider for advancing through frames or a gallery corresponding to the first plurality of frames.

18. The device according to claim 15, in, Displaying the interface for selecting the at least one selected frame includes displaying one or more frames of the first plurality of frames.

19. The device according to claim 15, further comprising: include: determining that an amount of motion in the at least one frame exceeds a motion threshold; as well as Based on determining that the amount of motion in the at least one frame exceeds the motion threshold, the at least one suggested frame is displayed.

20. The device according to claim 15, in, The at least one proposal frame is generated using a deep learning neural network.

21. The device according to claim 15, in, The at least one selected frame comprises a frame selected from the first plurality of frames.

22. The device according to claim 15, in, The at least one additional frame is a key frame.

23. A method for processing one or more frames, include: obtaining a first frame from an image capture system, wherein the first frame is captured prior to obtaining a capture input, the first frame being associated with a first setup domain; obtaining a second frame in response to the capture input, wherein the second frame is associated with a second settings domain, the second settings domain being different from the first settings domain; and A third frame is generated based on the first frame and the second frame through a domain transformation model, the third frame being associated with the second setting domain.

24. The method according to claim 23, in, The domain transformation model includes a deep learning neural network.

25. The method according to claim 23, in, The domain transformation model includes a generator of a generative adversarial network.

26. The method according to claim 23, in, Generating the third frame includes transforming the first frame from the first setting domain to the second setting domain.

27. The method according to claim 26, in: The first setting domain includes a first resolution; The second setting field includes a second resolution; as well as Transforming the first frame from the first setting domain to the second setting domain includes upscaling the first frame from the first resolution to the second resolution.

28. The method according to claim 26, in: The first setting field includes a first frame rate; The second setting field includes a second frame rate; as well as Transforming a first frame from the first setting domain to the second setting domain includes frame rate converting the first frame from the first frame rate to the second frame rate.

29. The method according to claim 26, in: The first setting domain includes a first resolution and a first frame rate; The second setting field includes a second resolution and a second frame rate; as well as Converting the first frame from the first setting domain to the second setting domain includes upscaling the first frame from the first resolution to the second resolution and frame rate converting the first frame from the first frame rate to the second frame rate.

30. The method according to claim 23, in, The second frame is obtained from an additional image capture system different from the image capture system.

31. The method according to claim 23, in, The first frame is stored in a frame buffer.

32. The method according to claim 23, in, The second frame is a key frame.

33. An apparatus for processing one or more frames, include: Memory; as well as one or more processors coupled to the memory and configured to: obtaining a first frame from an image capture system, wherein the first frame is captured prior to obtaining a capture input, the first frame being associated with a first setup domain; obtaining a second frame in response to the capture input, wherein the second frame is associated with a second settings domain, the second settings domain being different from the first settings domain; and A third frame is generated based on the first frame and the second frame through a domain transformation model, the third frame being associated with the second setting domain.

34. The device according to claim 33, in, The domain transformation model includes a deep learning neural network.

35. The device according to claim 33, in, The domain transformation model includes a generator of a generative adversarial network.

36. The device according to claim 33, in, To generate the third frame, the one or more processors are configured to transform the first frame from the first setting domain to the second setting domain.

37. The device according to claim 36, in: The first setting domain includes a first resolution; The second setting field includes a second resolution; as well as To transform the first frame from the first setting domain to the second setting domain, the one or more processors are configured to upscale the first frame from the first resolution to the second resolution.

38. The device according to claim 36, in: The first setting field includes a first frame rate; The second setting field includes a second frame rate; and To transform a first frame from the first setting domain to the second setting domain, the one or more processors are configured to frame rate convert the first frame from the first frame rate to the second frame rate.

39. The device according to claim 36, in: The first setting domain includes a first resolution and a first frame rate; The second setting field includes a second resolution and a second frame rate; as well as To transform the first frame from the first setting domain to the second setting domain, the one or more processors are configured to: upscale the first frame from the first resolution to the second resolution and frame rate convert the first frame from the first frame rate to the second frame rate.

40. The device according to claim 33, in, The second frame is obtained from an additional image capture system different from the image capture system.

41. The device according to claim 33, in, The first frame is stored in a frame buffer.

42. The device according to claim 33, in, The second frame is a key frame.