Camera initialization to reduce latency

JP2024535666A5Pending Publication Date: 2025-08-13QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024505601
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-25
Filing Date
2022-08-18
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Existing camera systems experience significant latency due to the need to initialize high-power cameras, which consume substantial power and resources, leading to delayed availability and potential missed capture opportunities.

Method used

Implement predictive camera initialization using a low-power camera to detect events and trigger high-power camera initialization before the event occurs, optimizing power modes and reducing latency through pre-focusing, pre-convergence of exposure and focus values, and buffering raw data for immediate high-power camera processing.

Benefits of technology

Reduces camera latency by ensuring high-power cameras are ready for immediate use when needed, minimizing power consumption and resource utilization while maintaining image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems, methods, and non-transitory media are provided for predictive camera initialization. An example method can include obtaining image data from a first image capture device depicting a scene, classifying the scene based on the image data, predicting a camera use event based on the scene classification, and adjusting a power mode of at least one of the first image capture device and a second image capture device based on the predicted camera use event.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] For example, aspects of the present disclosure relate to predictive camera initialization to reduce latency in camera usage. [Background technology]

[0002] Electronic devices are increasingly equipped with camera hardware to capture images and / or videos for consumption. For example, a computing device may include a camera (e.g., a mobile device such as a mobile phone or smartphone including one or more cameras) to enable the computing device to capture videos or images of a scene, person, object, etc. The images or videos may be captured and processed by the computing device (e.g., a mobile device, an IP camera, an extended reality device, a connected device, a security system, etc.) and stored and / or output (e.g., displayed on the device and / or another device) for consumption. In some cases, the images or videos may be further processed for effects (e.g., compression, image enhancement, image restoration, scaling, frame rate conversion, noise reduction, etc.) and / or for specific applications such as computer vision, extended reality (e.g., augmented reality, virtual reality, etc.), object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, and automation, among others.

[0003] In some cases, the electronic device may process the image to detect objects, faces, events, and / or any other items captured by the image. Object detection may be useful for various applications, such as, for example, authentication, automation, gesture recognition, surveillance, extended reality, computer vision, among others. In some examples, the electronic device may implement a low-power or "always-on" (AON) camera that operates persistently or periodically to automatically detect certain objects in an environment. The low-power camera may be implemented for various use cases, such as, for example, persistent gesture detection, persistent object (e.g., face / person, animal, vehicle, device, plane, event, etc.) detection, persistent object scanning (e.g., quick response (QR) code scanning, barcode scanning, etc.), persistent face recognition for authentication, etc. In many cases, the imaging, processing, and / or performance capabilities / results of the low-power camera may be limited. Thus, in some cases, the electronic device may also implement a higher power camera with higher imaging, processing, and / or performance capabilities / results, which the electronic device may use at particular times and / or in particular scenarios where the higher imaging, processing, and / or performance capabilities / results are desired. Summary of the Invention

[0004] Systems and techniques for predictive camera initialization to reduce latency are described herein. According to at least one example, a method of predictive camera initialization to reduce latency is provided. The method can include obtaining image data describing a scene from a first image capture device, classifying the scene based on the image data, predicting a camera usage event based on the scene classification, and adjusting a power mode of at least one of the first image capture device and a second image capture device based on the predicted camera usage event.

[0005] According to at least one example, a non-transitory computer-readable medium for predictive camera initialization to reduce latency is provided. The non-transitory computer-readable medium can include instructions that, when executed by one or more processors, cause the one or more processors to obtain image data depicting a scene from a first image capture device, classify the scene based on the image data, predict a camera usage event based on the classification of the scene, and adjust a power mode of at least one of the first image capture device and a second image capture device based on the predicted camera usage event.

[0006] According to at least one example, an apparatus for predictive camera initialization to reduce latency is provided. The apparatus may include a memory and one or more processors configured to: obtain image data depicting a scene from a first image capture device, classify the scene based on the image data, predict a camera usage event based on the classification of the scene, and adjust a power mode of at least one of the first image capture device and a second image capture device based on the predicted camera usage event.

[0007] According to at least one example, another apparatus is provided for predictive camera initialization to reduce latency, the apparatus may be meant to: obtain image data describing a scene from a first image capture device, classify the scene based on the image data, predict a camera usage event based on the scene classification, and adjust a power mode of at least one of the first image capture device and a second image capture device based on the predicted camera usage event.

[0008] In some aspects, the methods, non-transitory computer-readable media, and apparatus described above may pre-converge exposure and / or focus values ​​based on one or more images from a first image capture device.

[0009] In some aspects, the above-described methods, non-transitory computer-readable media, and apparatus can obtain at least one of location data indicating a location of the electronic device associated with a first image capture device and sensor data from one or more sensors associated with the electronic device, the sensor data including at least one of motion measurements indicative of movement associated with the electronic device, audio data captured by the one or more sensors, and position measurements indicative of a position of the electronic device, and classify a scene based on the image data and at least one of the location data, the audio data, and the sensor data.

[0010] In some examples, the predicted camera usage event may include a user input configured to trigger an electronic device associated with the first image capture device to capture additional image data.

[0011] In some examples, adjusting the power mode of at least one of the first image capture device and the second image capture device can include initializing the second image capture device, hi some examples, the second image capture device is initialized to a power mode that consumes more power than a respective power mode of the first image capture device used to capture the image data.

[0012] In some aspects, the methods, non-transitory computer-readable media, and apparatus described above can capture additional image data using a second image capture device in a power mode that consumes more power than the respective power mode of the first image capture device. In some instances, the power modes include at least one of a higher resolution than a resolution associated with the respective power mode of the first image capture device used to capture the image data, a higher frame rate than a frame rate associated with the respective power mode of the first image capture device used to capture the image data, a number of image sensors that is higher than a number of image sensors associated with the respective power mode of the first image capture device used to capture the image data, and the first image sensor supporting a particular power mode that consumes more power than a different power mode supported by a second image sensor associated with the first image capture device.

[0013] In some cases, the second image capture device is associated with the first camera pipeline that consumes more power than the second camera pipeline associated with the first image capture device. In some examples, the first camera pipeline includes at least one of more image processing capabilities than the second camera pipeline and one or more hardware components having higher processing performance than the second camera pipeline.

[0014] In some examples, initializing the second image capture device includes increasing a power mode of the second image capture device.

[0015] In some examples, initializing the second image capture device includes increasing a power mode of the second image capture device. In some cases, increasing the power mode of the second image capture device includes at least one of increasing power of at least one of the second image capture device and one or more hardware components associated with a camera pipeline of the second image capture device.

[0016] In some aspects, the methods, non-transitory computer-readable media, and apparatus described above may include storing image data from a first image capture device in a buffer, where at least a portion of the image data is stored in the buffer at least one of during initialization of a second image capture device and before initialization of the second image capture device is completed, and processing at least a portion of the image data stored in the buffer via a camera pipeline associated with the second image capture device, where at least a portion of the image data is processed after the second image capture device is initialized.

[0017] In some examples, adjusting the power mode of at least one of the first image capture device and the second image capture device includes increasing a frequency of at least one of a processor associated with a camera pipeline and a memory associated with the camera pipeline of the second image capture device.

[0018] In some cases, adjusting the power mode of at least one of the first image capture device and the second image capture device includes pre-allocating memory for a camera application associated with the second image capture device.

[0019] In some examples, classifying the scene includes detecting an event associated with the image data. In some cases, the event includes at least one of a scene depicted in the image data, a particular movement of an electronic device associated with the first image capture device, a position of the electronic device relative to a user associated with the electronic device, a crowd detected in the image data, a gesture associated with one or more users, a pattern displayed on an object, and a position of a group of people relative to one another.

[0020] In some examples, adjusting the power mode of at least one of the first image capture device and the second image capture device includes determining one or more initialization settings associated with initialization of the second image capture device based on a type of event associated with the scene, and initializing the second image capture device according to the one or more initialization settings.

[0021] In some aspects, the methods, non-transitory computer-readable media, and apparatus described above may be capable of initializing a timer associated with an expiration value in response to classifying the scene, determining that the value of the timer has reached the expiration value prior to the occurrence of the predicted camera use event, and reducing a power mode of the second image capture device based on the value of the timer having reached the expiration value prior to the occurrence of the predicted camera use event. In some examples, reducing the power mode includes at least one of turning off the second image capture device and reducing one or more power settings associated with at least one of the second image capture device and a camera pipeline associated with the second image capture device.

[0022] In some examples, adjusting the power mode of the second image capture device includes turning on or implementing at least one of a flood illuminator, a depth sensor device, a dual image capture device system, a structured light system, a time-of-flight system, an audio algorithm, location services, and a camera pipeline that is different from the camera pipeline associated with the first image capture device.

[0023] In some examples, adjusting a power mode of at least one of the first image capture device and the second image capture device includes decreasing a power mode of the first image capture device. In some cases, decreasing a power mode of the first image capture device includes at least one of decreasing power of at least one of the first image capture device and one or more hardware components associated with a camera pipeline of the first image capture device.

[0024] In some examples, adjusting a power mode of at least one of the first image capture device and the second image capture device includes increasing a power mode of the first image capture device. In some cases, increasing the power mode of the first image capture device includes at least one of increasing power of at least one of the first image capture device and one or more hardware components associated with a camera pipeline of the first image capture device.

[0025] In some aspects, each of the above-mentioned devices may be, be part of, or include a mobile, device, smart or connected device, camera system, and / or extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device). In some examples, the device may include or be part of a vehicle, a mobile device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device, a personal computer, a laptop computer, a tablet computer, a server computer, a robotics device or system, an aviation system, or other device. In some aspects, the device includes an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, the device includes one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device includes one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, the above-mentioned devices may include one or more sensors. In some cases, one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level, and / or other state), and / or for other purposes.

[0026] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used independently to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0027] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0028] Illustrative examples of the present application are described in detail below with reference to the following figures: [Brief description of the drawings]

[0029] [Figure 1] FIG. 1 illustrates an example of an electronic device that can implement predictive camera initialization, according to some examples of the present disclosure. [Figure 2A] FIG. 2 illustrates an example system process for predictive camera initialization, according to some examples of the present disclosure. [Figure 2B] FIG. 2 illustrates an example system process for predictive camera initialization, according to some examples of the present disclosure. [Diagram 3] 1 is a flowchart illustrating an example process for predictive camera initialization, according to some examples of this disclosure. [Figure 4] 1A-1C illustrate example camera initialization states adjusted at different times based on a particular stimulus, in accordance with some examples of the present disclosure. [Diagram 5] FIG. 13 illustrates an example of different initialization states implemented based on different predictive camera event determinations, according to some examples of the present disclosure. [Figure 6] 1 is a flowchart illustrating an example process for predictive camera initialization, according to some examples of this disclosure. [Figure 7] FIG. 2 illustrates an example of a computing device architecture, according to some examples of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0030] Specific aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments can be practiced without these specific details. The figures and descriptions are not intended to be limiting.

[0031] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0032] Electronic devices (e.g., mobile phones, wearable devices (e.g., smart watches, smart bracelets, smart glasses, etc.), tablet computers, extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, etc.), connected devices, laptop computers, etc.) can implement cameras to detect and / or recognize events of interest. For example, electronic devices can implement low-power cameras to detect and / or recognize events of interest, such as on-demand, continuously, or periodically. Examples of events of interest can include gestures (e.g., hand gestures, smiles, etc.), actions (e.g., by a device, a person, and / or an animal), the presence or occurrence of one or more objects, etc. Objects associated with events of interest can include and / or point to, for example, without limitation, a face, a code (e.g., a quick response (QR) code, a barcode, etc.), a document, a scene or environment, a link, a machine-readable code, a crowd, etc. Low-power cameras can implement low-power hardware and / or energy-efficient image processing software / pipelines used to capture image data, detect events of interest, etc. Low-power cameras can remain on, or "wake up," to monitor movement and / or objects in a scene and detect events in the scene while using less battery power than other devices, such as high-power cameras.

[0033] For example, a low-power camera can monitor movement and / or activity in a scene to discover objects. To illustrate, an XR device can implement a low-power camera that periodically discovers an XR controller and / or other tracked objects, a mobile phone can implement a low-power camera that periodically checks for objects (e.g., codes, documents, faces, etc.) or events, a smart home assistant can implement a low-power camera that periodically checks for user presence, etc. Upon discovering an event (e.g., an object, gesture, activity, scene, etc.), the low-power camera can trigger one or more actions, such as, for example, object detection, object recognition, authentication (e.g., face recognition, etc.), one or more image processing tasks, among other actions. In some cases, the low-power camera can "wake up" other devices and / or components, such as other cameras, sensors, processing hardware, etc.

[0034] In some examples, low-power cameras (sometimes referred to as "always-on" (AON) cameras) can operate persistently or periodically to automatically detect certain objects / events in the environment. Additionally, low-power cameras can be configured to consume less power and computational resources than high-power or "main" cameras. For example, a low-power camera pipeline can enable persistent or periodic imaging with limited / reduced power consumption compared to a high-power or "main" camera pipeline using lower / reduced resolution, low-power image sensors, low-power memory resources (e.g., on-chip static random-access memory (SRAM) as opposed to dynamic random-access memory (DRAM)), island voltage rails to reduce leakage, ring oscillators for clock sources (e.g., as opposed to phase-locked loops), low-power physical interfaces, low-power image processing operations, etc. In some cases, to further reduce power consumption and / or resource utilization, the low power camera pipeline may not implement certain operations (e.g., noise reduction, image warping, image enhancement, etc.), may not process certain types of data (e.g., color image data as opposed to mono / luminance data), may not employ certain hardware (e.g., downscalers, color converters, lens aberration correction hardware, digital signal processors, neural processors, neural network accelerators, high power physical interfaces such as mobile industry processor interface (MIPI) camera serial interface (CSI), certain computer vision blocks, etc.).

[0035] In general, the imaging, processing, and / or performance capabilities and results of a low-power camera may be lower than those of a high-power camera. For example, a low-power camera may generate lower quality images / videos than a high-power camera and / or may provide more limited features and / or effects than a high-power camera. Thus, in some cases, in addition to implementing a low-power camera, an electronic device may also implement a high-power camera that supports higher imaging, processing, and / or performance capabilities / results than a low-power camera. In some examples, an electronic device may use such a high-power camera at a particular time and / or in a particular scenario where higher imaging, processing, and / or performance capabilities / results are desired.

[0036] In many cases, it may be desirable for a high-power camera and / or a camera application associated with the high-power camera to be available as soon as possible when a user attempts to obtain an image, video, or preview. However, maintaining a high-power camera in an initialized state (e.g., ready, powered on, etc.) consumes a significant amount of power and computational resources. To reduce power consumption and resource utilization, high-power cameras are generally maintained in an off or uninitialized state. This may delay the availability and initial operation of the high-power camera and / or camera application when a user attempts to use the high-power camera and / or camera application. For example, initialization of a high-power camera may take a certain amount of time due to allocation of memory for the camera application, powering on / up the image sensor of the high-power camera, powering on / up various hardware blocks associated with the high-power camera and / or camera pipeline of the high-power camera, convergence of exposure and / or focus values, etc. These processes result in camera initialization / startup latency and user annoyance at not being able to immediately capture an image / video of a subject. In some cases, these processes can even result in missed opportunities when attempting to capture fleeting events.

[0037] Described herein are systems, apparatus, methods (also referred to as processes), and computer-readable media (collectively referred to herein as "systems and techniques") for predictive camera initialization to reduce latency. In some examples, the systems and techniques described herein may implement a prediction engine that predictively initializes a high-power camera prior to an estimated / anticipated high-power camera usage event, such as a predicted user attempt / input to use a high-power camera and / or camera application. In some examples, the prediction engine may trigger initialization of the high-power camera based on a particular event / stimuli that is at least partially observed by a low-power camera and that is estimated to indicate a future / imminent camera usage event, such as a scene, condition, and / or activity that is known and / or estimated to often occur prior to a user attempt (e.g., user input) to capture an image / video.

[0038] For example, in some cases, the low-power camera may observe a group of people gathering and / or smiling. A classifier associated with the predictor may use image data captured by the low-power camera to determine the likelihood that a user may activate the high-power camera to capture an image / video of the group. As another example, the low-power camera may output images to the classifier showing commonly captured scenes (e.g., sunsets, beaches, animals, landmarks, etc.) or events (e.g., concerts, user activities, gestures (e.g., camera snaps, electronic device pose relative to the horizon, one or more user poses relative to the electronic device, hand gestures, etc.), games, weddings, etc.). The predictor may then trigger initialization of the high-power camera based on the predicted output from the classifier. In some cases, the predictor's decision to trigger initialization of the high-power camera may be further based on other data, such as, for example, other sensor data (e.g., inertial sensor data, audio data, etc.), global navigation satellite system (GNSS) or global positioning system (GPS) data.

[0039] In some cases, the predictor can also use data from the low-power camera to determine the amount of time available / utilized to initialize the high-power camera. For example, if the predictor determines that a group of people is likely preparing to take a group image, the predictor can determine that more time may be available to initialize the high-power camera before the group image (e.g., compared to another image event such as a sports image or a selfie) because the group may take longer to assemble (e.g., compared to another image event). The predictor may determine that the additional time may allow for longer exposure and / or focus convergence to optimize the image. Based on the additional time, the predictor can trigger a camera pipeline associated with the low-power camera or the high-power camera to converge or pre-converge auto-exposure and / or focus values.

[0040] In some examples, the predictor can determine which initialization steps and / or camera processes it triggers based on data from the low-power camera. For example, the predictor can trigger high-dynamic range (HDR) mode exposure control when data from the low-power camera is capturing a sunset and trigger non-HDR mode when data from the low-power camera is not capturing a sunset. As another example, the predictor can allocate a larger buffer when data from the low-power camera is capturing a user's smile than when data from the low-power camera is capturing a quick response (QR) code. In some cases, in addition to triggering initialization of the high-power camera, the predictor can trigger the low-power camera pipeline to store raw data from the low-power camera in a buffer for processing by the high-power camera once the high-power camera is initialized. In some examples, the raw data can include data captured by the low-power camera while the high-power camera is initialized. In some examples, the high-power camera can use the raw data in the buffer to achieve zero shutter lag (ZSL), in which case the buffer can act as a ZSL buffer.

[0041] In some cases, the predictor may use certain criteria to determine when to deinitialize and / or limit the initialization of a high-power camera to reduce the time the high-power camera remains initialized and the amount of power and resource consumption by the high-power camera. For example, the predictor may deinitialize a high-power camera after a preset period of time if no images / videos have been captured by the high-power camera within that period. As another example, the predictor may set a maximum number of initializations allowed within a preset period of time and skip initializing a new high-power camera if the maximum number of initializations is reached within the preset period of time.

[0042] The systems and techniques described herein can be implemented for various electronic devices to intelligently power up / down different camera sensors with minimal or reduced latency channel sensor feeds to a fewer number of processing devices such as image signal processors. For example, the systems and techniques described herein can be implemented for mobile computing devices (e.g., smartphones, tablets, laptops, cameras, etc.), smart wearable devices (e.g., head mounted displays, extended reality (e.g., virtual reality, augmented reality, etc.) glasses, etc.), connected devices or Internet-of-Things (IoT) devices (e.g., smart TVs, smart security cameras, smart home appliances, etc.), autonomous robotic devices, autonomous driving systems, and / or any other devices with camera hardware.

[0043] An electronic device (and / or a low-power camera thereon) can monitor (and implement the systems and techniques described herein for) various types of events. Non-limiting examples of detected events can include face detection, scene detection (e.g., sunset, document scan, landmarks, etc.), people crowd detection, animal / pet detection, pattern detection (e.g., QR codes, barcodes, etc.), text detection, object detection, gesture detection (e.g., smile detection, emotion detection, hand gestures, etc.), posture detection, etc.

[0044] Various aspects of the application are described with respect to the figures.

[0045] 1 is a diagram illustrating an example of an electronic device 100 in which predictive camera initialization (e.g., camera pre-initialization) and other techniques described herein can be implemented. In some examples, the electronic device 100 can include an electronic device configured to provide one or more functions, such as, for example, imaging functions, extended reality (XR) functions (e.g., locating / tracking, detection, classification, mapping, content rendering, etc.), video functions, image processing functions, device management and / or control functions, gaming functions, autonomous driving or navigation functions, computer vision functions, robotics functions, automation, computer vision, electronic communication functions (e.g., audio / video calling, electronic messaging, etc.), web browsing functions, etc.

[0046] For example, in some cases, electronic device 100 may be an XR device (e.g., a head-mounted display, a head-up display device, smart glasses, etc.) configured to provide XR functionality and perform predictive camera initialization. In some cases, electronic device 100 may implement one or more applications, such as, but not limited to, an XR application, a camera application, an application for managing and / or controlling components and / or operations of electronic device 100, a smart home application, a video game application, a device control application, an autonomous driving application, a navigation application, a productivity application, a social media application, a communication application, a modeling application, a media application, an e-commerce application, a browser application, a design application, a map application, and / or any other application. As another example, electronic device 100 may be a smartphone configured to perform predictive camera initialization as described herein.

[0047] In the illustrative example shown in FIG. 1 , electronic device 100 may include one or more image sensors, such as image sensor 102 and image sensor 104, audio sensor 106 (e.g., ultrasonic sensor, microphone, etc.), inertial measurement unit (IMU) 108, and one or more computational components 110. In some cases, electronic device 100 may optionally include one or more other / additional sensors, such as, for example, but not limited to, radar, light detection and ranging (LIDAR) sensors, touch sensors, pressure sensors (e.g., air pressure sensors and / or any other pressure sensors), gyroscopes, accelerometers, magnetometers, and / or any other sensors. In some examples, electronic device 100 may include additional components, such as, for example, light-emitting diode (LED) devices, storage devices, caches, GNSS / GPS receivers, communication interfaces, displays, memory devices, etc. An exemplary architecture and exemplary hardware components that may be implemented by electronic device 100 are further described below with respect to FIG. 7 .

[0048] Electronic device 100 may be part of or implemented by a single computing device or multiple computing devices. In some examples, electronic device 100 may be part of electronic device(s), such as a camera system (e.g., digital camera, IP camera, video camera, security camera, etc.), a telephone system (e.g., smartphone, mobile phone, conferencing system, etc.), a laptop or notebook computer, a tablet computer, a set-top box, a smart television, a display device, a game console, an XR device such as an HMD, a drone, an in-vehicle computer, an IoT (Internet of Things) device, a smart wearable device, or any other suitable electronic device(s).

[0049] In some implementations, the image sensor 102, the image sensor 104, the audio sensor 106, the IMU 108, and / or one or more computational components 110 may be part of the same computing device. For example, in some cases, the image sensor 102, the image sensor 104, the audio sensor 106, the IMU 108, and / or one or more computational components 110 may be integrated with or into a camera system, a smartphone, a laptop, a tablet computer, a smart wearable device, an XR device such as an HMD, an IoT device, a gaming system, and / or any other computing device. In other implementations, the image sensor 102, the image sensor 104, the audio sensor 106, the IMU 108, and / or one or more computational components 110 may be part of or implemented by two or more separate computing devices.

[0050] The one or more computational components 110 of the electronic device 100 may include, for example, without limitation, a central processing unit (CPU) 112, a graphics processing unit (GPU) 114, a digital signal processor (DSP) 116, and / or an image signal processor (ISP) 118. In some examples, the electronic device 100 may include other processors, such as, for example, a computer vision (CV) processor, a neural network processor (NNP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or the like. The electronic device 100 can use one or more computing components 110 to perform various computing operations, such as, for example, extended reality operations (e.g., tracking, localization, object detection, classification, pose estimation, mapping, content anchoring, content rendering, etc.), device control operations, image / video processing, graphics rendering, machine learning, data processing, modeling, computation, computer vision, messaging, and / or any other operation.

[0051] In some cases, one or more of the computing components 110 may include other electronic circuitry or hardware, computer software, firmware, or any combination thereof for performing any of the various operations described herein. In some examples, one or more of the computing components 110 may include more or fewer computing components than those illustrated in Figure 1. Additionally, the CPU 112, GPU 114, DSP 116, and ISP 118 are merely illustrative examples of computing components provided for purposes of explanation.

[0052] Image sensor 102 and / or image sensor 104 may include any image and / or video sensor or capture device, such as a digital camera sensor, a video camera sensor, a smartphone camera sensor, an image / video capture device on an electronic device such as a television or computer, a camera, etc. In some cases, image sensor 102 and image sensor 104 may be part of a camera or computing device, such as a digital camera, a video camera, an IP camera, a smartphone, a smart television, a gaming system, etc. Additionally, in some cases, image sensor 102 and image sensor 104 may include multiple image sensors, such as rear and front sensor devices, and may be part of a dual camera or other multi-camera assembly (e.g., including two cameras, three cameras, four cameras, or other number of cameras).

[0053] In some examples, image sensor 102 can include or be part of a low-power or "always-on" camera, and image sensor 104 can include or be part of a high-power or "main" camera. In some examples, the low-power camera can implement low-power hardware and / or more energy-efficient image processing software (than the high-power camera) to detect events, process captured image data, etc. In some cases, the low-power camera can implement lower power settings and / or modes than the high-power camera (e.g., image sensor 104), such as, for example, a lower frame rate, lower resolution, fewer image sensors, a low-power mode, a low-power camera pipeline (including software and / or hardware), etc. In some examples, low-power cameras may implement fewer and / or lower power image sensors than high-power cameras, may use low-power memory such as on-chip static random access memory (SRAM) rather than dynamic random access memory (DRAM), may use island voltage rails to reduce leakage, may use ring oscillators rather than phased-locked loops (PLLs) as clock sources, and / or other low-power processing hardware / components. In some examples, low-power cameras may not handle high-power and / or complexity sensor technologies (e.g., phase-detection autofocus, dual photodiode (2PD) pixels, red-green-blue-clear (RGBC) color sensing, etc.) and / or data (e.g., mono / luminance data rather than full-color image data).

[0054] In some cases, a low-power camera can remain on or "wake up" to monitor movement and / or events in a scene and / or detect events in a scene while using less battery power than other devices, such as higher power / resolution cameras. For example, a low-power camera can continually observe or wake up to observe movement and / or activity in a scene to discover objects in the scene. In some cases, upon discovering an event, the low-power camera can trigger one or more actions, such as, for example, object detection, object recognition, face recognition, image processing tasks, among other actions. In some cases, the low-power camera can also "wake up" other devices, such as other sensors, processing hardware, etc.

[0055] In some examples, each image sensor (e.g., image sensor 102, image sensor 104) can capture image data, generate frames based on the image data, and / or provide image data or frames to one or more computational components 110 for processing. A frame can include a video frame of a video sequence or a still image. A frame can include a pixel array representing a scene. For example, a frame can be a Red-Green-Blue (RGB) frame with red, green, and blue color components per pixel, a Luminance, Red Difference, Blue Difference (YCbCr) frame with one Luminance component and two Chrominance (Color) components (Red Difference and Blue Difference) per pixel, or any other suitable type of color or monochrome image.

[0056] In some examples, one or more of the computational components 110 can use data from the image sensor 102, the image sensor 104, the audio sensor 106, the IMU 108, and / or any other sensors and / or components to perform image / video processing, predictive camera initialization, XR processing, device management / control, and / or other operations described herein. For example, in some cases, the one or more computational components 110 can perform predictive camera initialization, device control / management, tracking, localization, object detection, object classification, pose estimation, shape estimation, scene mapping, content anchoring, content rendering, image processing, modeling, content generation, gesture detection, gesture recognition, and / or other operations based on data from the image sensor 102, the image sensor 104, the audio sensor 106, the IMU 108, and / or any other components.

[0057] In some examples, one or more of the computational components 110 may implement one or more software engines and / or algorithms, such as, for example, a predictor 120 described herein. In some cases, one or more of the computational components 110 may implement one or more additional components and / or algorithms, such as a machine learning model(s), a computer vision algorithm(s), a neural network(s), and / or any other algorithms and / or components. For example, in some cases, the predictor 120 implemented by one or more of the computational components 110 may implement a machine learning engine 122.

[0058] In some examples, the predictor 120 can predict a camera event (e.g., a user input to capture an image / video, a user input to use a camera application, a user input to trigger a camera preview, etc.) based on sensor data from the image sensor 102 and, optionally, the audio sensor 106, the IMU 108, and / or any other components. The predictor 120 can then trigger a predictive initialization (e.g., pre-initialization) of the high-power camera (e.g., the image sensor 104) and / or the high-power camera pipeline before the predicted camera event occurs. In some examples, the predictive camera initialization can reduce the latency / delay between the time of a user input to capture an image / video and / or access a camera application / preview and the capture of the image / video and / or the access of the camera application / preview.

[0059] In some cases, the predictor 120 can trigger some predictive initialization steps in response to predicting a camera event. For example, the predictor 120 can burst the CPU (e.g., CPU 112) and / or memory to a higher frequency to reduce camera initialization and / or camera pipeline latency. As another example, the predictor 120 can pre-allocate memory for a camera application, trigger low-resolution data capture by the image sensor 102 to pre-converge exposure and / or focus values, etc. In some cases, predictor 120 may power on specific hardware and / or software blocks for image sensor 104 and / or camera processing, such as, for example, but not limited to, a high power physical interface (e.g., MIPI CSI, etc.), a PLL for generating clocks, on-chip infrastructure, a downscaler, a color converter, a DSP (e.g., DSP 116), an ISP (e.g., ISP 118), a neural network accelerator, a neural processor, one or more blocks for computer vision (e.g., feature extraction, feature description, object detection, recognition, etc.), image warping, noise reduction, PDAF, 2PD pixels, RGBC color sensing, HDR, etc. In some cases, predictor 120 may adjust one or more settings as part of the predictive initialization, such as, for example, but not limited to, increasing / decreasing resolution, increasing / decreasing frame rate, increasing / decreasing buffer size pre-allocation, activating more / less image sensors, increasing / decreasing power modes, etc.

[0060] In some cases, predictor 120 may extract / detect visual features from one or more frames captured by a camera device, such as a low-power camera device (e.g., image sensor 102). In some examples, predictor 120 may implement detectors and / or algorithms for extracting visual features, such as, but not limited to, scale-invariant feature transform (SIFT), speeded up robust features (SURF), Oriented FAST and rotated BRIEF (ORB), and / or any other detectors / algorithms.

[0061] In some cases, the predictor 120 can implement a machine learning engine 122 to generate machine learning classifications used to predict camera events. The machine learning engine 122 can generate the classifications based on data from the image sensor 102, the audio sensor 106, the IMU 108, and / or one or more other components, such as a GNSS / GPS receiver, a radar, a wireless interface, etc. The predictor 120 can predict camera events using output from the machine learning engine 122. For example, the predictor 120 can use inertial sensor data (e.g., from the IMU 108) that indicates that the electronic device 100 has been lifted by a user (which can suggest that the user may use the electronic device 100 to capture an image / video) based on the output of a machine learning classifier (e.g., the machine learning engine 122) trained on positive and negative examples of the inertial sensing data. In some examples, the predictor 120 can use inertial sensor data (e.g., from the IMU 108) that indicates that the electronic device 100 is being held in movements / positions commonly used when taking photos / videos, based on the output of a machine learning classifier (e.g., the machine learning engine 122) trained on positive and negative examples of inertial sensing data.

[0062] In some examples, the predictor 120 can use audio sensor data (e.g., from the audio sensor 106) that indicates the electronic device 100 is in a crowded area (e.g., a concert, sporting event, party, play, wedding, etc.) or indicates speech associated with taking an image / video based on the output of a machine learning classifier (e.g., the machine learning engine 122) trained on positive and negative examples of the audio sensor data. In other examples, the predictor 120 can use GNSS / location data that indicates the electronic device 100 is in a commonly photographed area (e.g., a landmark, stage, etc.) based on a set of location-tagged photos.

[0063] In some cases, the predictor 120 can use camera data from a low-power camera (e.g., image sensor 102) based on the output of a machine learning classifier (e.g., machine learning engine 122) trained on positive and negative examples of camera data (e.g., images). In some examples, the camera data can indicate that a group of people is gathered together (e.g., as if preparing for a group photo), that users or a group of people are smiling or posing (e.g., as if posing for a photo), the presence of a commonly photographed scene (e.g., a sunset, a beach, an animal, a landmark, a landscape, etc.), the presence of a commonly scanned pattern (e.g., a QR code), the presence of a gesture (e.g., waving, pointing, etc.), that a commonly photographed activity (e.g., a sporting event, a show, a play, etc.) is taking place, etc. In some cases, the predictor 120 can use any combination of camera data, audio sensor data, GNSS / location data, inertial sensor data, etc.

[0064] In some cases, the predictor 120 can track the usage of the camera devices (e.g., image sensor 102, image sensor 104) and use the camera usage statistics when predicting camera events. For example, if the camera usage statistics indicate that the user of the electronic device 100 often connects the image sensor 104 of the electronic device 100 to a telescope to capture astrophotography images, the predictor 120 can use such camera usage statistics to predict the usage of the image sensor 104 when the image data from the image sensor 102 indicates that the electronic device 100 is near or connected / connected to a telescope. As another example, if the audio sensor data indicates that the electronic device 100 is often in a noisy area and the camera usage statistics indicate that the user often captures images / videos in the presence of certain types of noise, the predictor 120 can use the camera usage statistics to predict the usage of the image sensor 104 when the audio sensor data indicates that the electronic device 100 is around those types of noise and not when the audio sensor data indicates that the electronic device 100 is around other types of noise.

[0065] In some examples, the predictor 120 can monitor various types of events to predict camera events. Non-limiting examples of events can include face detection, scene detection (e.g., sunset, room, landmark, landscape, concert, etc.), crowd detection, animal / pet detection, pattern (e.g., QR code, etc.) detection, document detection, text detection, object detection, gesture detection (e.g., smile detection, emotion detection, hand gestures, etc.), posture detection, noise detection, location detection, motion detection, activity detection, among others.

[0066] The components shown in FIG. 1 for electronic device 100 are illustrative examples provided for purposes of explanation. In other examples, electronic device 100 may include more or less components than those shown in FIG. 1. Although electronic device 100 is shown as including certain components, one of ordinary skill in the art will understand that electronic device 100 may include more or less components than those shown in FIG. 1. For example, electronic device 100 may, in some examples, include one or more memory devices (e.g., RAM, ROM, cache, etc.), one or more networking interfaces (e.g., wired and / or wireless communication interfaces, etc.), one or more display devices, cache, storage devices, and / or other hardware or processing devices not shown in FIG. 1. Illustrative examples of computing devices and hardware components that may be implemented by electronic device 100 are described below with respect to FIG. 7.

[0067] 2A illustrates an example system process 200 for predictive camera initialization. In this example, image sensor 102 represents a low power camera. Image sensor 102 captures image data of a scene and provides sensor data 220 including the captured image data to low power camera pipeline 202 for processing. Low power camera pipeline 202 may include image processing operations and / or hardware used to process sensor data 220 including image data captured by image sensor 102. In this example, low power camera pipeline 202 includes image sensor 102.

[0068] In some examples, the low power camera pipeline 202 includes pre-processing (e.g., image resizing, denoising, segmentation, edge smoothing, color correction / conversion, debayering, scaling, gamma correction, etc.). In some cases, the low power camera pipeline 202 can include one or more image post-processing operations. In some examples, the low power camera pipeline 202 can activate / include low power hardware, configurations, and / or processes, such as, for example, lower / reduced resolution, lower / reduced frame rate, lower power sensors, on-chip SRAM (e.g., rather than DRAM), island voltage rails, ring oscillators for clock sources (e.g., rather than PLLs), fewer / reduced number of image sensors, etc.

[0069] The low power camera pipeline 202 can output processed sensor data (e.g., processed image data from sensor data 220) to the predictor 120. The predictor 120 can use the sensor data from the low power camera pipeline 202 to predict camera events based on events captured and detected in the sensor data (e.g., image data in sensor data 220 from image sensor 102). In some examples, the predictor 120 can also, optionally, use sensor data 222 from non-camera sensors 204 to predict camera events. The non-camera sensors 204 can include or represent one or more sensors / devices, such as, for example, an audio sensor (e.g., audio sensor 106), an IMU (e.g., IMU 108), a GNSS / GPS receiver, and / or any other sensor (e.g., radar, LIDAR, pressure sensor, etc.). The sensor data 222 may include inertial sensor measurements (e.g., acceleration, velocity, angular velocity, orientation, heading, etc.), acoustic data (e.g., noise / sound, speech, utterance, voice, tone, pitch, sound level, noise / sound patterns, ultrasound, infrasound, etc.), location measurements, altitude measurements, distance measurements, any other sensor data, and / or any combination thereof.

[0070] Events captured in the sensor data (e.g., sensor data 220 and / or sensor data 222) may include, for example, without limitation, objects detected in an image, faces detected in an image, scenes detected in an image (e.g., a sunset, a room, a landmark, a landscape, a concert, etc.), groups of people detected in an image, animals detected in an image, patterns detected in an image (e.g., a QR code, etc.), documents detected in an image, text detected in an image, gestures detected in an image (e.g., smile detection, emotion detection, hand gestures, etc.), postures detected in an image, acoustic data captured in the audio sensor data (e.g., noise / sound, speech or sound, sound / noise patterns, voice, tone, pitch, specific ultrasound, specific infrasound, sound levels, etc.), measured locations, motion detection, activity detection, among other events (e.g., detected / measured objects, activities, movements, etc.).

[0071] The camera events being predicted by the predictor 120 may include one or more events associated with a user interaction with and / or use of a high-power camera (e.g., image sensor 104) and / or a camera application. For example, the predicted camera events may include a user input to open a camera application through which the user can get a preview and / or capture an image / video, a user input to trigger capture, an image / video, a hotkey or shortcut pressed by a user on the electronic device 100 to trigger the capture of an image / video, etc. For example, predicting a camera event may include predicting that a user of the electronic device 100 is about to use the electronic device 100 to take a picture / video (e.g., via a camera application that triggers the image sensor 104 or via selection / activation of a hotkey or shortcut that triggers the image sensor 104 to capture an image / video) or is about to open a camera application to view a preview feed and / or trigger the image sensor 104.

[0072] The predictor 120 can detect events captured in the sensor data (e.g., sensor data 220, sensor data 222) and determine whether the detected event is commonly observed before (immediately before, within a threshold period of time prior, within some events / actions prior, etc.) the camera event(s) (e.g., the detected event led to or preceded the camera event with at least a threshold frequency, the detected event is estimated to have a threshold likelihood of leading to or preceding the camera event, and / or the detected event has a threshold number of occurrences that previously led to or preceded the camera event) and / or whether the detected event indicates that it is likely (e.g., within a threshold likelihood) that the camera event will follow the detected event (e.g., immediately after the detected event, within a period of time after the detected event, within some events / actions after the detected event, etc.). In some examples, the predictor 120 can predict camera events based on detected events and event prior probabilities, event statistics (such as detected event statistics, camera event statistics, statistical correlation / causality of detected events and camera events), user input(s) providing one or more preferences and / or configurations of detected events selected to trigger a camera event prediction, a predictive model, camera usage analysis, and / or any combination thereof.

[0073] In some cases, predictor 120 can implement a machine learning classifier (e.g., machine learning engine 122) to predict camera events based on sensor data (e.g., sensor data 220, sensor data 222). In some examples, the machine learning classifier can be trained on positive and negative examples of sensor data, such as, for example, image sensor data, inertial sensor data, audio sensor data, GNSS / GPS / location data, other sensor data, and / or any combination thereof.

[0074] If the predictor 120 determines that the detected event indicates that the detected event is commonly observed before the camera event(s) (e.g., the detected event led to or preceded the camera event at least a threshold frequency, the detected event is estimated to have a threshold likelihood of leading to or preceding the camera event, and / or the detected event has a threshold number of occurrences that previously led to or preceded the camera event) and / or that the camera event is likely to follow the detected event (e.g., within a threshold likelihood), the predictor 120 may generate a positive camera event prediction result (e.g., the predictor 120 may determine that the camera event is predicted to occur). Based on the positive camera event prediction result (e.g., based on a determination that the camera event is predicted to occur), the predictor 120 may trigger a predictive initialization 224 of the high power camera pipeline 212.

[0075] The high-power camera pipeline 212 may include one or more operations and / or hardware used to capture images / video and / or process the captured images / video. In some cases, the high-power camera pipeline 212 may be the same as or include the low-power camera pipeline 202 with one or more adjusted settings to produce higher image quality, produce additional and / or more complex image effects, and / or achieve higher processing / output performance. For example, in some cases, the high-power camera pipeline 212 may include the low-power camera pipeline 202 with one or more settings, such as increasing image resolution, increasing frame rate, utilizing full color image data (e.g., as opposed to only mono / luminance data), etc. In other cases, the high-power camera pipeline 212 may include one or more image sensors, settings, operations, and / or hardware blocks that are different from the low-power camera pipeline 202.

[0076] In some examples, the low power camera pipeline 202 can include a low power image sensor (e.g., image sensor 102) and the high power camera pipeline 212 can include a high power image sensor (e.g., image sensor 104). In other examples, the low power camera pipeline 202 and the high power camera pipeline 212 can include the same image sensor(s) (e.g., image sensor 102 and / or image sensor 104). In some cases, the low power camera pipeline 202 and the high power camera pipeline 212 can include the same image sensor(s), but the high power camera pipeline 212 can implement the image sensor(s) with or in a high power mode (e.g., higher resolution, higher frame rate, etc.).

[0077] In some examples, the high power camera pipeline 212 includes one or more image pre-processing, one or more post-processing operations, and / or any other image processing operations. For example, the high power camera pipeline 212 can include image resizing, noise removal, segmentation, edge smoothing, color correction / conversion, debayering, scaling, gamma correction, tone mapping, color sensing, sharpening, compression, demosaicing, noise reduction (e.g., chroma noise reduction, luminance noise reduction, temporal noise reduction, etc.), feature extraction, feature recognition, computer vision, auto exposure, auto white balance, auto focus, depth sensing, image stabilization, sensor fusion, HDR, and / or any other operations. In some examples, the high power camera pipeline 212 can activate / include high power hardware, configurations, and / or processes, such as, for example, higher / increased resolution, higher / increased frame rate, high power image sensor (e.g., image sensor 104), DRAM usage / allocation, PLL for clock source, more / increased number of image sensors, etc.

[0078] In some examples, predictive initialization 224 may include powering up / on and / or initiating / preparing one or more hardware components / devices and / or one or more operations associated with the high power camera pipeline 212. In some cases, predictive initialization 224 may include powering up / on one or more hardware blocks associated with the image sensor 104 and camera processing. In some cases, predictive initialization 224 may include bursting a processor (e.g., CPU 112) and / or memory to a higher frequency to enable reduced latency. In some cases, predictive initialization 224 may include pre-allocating memory (e.g., DRAM, etc.) for the camera application. In some cases, predictive initialization 224 may include performing a low resolution image capture (e.g., via image sensor 102) to pre-converge exposure and / or focus values. In some cases, predictive initialization 224 may include any combination of the aforementioned operations, settings, and / or hardware blocks. In some cases, predictive initialization 224 can start / prepare the high power camera pipeline 212 to capture images / video with zero or reduced latency in response to user input.

[0079] In some cases, the low power camera pipeline 202 can store raw sensor data from the image sensor 102 in a buffer for offline (e.g., later) processing by the high power camera pipeline 212. For example, FIG. 2B shows an example system process 230 for predictive camera initialization involving offline processing of raw sensor data from the image sensor 102 by the high power camera pipeline 212.

[0080] In this example, in addition to leveraging the data, processing, and hardware used in the system process 200 shown in FIG. 2A, the system process 230 can implement a multi-frame buffer 232 for storing sensor data 220 from the low power camera pipeline 202. The low power camera pipeline 202 can receive sensor data 220 from the image sensor 102 and store the sensor data 220 in the multi-frame buffer 232. In some examples, the sensor data 220 can include raw sensor data from the image sensor 102. In some cases, the multi-frame buffer 232 can function as a ZSL buffer for the high power camera pipeline 212.

[0081] When the predictor 120 triggers the high power camera pipeline 212 via predictive initialization 224, the high power camera pipeline 212 can retrieve sensor data 220 from the multi-frame buffer 232 for processing. The sensor data 220 can provide the high power camera pipeline 212 with image data captured before and / or while the high power camera pipeline 212 is initialized, which the high power camera pipeline 212 can leverage to reduce image / video capture latency and / or improve the processing, quality, and / or performance of the captured images / videos.

[0082] In some examples, the low power camera pipeline 202 may not be able to process images at as high a quality as images processed by the high power camera pipeline 212. Nevertheless, the low power camera pipeline 202 may store raw sensor data in the multi-frame buffer 232 (or a different memory, such as DRAM) for (later) offline processing by the high power camera pipeline 212. Making this raw sensor data available to the high power camera pipeline 212 may provide various advantages, as the low power camera pipeline 202 may be operational while the high power camera pipeline 212 is being initialized.

[0083] Once the high power camera pipeline 212 is fully initialized, the high power camera pipeline 212 can then read back and process the raw sensor data offline in the multi-frame buffer 232. In some examples, the system process 230 can reduce or even completely eliminate the start-up latency of the high power camera pipeline 212 in the final capture of an image / video performed by the high power camera pipeline 212 (e.g., including the image sensor 104) in response to user input.

[0084] 3 is a flow chart illustrating an example process 300 for predictive camera initialization. In this example, at block 304, the process 300 may include performing camera usage prediction (e.g., via predictor 120) using camera data 302. The camera data 302 may include image data from a low power image sensor, such as image sensor 102. In some examples, the camera data 302 may be the same as the sensor data 220 described above with respect to FIG. 2A.

[0085] Camera usage prediction can include and / or represent camera event prediction, as described above. For example, camera usage prediction can include predicting whether a user of electronic device 100 is going to use electronic device 100 to take a photo / video (e.g., via a camera application that triggers the high-power camera pipeline or via selecting / activating a hotkey or shortcut that triggers the high-power camera pipeline to capture an image / video) or to open a camera application, view a preview feed, and / or trigger the high-power camera pipeline.

[0086] At block 306, if the camera usage prediction results in a predicted camera event (e.g., the predictor 120 predicts that a user will trigger an image / video capture (e.g., via user input) and / or use a camera application to view a preview feed and / or attempt to trigger an image / video capture), the process 300 may include initializing a high-power camera pipeline (e.g., the high-power camera pipeline 212) and a configurable timer (e.g., via the predictor 120). The configurable timer may provide a time limit for the user to trigger an image / video capture and / or start / use a camera application before the initialized camera pipeline is deinitialized to save power (e.g., from continuing to run the high-power camera pipeline when it is not needed or being used).

[0087] At block 308, the process 300 may include determining whether a configurable timer has expired. At block 314, if the configurable timer has expired, the process 300 may include deinitializing (e.g., powering down, disabling, stopping, turning off) the high power camera pipeline (e.g., via the predictor 120).

[0088] If the configurable timer has not expired, the high-power camera pipeline may remain initialized. At block 310, the process 300 may include determining (e.g., via the predictor 120) whether a user-initiated camera event has occurred. The user-initiated camera event may include a user input to trigger image / video capture by the high-power camera pipeline and / or to access a preview feed using a camera application and / or to trigger image / video capture by the high-power camera pipeline. For example, in some cases, the user-initiated camera event may include a user selection of a hotkey or shortcut to trigger image / video capture. In other cases, the user-initiated camera event may include a voice command from the user requesting the electronic device 100 to trigger the high-power camera pipeline to capture an image / video. In other cases, the user-initiated camera event may include a user input received via a camera application to trigger the camera application (and the high-power camera pipeline) to start / perform image / video capture.

[0089] If process 300 determines that a user-initiated camera event has not occurred (e.g., if a user-initiated camera event is not detected), process 300 may return to block 308 to determine whether the configurable timer has expired. Process 300 may either keep the high-power camera pipeline initialized if the configurable timer has not expired or deinitialize the high-power camera pipeline if the configurable timer has expired (e.g., as described with respect to block 314).

[0090] If the process 300 determines that a user-initiated camera event has occurred (e.g., if a user-initiated camera event is detected), then in block 312, the process 300 may include performing a high-speed camera start. The high-speed camera start may include triggering an already initialized high-power camera pipeline to capture an image / video in response to the user-initiated camera event. Because the high-power camera pipeline has already been initialized, the high-speed camera start may be performed without latency (or with reduced / minimal latency) from the time of the user-initiated camera event. For example, because the high-power camera pipeline has already been initialized, in response to the user-initiated camera event, the high-power camera pipeline may capture an image / video without any latency (or with reduced / minimal latency).

[0091] 4 is a diagram illustrating example camera initialization states that are modified / adjusted at different times based on certain stimuli, such as, for example, predictive initialization, timer expiration, and / or state change triggers. In this example, a high-power camera pipeline (e.g., high-power camera pipeline 212) is in an uninitialized state 402 at time t1.

[0092] At time t2, the predictor 120 triggers predictive initialization 410, as described above with respect to Figures 2A, 2B, and 3. Based on the predictive initialization 410, the high-power camera pipeline changes from the uninitialized state 402 to the initialized state 404.

[0093] At time t3, a predictor (e.g., predictor 120) detects timer expiration 412 and triggers camera deinitialization. Based on the camera deinitialization, the high power camera pipeline changes from the initialized state 404 to the deinitialized state 406. In some examples, the deinitialized state 406 may be the same as the uninitialized state 402. In some examples, the deinitialized state 406 may be a state in which the power mode of the high power camera pipeline is reduced to a low power mode.

[0094] At time t4, the predictor triggers another predictive initialization 414. Based on the predictive initialization 414, the high-power camera pipeline transitions from the deinitialized state 406 back to the initialized state 404.

[0095] At time t5, a user initiated camera event 416 may trigger a high speed camera start 408, such as the high speed camera start described with respect to Figure 3. At the high speed camera start 408, the high power camera pipeline may capture images / video with no latency (or reduced / minimal latency) in response to the user initiated camera event 416.

[0096] At time t6, the predictor detects a state change trigger 418 configured to change the state of the high power camera pipeline from an initialized state (initialized state 404) to a deinitialized state 406. In some examples, the state change trigger 418 can include the expiration of a timer, such as timer expiration 412. In some examples, the state change trigger 418 can be based on a user input from a user of the electronic device 100. For example, the state change trigger 418 can be based on closing a camera application, powering down / off the high power camera pipeline, or user input (e.g., via a user interface, via a verbal command, via one or more buttons or keys, etc.) to trigger deinitialization of the high power camera pipeline.

[0097] 5 is a diagram illustrating an example of different initialization states implemented based on different predictive camera event determinations. In this example, the high-power camera pipeline of electronic device 100 is in an uninitialized state 502 during the absence of a camera event prediction 500.

[0098] The electronic device 100 then detects an object 530 in a scene of the electronic device 100 (e.g., via a low-power camera sensor such as image sensor 102). Based on the detection of the object 530, a predictor (e.g., predictor 120) can determine a camera event prediction 520. The camera event prediction 520 can predict a user-initiated camera event, as described above.

[0099] In some examples, the object 530 can be configured to trigger a camera event prediction. In some examples, the predictor can learn from previous examples that there is a threshold likelihood of a user-initiated event occurring after detection of the object 530. For example, the predictor can determine the camera event prediction 520 based on an output from a machine learning classifier (e.g., the machine learning engine 122) that indicates that a user-initiated camera event is predicted to occur after detection of the object 530.

[0100] Based on the camera event prediction 520, the predictor can trigger (eg, via an instruction / command) the high-power camera pipeline to change from the uninitialized state 502 to the initialized state 522.

[0101] FIG. 6 is a flow chart illustrating an example process 600 for predictive camera initialization. In this example, at block 602, the process 600 may include acquiring image data (e.g., sensor data 220) depicting a scene (e.g., a sunset, a landmark, a landscape, an environment, an event, etc.) from a first image capture device (e.g., image sensor 102). In some examples, the first image capture device may automatically capture the image data without user input to trigger the image data to be captured. In some cases, the first image capture device may capture the image data using a low power mode and / or a low power camera pipeline (e.g., low power camera pipeline 202) (e.g., relative to a high power mode supported by the first image capture device and / or a second image capture device on the electronic device associated with the first image capture device).

[0102] At block 604, process 600 may include classifying a scene based on the image data. In some examples, classifying the scene may include detecting an event associated with (e.g., depicted in) the image data. In some cases, the detected event may include a scene depicted in the image data. In some cases, the scene and / or the detected event may include a particular movement of the electronic device, a position of the electronic device relative to a user associated with the electronic device, a crowd detected in the image data, a gesture associated with one or more users, a pattern displayed on an object (e.g., a QR code, a link, etc.), and / or a position of a group of people relative to one another.

[0103] In some cases, the scene may be classified using machine learning or image processing algorithms.

[0104] At block 606, process 600 may include predicting a camera use event based on the classification of the scene. In some examples, the camera use event may include a user input configured to trigger an electronic device (e.g., electronic device 100) associated with the first image capture device to capture additional image data.

[0105] In some cases, camera use events can be detected using a machine learning classifier trained on positive and negative examples of image data, such as positive and negative examples of image data that capture the same type of scene as the classified scene.

[0106] At block 608, process 600 may include adjusting a power mode of the first image capture device and / or the second image capture device (e.g., image sensor 104) based on the predicted camera usage event. In some examples, the second image capture device includes a higher power camera device relative to the first image capture device. For example, the second image capture device may support a power mode that consumes more power than a power mode supported by the first image capture device. In some cases, the higher power camera device may include a first image sensor that supports a higher resolution than the first image capture device, a higher frame rate than the first image capture device, a greater number of image sensors than the first image capture device, and / or a higher power mode than a second image sensor associated with the first image capture device.

[0107] In some cases, the second image capture device is associated with (e.g., uses and / or is part of) a high power camera pipeline (e.g., high power camera pipeline 212) versus a low power camera pipeline (e.g., low power camera pipeline 202) associated with the first image capture device. In some examples, the high power camera pipeline includes one or more hardware components that have more image processing capability and / or higher processing performance than the low power camera pipeline. For example, in some cases, the second image capture device is associated with a first camera pipeline that consumes more power than the second camera pipeline associated with the first image capture device. In some cases, the first camera pipeline includes one or more hardware components that have more image processing capability and / or higher processing performance than the second camera pipeline.

[0108] In some examples, adjusting the power mode of at least one of the first image capture device and the second image capture device can include initializing the second image capture device. In some examples, the second image capture device is initialized to a power mode (e.g., a high power mode) that consumes more power than a respective power mode associated with the first image capture device (e.g., than a respective power mode of the first image capture device used to capture image data). In some cases, initializing the second image capture device can include increasing the power mode of the second image capture device.

[0109] In some aspects, process 600 may include capturing the additional image data using a second image capture device in a power mode that consumes more power than a respective power mode of the first image capture device used to capture the image data. In some examples, the power mode of the second image capture device may include a higher resolution than a resolution associated with a respective power mode of the first image capture device used to capture the image data, a higher frame rate than a frame rate associated with a respective power mode of the first image capture device, a number of image sensors that is greater than a number of image sensors associated with a respective power mode of the first image capture device, and / or a first image sensor supporting a particular power mode that consumes more power than a different power mode supported by the second image sensor associated with the first image capture device.

[0110] In some cases, initializing the second image capture device can include increasing a power mode of the second image capture device. In some examples, increasing a power mode of the second image capture device can include increasing power of the second image capture device and / or one or more hardware components associated with a camera pipeline of the second image capture device.

[0111] In some aspects, process 600 may include storing image data from a first image capture device in a buffer (e.g., multi-frame buffer 232) and processing at least a portion of the image data stored in the buffer via a camera pipeline associated with a second image capture device (e.g., high power camera pipeline 212). In some cases, at least a portion of the image data is stored in the buffer during initialization of the second image capture device and / or before initialization of the second image capture device is complete. In some cases, at least a portion of the image data is processed after the second image capture device is initialized.

[0112] In some aspects, adjusting the power mode of at least one of the first image capture device and the second image capture device includes increasing a frequency of at least one of a processor associated with a camera pipeline and / or a memory associated with the camera pipeline of the second image capture device (e.g., bursting the processor and / or memory).

[0113] In some aspects, adjusting the power mode of at least one of the first image capture device and the second image capture device includes pre-allocating memory for a camera application associated with the second image capture device.

[0114] In some aspects, the process 600 includes pre-converging exposure and / or focus values ​​based on one or more images from the first image capture device.

[0115] In some aspects, the process 600 includes obtaining location data indicative of a location of the electronic device associated with the first image capture device and / or sensor data from one or more sensors associated with the electronic device, and classifying the scene based on the image data and at least one of the location data, audio data, and / or sensor data. In some examples, the additional camera use event can include an additional user input configured to trigger the electronic device to capture further image data. In some examples, the sensor data includes motion measurements indicative of motion associated with the electronic device, audio data captured by the one or more sensors, and / or position measurements indicative of a location of the electronic device.

[0116] In some cases, the predicted camera usage event may include a user input configured to trigger the electronic device to capture additional image data.

[0117] In some cases, predicting the additional camera use event can include detecting the additional event based on the location data, the audio data, and / or the sensor data. In some cases, the additional camera use event is predicted further based on the additional event. In some examples, the additional event can include one or more sounds in the audio data, a particular movement of the electronic device, and / or a particular pose of the electronic device.

[0118] In some aspects, adjusting the power mode of at least one of the first image capture device and the second image capture device includes determining one or more initialization settings associated with initializing the second image capture device and initializing the second image capture device according to the one or more initialization settings. In some examples, the one or more initialization settings are determined based on a type of scene associated with the classified scene. For example, the initialization settings may include an HDR mode if the scene includes a particular scene (e.g., a sunset, a landscape, etc.) and a non-HDR mode if the scene instead includes a user's smiling face.

[0119] In some aspects, adjusting the power mode of at least one of the first image capture device and the second image capture device includes turning on or implementing a flood illuminator, a depth sensor device, a dual image capture device system, a structured light system, a time-of-flight system, an audio algorithm, location services, and / or a camera pipeline different from the camera pipeline associated with the first image capture device.

[0120] In some aspects, the process 600 may include initializing a timer associated with an expiration value in response to classifying the scene, determining that the value of the timer reaches the expiration value prior to the occurrence of the predicted camera use event, and reducing a power mode of the second image capture device based on the value of the timer reaching the expiration value prior to the occurrence of the predicted camera use event. In some examples, reducing the power mode may include turning off the second image capture device and / or reducing one or more power settings associated with the second image capture device and / or a camera pipeline associated with the second image capture device. In some cases, the expiration value may vary based on the type of the classified scene. For example, if the scene includes a group of people gathering in a manner that is presumed to indicate that the group is preparing to take a group photo, the expiration value may be increased (relative to expiration values ​​in other types of scenes) to give the group more time to gather and coordinate for the group photo.

[0121] In some aspects, adjusting a power mode of at least one of the first image capture device and the second image capture device can include decreasing a power mode of the first image capture device. In some cases, decreasing a power mode of the first image capture device can include at least one of decreasing power of the first image capture device and at least one of one or more hardware components associated with a camera pipeline of the first image capture device.

[0122] In some aspects, adjusting a power mode of at least one of the first image capture device and the second image capture device can include increasing a power mode of the first image capture device. In some cases, increasing the power mode of the first image capture device can include at least one of increasing power of at least one of the first image capture device and one or more hardware components associated with a camera pipeline of the first image capture device.

[0123] In some examples, the process 600 may be performed by one or more computing devices or apparatuses. In one illustrative example, the process 600 may be performed by the electronic device 100 shown in FIG. 1. In some examples, the process 600 may be performed by one or more computing devices having a computing device architecture 700 shown in FIG. 7. In some cases, such a computing device or apparatus may include a processor, microprocessor, microcomputer, or other components of a device configured to perform the steps of the process 600. In some examples, such a computing device or apparatus may include one or more sensors configured to capture image data and / or other sensor measurements. For example, the computing device may include a smartphone, a head-mounted display, a mobile device, or other suitable device. In some examples, such a computing device or apparatus may include a camera configured to capture one or more images or videos. In some cases, such a computing device may include a display for displaying the images. In some examples, the one or more sensors and / or the camera are separate from the computing device, in which case the computing device receives the sensed data. Such a computing device may further include a network interface configured to communicate data.

[0124] The components of a computing device may be implemented in circuitry. For example, the components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform various operations described herein. The computing device may further include a display (as an example of an output device or in addition to an output device), a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0125] Process 600 is illustrated as a logical flow diagram, whose operations represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement a process.

[0126] Additionally, process 600 may be executed under the control of one or more computer systems comprised of executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a number of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0127] 7 illustrates an exemplary computing device architecture 700 of an exemplary computing device that may implement various techniques described herein. For example, the computing device architecture 700 may implement at least some portions of the electronic device 100 illustrated in FIG. 1. The components of the computing device architecture 700 are shown in electrical communication with each other using a connection 705, such as a bus. The exemplary computing device architecture 700 includes a processing unit (CPU or processor) 710 and computing device connections 705 that couple various computing device components to the processor 710, including a computing device memory 715, such as a read only memory (ROM) 720 and a random access memory (RAM) 725.

[0128] The computing device architecture 700 may include a cache of high speed memory that is directly connected to the processor 710, near the processor 710, or integrated as part of the processor 710. The computing device architecture 700 may copy data from the memory 715 and / or the storage device 730 to the cache 712 for faster access by the processor 710. In this manner, the cache may provide a performance boost that avoids delays to the processor 710 while waiting for data. These and other modules may control or be configured to control the processor 710 to perform various actions. Other computing device memories 715 may also be available. The memory 715 may include multiple different types of memory with different performance characteristics. The processor 710 may include any general purpose processor, as well as hardware or software services stored in the storage device 730 and configured to control the processor 710, as well as special purpose processors with software instructions built into the processor design. The processor 710 may be a self-contained system that includes multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetrical or asymmetrical.

[0129] To enable user interaction with the computing device architecture 700, the input device 745 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. The output device 735 can also be one or more of several output mechanisms known to those skilled in the art, such as a display, projector, television, speaker device, etc. In some cases, a multimodal computing device may enable a user to provide multiple types of input to communicate with the computing device architecture 700. The communication interface 740 can generally govern and manage user input and computing device output. There is no constraint to operate on any particular hardware configuration, and therefore the basic functions herein may be easily replaced with improved hardware or firmware configurations as they are developed.

[0130] The storage device 730 is a non-volatile memory and may be a hard disk or other type of computer readable medium capable of storing data accessible by a computer, such as a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a random access memory (RAM) 725, a read only memory (ROM) 720, and hybrids thereof. The storage device 730 may include software, code, firmware, etc. for controlling the processor 710. Other hardware or software modules are contemplated. The storage device 730 may be connected to the computing device connection 705. In one aspect, a hardware module that performs a particular function may include a software component stored in a computer readable medium that connects with the necessary hardware components, such as the processor 710, the connection 705, the output device 735, etc., to perform the function.

[0131] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media that can store, store, or transport instructions and / or data. Computer-readable media may include non-transitory media on which data is stored and does not include carrier waves and / or transitory electronic signals propagating over wireless or wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memory, or memory devices. Computer-readable media may have code and / or machine-executable instructions stored on the computer-readable medium, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0132] In some embodiments, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when mentioned, non-transitory computer readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0133] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the embodiments may be practiced without these specific details. For clarity of explanation, in some cases, the technology may be presented as including individual functional blocks comprising devices, device components, and steps or routines in a method embodied in software or a combination of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.

[0134] Particular embodiments may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although the flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of steps may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.

[0135] The processes and methods according to the examples described above may be implemented using computer executable instructions stored or otherwise available from a computer readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general purpose computer, a special purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessible over a network. The computer executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, network attached storage devices, etc.

[0136] Devices implementing processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program product) to perform the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, and the like. The functionality described herein may also be embodied in a peripheral device or an add-in card. Such functionality may also be implemented on a circuit board among different chips, or on different processes executing in a single device, as further examples.

[0137] The instructions, media for carrying such instructions, computing resources for executing such instructions, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.

[0138] In the above description, aspects of the present application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while exemplary embodiments of the present application have been described in detail herein, it should be understood that the inventive concepts may be embodied and employed in various other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. The various features and aspects of the present application described above may be used individually or jointly. Moreover, the embodiments may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present specification. Thus, the present specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, the methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in an order different from that described.

[0139] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terminology used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of the present specification.

[0140] When a component is described as being "configured to" perform a particular operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming a programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof.

[0141] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or that is in communication, either directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).

[0142] Claim language or other language in this disclosure reciting "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and can additionally include unrecited items in the set of A and B.

[0143] The various exemplary logic blocks, modules, circuits, and algorithm steps described with respect to the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0144] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium including program code including instructions, which, when executed, perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium, such as a random access memory (RAM), such as a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a FLASH memory, a magnetic or optical data storage medium, etc. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.

[0145] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein.

[0146] Illustrative examples of the present disclosure include the following: Aspect 1. An apparatus for predictive camera initialization, comprising: a memory; and one or more processors coupled to the memory, wherein the one or more processors are configured to: obtain image data depicting a scene from a first image capture device; classify the scene based on the image data; predict a camera usage event based on the scene classification; and adjust a power mode of at least one of the first image capture device and a second image capture device based on the predicted camera usage event.

[0147] Aspect 2. The apparatus of aspect 1, wherein to adjust the power mode of at least one of the first image capture device and the second image capture device, the one or more processors are configured to initialize the second image capture device, and the second image capture device is initialized to a power mode that consumes more power than the respective power mode of the first image capture device used to capture image data.

[0148] Aspect 3. The apparatus of aspect 2, wherein the one or more processors are configured to capture additional image data using a second image capture device in a power mode that consumes more power than a respective power mode of the first image capture device, the power mode being associated with at least one of the first image sensors supporting a higher resolution than a respective power mode of the first image capture device used to capture the image data, a higher frame rate than a respective power mode of the first image capture device used to capture the image data, a number of image sensors that is greater than a number of image sensors associated with the respective power mode of the first image capture device used to capture the image data, and a particular power mode that consumes more power than a different power mode supported by the second image sensor associated with the first image capture device.

[0149] Aspect 4. The apparatus of aspect 2 or 3, wherein the second image capture device is associated with a first camera pipeline that consumes more power than a second camera pipeline associated with the first image capture device, and the first camera pipeline includes at least one of one or more hardware components having more image processing capability than the second camera pipeline and higher processing performance than the second camera pipeline.

[0150] Aspect 5. The apparatus of any one of Aspects 2 to 4, wherein, to initialize the second image capture device, the one or more processors are configured to increase a power mode of the second image capture device.

[0151] Aspect 6. The apparatus of aspect 5, wherein to increase a power mode of the second image capture device, the one or more processors are configured to increase power of at least one of the second image capture device and one or more hardware components associated with a camera pipeline of the second image capture device.

[0152] Aspect 7. The apparatus of any one of aspects 2 to 6, wherein the one or more processors are configured to store image data from a first image capture device in a buffer, where at least a portion of the image data is stored in the buffer at least one of during initialization of a second image capture device and before initialization of the second image capture device is completed, and process at least a portion of the image data stored in the buffer via a camera pipeline associated with the second image capture device, where at least a portion of the image data is processed after the second image capture device is initialized.

[0153] Aspect 8. The apparatus of any one of aspects 1 to 7, wherein to adjust a power mode of at least one of the first image capture device and the second image capture device, the one or more processors are configured to increase a frequency of at least one of a processor associated with a camera pipeline and a memory associated with the camera pipeline of the second image capture device.

[0154] Aspect 9. The apparatus of any one of aspects 1 to 8, wherein to adjust a power mode of at least one of the first image capture device and the second image capture device, the one or more processors are configured to pre-allocate memory to a camera application associated with the second image capture device.

[0155] Aspect 10. An apparatus described in any one of aspects 1 to 9, wherein one or more processors are configured to pre-converge at least one of exposure values ​​and focus values ​​based on one or more images from a first image capture device.

[0156] Aspect 11. A device as described in any one of aspects 1 to 10, wherein one or more processors are configured to obtain at least one of location data indicating a location of the device and sensor data from one or more sensors associated with the device, the sensor data including at least one of motion measurements indicating motion associated with the device, audio data captured by the one or more sensors, and position measurements indicating a position of the device, and classify a scene based on the image data and at least one of the location data, audio data, and sensor data.

[0157] Aspect 12. The device of any one of aspects 1 to 11, wherein the predicted camera usage event includes a user input configured to trigger the device to capture additional image data.

[0158] Aspect 13. The device of any one of aspects 1 to 12, wherein to classify the scene, the one or more processors are configured to detect events associated with the image data, the events including at least one of a scene depicted in the image data, a specific movement of the device, a position of the device relative to a user associated with the device, a crowd detected in the image data, a gesture associated with one or more users, a pattern displayed on an object, and a position of a group of people relative to one another.

[0159] Aspect 14. The apparatus of any one of aspects 1 to 13, wherein to adjust a power mode of at least one of the first image capture device and the second image capture device, the one or more processors are configured to determine one or more initialization settings associated with initialization of the second image capture device based on a type of event associated with the scene, and initialize the second image capture device according to the one or more initialization settings.

[0160] Aspect 15. The apparatus of any one of aspects 1 to 14, wherein one or more processors are configured to initialize a timer associated with an expiration value in response to classifying the scene, determine that the value of the timer has reached the expiration value before the occurrence of the predicted camera use event, and reduce a power mode of a second image capture device based on the value of the timer having reached the expiration value before the occurrence of the predicted camera use event, wherein reducing the power mode includes at least one of turning off the second image capture device and reducing one or more power settings associated with the second image capture device and at least one of the camera pipelines associated with the second image capture device.

[0161] Aspect 16. The apparatus of any one of aspects 1 to 15, wherein to adjust a power mode of the second image capture device, the one or more processors are configured to turn on or implement at least one of a flood illuminator, a depth sensor device, a dual image capture device system, a structured light system, a time-of-flight system, an audio algorithm, location services, and a camera pipeline that is different from the camera pipeline associated with the first image capture device.

[0162] Aspect 17. The apparatus of any one of aspects 1 to 16, wherein to adjust a power mode of at least one of the first image capture device and the second image capture device, the one or more processors are configured to reduce a power mode of the first image capture device, and reducing the power mode of the first image capture device includes at least one of reducing power of the first image capture device and at least one of one or more hardware components associated with a camera pipeline of the first image capture device.

[0163] Aspect 18. The apparatus of any one of aspects 1 to 17, wherein to adjust a power mode of at least one of the first image capture device and the second image capture device, the one or more processors are configured to increase a power mode of the first image capture device, and increasing the power mode of the first image capture device includes at least one of increasing power of at least one of the first image capture device and one or more hardware components associated with a camera pipeline of the first image capture device.

[0164] Example 19. The apparatus of any one of examples 1 to 18, wherein the apparatus comprises a mobile device.

[0165] Aspect 20. The apparatus of aspect 19, wherein the apparatus includes an extended reality device.

[0166] Aspect 21. A method of predictive camera initialization comprising: obtaining image data depicting a scene from a first image capture device; classifying the scene based on the image data; predicting a camera use event based on the scene classification; and adjusting a power mode of at least one of the first image capture device and a second image capture device based on the predicted camera use event.

[0167] Aspect 22. The method of aspect 21, wherein adjusting the power mode of at least one of the first image capture device and the second image capture device includes initializing the second image capture device, wherein the second image capture device is initialized to a power mode that consumes more power than the respective power mode of the first image capture device used to capture image data.

[0168] Aspect 23. The method of aspect 22, further comprising capturing additional image data using the second image capture device in a power mode that consumes more power than a respective power mode of the first image capture device, wherein the power mode is associated with at least one of the first image sensor supporting a higher resolution than a resolution associated with the respective power mode of the first image capture device used to capture the image data, a higher frame rate than a frame rate associated with the respective power mode of the first image capture device used to capture the image data, a number of image sensors that is greater than a number of image sensors associated with the respective power mode of the first image capture device used to capture the image data, and a particular power mode that consumes more power than a different power mode supported by the second image sensor associated with the first image capture device.

[0169] Aspect 24. The method of aspect 22 or 23, wherein a second image capture device is associated with a first camera pipeline that consumes more power than a second camera pipeline associated with the first image capture device, and the first camera pipeline includes at least one of one or more hardware components having more image processing capability than the second camera pipeline and higher processing performance than the second camera pipeline.

[0170] Example 25. The method of any one of examples 22 to 24, wherein initializing the second image capture device includes increasing a power mode of the second image capture device.

[0171] Aspect 26. The method of aspect 25, wherein increasing the power mode of the second image capture device includes increasing power of at least one of the second image capture device and one or more hardware components associated with the second image capture device and a camera pipeline of the second image capture device.

[0172] Aspect 27. The method of any one of aspects 22 to 26, further comprising: storing image data from a first image capture device in a buffer, where at least a portion of the image data is stored in the buffer at least one of during initialization of a second image capture device and before initialization of the second image capture device is completed; and processing at least a portion of the image data stored in the buffer via a camera pipeline associated with the second image capture device, where at least a portion of the image data is processed after the second image capture device is initialized.

[0173] Aspect 28. The method of any one of aspects 21 to 27, wherein adjusting the power mode of at least one of the first image capture device and the second image capture device includes increasing a frequency of at least one of a processor associated with a camera pipeline and a memory associated with the camera pipeline of the second image capture device.

[0174] Aspect 29. The method of any one of aspects 21 to 28, wherein adjusting the power mode of at least one of the first image capture device and the second image capture device includes pre-allocating memory to a camera application associated with the second image capture device.

[0175] Aspect 30. A method according to any one of aspects 21 to 29, further comprising pre-converging at least one of exposure and focus values ​​based on one or more images from a first image capture device.

[0176] Aspect 31. The method of any one of aspects 21 to 30, further comprising: acquiring at least one of location data indicating a location of the electronic device associated with the first image capture device; and sensor data from one or more sensors associated with the electronic device, the sensor data including at least one of motion measurements indicating movement associated with the electronic device, audio data captured by the one or more sensors, and position measurements indicating a position of the electronic device; and classifying the scene based on the image data and at least one of the location data, the audio data, and the sensor data.

[0177] Aspect 32. The method of any one of aspects 21 to 31, wherein the predicted camera usage event includes a user input configured to trigger an electronic device associated with the first image capture device to capture additional image data.

[0178] Aspect 33. The method of any one of aspects 21 to 32, wherein classifying the scene includes detecting an event associated with the image data, the event including at least one of a scene depicted in the image data, a particular movement of an electronic device associated with the first image capture device, a position of the electronic device relative to a user associated with the electronic device, a crowd detected in the image data, a gesture associated with one or more users, a pattern displayed on an object, and a position of a group of people relative to each other.

[0179] Aspect 34. The method of any one of aspects 21 to 33, wherein adjusting the power mode of at least one of the first image capture device and the second image capture device includes determining one or more initialization settings associated with initialization of the second image capture device based on a type of event associated with the scene, and initializing the second image capture device according to the one or more initialization settings.

[0180] Aspect 35. The method of any one of aspects 21 to 34, further comprising: initializing a timer associated with an expiration value in response to classifying the scene; determining that the value of the timer has reached the expiration value before the occurrence of the predicted camera use event; and reducing a power mode of a second image capture device based on the value of the timer having reached the expiration value before the occurrence of the predicted camera use event, wherein reducing the power mode comprises at least one of turning off the second image capture device and reducing one or more power settings associated with the second image capture device and at least one of the camera pipelines associated with the second image capture device.

[0181] Aspect 36. The method of any one of aspects 21 to 35, wherein adjusting the power mode of the second image capture device includes turning on or implementing at least one of a flood illuminator, a depth sensor device, a dual image capture device system, a structured light system, a time-of-flight system, an audio algorithm, location services, and a camera pipeline that is different from the camera pipeline associated with the first image capture device.

[0182] Aspect 37. The method of any one of aspects 21 to 36, wherein adjusting the power mode of at least one of the first image capture device and the second image capture device includes reducing a power mode of the first image capture device, and reducing the power mode of the first image capture device includes at least one of reducing power of at least one of the first image capture device and one or more hardware components associated with a camera pipeline of the first image capture device.

[0183] Aspect 38. The method of any one of aspects 21 to 37, wherein adjusting the power mode of at least one of the first image capture device and the second image capture device includes increasing the power mode of the first image capture device, and increasing the power mode of the first image capture device includes at least one of increasing power of at least one of the first image capture device and one or more hardware components associated with a camera pipeline of the first image capture device.

[0184] Aspect 39. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform a method as recited in any one of aspects 21 to 38.

[0185] Embodiment 40. An apparatus comprising means for performing the method according to any one of embodiments 21 to 38.

[0186] Aspect 41. The apparatus of aspect 40, wherein the apparatus comprises a mobile device.

[0187] Aspect 42. The apparatus of aspect 41, wherein the mobile device includes an extended reality device.

Claims

1. 1. An apparatus for predictive camera initialization, comprising: Memory and one or more processors coupled to the memory, acquiring image data describing a scene from a first image capture device; classifying the scene based on the image data; predicting a camera use event based on the classification of the scene; adjusting a power mode of at least one of the first image capture device and the second image capture device based on the predicted camera usage event, and pre-allocating memory for a camera application associated with the second image capture device; one or more processors configured to perform An apparatus comprising:

2. To adjust the power mode of at least one of the first image capture device and the second image capture device, the one or more processors:

10. The apparatus of claim 1, configured to initialize the second image capture device, wherein the second image capture device is initialized to a power mode that consumes more power than a respective power mode of the first image capture device used to capture the image data.

3. the one or more processors: configured to capture additional image data using the second image capture device in the power mode that consumes more power than the respective power mode of the first image capture device, the power mode comprising: a resolution greater than the resolution associated with the respective power mode of the first image capture device used to capture the image data; a frame rate that is higher than the frame rate associated with the respective power mode of the first image capture device used to capture the image data; a number of image sensors greater than the number of image sensors associated with the respective power modes of the first image capture device used to capture the image data; and a first image sensor that supports a particular power mode that consumes more power than a different power mode supported by a second image sensor associated with the first image capture device; The device of claim 2 , wherein the device is associated with at least one of:

4. the second image capture device is associated with a first camera pipeline that consumes more power than a second camera pipeline associated with the first image capture device, the first camera pipeline including at least one of one or more hardware components having more image processing capability and higher processing performance than the second camera pipeline; to initialize the second image capture device, the one or more processors are configured to increase the power mode of the second image capture device; and / or to increase the power mode of the second image capture device, the one or more processors are configured to increase power of at least one of the second image capture device and one or more hardware components associated with a camera pipeline of the second image capture device; 3. The apparatus of claim 2.

5. the one or more processors: storing the image data from the first image capture device in a buffer, wherein at least a portion of the image data is stored in the buffer at least one of during the initialization of the second image capture device and before the initialization of the second image capture device is completed; processing at least the portion of the image data stored in the buffer through a camera pipeline associated with the second image capture device, wherein the at least the portion of the image data is processed after the second image capture device is initialized; The apparatus of claim 2 configured to:

6. To adjust the power mode of at least one of the first image capture device and the second image capture device, the processor: The apparatus of claim 1 , configured to increase a frequency of at least one of a processor associated with a camera pipeline of the second image capture device and a memory associated with the camera pipeline.

7. the one or more processors: configured to pre-converge at least one of an exposure value and a focus value based on one or more images from the first image capture device; and / or the one or more processors: acquiring at least one of location data indicative of a location of the device and sensor data from one or more sensors associated with the device, the sensor data including at least one of motion measurements indicative of movement associated with the device, audio data captured by the one or more sensors, and position measurements indicative of a position of the device; classifying the scene based on the image data and at least one of the location data, the audio data, and the sensor data; configured to:

10. The apparatus of claim 1.

8. the predicted camera usage event includes a user input configured to trigger the device to capture additional image data; and / or To classify the scene, the one or more processors are configured to detect events associated with the image data, the events including at least one of the scene depicted in the image data, a particular movement of the device, a position of the device relative to a user associated with the device, a crowd detected in the image data, a gesture associated with one or more users, a pattern displayed on an object, and a position of a group of people relative to one another.

10. The apparatus of claim 1.

9. To adjust the power mode of at least one of the first image capture device and the second image capture device, the one or more processors: determining one or more initialization settings associated with initializing the second image capture device, the one or more initialization settings based on a type of event associated with the scene; initializing the second image capture device according to the one or more initialization settings; The apparatus of claim 1 configured to:

10. the one or more processors: In response to classifying the scene, initializing a timer associated with an expiration value; determining that the timer has reached the expiration value prior to the occurrence of the predicted camera use event; reducing the power mode of the second image capture device based on the value of the timer reaching the expiration value prior to the occurrence of the predicted camera usage event, comprising at least one of turning off the second image capture device and reducing one or more power settings associated with at least one of the second image capture device and a camera pipeline associated with the second image capture device; The apparatus of claim 1 configured to:

11. 10. The apparatus of claim 1, wherein to adjust the power mode of the second image capture device, the one or more processors are configured to turn on or implement at least one of a flood illuminator, a depth sensor device, a dual image capture device system, a structured light system, a time-of-flight system, an audio algorithm, location services, and a camera pipeline different from a camera pipeline associated with the first image capture device.

12. To adjust the power mode of at least one of the first image capture device and the second image capture device, the one or more processors: reducing the power mode of the first image capture device, the reducing the power mode including at least one of reducing power to at least one of the first image capture device and one or more hardware components associated with a camera pipeline of the first image capture device; or increasing the power mode of the first image capture device, the increasing power mode including at least one of increasing power of the first image capture device and at least one of one or more hardware components associated with a camera pipeline of the first image capture device. The apparatus of claim 1 configured to:

13. The apparatus of claim 1 , wherein the apparatus comprises a mobile device, and optionally the apparatus comprises an extended reality device.

14. 1. A method for predictive camera initialization, comprising: acquiring image data describing a scene from a first image capture device; classifying the scene based on the image data; predicting a camera use event based on the classification of the scene; adjusting a power mode of at least one of the first image capture device and the second image capture device based on the predicted camera usage event, the power mode including pre-allocating memory for a camera application associated with the second image capture device; A method comprising:

15. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 14.