Reducing image capture latency in head-mounted devices
Patent Information
- Application Number
- JP2026023051
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-12-04
- Filing Date
- 2026-02-16
- Publication Date
- 2026-09-07
Smart Images

Figure 2026142550000001_ABST
Abstract
Description
[Technical Field]
[0001] The present application claims the priority benefit of U.S. Patent Application No. 19 / 408,566 filed on December 4, 2025 and U.S. Provisional Patent Application No. 63 / 763,536 filed on February 26, 2025, the entire contents of which are incorporated herein by reference. The present disclosure relates generally to electronic devices, and more specifically to head-mounted devices having one or more cameras. [Background Art]
[0002] Some electronic devices can be worn on a user's head. This type of electronic device is sometimes referred to as a head-mounted device. A head-mounted device may include a camera for capturing images of the surrounding physical environment. The embodiments described herein are developed against this background. [Summary of the Invention]
[0003] One aspect of the present disclosure provides a head-mounted device, including: one or more image sensors; a first processing circuit configured to execute an operating system for the head-mounted device, wherein the first processing circuit is an application processor configured to execute one or more applications on the operating system and is operable between a sleep state and a wake state; a second processing circuit configured to instruct the one or more image sensors to capture images, wherein the second processing circuit is operable between a sleep state and a wake state, and may include a camera driver for controlling an image signal processing (ISP) circuit configured to receive and process captured images output from the one or more image sensors; and a third processing circuit configured to detect a user input and wake up the first processing circuit and the second processing circuit simultaneously from the sleep state in response to detection of the user input. The third processor may be a processor that remains in a wake state continuously.
[0004] One aspect of the present disclosure provides a method for operating a head-mounted device having a first processor, a second processor, and a third processor. The method may include using the third processor to detect user input; using the third processor to simultaneously wake up the first processor and the second processor, wherein the first processor has a first wake time and the second processor has a second wake time shorter than the first wake time, in response to the detection of user input; and initiating image capture using a camera driver running on the second processor, in response to the second processor waking up from a sleep state to a wake state. The first processor may be operable between a sleep state and a wake state and consume a first amount of power in the wake state; the second processor may consume a second amount of power less than or equal to the first amount in the wake state; and the third processor may consume a third amount of power less than the first amount.
[0005] One aspect of the present disclosure provides an electronic device comprising: a first processor on which an operating system for the electronic device is executed; a second processor on which a camera driver is executed, the camera driver being configured to initiate image capture while the first processor is transitioning from a sleep state to a wake state; and a third processor configured to detect user input, the first processor being configured to transition from a sleep state to a wake state based on the second processor detecting user input. The third processor may be configured to wake up the first and second processors simultaneously in response to detecting user input. The electronic device may further include one or more cameras configured to capture images; an image signal processing circuit configured to receive captured images and output corresponding processed images; and a memory configured to store the processed images. The image signal processing circuit may include a computer vision processing circuit having a plurality of subsystems configured to receive captured images and to operate in a first power domain, and a backend image signal processing pipeline coupled to a computer vision processing circuit and configured to operate in a second power domain different from the first power domain. [Brief explanation of the drawing]
[0006] [Figure 1] This is a diagram of an exemplary system having a transparent display, according to several embodiments.
[0007] [Figure 2] This figure shows exemplary hardware components that may be included in a system of the type shown in Figure 1, according to several embodiments.
[0008] [Figure 3]This figure shows how exemplary systems, according to several embodiments, may include multiple processors for coordinating low-latency image capture.
[0009] [Figure 4] These are flowcharts illustrating the steps required to operate the type of system shown in Figures 1 to 3, according to several embodiments. [Modes for carrying out the invention]
[0010] The physical environment can refer to the physical world that people can perceive and / or interact with without the help of electronic devices. The physical environment may include physical features such as physical surfaces or physical objects. For example, the physical environment corresponds to a physical park, which includes physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.
[0011] In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through electronic devices. For example, an XR environment may include augmented reality (AR) content, mixed reality (MR) content, and virtual reality (VR) content. In an XR system, a subset of a person's bodily movements or their representation is tracked, and accordingly, one or more properties of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics.
[0012] As an example, an XR system can detect a person's head movements and adjust the graphic content and sound field presented to that person accordingly, in the same way that such views and sounds would change in the physical environment. As another example, an XR system can detect the movement of an electronic device presenting an XR environment (e.g., a mobile phone, tablet, laptop) and adjust the graphic content and sound field presented to that person accordingly, in the same way that such views and sounds would change in the physical environment. In some situations (e.g., for accessibility reasons), an XR system can adjust the characteristics of the graphic content within the XR environment in response to a representation of bodily movement (e.g., a voice command).
[0013] The existence of a wide variety of electronic systems enables people to perceive and / or interact with various XR environments. Examples include head-wearable systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-wearable system may have one or more speakers and an integrated opaque display. Alternatively, a head-wearable system may be configured to accept an external opaque display (e.g., a smartphone). A head-wearable system may incorporate one or more imaging sensors for capturing images or videos of the physical environment, and / or one or more microphones for capturing audio of the physical environment.
[0014] The head-worn system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed towards the person's eyes. The display can utilize digital light projection, organic light-emitting diodes (OLEDs), LEDs, micro-light-emitting diodes (uLEDs), liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In some implementations, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface. The display of device 10 is optional and may be omitted if necessary.
[0015] The system 10 in Figure 1 (sometimes referred to as electronic device 10, head-mounted device 10, etc.) may be a head-mounted device (HMD) having one or more displays. The displays in system 10 may include a display 20, sometimes referred to as a near-eye display, mounted within a support structure (housing) 8. The support structure 8 may have the shape of a pair of glasses or goggles (e.g., a support frame), may form a helmet-shaped housing, or may have other configurations that help mount and secure components of the near-eye display 20 near the user's head or eyes. The near-eye display 20 may include one or more display modules, such as display module 20A, and one or more optical systems, such as optical system 20B. The display module 20A may be mounted on a support structure such as the support structure 8. Each display module 20A may emit light 38 (image light) that is redirected towards the user's eyes in the eye box 24 using an associated optical system 20B. The display 20 is optional and can be omitted from device 10.
[0016] The operation of system 10 may be controlled using a control circuit 16. Processing circuits within the control circuit 16 may be used to control the operation of device 10. The processing circuits may be based on one or more microprocessors, microcontrollers, digital signal processors, baseband processors, power management units, audio chips, application-specific integrated circuits, etc. The control circuit 16 may be configured to perform operations within system 10 using hardware (e.g., dedicated hardware or circuitry), firmware, and / or software. Software code and other data for performing operations within system 10 may be stored on a non-temporary computer-readable storage medium (e.g., a tangible computer-readable storage medium) within the control circuit 16. Software code may be referred to as software, data, program instructions, instructions, or code. The non-temporary computer-readable storage medium (sometimes commonly referred to as memory) may include non-volatile memory such as non-volatile random-access memory (NVRAM), one or more hard drives (e.g., magnetic drives or solid-state drives), one or more removable flash drives, or other removable media. Software stored in a non-temporary computer-readable storage medium may be executed on the processing circuit of the control circuit 16. The control circuit 16, which has both a storage circuit and a processing circuit, may also be collectively referred to as the storage and processing circuit.
[0017] System 10 may include input / output circuit components such as input / output devices 12. The input / output devices 12 may be used to enable System 10 to receive data from external devices (e.g., portable devices such as tethered computers, handheld devices or laptop computers, or other electronic devices) and to enable a user to provide user input to the head-mounted device 10. The input / output devices 12 may also be used to collect information about the environment in which System 10 (e.g., the head-mounted device 10) is operating. Output components within the device 12 may enable System 10 to provide output to a user and may also be used for communication with external electrical equipment. The input / output devices 12 may include one or more cameras 14, sometimes referred to as image sensors. The cameras 14 may be used to collect images of physical objects that are optionally digitally merged with virtual objects on the display of System 10. The input / output devices 12 may include sensors and other components 18 (e.g., accelerometers, gyroscopes, depth sensors, light sensors, haptic output devices, speakers, batteries, wireless communication circuits for communication between System 10 and external electronic equipment).
[0018] A camera 14 mounted on the front of system 10 and facing outwards (towards the front of system 10, away from the user) may be referred to herein as an outward-facing (external-facing) camera or a forward-facing (front-facing) camera. Camera 14 can capture visual odometry information, image information processed to locate objects within the user's field of view (e.g., so that virtual content can be properly registered to real-world objects), image content displayed to the user of system 10 in real time, and / or other appropriate image data. For example, an outward-facing camera can enable system 10 to monitor the movement of system 10 relative to the environment surrounding system 10 (e.g., the camera may be used when forming part of a visual odometry system or a visual inertial odometry system). An outward-facing camera may also be used to capture images of the environment displayed to the user of system 10. If necessary, images from multiple outward-facing cameras may be merged with each other for the user, and / or outward-facing camera content may be merged with computer-generated content.
[0019] The display module 20A may be a liquid crystal display, an organic light-emitting diode display, a laser-based display, or another type of display. The optical system 20B can form lenses that enable a viewer (e.g., the viewer's eyes in the eyebox 24) to view images on the display(s) 20. There may be two optical systems 20B associated with each of the user's left and right eyes (e.g., for forming the left lens and the right lens). It is also possible to generate images for both eyes with a single display 20, or to display images using a pair of displays 20. In a configuration with multiple displays (e.g., a display for the left eye and a display for the right eye), the focal length and position of the lenses formed by the system 20B may be selected so that no gaps between the displays are visible to the user (i.e., the images on the left display and the right display overlap or merge seamlessly).
[0020] If necessary, the optical system 20B may include a transparent structure (e.g., an optical coupler) that enables optical combination of image light from a physical object 28 with a virtual (computer-generated) image, such as a virtual image, within the image light 38. The light from a physical object 28 in the physical environment or scene may be referred to and defined herein as world light, scene light, ambient light, external light, or environmental light. In this type of system, a user of system 10 can see both the physical environment around the user and computer-generated content superimposed on the physical environment. The camera 14 may also be used in device 10 (for example, in a configuration where the camera captures an image of a physical object 28, and this content is modified and presented as virtual content in the optical system 20B).
[0021] System 10 may optionally include wireless circuit components and / or other circuit components to support communication with a computer or other external device (e.g., a computer that supplies image content to display 20). During operation, control circuit 16 may supply image content to display 20. The content may be remotely received (e.g., from a computer or other content source coupled to system 10) and / or generated by control circuit 16 (e.g., text, other computer-generated content, etc.). A viewer may view, in eyebox 24, the content supplied to display 20 by control circuit 16.
[0022] FIG. 2 is a diagram showing illustrative hardware components that may be included in a system of the type described in connection with FIG. 1 (e.g., device 10). As shown in FIG. 2, device 10 may include one or more hardware and / or software subsystems including one or more outward-facing image sensing subsystems such as outward-facing camera 50, one or more tracking subsystems such as tracking sensor 54, computer vision processing (CVP) circuit such as CVP circuit 60, a separate image signal processing pipeline such as high-quality (back-end) pipeline 72, and one or more display(s) 20.
[0023] One or more cameras 50 may be used to collect information about the external real-world environment or scene surrounding device 10. Camera 50 may include one or more of forward-facing cameras 14 in FIG. 1. At least some of cameras 50 may be configured to capture one or more images of a scene, which may optionally be presented as a live video pass-through feed to a user using display 20. Camera 50 may include a color image sensor and / or optionally a monochrome (black and white) image sensor.
[0024] Cameras 50 can have different fields of view. Some cameras 50 can have a wide or very wide field of view, while some cameras 50 can have a relatively narrow field of view. Not all cameras 50 need to be used to capture passthrough content. Some cameras 50 may be forward-facing (e.g., pointed towards the scene in front of the user), some may be downward-facing (e.g., pointed towards the user's torso, hands, or other parts of the user), some may be side-facing (e.g., pointed towards the user's left and right sides), and some may be pointed in other directions relative to the front of device 10. All of these cameras 50 configured to collect information about the external physical environment surrounding device 10 may be collectively referred to as “external-facing” cameras.
[0025] The tracking sensor 54 may include an eye-tracking subsystem, sometimes referred to as an eye-tracker, configured to collect eye-line information or eye-point information. The eye-tracker may monitor the user's eyes using one or more "inward-facing" cameras and / or other eye-tracking components (e.g., eye-facing components and / or other light sources that emit a light beam so that a reflection of the beam from the user's eyes can be detected). One or more eye-tracking sensors 54 may face the user's eyes and track the user's gaze. The camera in the eye-tracking subsystem may determine the location of the user's eyes (e.g., the center of the user's pupil), the direction the user's eyes are facing (the direction of the user's gaze), the size of the user's pupil (e.g., the amount of gradual spatial adjustment of light modulation and / or other optical parameters, and / or one or more of these parameters, and / or an area where one or more of these optical parameters are adjusted based on pupil size), the current focus of the user's eye lens (e.g., whether the user is focused on near or far vision, which may be used to assess whether the user is daydreaming or thinking strategically or tactically), and / or other gaze information. Eye-tracking cameras may also be referred to as inward-facing cameras, gaze detection cameras, eye-tracking cameras, or eye-monitoring cameras. If necessary, other types of optical sensors (e.g., infrared and / or visible light-emitting diodes and photodetectors) may also be used to monitor the user's gaze.
[0026] The tracking sensor 54 may also include a face and body tracking subsystem configured to perform face tracking (e.g., for capturing images of the user's jaw, mouth, etc. while the device is worn on the user's head) and body tracking (e.g., by capturing images of the user's torso, arms, hands, legs, etc. while the device is worn on the user's head). Optionally, the face and body tracking subsystem may also track the user's head posture by directly determining any movement, yaw, pitch, roll, etc. of the head-mounted device 10. The yaw, roll, and pitch of the user's head may collectively define the user's "head posture". For example, the tracking sensor 54 may include an inertial measurement unit (IMU). The inertial measurement unit may include one or more gyroscopes, gyrocompasses, accelerometers, magnetometers, other inertial sensors, and other position and motion sensors. These position sensors and motion sensors may assume that the head-mounted device 10 is worn on the user's head. Therefore, references herein to head posture, head movement, user head yaw (e.g., rotation about a vertical axis), user head pitch (e.g., rotation about a left-right axis), user head roll (e.g., rotation about a front-back axis), and the like, may be considered interchangeable with references to device posture, device movement, device yaw, device pitch, device roll, and the like. In certain embodiments, the tracking sensor 54 may also include a six degrees of freedom (DoF) tracking subsystem. Six DoF tracking subsystems or sensors may be used to monitor both rotational movements such as roll, pitch, and yaw, and position / translational movement in a 3D environment.
[0027] The tracking sensor 54 may optionally further include a hand tracking subsystem, sometimes referred to as a hand tracker, configured to monitor the user's hand movements / gestures and acquire hand gesture data. For example, the hand tracker may include a camera and / or other gesture tracking components (e.g., an outward-facing component and / or light source that emits a light beam so that the reflection of the beam from the user's hand can be detected) to monitor the user's hand(s). One or more hand tracking sensors may be directed at the user's hand, track movements associated with the user's hand, determine whether the user is performing a tap or swipe action with their fingertip or hand, determine whether the user is performing a contactless button press or object selection action with their hand, determine whether the user is performing a grasp or grip action with their hand, determine whether the user is pointing or pinching a given object presented on the display 20 with their hand or fingers, determine whether the user is performing a wave or thrusting action with their hand, or generally measure / monitor three-dimensional contactless gestures ("air gestures") associated with the user's hand. Tracking sensors 54 that are operable to acquire gaze, posture, hand gestures, and other information about the user's movements in device 10 may be collectively referred to as "user tracking" sensors.
[0028] The embodiment in Figure 2 is illustrative, in which the outward-facing camera 50 and tracking sensor 54 (e.g., an optical sensor employed to acquire gaze, posture, and / or other user-related data) are shown as separate, independent subsystems. In some embodiments, one or more outward-facing cameras 50 may also be used to acquire posture information, location information, and / or other motion / position information associated with the device 10. To help protect user privacy, any personal user information collected by the sensors may be processed using best practices. These best practices include meeting or exceeding any applicable privacy regulations. Opt-in and opt-out options and / or other options may be provided to allow users to control the use of their personal data.
[0029] The electronic device 10 can be configured to collect contextual information of the surrounding real-world (physical) environment or scene. Collecting contextual information may include, for example, identifying one or more objects of interest in the environment, detecting when a user enters a particular room or environment, detecting when a user is engaged in a particular activity, detecting the current location of device 10, detecting the current user context or usage scenario (e.g., detecting whether a user is currently watching a movie, playing a video game, or talking to another person or avatar), and / or determining other contextual information related to the operation of device 10. Collecting contextual information may involve capturing one or more images using the outward-facing camera 50 and / or acquiring data from tracking sensors 54. Such images captured for contextual purposes do not need to be output by the display 20 for human consumption. Therefore, the processing requirements and complexity for handling such images may be fewer than the conventional image signal processing steps required to process images output by a display for human consumption (viewing).
[0030] According to one embodiment, the image signal processing circuit on device 10 may be segmented into a first part including a computer vision processing (CVP) circuit 60 and a separate second part including a high-quality (HQ) pipeline 72. In other words, the CVP circuit 60 and the HQ pipeline 72 may be collectively referred to and defined herein as the image signal processing (ISP) circuit. Image and / or data outputs from sensors 50 and 54 that only need to be analyzed for contextual purposes may be processed using only the CVP circuit 60 (e.g., without being processed by the high-quality pipeline 72), while image and / or data outputs from sensors 50 and 54 that will be output onto the display 20 for human viewing may be processed by both the CVP circuit 60 and the high-quality pipeline 72. Components within the CVP circuit 60 may be operated in a first power domain, and components within the HQ pipeline 72 may be operated in a second power domain different from the first power domain (e.g., the CVP circuit 60 and the HQ pipeline 72 may be configured to operate in different power domains).
[0031] The components within the CVP circuit 60 can generally operate in a lower power range than the components within the HQ pipeline 72. The high-quality pipeline 72 can be power-gate controlled (e.g., the HQ pipeline 72 can be selectively activated and deactivated to reduce overall power consumption). When processing images to be output onto the display 20 for human consumption, the high-quality pipeline 72 can be selectively activated to perform some or all of the image processing functions provided by the HQ pipeline 72 (e.g., it can be powered on). When processing images solely for contextual purposes (e.g., to support one or more computer vision algorithms running on the device 10) without the need to display such images, the HQ pipeline 72 can be selectively deactivated to conserve power (e.g., it can be powered off or idled). In other words, when the CVP circuit 60 is activated, it consumes a first amount of power, and when the HQ pipeline 72 is activated, it consumes a second amount of power greater than the first amount. Operating the image signal processing circuit on device 10 in this manner may be technically advantageous in minimizing power consumption on device 10. Such reduced power operation may be beneficial for a small, lightweight device 10 that can be powered by a battery for all-day use.
[0032] As shown in Figure 2, the CVP circuit 60 may include one or more hardware and / or software subsystems, such as a sensor interface 62, a front-end (FE) processor 64, a statistical front-end (FE) processor 66, a statistical back-end (BE) processor 68, a central processing unit (CPU) such as a computer vision processing (CVP) CPU 70, and / or other image signal processing components. The sensor interface 62 may be configured to receive images (e.g., raw pixel data) from the camera 50, the tracking sensor 54, and / or other image sensors in the device 10. The front-end processor 64 may be configured to perform faulty / defective pixel correction, image scaling or binning operations, image cropping or resizing, and / or other front-end or image preprocessing operations. The statistical FE processor 66 may be configured to collect pixel statistics such as minimum pixel value, maximum pixel value, average pixel value, color plane information (e.g., red, green, and blue color planes), color and / or luminance histograms, and other front-end image statistics. The statistical BE processor 68 may be configured to convert an image from a raw Bayer region into a color image and generate additional statistical information.
[0033] The color image output from the statistical BE processor 68 may be provided to one or more downstream computer vision processing algorithms or tasks running on device 10 (for example, processor 68 may output an image to one or more client processors). The statistical FE processor 66 and BE processor 68 may be collectively referred to as the CVP statistical pipeline. Although the CVP circuit 60 is shown as including a single instance of interface 62, processor 64, processor 66, and processor 68, the CVP circuit 60 may include multiple sensor interface blocks 62 for interfacing with multiple sensors, multiple front-end processors 64 for performing image preprocessing operations in parallel, multiple processors 66 for performing front-end statistical calculations in parallel, and / or multiple processors 68 for performing back-end statistical calculations in parallel. The computer vision processing CPU 70 may be configured to manage and coordinate the operation of blocks 62, 64, 66, and 68 to process each input image frame.
[0034] The computer vision processing circuit 60 primarily comprises components for performing front-end image signal processing operations. Therefore, the computer vision processing circuit 60 may be referred to and defined herein as the “front-end” image signal processing (ISP) circuit. In contrast, the HQ pipeline 72 primarily comprises components configured to perform back-end image signal processing operations. Therefore, the high-quality pipeline 72 may be referred to and defined herein as the “back-end” image signal processing (ISP) circuit. The high-quality (back-end) pipeline 72 may be a more complex and power-intensive version of the statistical back-end processor 68 of the CVP circuit 60. For example, the HQ pipeline 72 may include components configured to perform faulty / defective pixel correction, noise reduction, white balance, demosaicing, color space conversion, tone mapping (including, for example, global and local tone mapping), color correction, gamma correction, shading correction, image sharpening, high dynamic range (HDR) correction, edge-recognition local image adjustment, image fusion (for example, merging multiple image frames together for noise reduction and high dynamic range), image signal processing operations that are not present at all in the CVP circuit 60, and / or other image signal processing functions for outputting a corresponding image for display.
[0035] The image(s) output by the backend processor 68 of the CVP circuit 60 may be processed according to a first set of image processing requirements that can optionally generate images of lower fidelity (quality) for computer vision consumption, while the image(s) output by the HQ pipeline 72 may be processed according to a second set of image processing requirements different from the first set of image processing requirements that can optionally generate images of relatively high fidelity (quality) to be displayed for human consumption. In some embodiments, the CVP circuit 60 may be configured to output processed images having a first quality and / or using a first amount of power, while the HQ pipeline 72 may be configured to output processed images having a second quality higher than the first quality and / or using a second amount of power higher than the first amount. In some embodiments, the CVP circuit 60 may be configured to output processed images by performing a first set of image processing operations, while the HQ pipeline 72 may be configured to output processed images by performing additional image processing operations different from the first set of image processing operations. The image output by processor 68 may consequently be provided to one or more client processors. The example in Figure 2, where the HQ pipeline 72 can output content for human consumption via one or more displays 20, is illustrative. The displays 20 are optional and can be omitted from device 10. If necessary, the content output from the HQ pipeline 72 can be stored in memory for later processing.
[0036] In other types of mobile electronic devices, such as smartphones with cameras, the user is typically presented with a preview of the image to be captured before opening the camera application and pressing the capture button. In such scenarios, the smartphone can determine with a high probability that the user is about to press the capture button and can prepare itself for image capture by proactively waking up all the necessary hardware and / or software subsystems required for image capture.
[0037] In contrast to capturing images on a smartphone, a user operating device 10 can initiate or trigger image capture without necessarily opening a camera (image capture) application. In other words, device 10 may not know in advance when the user will press the capture button. As mentioned above, device 10, which may be a lightweight head-mounted device with low power consumption for all-day use, may include one or more processors, at least some of which may operate in sleep mode to reduce active power consumption.
[0038] For example, device 10 may include an application processor on which the operating system of device 10 runs. The application processor should operate in sleep mode most of the time to conserve power. Pressing the capture button, which may occur at any time based on the user's willpower, may trigger the application processor to wake up. However, waiting for the application processor to fully wake up before allowing the camera to capture an image may introduce a considerable shutter lag. "Shutter lag" refers to and may be defined herein as the delay between pressing the capture (shutter) button and the moment when the image is actually captured by the camera.
[0039] According to one embodiment, a method is provided for operating a device 10 that reduces shutter lag. The device 10 utilizes an always-on processor to monitor button presses, and in response to detecting a button press, it can wake up the application processor and, in parallel with the wake-up of the application processor, prepare a camera pipeline for image capture even before the application processor is fully awake. For example, the camera pipeline can start capturing one or more images and start processing the captured images before the application processor is ready to process the images. The terms “camera pipeline” or “camera stack” can refer to all subsystems involved in capturing one or more images and may include at least an outward-facing camera 50, a computer vision processing circuit 60, a high-quality pipeline 72, associated memory devices for storing the captured images (one or more), and a camera driver (e.g., a software subsystem configured to coordinate the operation of the image signal processing circuit).
[0040] Figure 3 shows how, according to several embodiments, device 10 may include multiple processors for coordinating low-latency image capture. As shown in Figure 3, device 10 may include one or more cameras 50, a computer vision processing circuit 60, a high-quality pipeline 72, a memory device 76, and one or more processing circuits such as processors 100, 102, and 106. The camera 50 may be an outward-facing image sensor configured to capture one or more images of a scene or physical environment. The captured images can be processed by the computer vision processing (CVP) circuit 60 and then by the high-quality pipeline 72. The computer vision processing circuit 60 and the high-quality pipeline 72 can therefore receive input (raw) images and output corresponding “processed” images. The computer vision processing circuit 60 and the high-quality pipeline 72 configured to produce processed images may be collectively referred to as an image signal processing (ISP) circuit 74. The processed images output from the ISP circuit 74 can be stored in memory 76. The memory device 76 may be part of the storage subsystem within the control circuit 16 (see Figure 1). The memory device 76 may be implemented as volatile memory such as random access memory (e.g., dynamic RAM or DRAM), non-volatile (persistent) memory such as flash memory, magnetic drives, optical drives, or solid-state drives, or as other types of storage devices.
[0041] Processor 106 may represent the application processor of device 10. The application processor 106 may also be referred to as the application processing circuit 106. The application processor 106 may be configured to run an operating system (OS), such as the operating system 108 for device 10. The operating system 108 may be used to manage multiple applications running on device 10, such as enabling the user to switch between different applications (e.g., photo / video organizing and editing applications, media streaming applications, game applications, social media applications, map / navigation applications, health and fitness applications, automated assistant applications, information retrieval applications, taxi dispatch applications, online banking applications, etc.), to manage security features on device 10, such as performing biometric authentication for secure access and data encryption to protect user data, and / or to manage productivity features such as the user's calendar, reminders, notes, and files, to name a few. In general, the operating system 108 may be designed to provide a secure and intuitive platform that prioritizes user experience, privacy, and seamless functionality across a wide variety of services and applications.
[0042] As described above, the application processor 106, which runs the full OS stack, can consume a considerable amount of power if kept constantly active. To help extend the battery life of device 10, the application processor 106 may be configured to sleep when applications running on device 10 are idle. The application processor 106 can therefore toggle between sleep and wake states. When one or more applications managed by the operating system 108 are required or activated by the user, the application processor 106 can be woken up by transitioning from sleep to wake state. The amount of time it takes for the application processor 106 to transition from sleep to wake state may be referred to herein as the application processor “wake time”.
[0043] In contrast to the application processor 106, the processor 100 may always be powered on. For example, the processor 100 may be a specialized low-power processing subsystem designed to continuously handle a specific take without consuming much battery power. The processor 100 remains active even when the application processor 106 is in sleep or idle state. A processor 100 operating at a minimum power level may be configured to handle lightweight tasks such as monitoring sensors (e.g., sensor 18 in Figure 1, image sensor 50 and tracking sensor 54 in Figure 2, and / or other sensors), monitoring voice commands (e.g., "Hey Siri" or other voice commands), managing notifications, and / or maintaining wireless connectivity for a specific application. Such a type of processor 100 may be referred to and defined herein as an “always-on” processor (AOP) or an “always-awake” processor. An always-on processor 100 may also be referred to as an always-on processing circuit 100. Having an always-on processor 100 that is always (continuously or constantly) active as long as the battery is not completely depleted ensures a fast response to user input or certain trigger events that would otherwise require the attention of processor 106, by eliminating the latency of waking up the main processor 106 (for example, to bypass the application processor wake time). Offloading lightweight tasks from the main application processor 106 to the always-on processor 100 is also technically advantageous and beneficial in helping to conserve energy while optimizing the overall performance of the system 10.
[0044] Device 10 may further comprise low-power computing blocks, such as a low-power computing processor 102. The low-power computing processor 102 may also be referred to as a low-power computing circuit 102. The low-power computing processor 102 may include a camera driver, such as a camera driver 104 configured to control the camera pipeline. The camera driver 104 may also be referred to as an image sensor or image signal processing driver, and is a software subsystem configured to coordinate the operation of the ISP circuit 74. The processor 102 can directly access memory 76, which may be referred to herein as an image storage circuit. The application processor 106 can access or retrieve images stored in memory 76 via the low-power computing processor 102. Alternatively or additionally, the application processor 106 may also directly access images stored on memory 76. Unlike processors 102 and 106, the always-on processor 100 cannot access memory 76 (for example, processor 100 should not be able to access stored images). Processors 100, 102, and 106, configured to operate in this manner, can be considered to have different memory access privileges. For example, the always-on processor 100 may have a first memory access privilege, the application processor 106 may have a second memory access privilege that is higher than or equal to the first memory access privilege, and the low-power computing processor 102 may have a third memory access privilege that is higher than or equal to the second memory access privilege.
[0045] The low-power computing processor 102 may operate in a wake state or a sleep state. A wake-state processor 102 may consume less power than a wake-state application processor 106. A wake-state processor 102 may consume more power than an always-on processor 100. Generally, the application processor 106 may consume a first amount of power, the low-power computing processor 102 may consume a second amount of power less than the first, and the always-on processor may consume a third amount of power less than the second. In some embodiments, when operating in a wake state, the low-power computing processor 102 may consume a similar amount of power as the always-on processor 100. The processor 102 can toggle between a sleep (idle) state and a wake state, while the processor 100 is always active (e.g., the processor 100 is always awake but consumes a small amount of power).
[0046] Figure 4 is a flowchart of exemplary steps for operating the type of device 10 described in relation to Figures 1 to 3. During the operation of block 200, device 10 may detect or predict user input for capturing an image. For example, an always-on processor 100 may be configured to monitor user input 110. User input 110 may include, to name a few, button presses (e.g., the user pressing or pressing a physical button on the housing or frame of device 10), touches (e.g., a physical tap or pressure on a part of the housing or frame of device 10), voice commands (e.g., asking Siri or another automated assistant to take a picture), hand gestures (e.g., a gesture from the user's finger or hand(s), or other movement to trigger image capture), and / or remote triggers (e.g., using a remote controller that communicates wirelessly with device 10). In the example above, such user input(s) 110 may be detected via a physical button, virtual button, touch sensor, microphone, motion sensor, or other type of sensor, which can then alert the always-on processor 100 to the user's intention to capture an image (as indicated by arrow 112 in Figure 3).
[0047] During the operation of block 202, the always-on processor 100 may be configured to wake up the low-power computing processor 102 (as indicated by arrow 114-1 in Figure 3) and simultaneously wake up the application processor 106 (as indicated by arrow 114-2 in Figure 3). In other words, the always-on processor 100 can wake up processors 102 and 106 from sleep states in parallel. Before block 202, both the low-power processor 102 and the application processor 106 may be in a sleep (idle) state. After receiving a wake signal from processor 100, the low-power computing processor 102 may wake up and begin transitioning to wake (active) mode. Similarly, after receiving a wake signal from processor 100, the application processor 106 may wake up and begin transitioning to wake (active) mode.
[0048] As used herein, the term “simultaneous” means that events overlap at least partially in time. In other words, if at least a portion of the first event occurs simultaneously with at least a portion of the second event (for example, if at least a portion of the first event occurs while, during, or when at least a portion of the second event occurs), the first and second events are referred to as “simultaneous” with respect to each other. The first and second events may be simultaneous if they are simultaneous (for example, if the entire duration of the first event overlaps in time with the entire duration of the second event), but they may also be simultaneous if they are not simultaneous (for example, if the first event begins before or after the start of the second event, if the first event ends before or after the end of the second event, or if the first and second events do not overlap in time). As used herein, the term “while” is synonymous with “simultaneous.”
[0049] The application processor 106 may have a first wake time, and the low-power computing processor 102 may have a second wake time shorter than the first wake time. In other words, the low-power computing processor 102 can wake up faster than the application processor 106. During the operation of block 204, the processor 102 can transition completely to the wake state (before the application processor 106 transitions to the wake state), and then instruct the image signal processing circuit 74 to start streaming images from one or more cameras 50. To achieve this, the camera driver 104 running on the processor 102 can simultaneously activate the CVP circuit 60 (as indicated by arrow 116-1 in Figure 3) and the HQ pipeline 72 (as indicated by arrow 116-2 in Figure 3). The image signal processing circuit 74 can then send a signal to the camera 50 instructing it to capture one or more images (as indicated by arrow 118 in Figure 3). The amount of time elapsed between the detection of user input in block 200 and the actual image capture in block 204 is referred to and defined herein as "capture latency." Camera 50 can output a raw image to the ISP circuit 74, as indicated by arrow 120 in Figure 3.
[0050] After receiving a raw image from camera 50, the CVP circuit 60 can utilize at least some of its components, such as the statistical FE processor 66 and / or the statistical BE processor 68, to perform camera adjustments, including exposure (sometimes referred to as auto exposure), white balance (sometimes referred to as auto white balance), tone mapping, lens shading, lens correction, and / or other types of adjustments that may affect the final processed image. Using the CVP circuit 60 to initiate the analysis of the captured image and perform operations such as auto exposure (AE) and auto white balance (AWB) can be technically advantageous and beneficial in obtaining appropriate camera settings, enabling the overall camera pipeline to obtain a properly exposed, more aesthetically pleasing image. The image processed by the ISP circuit 74 can be stored in memory 76, as indicated by arrow 122 in Figure 3 (see also the operation of block 206 in Figure 4).
[0051] In block 208, the application processor 106 can be fully awakened (for example, the processor 106 completes the transition to the awakened state). At this point, the main operating system 108 running on the processor 106 may be fully operational and ready to handle desired tasks and workloads. Once the application processor 106 is active, it may be configured to immediately ping the low-power computing processor 102 to check for images (see operation in block 210). As indicated by arrow 130 in Figure 3, the application processor 106 may send one or more pings to the low-power computing processor 102 to retrieve or request one or more images from memory 76. Such pings sent from the application processor 106 may be referred to as image requests.
[0052] During the operation of block 212, the low-power computing processor 102 can send the addresses of one or more images stored in memory 76 during block 206 in response to an image request ping, as indicated by arrow 132 in Figure 3. During the operation of block 214, the application processor 106 can access the corresponding images from memory 76 using the addresses received from processor 102 (for example, memory 76 can output the stored images to processor 106, as indicated by arrow 134 in Figure 3). This example, in which application processor 106 retrieves address information from low-power computing processor 102 and then uses the retrieved address information to access memory 76, is illustrative. In other embodiments, in response to receiving a ping 130 from application processor 106, low-power computing processor 102 can retrieve the stored images from memory 76 and then forward the retrieved message to application processor 106. Other methods can be used to retrieve and transmit captured images to application processor 106 as needed.
[0053] During the operation of block 216, device 10 may optionally display the captured images using the display 20 shown in Figure 1. For example, application processor 106 may receive one or more captured images from memory 76 and then transmit the captured images(s) to display 20 for output. If necessary, the captured images(s) may be stored in memory 76 for subsequent (downstream) processing and / or transmitted to other external devices or the cloud for online storage. Operating device 10 in this way to capture images, process the captured images, and then store the processed images (and optionally display the stored images) may be technically advantageous and beneficial in minimizing shutter lag. Shutter lag reduction may be achieved by using an always-on processor 100 to wake up the application processor 106 and the low-power computing processor 102 simultaneously, with the low-power computing processor 102 preemptively kickstarting the image capture process even before processor 106 is fully awake. Displaying one or more images captured locally on device 10 is illustrative. If necessary, one or more images captured using the operation shown in Figure 4 may be shared with or otherwise transmitted to other computing devices (e.g., smartphones, tablets, laptop computers, desktop computers, watches, other head-mounted devices, etc.) and viewed on those other computing devices.
[0054] The operation in Figure 4 is illustrative. In some embodiments, one or more of the operations described may be modified, replaced, or omitted. In some embodiments, one or more of the operations described may be performed in parallel. In some embodiments, additional processes may be added or inserted between the operations described. The order of some operations may be reversed or changed as needed, and / or the timing of the operations described may be adjusted so that they occur at slightly different times. In some embodiments, the operations described may be distributed across a larger system.
[0055] According to one embodiment, the head-mounted device includes one or more image sensors, a first processing circuit configured to run an operating system for the head-mounted device, a second processing circuit configured to instruct one or more image sensors to capture images, and a third processing circuit configured to detect user input and simultaneously wake up the first and second processing circuits from a sleep state in response to the detection of user input.
[0056] According to another embodiment, the first processing circuit optionally includes an application processor configured to run one or more applications using an operating system, and is operable between a sleep state and a wake state.
[0057] According to another embodiment, the third processing circuit optionally includes a processor that is continuously in a wake state.
[0058] According to another embodiment, the head-mounted device further includes an image signal processing (ISP) circuit optionally configured to receive and process captured images output from one or more image sensors, the second processing circuit being operable between sleep and wake states and including a camera driver for controlling the image signal processing circuit.
[0059] According to another embodiment, the first processing circuit consumes a first amount of power in the wake state, the second processing circuit optionally consumes a second amount of power less than or equal to the first amount of power in the wake state, and the third processing circuit optionally consumes a third amount of power less than the first amount of power.
[0060] According to another embodiment, the image signal processing circuit optionally includes a plurality of subsystems configured to receive captured images and to operate in a first power domain, and a backend image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
[0061] According to another embodiment, the head-mounted device further optionally includes a memory device configured to receive and store processed images output from a backend image signal processing pipeline, wherein a second processing circuit is operable to access the memory device, and a third processing circuit is unable to access the memory device.
[0062] According to another embodiment, the second processing circuit is optionally configured to instruct one or more image sensors to begin capturing images in response to waking up from a sleep state to a wake state.
[0063] According to another embodiment, the first processing circuit is optionally configured to send a ping to the second processing circuit about the captured image in response to waking up from a sleep state to a wake state.
[0064] According to another embodiment, the head-mounted device optionally further includes one or more displays configured such that a first processing circuit pings a second processing circuit about the captured image, transmits the captured image to one or more displays, and then displays the captured image.
[0065] According to one embodiment, a method for operating a head-mounted device having a first processor, a second processor, and a third processor includes: using the third processor to detect user input; using the third processor to simultaneously wake up the first processor and the second processor, wherein the first processor has a first wake time and the second processor has a second wake time shorter than the first wake time, in response to the detection of user input; and initiating image capture using a camera driver running on the second processor in response to the second processor waking up from a sleep state to a wake state.
[0066] According to another embodiment, the first processor is capable of selectively operating between a sleep state and a wake state, consuming a first amount of power in the wake state; the second processor selectively consumes a second amount of power less than or equal to the first amount of power in the wake state; and the third processor selectively consumes a third amount of power less than the first amount of power.
[0067] According to another embodiment, initiating image capture using a camera driver running on a second processor optionally includes activating an image signal processing circuit, instructing one or more image sensors to start capturing images and output the captured images to the image signal processing circuit, and using the image signal processing circuit to output the processed images to a memory device.
[0068] According to another embodiment, the method further optionally includes using the first processor to send a request to the second processor to retrieve an image from a memory device, in response to the first processor waking up from a sleep state to a wake state.
[0069] In another embodiment, the method optionally further includes using a second processor to transmit address information to the first processor in response to receiving a request from the first processor, and using the first processor to access a memory device using the address information obtained from the second processor.
[0070] According to one embodiment, the electronic device includes a first processor on which the electronic device's operating system is executed, a second processor on which a camera driver is executed, the camera driver being configured to initiate image capture while the first processor is transitioning from a sleep state to a wake state, and a third processor configured to detect user input, wherein the first processor is configured to transition from a sleep state to a wake state based on the second processor detecting user input.
[0071] According to another embodiment, the third processor is optionally configured to wake up the first and second processors simultaneously in response to detecting user input.
[0072] According to another embodiment, the electronic device further includes one or more cameras optionally configured to capture images, an image signal processing circuit configured to receive captured images and output corresponding processed images, and a memory configured to store the processed images.
[0073] According to another embodiment, the image signal processing circuit optionally includes a plurality of subsystems configured to receive captured images and to operate in a first power domain, and a backend image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
[0074] According to another embodiment, the first processor is optionally configured to send an image request to the second processor in response to a transition to a wake state, and the second processor is optionally configured to respond to the image request by sending address information associated with a stored image in memory to the first processor.
[0075] The above is merely illustrative, and various modifications may be made to the described embodiments. The aforementioned embodiments may be implemented individually or in any combination.
Claims
1. It is a head-mounted device, One or more image sensors, A first processing circuit configured to run the operating system of the head-mounted device, A second processing circuit configured to instruct one or more image sensors to capture an image, A head-mounted device comprising: a third processing circuit configured to detect user input and, in response to the detection of said user input, simultaneously wake up the first processing circuit and the second processing circuit from a sleep state.
2. The head-mounted device according to claim 1, wherein the first processing circuit comprises an application processor configured to execute one or more applications using the operating system, and is operable between a sleep state and a wake state.
3. The head-mounted device according to claim 2, wherein the third processing circuit comprises a processor that is continuously in the wake state.
4. The second processing circuit further comprises an image signal processing (ISP) circuit configured to receive and process the captured images output from one or more image sensors, wherein the second processing circuit is operable between the sleep state and the wake state and includes a camera driver for controlling the image signal processing circuit. The head-mounted device according to claim 2.
5. The first processing circuit consumes a first amount of power in the wake state, The second processing circuit consumes a second amount of power less than or equal to the first amount of power in the wake state. The third processing circuit consumes a third amount of power that is less than the first amount of power. The head-mounted device according to claim 4.
6. The aforementioned image signal processing circuit is A computer vision processing circuit having a plurality of subsystems configured to receive the captured image and to operate in a first power domain, The head-mounted device according to claim 4, comprising: a backend image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
7. A memory device configured to receive and store processed images output from the backend image signal processing pipeline, wherein the second processing circuit is operable to access the memory device, and the third processing circuit is unable to access the memory device. The head-mounted device according to claim 6, further comprising the following:
8. The head-mounted device according to claim 4, wherein the second processing circuit is configured to instruct one or more image sensors to start capturing the image in response to waking up from the sleep state to the wake state.
9. The head-mounted device according to claim 4, wherein the first processing circuit is further configured to send a ping to the second processing circuit regarding the captured image in response to waking up from the sleep state to the wake state.
10. The first processing circuit sends a ping to the second processing circuit regarding the captured image, and after transmitting the captured image to one or more displays, one or more displays configured to display the captured image, The head-mounted device according to claim 9, further comprising the following:
11. A method for operating a head-mounted device having a first processor, a second processor, and a third processor, The third processor described above is used to detect user input, In response to detecting the user input, the third processor is used to simultaneously wake up the first processor and the second processor, wherein the first processor has a first wake time and the second processor has a second wake time shorter than the first wake time. A method comprising: initiating image capture using a camera driver running on the second processor in response to the second processor waking up from a sleep state to a wake state.
12. The first processor is capable of operating between the sleep state and the wake state, and consumes a first amount of power in the wake state. The second processor consumes a second amount of power that is less than or equal to the first amount of power in the wake state. The third processor consumes a third amount of power that is less than the first amount of power. The method according to claim 11.
13. Initiating the image capture using the camera driver running on the second processor is: Activating the image signal processing circuit, An instruction is given to one or more image sensors to start capturing an image and to output the captured image to the image signal processing circuit. The method according to claim 11, further comprising outputting the image processed by the image signal processing circuit to a memory device.
14. In response to the first processor waking up from the sleep state to the wake state, the first processor is used to send a request to the second processor to retrieve an image from the memory device. The method according to claim 13, further comprising:
15. Using the second processor, in response to receiving the request from the first processor, the address information is transmitted to the first processor, The first processor accesses the memory device using the address information obtained from the second processor, The method according to claim 14, further comprising:
16. It is an electronic device, A first processor on which the operating system of the aforementioned electronic device is executed, A camera driver, wherein the camera driver is configured to start image capture while the first processor is transitioning from a sleep state to a wake state, and the camera driver is executed on a second processor, An electronic device comprising: a third processor configured to detect user input, wherein the first processor is configured to transition from the sleep state to the wake state based on the second processor detecting the user input.
17. The electronic device according to claim 16, wherein the third processor is configured to wake up the first processor and the second processor simultaneously in response to detecting the user input.
18. One or more cameras configured to capture images, An image signal processing circuit configured to receive the captured image and output a corresponding processed image, A memory configured to store the processed image, The electronic device according to claim 16, further comprising the above.
19. The aforementioned image signal processing circuit is A computer vision processing circuit having a plurality of subsystems configured to receive the captured image and to operate in a first power domain, The electronic device according to claim 18, comprising: a backend image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
20. The electronic device according to claim 18, wherein the first processor is configured to transmit an image request to the second processor in response to a transition to the wake state, and the second processor is configured to respond to the image request by transmitting address information related to the stored image in the memory to the first processor.