Reducing image capture latency in a head-mounted device
Patent Information
- Application Number
- CN202610198779.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-12-04
- Filing Date
- 2026-02-11
- Publication Date
- 2026-08-28
Smart Images

Figure CN122661584A_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 19 / 408,566, filed December 4, 2025, and U.S. Provisional Patent Application No. 63 / 763,536, filed February 26, 2025, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates generally to electronic devices, and more specifically to head-mounted devices having one or more cameras. Background Technology
[0003] Some electronic devices can be mounted on a user's head. This type of electronic device can be referred to as a head-mounted device. A head-mounted device may include a camera for capturing images of the surrounding physical environment. It is in this context that the embodiments described herein came into being. Summary of the Invention
[0004] One aspect of this disclosure provides a head-mounted device comprising: one or more image sensors; a first processing circuit configured to run an operating system for the head-mounted device; a second processing circuit configured to direct the one or more image sensors to capture images; and a third processing circuit configured to detect user input and, in response to detecting user input, concurrently wake the first and second processing circuits from a sleep state. The first processing circuit may be an application processor configured to run one or more applications using the operating system and capable of operating between a sleep state and a wake state. The third processor may be a processor that remains continuously awake. The second processing circuit may be capable of operating between a sleep state and a wake state and may include a camera driver for controlling image signal processing (ISP) circuitry configured to receive and process captured images output from the one or more image sensors.
[0005] One aspect of this disclosure provides a method for operating a head-mounted device having a first processor, a second processor, and a third processor. The method may include: detecting user input using the third processor; concurrently waking up the first and second processors using the third processor in response to detecting user input, wherein the first processor has a first wake-up time, and wherein the second processor has a second wake-up time less than the first wake-up time; and initiating image capture using a camera driver running on the second processor in response to the second processor waking from a sleep state to a wake-up state. The first processor may be able to operate between a sleep state and a wake-up state, and consumes a first power amount in the wake-up state; the second processor may consume a second power amount less than or equal to the first power amount in the wake-up state; and the third processor may consume a third power amount less than the first power amount.
[0006] One aspect of this disclosure provides an electronic device comprising: a first processor on which an operating system of the electronic device executes; a second processor on which a camera driver executes, wherein the camera driver is configured to initiate image capture when the first processor transitions from a sleep state to a wake state; and a third processor configured to detect user input, wherein the first processor is configured to transition from a sleep state to a wake state based on the user input detected by the second processor. The third processor may be configured to concurrently wake up the first and second processors in response to the detection of user input. The electronic device may further comprise: one or more cameras configured to capture images; an image signal processing circuit configured to receive the captured images and output corresponding processed images; and a memory configured to store the processed images. The image signal processing circuit may include: a computer vision processing circuit configured to receive the captured images and having a plurality of subsystems configured to operate in a first power domain; and a back-end image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain. Attached Figure Description
[0007] Figure 1 This is an illustration of an exemplary system with a transparent display according to some implementation schemes.
[0008] Figure 2 It is shown that, according to some implementation schemes, it can be included Figure 1 A diagram illustrating an exemplary hardware component within a system of the type shown.
[0009] Figure 3This is an illustration showing how an exemplary system according to some implementations may include multiple processors for orchestrating low-latency image capture.
[0010] Figure 4 It is based on some implementation schemes for operation Figures 1 to 3 The flowchart illustrates the exemplary steps of a system of the type shown. Detailed Implementation
[0011] The physical environment can refer to the physical world that people can sense and / or interact with without the aid of electronic devices. The physical environment can include physical features, such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.
[0012] Conversely, extended reality (XR) environments refer to fully or partially simulated environments that people sense and / or interact with via electronic devices. For example, XR environments may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, etc. In the case of an XR system, a subset of a person's physical motion or a representation thereof is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law.
[0013] As an example, an XR system can detect head movement and, in response, adjust the graphical content and sound field presented to that person in a manner similar to how such views and sounds would change in the physical environment. As another example, an XR system can detect movement of an electronic device (e.g., a mobile phone, tablet, laptop, etc.) presenting the XR environment and, in response, adjust the graphical content and sound field presented to that person in a manner similar to how such views and sounds would change in the physical environment. In some cases (e.g., for accessibility reasons), an XR system may adjust the characteristics of the graphical content in the XR environment in response to a representation of physical movement (e.g., a voice command).
[0014] Many different types of electronic systems enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped like lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may combine one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment.
[0015] The head-mounted system may have a transparent or semi-transparent display, rather than an opaque display. The transparent or semi-transparent display may have a medium through which light representing the image is directed to the viewer's eyes. The display may utilize digital light projection, organic light-emitting diodes (OLEDs), LEDs, micro-light-emitting diodes (uLEDs), liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In some specific instances, the transparent or semi-transparent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection technology, which projects graphic images onto the viewer's retina. The projection system may also be configured to project virtual objects onto a physical environment, such as as a hologram or onto a physical surface. The display in device 10 is optional and may be omitted if desired.
[0016] Figure 1System 10 (sometimes referred to as electronic device 10, head-mounted device 10, etc.) may be a head-mounted device having one or more displays. The displays in system 10 may include a display 20 (sometimes referred to as a near-eye display) mounted within a support structure (housing) 8. The support structure 8 may be shaped like a pair of glasses or goggles (e.g., a support frame), may be formed with a helmet-shaped housing, or may have other configurations to help mount and secure components of the near-eye display 20 to the user's head or near their eyes. The near-eye display 20 may include one or more display modules such as display module 20A, and one or more optical systems such as optical system 20B. Display module 20A may be mounted in the support structure such as support structure 8. Each display module 20A may emit light 38 (image light), which is redirected towards the user's eye at the eye-adaptive zone 24 using an associated optical system in optical system 20B. Display 20 is optional and may be omitted from device 10.
[0017] The control circuit 16 can be used to control the operation of the system 10. The processing circuitry within the control circuit 16 can be used to control the operation of the device 10. The processing circuitry may be based on one or more microprocessors, microcontrollers, digital signal processors, baseband processors, power management units, audio chips, application-specific integrated circuits (ASICs), etc. The control circuit 16 can be configured to perform operations within the system 10 using hardware (e.g., dedicated hardware or circuitry), firmware, and / or software. Software code and other data used to perform operations within the system 10 can be stored on a non-transitory computer-readable storage medium (e.g., a tangible computer-readable storage medium) within the control circuit 16. Software code may sometimes be referred to as software, data, program instructions, commands, or code. The non-transitory computer-readable storage medium (sometimes commonly referred to as memory) may include non-volatile memory such as non-volatile random access memory (NVRAM), one or more hard disk drives (e.g., disk drives or solid-state drives), one or more removable flash drives, or other removable media, etc. Software stored on the non-transitory computer-readable storage medium can be executed on the processing circuitry of the control circuit 16. The control circuit 16, which has both storage and processing circuitry, is sometimes collectively referred to as the storage and processing circuitry.
[0018] System 10 may include input / output circuitry such as input-output device 12. Input-output device 12 may be used to allow system 10 to receive data from external equipment (e.g., a tethered computer, portable device such as a handheld device or laptop computer), or other electrical equipment, and to allow user input to head-mounted device 10. Input-output device 12 may also be used to collect information about the environment in which system 10 (e.g., head-mounted device 10) operates. Output components in device 12 may allow system 10 to provide output to a user and may be used to communicate with external electronic equipment. Input-output device 12 may include one or more cameras 14 (sometimes referred to as image sensors). Cameras 14 may be used to collect images of physical objects, which may optionally be digitally merged with virtual objects on a display in system 10. Input-output device 12 may include sensors and other components 18 (e.g., accelerometers, gyroscopes, depth sensors, light sensors, haptic output devices, speakers, batteries, wireless communication circuitry for communication between system 10 and external electronic devices, etc.).
[0019] A camera 14 mounted on the front of system 10 and facing outwards (towards the front of system 10 and away from the user) may be referred to herein as outward-facing, externally facing, forward-facing, or front-facing camera. Camera 14 can capture visual ranging information, image information processed to locate objects in the user's field of view (e.g., so that virtual content can be properly registered relative to real-world objects), image content displayed in real time to the user of system 10, and / or other suitable image data. For example, an outward-facing camera may allow system 10 to monitor the movement of system 10 relative to its surrounding environment (e.g., the camera may be used to form part of a visual ranging system or a visual-inertial ranging system). An outward-facing camera can also be used to capture images of the environment displayed to the user of system 10. If desired, images from multiple outward-facing cameras can be merged with each other and / or outward-facing camera content can be merged with computer-generated content for the user.
[0020] Display module 20A can be a liquid crystal display, an organic light-emitting diode display, a laser-based display, or other types of display. Optical system 20B can form lenses that allow a viewer (see, for example, the viewer's eye at eye-adaptation zone 24) to view an image on display 20. Two optical systems 20B may be associated with the user's respective left and right eyes (e.g., for forming a left lens and a right lens). A single display 20 can generate images for both eyes, or a pair of displays 20 can be used to display images. In a configuration with multiple displays (e.g., a left-eye display and a right-eye display), the focal length and positioning of the lenses formed by system 20B can be selected such that any gaps between the displays will be invisible to the user (e.g., allowing the images from the left and right displays to seamlessly overlap or merge).
[0021] If desired, the optical system 20B may include a transparent structure (e.g., an optical combiner, etc.) that allows image light from the physical object 28 to be optically combined with virtual (computer-generated) images, such as virtual images in image light 38. Light from the physical object 28 in the physical environment or scene may sometimes be referred to herein as and defined as world light, scene light, ambient light, external light, or ambient light. In this type of system, the user of system 10 can view both the physical environment surrounding the user and the computer-generated content overlaid on the physical environment. A camera 14 may also be used in device 10 (e.g., the camera captures an image of the physical object 28 and modifies that content, presenting it as virtual content at the optical system 20B).
[0022] If necessary, system 10 may include wireless circuitry and / or other circuitry to support communication with a computer or other external device (e.g., a computer that supplies image content to display 20). During operation, control circuitry 16 may provide image content to display 20. This content may be received remotely (e.g., from a computer or other content source coupled to system 10), and / or may be generated by control circuitry 16 (e.g., text, other computer-generated content, etc.). The content provided to display 20 by control circuitry 16 may be viewed by a viewer at eye level 24.
[0023] Figure 2 It shows that it can be included in the combination Figure 1 A diagram illustrating exemplary hardware components within a system of the described type (e.g., device 10). Figure 2As shown, device 10 may include one or more hardware and / or software subsystems, including one or more outward-facing image sensing subsystems (such as outward-facing camera 50), one or more tracking subsystems (such as tracking sensor 54), computer vision processing (CVP) circuitry (such as CVP circuitry 60), separate image signal processing pipelines (such as high-quality (back-end) pipeline 72), and one or more displays 20.
[0024] One or more cameras 50 can be used to collect information about the external real-world environment or scene surrounding the device 10. Cameras 50 may include... Figure 1 One or more front-facing cameras in front-facing camera 14. At least some of the cameras in camera 50 may be configured to capture one or more images of a scene, which may optionally be presented to a user as a live video pass-through feed using display 20. Camera 50 may include a color image sensor and / or optionally a monochrome (black and white) image sensor.
[0025] Cameras 50 may have different fields of view. Some cameras 50 may have wide or ultra-wide fields of view, while some cameras 50 may have relatively narrow fields of view. Not all cameras 50 need to be used to capture pass-through content. Some cameras 50 may be forward-facing (e.g., oriented towards the scene in front of the user); some cameras 50 may be downward-facing (e.g., oriented towards the user's torso, hands, or other parts of the user); some cameras 50 may be side-facing / lateral (e.g., oriented towards the user's left and right sides); and some cameras 50 may be oriented in other directions relative to the front of device 10. All these cameras 50 configured to collect information about the external physical environment surrounding device 10 are sometimes referred to as and defined as "outward-facing" or "externally-facing" cameras.
[0026] The tracking sensor 54 may include a gaze tracking subsystem, sometimes referred to as a gaze tracker, configured to acquire gaze information or gaze point information. The gaze tracker may employ one or more inward-facing cameras and / or other gaze tracking components (e.g., eye-facing components and / or other light sources emitting beams such that reflections of the beam from the user's eyes can be detected) to monitor the user's eyes. One or more gaze tracking sensors 54 may be oriented towards the user's eyes and track the user's gaze. The camera in the gaze tracking subsystem may determine the position of the user's eyes (e.g., the center of the user's pupil), the direction in which the user's eyes are oriented (the direction of the user's gaze), the size of the user's pupils (e.g., such that light modulation and / or other optical parameters are adjusted based on the pupil size, and / or the amount of sequence used to spatially adjust one or more of these parameters, and / or the area where one or more of these optical parameters are located), and may be used to monitor the current focal length of the lens in the user's eyes (e.g., whether the user is focusing on the near or far field, which can be used to assess whether the user is daydreaming or thinking strategically or tactically), and / or other gaze information. A gaze-tracking camera may sometimes be referred to as an inward-facing camera, gaze detection camera, eye-tracking camera, gaze tracking camera, or eye monitoring camera. Other types of optical sensors (e.g., infrared and / or visible light LEDs and photodetectors) may also be used to monitor the user's gaze if needed.
[0027] Tracking sensor 54 may also include a face and body tracking subsystem configured to perform face tracking (e.g., capturing images of the user's jaw, mouth, etc., when the device is worn on the user's head) and body tracking (e.g., capturing images of the user's torso, arms, hands, legs, etc., when the device is worn on the user's head). If desired, the face and body tracking subsystem may also track the user's head posture by directly determining any movement, yaw, pitch, roll, etc., of the head-mounted device 10. The yaw, roll, and pitch of the user's head can collectively define the user's "head posture." For example, tracking sensor 54 may include an inertial measurement unit (IMU). The IMU may include one or more gyroscopes, gyrocompasses, accelerometers, magnetometers, other inertial sensors, and other position and motion sensors. These position and motion sensors may assume that the head-mounted device 10 is mounted on the user's head. Therefore, in this document, references to head pose, head movement, user head yaw (e.g., rotation about a vertical axis), user head pitch (e.g., rotation about a left-right axis), user head roll (e.g., rotation about a front-back axis), etc., can be considered interchangeable with references to device pose, device movement, device yaw, device pitch, device roll, etc. In some embodiments, the tracking sensor 54 may also include a six-degree-of-freedom (DoF) tracking subsystem. A six-DoF tracking subsystem or sensor can be used to monitor rotational motions such as roll, pitch, and yaw, as well as positional / translational motions in a 3D environment.
[0028] The tracking sensor 54 may optionally also include a hand tracking subsystem, sometimes referred to as a hand tracker, configured to monitor the user's hand movements / gestures to obtain hand gesture data. For example, the hand tracker may include a camera and / or other gesture tracking components (e.g., outward-facing components and / or a light source emitting a beam of light such that reflections of the beam from the user's hand can be detected) to monitor the user's hand. One or more hand tracking sensors may be pointed at the user's hand and may track movements associated with the user's hand, determining whether the user is performing a tapping or swiping motion with his / her fingertips or hand, whether the user is performing a non-contact button press or object selection operation with his / her hand, whether the user is performing a grasping or gripping motion with his / her hand, whether the user is pointing or pinching a given object presented on display 20 with his / her hand or fingers, whether the user is performing a waving or striking motion with his / her hand, or may substantially measure / monitor three-dimensional non-contact gestures ("air gestures") associated with the user's hand. The tracking sensor 54, which is capable of operating to obtain information related to the user's movement, such as gaze, posture, hand gestures, and other information, is sometimes collectively referred to as a "user tracking" sensor.
[0029] Figure 2 Examples of the outward-facing camera 50 and tracking sensor 54 (e.g., optical sensors for acquiring gaze, posture, and / or other user-related data) shown are illustrated as separate, independent subsystems and are illustrative. In some embodiments, one or more of the outward-facing cameras 50 may also be employed to acquire posture information, location information, and / or other motion / position information associated with device 10. To help protect user privacy, best practices can be used to handle any personal user information acquired by the sensors. These best practices include meeting or exceeding any applicable privacy rules. Option-join and option-opt-out options and / or other options may be provided that allow users to control the use of their personal data.
[0030] Electronic device 10 can be configured to collect contextual information about the surrounding real-world (physical) environment or scene. Collecting contextual information may include, for example, identifying one or more objects of interest in the environment, detecting when a user enters a particular room or environment, detecting when a user engages in a particular activity, detecting the current location of device 10, detecting the current user context or usage scenario (e.g., detecting whether the user is currently watching a movie, playing a video game, or talking to another person or avatar), and / or determining other contextual information related to the operation of device 10. Collecting contextual information may involve capturing one or more images using an outward-facing camera 50 and / or obtaining data from a tracking sensor 54. Such images captured for contextual purposes do not need to be output by display 20 for human consumption. Therefore, the processing requirements and complexity for handling such images can be less than the conventional image signal processing steps required for processing images output by a display for human consumption (viewing).
[0031] According to one embodiment, the image signal processing circuitry on device 10 can be segmented into a first portion including computer vision processing (CVP) circuitry 60 and a separate second portion including a high-quality (HQ) pipeline 72. In other words, CVP circuitry 60 and HQ pipeline 72 may sometimes be collectively referred to herein and defined as image signal processing (ISP) circuitry. CVP circuitry 60 can be used alone to process images and / or data output from sensors 50 and 54 that only require analysis for contextual purposes (without needing processing by the high-quality pipeline 72), while images and / or data output from sensors 50 and 54 that will be output to display 20 for human viewing can be processed by both CVP circuitry 60 and the high-quality pipeline 72. Components within CVP circuitry 60 can operate in a first power domain, while components within HQ pipeline 72 can operate in a second power domain different from the first power domain (e.g., CVP circuitry 60 and HQ pipeline 72 can be configured to operate in different power domains).
[0032] The components in the CVP circuitry 60 typically operate in a lower power domain compared to the components in the HQ pipeline 72. The high-quality pipeline 72 can be power-gated (e.g., the HQ pipeline 72 can be selectively activated and deactivated to reduce overall power consumption). When processing images to be output on display 20 for human consumption, the high-quality pipeline 72 can be selectively activated (e.g., powered on) to perform some or all of the image processing functions provided by the HQ pipeline 72. When processing images only for contextual purposes (e.g., to support one or more computer vision algorithms running on device 10) and such images do not need to be displayed, the HQ pipeline 72 can be selectively deactivated (e.g., powered off or idle) to save power. In other words, the CVP circuitry 60 consumes a first amount of power when activated, while the HQ pipeline 72 consumes a second amount of power greater than the first amount when activated. Operating the image signal processing circuitry on device 10 in this manner can technically benefit from minimizing power consumption on device 10. Such power-saving operations could be beneficial for small and lightweight devices 10 that can be powered by batteries and used all day.
[0033] like Figure 2 As shown, the CVP circuit 60 may include one or more hardware and / or software subsystems, such as a sensor interface 62, a front-end (FE) processor 64, a statistical front-end (FE) processor 66, a statistical back-end (BE) processor 68, a central processing unit (CPU) (such as a computer vision processing (CVP) CPU 70), and / or other image signal processing components. The sensor interface 62 may be configured to receive images (e.g., raw pixel data) from the camera 50, the tracking sensor 54, and / or other image sensors within the device 10. The front-end processor 64 may be configured to perform defective / faulty pixel correction, image scaling or merging operations, image cropping or resizing, and / or other front-end or image preprocessing operations. The statistical FE processor 66 may be configured to collect pixel statistics, such as minimum pixel value, maximum pixel value, average pixel value, color plane information (e.g., red, green, and blue color planes), color and / or luminance histograms, and other front-end image statistics. The statistical BE processor 68 may be configured to convert images from the raw Bayer domain into color images and may generate additional statistical information.
[0034] The color image output from the statistical BE processor 68 can be provided to one or more downstream computer vision processing algorithms or tasks running on device 10 (e.g., processor 68 can output the image to one or more client processors). The statistical FE processor 66 and BE processor 68 can be collectively referred to as the CVP statistical pipeline. Although the CVP circuitry 60 is shown as a single instance including interface 62, processor 64, processor 66, and processor 68, the CVP circuitry 60 can include multiple sensor interface blocks 62 for interfacing with multiple sensors, multiple front-end processors 64 for performing image preprocessing operations in parallel, multiple processors 66 for performing front-end statistical computations in parallel, and multiple processors 68 for performing back-end statistical computations in parallel. The computer vision processing CPU 70 can be configured to manage and coordinate the operations of blocks 62, 64, 66, and 68 to process each incoming image frame.
[0035] The computer vision processing circuit 60 primarily includes components for performing front-end image signal processing operations. Therefore, the computer vision processing circuit 60 is sometimes referred to herein and defined as a "front-end" image signal processing (ISP) circuit. In contrast, the HQ pipeline 72 primarily includes components configured to perform back-end image signal processing operations. Therefore, the high-quality pipeline 72 is sometimes referred to herein and defined herein as a "back-end" image signal processing (ISP) circuit. The high-quality (back-end) pipeline 72 can be a more complex and higher-power-consuming version of the statistical back-end processor 68 of the CVP circuit 60. For example, the HQ pipeline 72 may include components configured to perform image signal processing operations in which the CVP circuit 60 is completely absent, such as poor / defective pixel correction, noise reduction, white balance, demosaicing, color space conversion, tone mapping (e.g., including global and local tone mapping), color correction, gamma correction, shadow correction, image sharpening, high dynamic range (HDR) correction, edge-aware local image adjustment, image fusion (e.g., merging multiple image frames together for noise reduction and high dynamic range), and other image signal processing functions for outputting corresponding images for display.
[0036] The image output by the back-end processor 68 of the CVP circuit 60 can be processed according to a first set of image processing requirements, which may optionally produce a lower fidelity (quality) image for computer vision consumption, while the image output by the HQ pipeline 72 can be processed according to a second set of image processing requirements, which may optionally produce a relatively higher fidelity (quality) image to be displayed for human consumption, different from the first set of image processing requirements. In some embodiments, the CVP circuit 60 can be configured to output a processed image with a first quality and / or using a first power level, while the HQ pipeline 72 can be configured to output a processed image with a second quality greater than the first quality and / or using a second power level greater than the first power level. In some embodiments, the CVP circuit 60 can be configured to output a processed image by performing the first set of image processing operations, while the HQ pipeline 72 can be configured to output a processed image by performing additional image processing operations different from the first set of image processing operations. The image output by the processor 68 can be provided as a result to one or more client processors. The HQ pipeline 72 can output content for human consumption via the display 20. Figure 2 The example is illustrative. Display 20 is optional and can be omitted from device 10. If needed, the output from HQ pipeline 72 can be stored in memory for later processing.
[0037] In other types of mobile electronic devices, such as smartphones with cameras, users typically open the camera app and are presented with a preview of the image to be captured before pressing the capture button. In such scenarios, the smartphone can determine with a high probability that the user is about to press the capture button and can prepare for image capture by preemptively waking up all the necessary hardware and / or software subsystems required for image capture.
[0038] Compared to capturing images on a smartphone, a user can initiate or trigger image capture by operating device 10 without having to open the camera (image capture) app. In other words, device 10 may not know in advance when the user will press the capture button. As described above, device 10 can be a lightweight head-mounted device with low power consumption for all-day use. This lightweight head-mounted device may include one or more processors, at least some of which can operate in sleep mode to reduce active power consumption.
[0039] For example, device 10 may include an application processor on which the operating system of device 10 is executed. The application processor should operate in sleep mode most of the time to save power. Pressing the capture button (which can occur at any time based on the user's whim) can trigger the application processor to wake up. However, waiting for the application processor to fully wake up before allowing the camera to capture an image can introduce a significant amount of shutter lag. "Shutter lag" herein may refer to and is defined as the delay between pressing the capture (shutter) button and the actual moment the image is captured at the camera.
[0040] According to one embodiment, a method for reducing shutter lag in an operating device 10 is provided. The device 10 may utilize an always-on processor to detect button presses and, in response to detecting a button press, wake up an application processor and prepare a camera pipeline for image capture in parallel with, and even before, the application processor is fully awakened. For example, the camera pipeline may begin capturing one or more images and begin processing the captured images before the application processor is ready to process them. The term "camera pipeline" or camera stack may refer to all subsystems involved in capturing one or more images, which may include at least an externally facing camera 50, computer vision processing circuitry 60, a high-quality pipeline 72, associated memory devices for storing the captured images, and a camera driver (e.g., a software subsystem configured to orchestrate the operation of image signal processing circuitry).
[0041] Figure 3 This is an illustration showing how a device 10, according to some embodiments, may include multiple processors for orchestrating low-latency image capture. (See diagram for details.) Figure 3 As shown, device 10 may include one or more cameras 50, computer vision processing circuitry 60, a high-quality pipeline 72, a memory device 76, and one or more processing circuits (such as processors 100, 102, and 106). Camera 50 may be an outward-facing image sensor configured to capture one or more images of a scene or physical environment. The captured images may be processed by computer vision processing (CVP) circuitry 60, and then by high-quality pipeline 72. Computer vision processing circuitry 60 and high-quality pipeline 72 can therefore receive incoming (raw) images and output corresponding "processed" images. Computer vision processing circuitry 60 and high-quality pipeline 72, configured to generate processed images, are sometimes collectively referred to as image signal processing (ISP) circuitry 74. The processed image output from ISP circuitry 74 may be stored in memory 76. Memory device 76 may be part of a storage subsystem within control circuitry 16 (see [link to relevant documentation]). Figure 1The memory device 76 may be implemented as volatile memory (such as random access memory (e.g., dynamic RAM or DRAM)), non-volatile (persistent) memory (such as flash memory), magnetic drive, optical drive or solid-state drive, or other types of storage devices.
[0042] Processor 106 may represent the application processor of device 10. Application processor 106 is sometimes referred to as application processing circuitry 106. Application processor 106 may be configured to run or execute an operating system (OS), such as operating system 108 for device 10. Operating system 108 may be used to manage multiple applications running on device 10, such as allowing users to switch between different applications (e.g., photo / video organization and editing applications, media streaming applications, gaming applications, social media applications, map / navigation applications, health and fitness applications, automation assistant applications, information search applications, taxi calling applications, online banking applications, etc.), to manage security features on device 10, such as performing biometric authentication for secure access and data encryption for protecting user data, and / or managing productivity features, such as the user's calendar, reminders, notes, and files, etc. Generally, operating system 108 may be designed to provide a secure and intuitive platform that prioritizes user experience, privacy, and seamless functionality across a wide variety of services and applications.
[0043] If constantly active, the application processor 106, running the full OS stack as described above, may consume significant power. To help extend the battery life of device 10, the application processor 106 can be configured to sleep when applications running on device 10 are idle. The application processor 106 can therefore switch between sleep and wake states. When a user needs or activates one or more applications managed by operating system 108, the application processor 106 can be woken up by transitioning from sleep to wake. The amount of time it takes for the application processor 106 to transition from sleep to wake is sometimes referred to herein as the application processor "wake-up time".
[0044] Compared to application processor 106, processor 100 can be always powered on. For example, processor 100 can be a dedicated low-power processing subsystem designed to continuously process specific tasks without consuming excessive battery power. Processor 100 remains active even when application processor 106 is in sleep or idle state. Processor 100, operating at minimum power levels, can be configured to handle lightweight tasks such as monitoring sensors (e.g., ...). Figure 1 Sensor 18 Figure 2The processor 100 includes an image sensor 50 and a tracking sensor 54 and / or other sensors, monitors voice commands (e.g., “Hey Siri” or other voice commands), manages notifications and / or maintains wireless connectivity for certain applications, etc. This type of processor 100 is sometimes referred to herein as and defined as an “always-on” processor (AOP) or an “always-wake” processor. The always-on processor 100 is sometimes referred to as the always-on processing circuit 100. As long as the battery is not completely depleted, the always-on processor 100, which is always (continuously or constantly) active, ensures a rapid response to certain triggering events that would otherwise require the attention of the processor 106 by eliminating the latency of waking up the main processor 106 (e.g., bypassing the application processor wake-up time). Offloading lightweight tasks from the main application processor 106 to the always-on processor 100 is also technically advantageous and beneficial in helping to conserve energy while optimizing the performance of the entire system 10.
[0045] Device 10 may also be provided with low-power computing blocks, such as low-power computing processor 102. Low-power computing processor 102 is sometimes referred to as low-power computing processing circuitry 102. Low-power computing processor 102 may include camera drivers, such as camera driver 104 configured to control a camera pipeline. Camera driver 104 (sometimes referred to as an image sensor or image signal processing driver) is a software subsystem configured to orchestrate the operation of ISP circuitry 74. Processor 102 can directly access memory 76, which is sometimes referred to herein as image storage circuitry. Application processor 106 can access or retrieve stored images from memory 76 via low-power computing processor 102. Alternatively or additionally, application processor 106 may also directly access images stored on memory 76. Unlike processors 102 and 106, processor 100 is always enabled and cannot access memory 76 (e.g., processor 100 should not be able to access stored images). Processors 100, 102, and 106 configured to operate in this way may be considered to have different memory access privileges. For example, always-on processor 100 may have a first memory access privilege, application processor 106 may have a second memory access privilege equal to or greater than the first memory access privilege, and low-power computing processor 102 may have a third memory access privilege equal to or greater than the second memory access privilege.
[0046] The low-power computing processor 102 can operate in either a wake-up or sleep state. A wake-up processor 102 consumes less power than a wake-up application processor 106. A wake-up processor 102 consumes more power than a always-on processor 100. Generally, the application processor 106 consumes a first amount of power; the low-power computing processor 102 consumes a second amount of power less than the first amount; and the always-on processor consumes a third amount of power less than the second amount. In some embodiments, when operating in a wake-up state, the low-power computing processor 102 may consume a similar amount of power as the always-on processor 100. Although the processor 102 can switch between sleep (idle) and wake-up states, the processor 100 is always active (e.g., the processor 100 is always awake but consumes a small amount of power).
[0047] Figure 4 It is used for operation combination Figures 1 to 3 A flowchart illustrating exemplary steps of device 10 of the described type. During operation of block 200, device 10 may detect or predict user input for capturing an image. For example, always-on processor 100 may be configured to monitor user input 110. User input 110 may be a button press (e.g., a user pressing or pushing a physical button on the housing or frame of device 10), a touch (e.g., a physical tap or pressure on a portion of the housing or frame of device 10), a voice command (e.g., asking Siri or another automated assistant to take a photo), a hand gesture (e.g., a gesture from the user's finger or hand or other movement used to trigger image capture), and / or a remote trigger (e.g., using a remote controller wirelessly connected to device 10), etc. In the example above, such user input 110 may be detected via a physical button, virtual button, touch sensor, microphone, motion sensor, or other type of sensor, which may then prompt the always-on processor 100 that the user intends to capture an image (e.g., ...). Figure 3 (As indicated by arrow 112 in the image).
[0048] During the operation of box 202, always-on processor 100 can be configured to wake up low-power computing processor 102 (e.g., Figure 3 (As shown by arrow 114-1 in the image) and concurrently wake up application processor 106 (as shown by arrow 114-1 in the image) Figure 3(As shown by arrow 114-2 in the diagram). In other words, always-on processor 100 can wake processors 102 and 106 from sleep in parallel. Before box 202, both low-power processor 102 and application processor 106 can be in sleep (idle) state. After receiving a wake-up signal from processor 100, low-power computing processor 102 can begin to wake up and transition to wake-up (active) mode. Similarly, after receiving a wake-up signal from processor 100, application processor 106 can begin to wake up and transition to wake-up (active) mode.
[0049] As used herein, the term "concurrent" means at least partially overlapping in time. In other words, the first and second events are referred to herein as "concurrent" if at least some of the first events occur simultaneously with at least some of the second events (e.g., if at least some of the first events occur during, concurrently with, or when at least some of the second events occur). The first and second events can be concurrent if they are synchronized (e.g., if the entire duration of the first event overlaps with the entire duration of the second event in time), but they can also be concurrent if they are asynchronous (e.g., if the first event begins before or after the second event, ends before or after the second event, or does not partially overlap in time). As used herein, the term "at the time of" is synonymous with "concurrent".
[0050] Application processor 106 may have a first wake-up time, while low-power computing processor 102 may have a second wake-up time shorter than the first wake-up time. In other words, low-power computing processor 102 may wake up faster than application processor 106. During the operation of block 204, processor 102 may fully transition to the wake-up state (before application processor 106 transitions to the wake-up state), and then image signal processing circuitry 74 may be directed to begin streaming images from one or more cameras 50. To achieve this, camera driver 104 running on processor 102 may concurrently activate CVP circuitry 60 (such as...). Figure 3 (as shown by arrow 116-1 in the image) and HQ pipeline 72 (as shown in the image) Figure 3 (As shown by arrow 116-2 in the image). The image signal processing circuit 74 can then transmit signals to the camera 50 (such as...). Figure 3(As indicated by arrow 118 in the diagram), this signal guides camera 50 to capture one or more images. The amount of time elapsed between detecting user input at box 200 and detecting actual image capture at box 204 is sometimes referred to herein as and defined as the "capture delay". Camera 50 can then output the raw image to ISP circuitry 74, as shown in the diagram. Figure 3 As indicated by arrow 120 in the image.
[0051] After receiving the raw image from camera 50, CVP circuit 60 can utilize at least some of its components (such as statistical FE processor 66 and / or statistical BE processor 68) to perform camera adjustments, including exposure adjustments (sometimes referred to as auto exposure), white balance (sometimes referred to as auto white balance), tone mapping, lens shading, lens correction, and / or other types of adjustments that may affect the final processed image. Using CVP circuit 60 to begin analyzing the captured image and performing operations such as auto exposure (AE) and auto white balance (AWB) can be technically advantageous and beneficial in obtaining appropriate camera settings, thereby enabling the entire camera pipeline to acquire properly exposed and more aesthetically pleasing images. Images processed by ISP circuit 74 can be stored in memory 76, such as... Figure 3 Arrow 122 is shown in the image (see also...) Figure 4 (The operation of box 206 in the middle).
[0052] At box 208, application processor 106 can be fully woken up (e.g., processor 106 completes the transition to a wake-up state). At this point, the host operating system 108 running on processor 106 can be fully operational and ready to handle the desired tasks and workloads. Once application processor 106 is active, it can be configured to immediately inspect low-power computing processor 102 for image examination (see operation at box 210). Figure 3 As illustrated by arrow 130, application processor 106 may output one or more checks to low-power computing processor 102 to retrieve or request one or more images from memory 76. Such checks output from application processor 106 are sometimes referred to as image requests.
[0053] During the operation of block 212, the low-power computing processor 102 can transmit the address of the image stored in memory 76 during block 206 in response to an image request verification, such as Figure 3 As indicated by arrow 132 in the diagram. During the operation of box 214, application processor 106 can use the address received from processor 102 to access the corresponding image from memory 76 (e.g., memory 76 can output a stored image to processor 106, such as...). Figure 3(As indicated by arrow 134 in the diagram). This example, in which application processor 106 retrieves address information from low-power computing processor 102 and then uses the retrieved address information to access memory 76, is illustrative. In other embodiments, in response to receiving verification 130 from application processor 106, low-power computing processor 102 may retrieve a stored image from memory 76 and then forward the retrieved message to application processor 106. Other methods may be employed, if desired, to retrieve a captured image and transmit it to application processor 106.
[0054] During the operation of box 216, device 10 may optionally use Figure 1 The display 20 shown is used to display the captured images. For example, the application processor 106 may receive one or more captured images from the memory 76 and then transmit the captured images to the display 20 for output. If needed, the captured images may be stored in the memory 76 for later (downstream) processing and / or transmitted to other external devices or the cloud for online storage. Operating the device 10 in this way to capture images, process the captured images, and then store the processed images (and optionally display the stored images) can be technically advantageous and beneficial in minimizing shutter lag. Reduction of shutter lag can be achieved by using an always-on processor 100 to concurrently wake up the application processor 106 and the low-power computing processor 102, which preemptively starts the image capture process even before the processor 106 is fully awake. Displaying the captured images locally at the device 10 is exemplary. If needed, using Figure 4 One or more images captured by the operation can be shared or otherwise sent to other computing devices (e.g., smartphones, tablets, laptops, desktops, watches, other head-mounted devices, etc.) and viewed on other computing devices.
[0055] Figure 4 The operations described are illustrative. In some embodiments, one or more of the described operations may be modified, replaced, or omitted. In some embodiments, one or more of the described operations may be performed in parallel. In some embodiments, additional processes may be added or inserted between the described operations. If necessary, the order of certain operations may be reversed or changed, and / or the timing of the described operations may be adjusted so that they occur at slightly different times. In some embodiments, the described operations may be distributed across a larger system.
[0056] According to one embodiment, the head-mounted device includes: one or more image sensors; a first processing circuit configured to run an operating system for the head-mounted device; a second processing circuit configured to direct the one or more image sensors to capture images; and a third processing circuit configured to detect user input and, in response to detecting user input, concurrently wake the first and second processing circuits from a sleep state.
[0057] According to another embodiment, the first processing circuit may optionally include an application processor configured to run one or more applications using an operating system and to operate between a sleep state and a wake state.
[0058] According to another embodiment, the third processing circuit may optionally include a processor that is constantly awake.
[0059] According to another embodiment, the head-mounted device may optionally further include an image signal processing (ISP) circuit configured to receive and process captured images output from one or more image sensors, wherein the second processing circuit is operable between a sleep state and a wake state and includes a camera driver for controlling the image signal processing circuit.
[0060] According to another embodiment, the first processing circuit consumes a first power amount in the wake-up state, the second processing circuit optionally consumes a second power amount less than or equal to the first power amount in the wake-up state, and the third processing circuit optionally consumes a third power amount less than the first power amount.
[0061] According to another embodiment, the image signal processing circuit may optionally include: a computer vision processing circuit configured to receive a captured image and having a plurality of subsystems configured to operate in a first power domain; and a back-end image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
[0062] According to another embodiment, the head-mounted device may optionally also include a memory device configured to receive and store processed images output from a back-end image signal processing pipeline, wherein a second processing circuitry is operable to access the memory device, and wherein a third processing circuitry is not operable to access the memory device.
[0063] According to another embodiment, the second processing circuit is optionally configured to instruct one or more image sensors to begin capturing images in response to waking from a sleep state to a wake state.
[0064] According to another embodiment, the first processing circuit may optionally be further configured to examine the second processing circuit for the captured image in response to waking from a sleep state to a wake state.
[0065] According to another embodiment, the head-mounted device may optionally include one or more displays configured to display the captured image after the first processing circuit verifies the captured image against the second processing circuit and transmits the captured image to the one or more displays.
[0066] According to one embodiment, a method of operating a head-mounted device having a first processor, a second processor, and a third processor includes: using the third processor to detect user input; in response to detecting user input, using the third processor to concurrently wake up the first processor and the second processor, wherein the first processor has a first wake-up time, and wherein the second processor has a second wake-up time less than the first wake-up time; and in response to the second processor waking up from a sleep state to a wake-up state, using a camera driver running on the second processor to initiate image capture.
[0067] According to another embodiment, the first processor may optionally be able to operate between a sleep state and a wake state and consume a first power amount in the wake state, the second processor may optionally consume a second power amount less than or equal to the first power amount in the wake state, and the third processor may optionally consume a third power amount less than the first power amount.
[0068] According to another embodiment, using a camera driver running on a second processor to initiate image capture may optionally include: activating an image signal processing circuit; instructing one or more image sensors to begin capturing images and outputting the captured images to the image signal processing circuit; and using the image signal processing circuit to output the processed images to a memory device.
[0069] According to another embodiment, the method may optionally further include: in response to the first processor waking from a sleep state to an awake state, using the first processor to transmit a request to the second processor to retrieve an image from a memory device.
[0070] According to another embodiment, the method may optionally further include: using a second processor to transmit address information to the first processor in response to receiving a request from the first processor; and using the first processor to access the memory device using the address information obtained from the second processor.
[0071] According to one embodiment, an electronic device includes: a first processor on which an operating system of the electronic device is executed; a second processor on which a camera driver is executed, wherein the camera driver is configured to initiate image capture when the first processor transitions from a sleep state to a wake state; and a third processor configured to detect user input, wherein the first processor is configured to transition from a sleep state to a wake state based on the user input detected by the second processor.
[0072] According to another embodiment, the third processor may optionally be configured to concurrently wake up the first and second processors in response to the detection of user input.
[0073] According to another embodiment, the electronic device may optionally further include: one or more cameras configured to capture images; an image signal processing circuit configured to receive the captured images and output corresponding processed images; and a memory configured to store the processed images.
[0074] According to another embodiment, the image signal processing circuit may optionally include: a computer vision processing circuit configured to receive a captured image and having a plurality of subsystems configured to operate in a first power domain; and a back-end image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
[0075] According to another embodiment, the first processor is optionally configured to transmit an image request to the second processor in response to transitioning to a wake-up state, and the second processor is optionally configured to respond to the image request by transmitting address information associated with an image stored in memory to the first processor.
[0076] The foregoing is merely illustrative and various modifications can be made to the described implementation scheme. The foregoing implementation scheme can be implemented individually or in any combination.
Claims
1. A head-mounted device, the head-mounted device comprising: One or more image sensors; A first processing circuit is configured to run an operating system for the head-mounted device; A second processing circuit is configured to guide the one or more image sensors to capture images; as well as A third processing circuit is configured to detect user input and, in response to detecting the user input, concurrently wake up the first processing circuit and the second processing circuit from a sleep state.
2. The head-mounted device of claim 1, wherein the first processing circuitry includes an application processor configured to run one or more applications using the operating system and to operate between the sleep state and the wake state.
3. The head-mounted device of claim 2, wherein the third processing circuitry includes a processor that remains in the wake-up state.
4. The head-mounted device according to claim 2, further comprising: An image signal processing (ISP) circuit configured to receive and process captured images output from one or more image sensors, wherein the second processing circuit is operable between the sleep state and the wake state and includes a camera driver for controlling the image signal processing circuit.
5. The head-mounted device according to claim 4, wherein: The first processing circuit consumes a first amount of power in the wake-up state; The second processing circuit consumes a second power amount less than or equal to the first power amount in the wake-up state; and The third processing circuit consumes a third amount of power that is less than the first amount of power.
6. The head-mounted device of claim 4, wherein the image signal processing circuit comprises: A computer vision processing circuit configured to receive a captured image and having multiple subsystems configured to operate in a first power domain; and A back-end image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
7. The head-mounted device according to claim 6, further comprising: A memory device configured to receive and store processed images output from the back-end image signal processing pipeline, wherein the second processing circuitry is operable to access the memory device, and wherein the third processing circuitry is not operable to access the memory device.
8. The head-mounted device of claim 4, wherein the second processing circuitry is configured to instruct the one or more image sensors to begin capturing the image in response to waking from the sleep state to the wake state.
9. The head-mounted device of claim 4, wherein the first processing circuitry is further configured to examine the second processing circuitry for the captured image in response to waking from the sleep state to the wake state.
10. The head-mounted device according to claim 9, further comprising: One or more displays are configured to display a captured image after the first processing circuit verifies the captured image against the second processing circuit and transmits the captured image to the one or more displays.
11. A method of operating a head-mounted device having a first processor, a second processor, and a third processor, the method comprising: The third processor is used to detect user input; In response to detecting the user input, the third processor is used to concurrently wake up the first processor and the second processor, wherein the first processor has a first wake-up time, and wherein the second processor has a second wake-up time that is less than the first wake-up time; as well as In response to the second processor waking up from sleep to wake, an image capture is initiated using the camera driver running on the second processor.
12. The method according to claim 11, wherein: The first processor is capable of operating between the sleep state and the wake state, and consumes a first amount of power in the wake state; The second processor consumes a second power amount less than or equal to the first power amount in the wake-up state; and The third processor consumes a third amount of power that is less than the first amount of power.
13. The method of claim 11, wherein initiating the image capture using the camera driver running on the second processor comprises: Activate the image signal processing circuit; Guide one or more image sensors to start capturing images and output the captured images to the image signal processing circuit; as well as The processed image is output to a memory device using the image signal processing circuit.
14. The method according to claim 13, further comprising: In response to the first processor waking up from the sleep state to the wake state, the first processor is used to transmit a request to the second processor to retrieve an image from the memory device.
15. The method of claim 14, further comprising: Using the second processor, in response to receiving the request from the first processor, address information is transmitted to the first processor; as well as The memory device is accessed using the address information obtained from the second processor, with the first processor in use.
16. An electronic device, the electronic device comprising: The first processor, on which the operating system of the electronic device is executed; A second processor, on which a camera driver is executed, wherein the camera driver is configured to initiate image capture when the first processor transitions from a sleep state to a wake state; and A third processor is configured to detect user input, wherein the first processor is configured to transition from the sleep state to the wake state based on the user input detected by the second processor.
17. The electronic device of claim 16, wherein the third processor is configured to concurrently wake up the first processor and the second processor in response to detecting the user input.
18. The electronic device according to claim 16, further comprising: One or more cameras, the one or more cameras being configured to capture images; An image signal processing circuit, configured to receive a captured image and output a corresponding processed image; and A memory configured to store the processed image.
19. The electronic device of claim 18, wherein the image signal processing circuit comprises: A computer vision processing circuit configured to receive a captured image and having multiple subsystems configured to operate in a first power domain; and A back-end image signal processing pipeline coupled to the computer vision processing circuit and configured to operate in a second power domain different from the first power domain.
20. The electronic device of claim 18, wherein the first processor is configured to transmit an image request to the second processor in response to transitioning to the wake-up state, and wherein the second processor is configured to respond to the image request by transmitting address information associated with an image stored in the memory to the first processor.