Low-power machine learning using real-time captured regions of interest

By identifying and processing regions of interest at lower resolutions, wearable devices efficiently perform computer vision tasks with reduced power and computational load, addressing the balance between performance and power consumption in mobile devices.

JP7756791B2Active Publication Date: 2025-10-20GOOGLE LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024508516
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-08-11
Filing Date
2022-08-12
Publication Date
2025-10-20
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

Computer vision techniques on mobile devices are costly in terms of processing and power consumption, necessitating a balance between performance and power that existing systems struggle to achieve.

Method used

A system for wearable computing devices that identifies regions of interest (ROIs) in image data, processing them at lower resolutions to reduce computational and power demands, using dual-stream image sensors and machine learning algorithms to perform tasks like object detection and OCR on-board without requiring additional resources.

Benefits of technology

Enables effective computer vision tasks with reduced computational and power consumption, allowing wearable devices to perform complex image processing without assistance from other devices, while maintaining accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007756791000001
    Figure 0007756791000001
  • Figure 0007756791000002
    Figure 0007756791000002
  • Figure 0007756791000003
    Figure 0007756791000003
Patent Text Reader

Abstract

Systems and methods for generating image content are described. The systems and methods can include detecting a first sensor data stream having a first image resolution and detecting a second sensor data stream having a second image resolution in response to receiving a request to have a sensor of a computing device identify image content associated with light data captured by the sensor. The systems and methods can also include identifying, by processing circuitry of the computing device, at least one region of interest in the first sensor data stream, determining cropping coordinates that define a first plurality of pixels of the at least one region of interest in the first sensor data stream, and generating a cropped image representative of the at least one region of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of and claims priority to U.S. Patent Application No. 17 / 819,170, filed August 11, 2022, and is a continuation of and claims priority to U.S. Patent Application No. 17 / 819,154, both of which in turn claim the benefit of U.S. Provisional Patent Application No. 63 / 260,206 and U.S. Provisional Patent Application No. 63 / 260,207, filed August 12, 2021. The disclosures of all of these applications are incorporated herein by reference in their entireties.

[0002] This application also claims priority to U.S. Provisional Patent Application No. 63 / 260,206, filed August 12, 2021, and U.S. Provisional Patent Application No. 63 / 260,207, filed August 12, 2021, the disclosures of which are incorporated herein by reference in their entireties.

[0003] Technical Field This description relates generally to methods, devices and algorithms used to process image content. [Background technology]

[0004] background Computer vision techniques enable computers to analyze images and extract information from them. Such computer vision techniques can be costly in terms of processing and power consumption. As mobile computing devices face increasing demands for a balance between performance and power consumption, device manufacturers strive to configure their devices to balance image degradation and device performance to avoid overtaxing the limitations of the mobile computing devices. Summary of the Invention

[0005] overview A system of one or more computers can be configured to perform particular operations or actions by installing software, firmware, hardware, or a combination thereof on the system that, during operation, causes the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by containing instructions that, when executed by a data processing device, cause the device to perform the actions.

[0006] In one general aspect, an image processing method is described. The method can include, in response to receiving a request to have a sensor of a wearable or non-wearable computing device identify image content associated with light data captured by the sensor, detecting a first sensor data stream based on the light data having a first image resolution and detecting a second sensor data stream based on the light data having a second image resolution.

[0007] The method may further include identifying, by processing circuitry of the wearable or non-wearable computing device, at least one region of interest in the first sensor data stream, determining, by the processing circuitry, cropping coordinates that define a first plurality of pixels of the at least one region of interest in the first sensor data stream, and generating, by the processing circuitry, a cropped image representing the at least one region of interest. The generating may include identifying a second plurality of pixels in the second sensor data stream using the cropping coordinates that define the first plurality of pixels of the at least one region of interest in the first sensor data stream and cropping the second sensor data stream to the second plurality of pixels.

[0008] Implementations may include any of the following features, singly or in combination: A method may also include performing, by processing circuitry, optical character resolution on the cropped image to generate a machine-readable version of the cropped image, performing, by processing circuitry, a search query using the machine-readable version of the cropped image to generate a plurality of search results, and displaying the search results on a display of a wearable or non-wearable computing device.

[0009] In some implementations, the second sensor data stream is stored in memory on the wearable or non-wearable computing device, and the method may further include, in response to identifying the at least one region of interest in the first sensor data stream, searching for a corresponding at least one region of interest in the second sensor data stored in memory, and restricting access to the first sensor data stream while continuing to detect and access the second sensor data stream.

[0010] In some implementations, the method may further include, in response to generating a cropped image representing the at least one region of interest, transmitting the generated cropped image to a mobile device in communication with the wearable or non-wearable computing device, receiving information about the at least one region of interest from the mobile device, and displaying the information on a display of the wearable or non-wearable computing device.

[0011] In some embodiments, the computing device is a battery-powered computing device, and the first image resolution has a low image resolution and the second image resolution has a high image resolution. In some embodiments, identifying the at least one region of interest further includes identifying text or at least one object represented in the first sensor data by a machine learning algorithm running on the wearable or non-wearable computing device and using the first sensor data as input.

[0012] In some embodiments, the processing circuitry includes at least a first image processor configured to perform image signal processing on the first sensor data stream and a second image processor configured to perform image signal processing on the second sensor data stream, wherein the first image resolution of the first sensor data stream is lower than the second image resolution of the second sensor data stream.

[0013] In some embodiments, generating the cropped image representative of the at least one region of interest is performed in response to detecting that the at least one region of interest meets a threshold condition, the threshold condition including detecting low blur of a second plurality of pixels in the second sensor data stream.

[0014] In a second general aspect, a wearable computing device is described that includes at least one processing device, at least one image sensor configured to capture light data, and a memory that stores instructions that, when executed, cause the wearable computing device to perform operations, including detecting a first sensor data stream based on the light data, having a first image resolution, and detecting a second sensor data stream based on the light data, having a second image resolution, in response to receiving a request to have the at least one image sensor identify image content associated with the light data.

[0015] The operations may further include identifying at least one region of interest in the first sensor data stream, determining cropping coordinates that define a first plurality of pixels of the at least one region of interest in the first sensor data stream, and generating a cropped image representing the at least one region of interest, wherein the generating includes identifying a second plurality of pixels in the second sensor data stream using the cropping coordinates that define the first plurality of pixels of the at least one region of interest in the first sensor data stream, and cropping the second sensor data stream to the second plurality of pixels.

[0016] Implementations may include any of the following features, alone or in combination: In some implementations, the at least one image sensor is a dual-stream image sensor configured to operate in a low image resolution mode until triggered to switch to operation in a high image resolution mode; In some implementations, the operations further include performing optical character resolution on the cropped image to generate a machine-readable version of the cropped image, performing a search query using the machine-readable version of the cropped image to generate a plurality of search results, and displaying the search results on a display of the wearable computing device; In some implementations, the wearable computing device may cause the performed optical character resolution to be output audibly through a speaker of the wearable computing device.

[0017] In some embodiments, the second sensor data stream is stored in memory on the wearable computing device, and the operations further include, in response to identifying the at least one region of interest in the first sensor data stream, searching for a corresponding at least one region of interest in the second sensor data stored in memory, and restricting access to the first sensor data stream while continuing to detect and access the second sensor data stream.

[0018] In some embodiments, the operations further include, in response to generating a cropped image representing the at least one region of interest, transmitting the generated cropped image to a mobile device in communication with the wearable computing device, receiving information about the at least one region of interest from the mobile device, and displaying the information on a display of the wearable computing device. In some embodiments, the operations include outputting information (e.g., audio, visual, tactile, etc.) at the wearable computing device.

[0019] In some implementations, the first image resolution has a low image resolution and the second image resolution has a high image resolution, and the operations can further include identifying at least one region of interest, wherein identifying the at least one region of interest further includes identifying a sentence or at least one object represented in the first sensor data by a machine learning algorithm running on the wearable computing device and using the first sensor data as input.

[0020] In some implementations, the at least one processing device includes a first image processor configured to perform image signal processing on the first sensor data stream and a second image processor configured to perform image signal processing on the second sensor data stream, wherein the first image resolution of the first sensor data stream is lower than the second image resolution of the second sensor data stream.

[0021] Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium. Details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 1 illustrates an example of a wearable computing device for generating and processing image content using limited computational and / or power resources, according to implementations described throughout this disclosure. [Figure 2] FIG. 1 illustrates a system for performing image processing on a wearable computing device, according to implementations described throughout this disclosure. [Figure 3A] FIG. 1 illustrates an example of a wearable computing device according to implementations described throughout this disclosure. [Figure 3B] FIG. 1 illustrates an example of a wearable computing device according to implementations described throughout this disclosure. [Figure 4A] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 4B] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 4C] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 5A] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 5B] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 6A] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 6B]1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 6C] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 7A] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 7B] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 7C] 1 is an exemplary flow diagram for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 8A] 10 is an exemplary flow diagram for emulating a dual resolution image sensor using a single image sensor to perform image processing tasks, according to implementations described throughout this disclosure. [Figure 8B] 10 is an exemplary flow diagram for emulating a dual resolution image sensor using a single image sensor to perform image processing tasks, according to implementations described throughout this disclosure. [Figure 9] 1 is a flowchart illustrating an example of a process for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. [Figure 10] 1A-1C illustrate examples of computing devices and mobile computing devices that can be used with the techniques described herein. DETAILED DESCRIPTION OF THE INVENTION

[0023] Detailed Description Like reference symbols in the various drawings indicate like elements.

[0024] This disclosure describes systems and methods for performing image processing that can enable wearable computing devices to effectively perform computer vision tasks. For example, the systems and methods described herein can utilize processors, sensors, neural networks, and / or image analysis algorithms to recognize text, symbols, and / or objects in images and / or extract information from images. In some implementations, the systems and methods described herein can perform such tasks while operating in reduced computational and / or low-power modes. For example, the systems and methods described herein can minimize the amount of image data used in performing computer vision and image analysis tasks, reducing the use of certain hardware and / or device resources (e.g., memory, processor, network bandwidth, etc.).

[0025] Configuring wearable computing devices to use fewer resources can provide the advantage of enabling such devices to perform relatively low-power computational tasks on-board the device without sending the tasks to other devices (e.g., servers, mobile devices, other computers, etc.). For example, instead of analyzing the entire image, the systems and methods described herein can identify regions of interest (ROIs) in the image data and analyze one or more of these regions on-board the wearable computing device. Analyzing a smaller amount of information while still maintaining accurate results can enable complex image processing tasks to be performed on the wearable computing device without the resource burden of analyzing the entire image.

[0026] In some implementations, the systems and methods described herein can be configured to run on devices to ensure that such devices can perform effective image processing tasks while reducing computational load and / or power. For example, the systems and methods described herein may enable wearable devices to perform object detection tasks, optical character recognition (OCR) tasks, and / or other image processing tasks while utilizing certain techniques to reduce power, memory, and / or processing consumption.

[0027] In some implementations, the systems and methods described herein can ensure that complex image processing computations can be performed on a wearable device without assistance from other resources and / or devices. For example, conventional systems may require the assistance of other communicatively coupled mobile devices, servers, and / or off-board systems to perform computationally intensive image processing tasks. The systems and methods described herein provide the advantage of generating portions of an image (e.g., ROIs, objects, cropped image content, etc.) that can be processed by devices with less computational power than, for example, a server, while still providing complete and accurate image processing capabilities.

[0028] The systems and methods described herein may enable wearable computing devices to use machine learning intelligence (e.g., neural networks, algorithms, etc.) to perform computer vision tasks such as object detection, motion tracking, face recognition, OCR tasks, etc., with low power consumption and / or low processing consumption.

[0029] In some implementations, the systems and methods described herein may use scalable levels of image processing to utilize fewer processing and / or power resources. For example, a low-resolution image signal processor may be used to perform low-level image processing tasks on a low-resolution image stream, while a high-resolution image signal processor may be used to perform high-level image processing tasks on a high-resolution image stream. In some implementations, output from the low-resolution image signal processor may be provided as input to additional sensors, detectors, machine learning networks, and / or processors. Such output may represent portions of an image (e.g., pixels, regions of interest, recognized image content, etc.) that can be used to identify content in a higher-resolution image, including image content representing corresponding pixels, regions of interest, and / or recognized image content from the low-resolution image.

[0030] In some implementations, the systems and methods described herein can utilize one or more sensors onboard the wearable computing device to divide detection and / or processing tasks. For example, the wearable computing device can include a dual-stream sensor (e.g., a camera) that can retrieve both high-resolution and low-resolution image streams from multiple images. For example, the low-resolution image stream can be captured by a sensor that functions as a scene camera. The high-resolution image stream can be captured by the same sensor that functions as a detail camera. One or both streams can be provided to different processors to perform image processing tasks. In some implementations, the individual streams can be used to generate different information for use in performing computer vision tasks, machine learning tasks, and / or other image processing tasks.

[0031] 1 is an example of a wearable computing device 100 for generating and processing image content using limited computational and / or power resources, according to implementations described throughout this disclosure. In this example, the wearable computing device 100 is shown in the form of AR smart glasses. However, battery-powered devices of any form factor can be substituted and combined with the systems and methods described herein. In some implementations, the wearable computing device 100 includes a system-on-chip (SOC) architecture (not shown) combined with any or all of one or more sensors, low-power island processors, high-power island processors, core processors, encoders, etc.

[0032] During operation, the wearable computing device 100 can capture a scene 102 via a camera, image sensor, etc. The wearable computing device 100 can be worn by and operated by a user 104. The scene 102 can include physical content as well as augmented reality (AR) content. In some implementations, the wearable computing device 100 can be communicatively coupled to other devices, such as a mobile computing device 106.

[0033] In operation, the mobile computing device 106 may include a dual-stream image sensor 108 (e.g., a dual-resolution image sensor). The dual-stream image sensor 108 may provide a low-resolution image stream 110 to a low-resolution (e.g., lightweight) image signal processor 112. The dual-stream image sensor 108 may simultaneously provide a high-resolution image stream 114 to a buffer 116. The high-resolution image signal processor 118 may obtain the high-resolution image stream 114 from the buffer 116 for further processing. The low-resolution image stream 110 may be provided to an ML model that identifies one or more ROIs (e.g., sentences, objects, hands, paragraphs, etc.) via, for example, an ROI detector 120. The processor 112 may determine ROI coordinates 122 (e.g., cropping coordinates) associated with the one or more ROIs.

[0034] The high-resolution image signal processor 112 can receive (or request) the ROI coordinates 122 and can use the ROI coordinates 122 to crop one or more high-resolution images from the high-resolution image stream 114. For example, the processor 112 can search the ROI coordinates 122 associated with the low-resolution image stream 110 to determine a frame and region in the high-resolution image stream 114 that matches the same capture time as the associated low-resolution image frame in which the ROI was identified. The high-resolution image signal processor 118 can then crop the image to the same cropping ROI coordinates 122. The cropped image 124 can be provided to an ML engine (e.g., ML processor 126, for example, to perform computer vision tasks including, but not limited to, OCR, object recognition, deep segmentation, hand detection, face detection, symbol detection, etc.).

[0035] In a non-limiting example, wearable computing device 100 can be triggered to begin real-time image processing in response to a request to identify image content associated with light data captured by dual-stream image sensor 108. The request may come from a user wearing wearable computing device 100. Real-time image processing can begin by detecting a first sensor data stream (e.g., low-resolution image stream 110) and a second sensor data stream (e.g., high-resolution image stream 114). Wearable computing device 100 can then identify at least one region of interest (e.g., ROI 128 and ROI 130). In this example, ROI 128 includes an analog clock, and ROI 130 includes presentation text and images. Device 100 can analyze the first stream (e.g., low-resolution image stream 110) to determine cropping coordinates for ROIs 128 and 130.

[0036] The wearable computing device 100 can obtain the high-resolution image stream 114 from the buffer 116 and identify the same corresponding (e.g., same capture time within a frame, same cropping coordinates) ROIs 128 and 130 in the high-resolution image stream 114 to crop one or more image frames of the high-resolution image stream 114 that correspond to the ROIs 128 and 130 from the low-resolution image stream 110. The cropped one or more image frames can be provided to the onboard ML processor 126 to perform additional processing of the cropped images. Cropping an image to provide less image content (e.g., less visual sensor data, a portion of the image, etc.) for image analysis can result in ML processing being performed using less power and lower latency than traditional ML processors that require analysis of the entire image. Thus, the wearable computing device 100 can perform ML processing in real time for tasks such as OCR, object recognition, deep paragraphing, hand detection, face detection, symbol detection, etc.

[0037] For example, wearable computing device 100 can identify ROI 130, perform image cropping, and use ML processing to determine that user 104's hand is pointing at the image and text in ROI 130. The processing can perform hand detection and OCR on ROI 130. Similarly, if audio content is available during a presentation, wearable computing device 100 can evaluate ROI 128 with respect to the timestamp at which user 104 pointed to ROI 130. The timestamp (e.g., 9:05) can be used to retrieve and analyze audio data (e.g., a transcript) associated with ROI 130. Such audio data can be converted into visual text for presentation on the display of wearable computing device 100. Thus, ML processor 126 can select the ROI but can also correlate the ROI as shown by output 132. Output 132 can be presented on the display of wearable computing device 100.

[0038] Although multiple processor blocks (e.g., processor 112, processor 118, processor 126, etc.) are shown, a single processor may be utilized to perform all processing tasks on wearable computing device 100. That is, individual processor blocks 112, 118, and 126 may represent different algorithms or code fragments that may run on a single processor on-board wearable computing device 100.

[0039] In some implementations, the object and / or ROI may be cropped by dual-stream image sensor 108 rather than being cropped by processor 112 or 118. In some implementations described in detail above, the object and / or ROI may be cropped by either image signal processor 112 or image signal processor 118. In some implementations, instead of cropping, the ROI may be prepared and sent to a mobile device in communication with wearable computing device 100. The mobile device may crop the ROI and provide the cropped image back to wearable computing device 100.

[0040] 2 illustrates a system 200 for performing image processing on a wearable computing device 100, according to implementations described throughout this disclosure. In some implementations, image processing is performed on the wearable computing device 100. In some implementations, image processing is shared among one or more devices. For example, image processing may be completed partially on the wearable computing device 100 and partially on the mobile computing device 202 (e.g., mobile computing device 106) and / or the server computing device 204. In some implementations, image processing is performed on the wearable computing device 100, and output from such processing is provided to the mobile computing device 202 and / or the server computing device 204.

[0041] In some implementations, wearable computing device 100 includes one or more computing devices, at least one of which is a display device that can be worn or placed near a person's skin. In some examples, wearable computing device 100 is or includes one or more wearable computing device components. In some implementations, wearable computing device 100 can include a head-mounted display (HMD) device, such as an optical head-mounted display (OHMD) device, a transparent head-up display (HUD) device, a virtual reality (VR) device, an AR device, or other devices, such as goggles or a headset, that have sensors, a display, and computing capabilities. In some implementations, wearable computing device 100 includes AR glasses (e.g., smart glasses). AR glasses refer to an optical head-mounted display device designed in the form of a pair of glasses. In some implementations, wearable computing device 100 is or includes a smartwatch. In some embodiments, wearable computing device 100 is or includes a piece of jewelry. In some embodiments, wearable computing device 100 is or includes a ring controller device or other wearable controller. In some embodiments, wearable computing device 100 is or includes earphones / headphones or smart earphones / headphones.

[0042] 2, system 200 includes wearable computing device 100 communicatively coupled to a mobile computing device 202 and optionally a server computing device 204. In some implementations, the communicative coupling may occur over a network 206. In some implementations, the communicative coupling may occur directly between wearable computing device 100, mobile computing device 202, and / or server computing device 204.

[0043] The wearable computing device 100 includes one or more processors 208, which may be formed in a substrate configured to execute one or more machine-executable instructions, or pieces of software, firmware, or a combination thereof. The processor 208 may be a semiconductor-based processor and may include semiconductor material capable of implementing digital logic. The processor 208 may include a CPU, a GPU, and / or a DSP, to name just a few examples.

[0044] The wearable computing device 100 may also include one or more memory devices 210. The memory device 210 may include any type of storage device that stores information in a format that can be read and / or executed by the processor 208. The memory device 210 may store applications and modules that, when executed by the processor 208, perform certain operations. In some examples, applications and modules may be stored on an external storage device and loaded into the memory device 210. The memory 210 may include or have access to a buffer 212, for example, for storing and retrieving image and / or audio content for the wearable computing device 100.

[0045] The wearable computing device 100 includes a sensor system 214. The sensor system 214 includes one or more image sensors 216 configured to detect and / or acquire image data. In some implementations, the sensor system 214 includes multiple image sensors 216. As shown, the sensor system 214 includes one or more image sensors 216. The image sensors 216 can capture and record images (e.g., pixels, frames, and / or portions of images) and video.

[0046] In some implementations, image sensor 216 is a red, green, and blue (RGB) camera. In some examples, image sensor 216 includes a pulsed laser sensor (e.g., a LiDAR sensor) and / or a depth camera. For example, image sensor 216 may be a camera configured to detect and convey information used to construct the image represented by image frame 226. Image sensor 216 can capture and record both images and video.

[0047] In operation, image sensor 216 is configured to continuously or periodically acquire (e.g., capture) image data (e.g., optical sensor data) while wearable computing device 100 is powered on. In some implementations, image sensor 216 is configured to operate as an always-on sensor. In some implementations, imaging sensor 216 can be powered on in response to detecting an object or region of interest.

[0048] In some implementations, image sensor 216 is a single sensor configured to function in a dual-resolution streaming mode. For example, image sensor 216 can operate in a low-resolution streaming mode to capture (e.g., receive, acquire, etc.) multiple images having a low image resolution. Sensor 216 can simultaneously operate in a high-resolution streaming mode to capture (e.g., receive, acquire, etc.) the same multiple images having a high image resolution. The dual mode can capture or retrieve multiple images as a first image stream having a low image resolution and a second image stream having a high image resolution.

[0049] In some implementations, wearable computing device 100 may use scalable levels of image processing using image sensor 216 to utilize fewer processing and / or power resources. For example, the systems and methods described herein may use low-level image processing, mid-level image processing, high-level image processing, and / or any combination between these levels, and / or any combination of these levels by using two or more different types of image signal processors. In some implementations, the systems and methods described herein may also vary the resolution of the image when implementing image processing techniques.

[0050] As used herein, high-level processing includes image processing on high-resolution image content. For example, computer vision techniques for performing content (e.g., object) recognition can be considered high-level processing. Mid-level image processing includes the processing task of deriving a scene description from image metadata, e.g., image metadata can indicate the location and / or shape of parts of a scene. Low-level image processing includes deriving a description from an image while ignoring scene objects and observer details.

[0051] The sensor system 214 may also include an inertial motion unit (IMU) sensor 218. The IMU sensor 218 may detect the motion, movement, and / or acceleration of the wearable computing device 100. The IMU sensor 218 may include a variety of different types of sensors, such as, for example, an accelerometer, a gyroscope, a magnetometer, and other such sensors.

[0052] In some implementations, sensor system 214 may also include an audio sensor 220 configured to detect audio received by wearable computing device 100. Sensor system 214 may include other types of sensors, such as optical sensors, distance and / or proximity sensors, contact sensors such as capacitive sensors, timers, and / or other sensors and / or different combinations of sensors. Sensor system 214 may be used to obtain information associated with the position and / or orientation of wearable computing device 100.

[0053] The wearable computing device 100 may also include a low-resolution image signal processor 222 for processing images from a low-resolution image stream (e.g., low-resolution image stream 110). The low-resolution image signal processor 222 may also be referred to as a low-power image signal processor. The wearable computing device 100 may also include a high-resolution image signal processor 224 for processing images from a high-resolution image stream (e.g., high-resolution image stream 114). The high-resolution image signal processor 224 may also be referred to as a high-power image signal processor (as opposed to the low-resolution image signal processor 222). The image stream may include multiple image frames 226, each having a particular image resolution 228. During operation, the image sensor 216 may be dual-functional to operate in either or both a low-power, low-resolution (LPLR) mode and / or a high-power, high-resolution (HPLR) mode.

[0054] The wearable computing device 100 may include a region of interest (ROI) detector 230 configured to detect an ROI 232 and / or an object 234. Additionally, the wearable computing device 100 may include a cropper 236 configured to crop image content (e.g., image frame 226) and the ROI 232 according to cropping coordinates 238. In some implementations, the wearable computing device 100 includes an encoder 240 configured to compress particular pixel blocks, image content, etc. The encoder 240 may receive image content for compression, for example, from a buffer 212.

[0055] The low-resolution image signal processor 222 can perform low-power computations to analyze the generated stream of images from the sensor 216 and detect objects and / or regions of interest within the images. Detection can include hand detection, object detection, paragraph detection, etc. Once an object and / or ROI of interest is detected, and the detection of such a region or object meets a threshold condition 256, a bounding box and / or other coordinates associated with the object and / or region can be identified. The threshold condition can represent a particular high level of quality for an object or region in an image frame of the stream of images. For example, the threshold condition may relate to any or all of the determinations of being focused on the object or region, remaining low-blur (e.g., low motion blur), and / or having adequate exposure, and / or having a particular object detected within the view of the computing device.

[0056] Wearable computing device 100 may also include one or more antennas 242 configured to communicate with other computing devices via wireless signals. For example, wearable computing device 100 may receive one or more wireless signals and use the wireless signals to communicate with other devices, such as mobile computing device 202 and / or server computing device 204, or other devices within range of antenna 242. The wireless signals may be triggered via a wireless connection, such as a short-range connection (e.g., a Bluetooth connection or a near-field communication (NFC) connection) or an Internet connection (e.g., Wi-Fi or a mobile network).

[0057] Wearable computing device 100 includes a display 244. Display 244 can include a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting display (OLED), an electrophoretic display (EPD), or a micro-projection display employing an LED light source. In some examples, display 244 is projected onto a user's field of view. In some examples, in the case of AR glasses, display 244 can provide a transparent or translucent display so that a user wearing the AR glasses can not only see the image provided by display 244 but also see information located in the field of view of the AR glasses behind the projected image.

[0058] Wearable computing device 100 also includes a control system 246 that includes various control system devices for facilitating operation of wearable computing device 100. Control system 246 may utilize a processor 248 (e.g., CPU, GPU, DSP, etc.) operably coupled to processor 208, sensor system 214, and / or components of wearable computing device 100.

[0059] Wearable computing device 100 also includes a UI renderer 250, which can work in conjunction with display 244 to render user interface objects or other content to a user of wearable computing device 100. For example, UI renderer 250 can receive images captured by wearable computing device 100 and generate and render additional user interface content on display 244.

[0060] Wearable computing device 100 also includes a communications module 252. Communications module 252 enables wearable computing device 100 to communicate to exchange information with another computing device within range of wearable computing device 100. For example, wearable computing device 100 may be operatively coupled to another computing device to facilitate communication, e.g., via a wired connection, a wireless connection, e.g., via Wi-Fi or Bluetooth, or other type of connection.

[0061] In some implementations, the wearable computing device 100 is configured to communicate with a server computing device 204 and / or the mobile computing device 202 over a network 206. The server computing device 204 can represent one or more computing devices in the form of several different devices, such as a standard server, a group of such servers, or a rack server system. In some implementations, the server computing device 204 is a single system that shares components such as a processor and memory. The network 206 can include the Internet and / or other types of data networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, or other types of data networks. The network 206 can also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within the network 206.

[0062] In some implementations, the wearable computing device 100 also includes a neural network NN254 that performs machine learning tasks onboard the wearable computing device 100. The neural network can be used in conjunction with machine learning models and / or operations to generate, modify, predict, and / or detect specific image content. In some implementations, image processing performed on sensor data acquired by the sensor 216 is referred to as a machine learning (ML) inference operation. An inference operation can be referred to as an image processing act, step, or sub-step that requires an ML model to make one or more predictions (or yield one or more predictions). Certain types of image processing performed by the wearable computing device 100 can be predicted using an ML model. For example, machine learning can use statistical algorithms that learn from existing data to render decisions about new data, a process called inference. In other words, inference refers to the process of taking an already trained model and using that trained model to make predictions. Some examples of inference may include image recognition (e.g., OCR text recognition, facial recognition, face, body or object tracking, etc.), and / or cognition (e.g., always-on voice perception, voice-prompt voice perception, etc.).

[0063] In some implementations, the ML model includes one or more NNs 254. The NNs 254 transform inputs received by an input layer through a series of hidden layers and generate outputs through an output layer. Each layer is constructed with a subset of the set of nodes. The nodes in a hidden layer are fully connected to all nodes in the preceding layer and provide their respective outputs to all nodes in the adjacent layer. The nodes in a single layer function independently of each other (i.e., do not share connections). The nodes in the output provide the transformed inputs to the requesting process. In some implementations, the NNs 254 utilized by the wearable computing device 100 are convolutional neural networks, which are not fully connected NNs.

[0064] Typically, the wearable computing device 100 can generate training data and perform ML tasks using a neural network (NN) and a classifier (not shown). For example, the processor 208 can be configured to execute a classifier that detects, for example, whether an ROI 232 is contained within an image frame 226 captured by the image sensor 216. The ROI 232 can also be referred to as an object of interest. The classifier can include or be defined by an ML model. The ML model can define several parameters used by the ML model to make a prediction (e.g., whether the ROI 232 is contained within the image frame 226). The ML model can be relatively small and can be configured to operate on an ROI rather than the entire image. Thus, the classifier can be configured to operate in a manner that conserves power and reduces latency through its relatively small ML model.

[0065] In some implementations, wearable computing device 100 can use an ROI classifier that executes an ML model to compute an ROI dataset (e.g., ROI 232). In some examples, ROI 232 represents an example of object location data and / or a bounding box dataset. ROI 232 can be data that defines where an ROI is located within image frame 226.

[0066] During operation, the wearable computing device 100 can receive or capture image content (e.g., multiple images / frames) with a dual resolution image sensor (e.g., image sensor 216). The image content may be low-resolution images that are sent to the low-resolution image signal processor 222 while also sending high-resolution images of the multiple images / frames to the high-resolution image signal processor 224. The low-resolution images can then be sent to an on-board machine learning model, such as a neural network 254. The NN 254 ​​can utilize an ROI detector 230 to identify one or more ROIs (e.g., text, objects, images, frames, pixels, etc.). The identified ROIs can each be associated with determined cropping coordinates 238. The wearable computing device 100 can transmit such cropping coordinates 238 to the high-resolution image signal processor 224 to crop the high-resolution image based on the crop coordinates of the low-resolution image corresponding to the detected ROI. The cropped image is smaller than the original high-resolution image, and therefore the wearable computing device 100 can use the NN 254 ​​to perform machine learning activities (e.g., OCR, object detection, etc.) on the cropped image.

[0067] In some implementations, wearable computing device 100 can perform other vision-based algorithms on the cropped region. For example, object recognition can be performed on the cropped region to identify and utilize information associated with the recognized object. For example, specific products, landmarks, storefronts, factories, animals, and / or various other general objects and / or object classifiers can be recognized and used as input to determine additional information for the user. In some implementations, barcode recognition and reading can be performed on the cropped region to identify and provide additional information on wearable computing device 100, which can be displayed, provided audibly, or other output. In some implementations, face detection can be performed on the cropped region to enable AR and / or VR content to be placed on a recognized face, capturing specific features of the recognized face, counting specific facial features, etc. In some implementations, feature tracking (e.g., feature point detection) can be performed on the cropped region to track object and / or user movement.

[0068] 3A and 3B illustrate various views of an example AR wearable computing device according to embodiments described throughout this disclosure. FIG. 3A is a front view of an example wearable computing device according to embodiments described throughout this disclosure. In this example, the wearable computing device is a pair of AR glasses 300A (e.g., wearable computing device 100 of FIG. 1). Typically, AR glasses 300A can include any or all of the components of system 200. AR glasses 300A can also be referred to as smart glasses, which represent an optical head-mounted display device designed in the shape of a pair of glasses. For example, smart glasses are glasses that add information (e.g., project a display) alongside what the wearer is looking at through the glasses.

[0069] Although AR glasses 300A are shown as the wearable computing device described herein, other types of wearable computing devices are possible. For example, a wearable computing device can include any battery-powered device, including, but not limited to, a head-mounted display (HMD) device such as an optical head-mounted display (OHMD) device, a transparent head-up display (HUD) device, an augmented reality (AR) device, or other devices such as goggles or headsets having sensors, displays, and computing capabilities. In some examples, the wearable computing device 300A can be a watch, a mobile device, a piece of jewelry, a ring controller, or other wearable controller.

[0070] 3A , the AR glasses 300A include a frame 302 having a display device 304 coupled thereto (or within the glass portion of the frame 302). The AR glasses 300A also include an audio output device 306, an illumination device 308, a perception system 310, a control system 312, at least one processor 314, and a camera 316.

[0071] The display device 304 can include a see-through near-eye display, such as a birdbath or a device used in waveguide optics. For example, such an optical design can project light from a display source onto a portion of teleprompter glass that acts as a beam splitter seated at a 45-degree angle. The beam splitter can allow for reflection and transmission values ​​that can partially reflect light from the display source and transmit the remaining light. Such an optical design allows the user to see both physical items in the world next to the digital images (e.g., UI elements, virtual content, etc.) generated by the display. In some implementations, waveguide optics can be used to depict content on the display device 304 of the AR glasses 300A.

[0072] An audio output device 306 (e.g., one or more speakers) can be coupled to the frame 302. The perception system 310 can include various perception devices and a control system 312, including various control system devices for facilitating operation of the AR glasses 300A. The control system 312 can include a processor 314 operably coupled to the components of the control system 312.

[0073] The camera 316 can capture still and / or video images. In some implementations, the camera 316 can be a depth camera that can collect data related to the distance of external objects from the camera 316. In some implementations, the camera 316 can be a point-tracking camera that can detect and follow, for example, one or more light markers on an external device, such as light markers on an input device, or a finger on a screen. In some implementations, the AR glasses 300A can include an illumination device 308 that can be selectively operated, for example, in conjunction with the camera 316, to detect objects (e.g., virtual and physical) within the field of view of the camera 316. The illumination device 308 can be selectively operated, for example, in conjunction with the camera 316, to detect objects within the field of view of the camera 316.

[0074] The AR glasses 300A may include a communications module (e.g., communications module 252) that communicates with the processor 314 and the control system 312. The communications module may provide for communications between devices housed within the AR glasses 300A as well as with external devices, such as a controller, a mobile device, a server, and / or other computing devices. The communications module may enable the AR glasses 300A to communicate and exchange information with another computing device and to authenticate other devices within range of the AR glasses 300A or other identifiable elements in the environment. For example, the AR glasses 300A may be operably coupled to another computing device to facilitate communications, for example, via a wired connection, a wireless connection, such as via Wi-Fi or Bluetooth, or another type of connection.

[0075] FIG. 3B is a rear view 300B of AR glasses 300A according to embodiments described throughout this disclosure. The AR glasses 300B may be an example of the wearable computing device 100 of FIG. 1. The AR glasses 300B are glasses that add information (e.g., project a display 320) alongside what the wearer is viewing through the glasses. In some embodiments, instead of projecting information, the display 320 is an in-lens microdisplay. In some embodiments, the AR glasses 300B (e.g., eyeglasses or spectacles) are vision aids that include lenses 322 (e.g., glass lenses or hard plastic lenses) attached to a frame 302 that holds the lenses in front of a person's eyes, typically utilizing a bridge 324 that rests on the nose and temples 326 (e.g., temples or temple pieces) that rest on the ears.

[0076] 4A-4C illustrate example flowcharts 400A, 400B, and 400C for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. For example, flowcharts 400A-400C may be implemented by wearable computing device 100 using one or more image sensors 216. In the example of FIGS. 400A-400C, sensor 402 may function in a dual-stream mode, where sensor 402 continuously and simultaneously processes both an incoming (i.e., captured) low-resolution image stream and an incoming (i.e., captured) high-resolution image stream.

[0077] 4A is a flow diagram 400A illustrating the use of an image sensor 402 installed in a wearable computing device and configured to capture or receive a stream of image data. The sensor 402 can capture and output two streams of image content. For example, the sensor 402 can acquire multiple high-resolution image frames captured using a full field of view (e.g., the maximum resolution of the camera / sensor 402). The multiple high-resolution image frames can be stored in a buffer 404. Typically, a high-power image signal processor 406 can optionally be utilized to perform analysis of the high-resolution image frames.

[0078] Additionally, the sensor 402 can acquire multiple low-resolution image frames (simultaneously) that are captured using the full field of view. The low-resolution image frames can be provided to a low-power image signal processor 408. The processor 408 processes the low-resolution image frames to generate a stream (e.g., a YUV-formatted image stream) of images 410 (e.g., one or more images), which are de-bayered, color-corrected, and shade-corrected images.

[0079] The low-power image signal processor 408 may perform low-power computations 412 to analyze the generated stream of images 410, thereby detecting objects and / or regions of interest 414 within the images 410. Detection may include hand detection, object detection, paragraph detection, etc. When an object and / or ROI of interest is detected, and the detection of such a region or object meets a threshold condition, a bounding box for the object and / or region is identified. The threshold condition may represent a particular high level of quality for the object or region in the image frames of the stream of images. For example, the threshold condition may relate to any or all determinations that the object or region is focused, remains low-blur, and / or has adequate exposure. The bounding box may represent cropping coordinates for each of the four corners of the box. In some implementations, the bounding box may not be rectangular, but may instead be triangular, circular, elliptical, or other shape.

[0080] 4B is a flow diagram 400B illustrating the use of an image sensor 402 installed in a wearable computing device and configured to capture or receive a stream of image data. Similar to FIG. 4A, the sensor 402 can capture and output two streams of image content: a plurality of high-resolution image frames and a plurality of low-resolution image frames. The plurality of high-resolution image frames can be provided to a buffer 404.

[0081] The low-resolution image frames may be provided to a low-power image signal processor 408. The processor 408 processes the low-resolution image frames to generate a stream (e.g., a YUV-formatted image stream) of images 410 (e.g., one or more images), which are de-bayered, color corrected, and shade corrected images.

[0082] The low-power image signal processor 408 may perform low-power computations 412 to analyze the generated stream of images 410, thereby detecting objects and / or regions of interest 414 within the images 410. Once an object and / or ROI of interest is detected, and the detection of such a region or object meets a threshold condition, a bounding box for the object and / or region is identified. The threshold condition may represent a certain high level of quality for the object or region in the image frames of the stream of images. For example, the threshold condition may relate to any or all determinations that the object or region is focused, remains low-blur, and / or has adequate exposure. The bounding box may represent cropping coordinates representing the region or object of interest 414. Once the object or region of interest 414 has been determined and the threshold condition has been met, the high-power image signal processor 406 retrieves (or receives) a high-resolution image frame from the buffer 404 (e.g., a raw frame that coincides with the same capture time as the low-resolution image frame in which the object or region of interest was identified).

[0083] For example, the cropping coordinates and the retrieved high-resolution image frame may be provided as input to a high-power image signal processor 406 (e.g., arrow 416). The high-power image signal processor 406 may output a processed (e.g., de-bayered, color corrected, shade corrected, etc.) image 418 that is cropped to the object and / or region of interest determined from the low-resolution image frame. This cropped full-resolution image 418 may be provided to additional on-board or off-board devices for further processing.

[0084] In some implementations, a single high-resolution image can be processed through a high-power image signal processor rather than multiple image frames, saving computational and / or device power. Because downstream processing can operate on a cropped image 418 rather than the entire high-resolution image, such additional downstream processing may still incur additional computational and / or device power.

[0085] In general, after the object and / or region of interest has been identified, the low-power image signal processor 408 may not continue processing the stream of low-resolution image frames. In some implementations, the low-power image signal processor 408 may be placed in a standby mode to await further quality metrics for previously provided / identified image frames. In some implementations, the sensor 402 may continue to trigger the performance of cognitive functions (i.e., ML processing) on ​​the low-resolution image stream, which may further affect the output to downstream functions.

[0086] FIG. 4C is a flow diagram 400C illustrating an example of additional downstream processing. As shown, a cropped image 418 represents a region of interest. The cropped image 418 is a high-resolution image that can be transmitted to an encoder (e.g., compression) block without buffering the image 418 (i.e., the image) in memory. The compressed output of the encoder is buffered in a buffer 422 and encrypted by an encryption block 424. The encrypted image 418 can be buffered in a buffer 426 and sent to a companion mobile computing device 202, for example, via Wi-Fi (e.g., network 206). The mobile computing device 202 can be configured to include computational power, memory, and battery capacity to support more complex processing than is supported on the wearable computing device 100, for example. Because the additional downstream processing utilizes a cropped version of the image content rather than the entire image, the downstream processing also has the advantage of saving computational, memory, and battery power. For example, the additional packets sent over Wi-Fi for a larger image, rather than the cropped image 418, would use a larger amount of power.

[0087] 5A and 5B illustrate example flowcharts 500A and 500B for performing image processing tasks on a wearable computing device, according to implementations described throughout this disclosure. For example, flowcharts 500A and 500B may be implemented by wearable computing device 100 using one or more image sensors 216. In the example of FIGS. 500A and 500B, sensor 402 may function in a dual-stream mode, where sensor 402 continuously and simultaneously processes both an incoming (i.e., captured) low-resolution image stream and an incoming (i.e., captured) high-resolution image stream.

[0088] FIG. 5A is a flow diagram 500A illustrating the use of an image sensor 402, for example, installed in the wearable computing device 100 and configured to capture or receive streams of image data. The sensor 402 can capture and output two streams of image content. For example, the sensor 402 can acquire multiple high-resolution image frames (e.g., the maximum resolution of the camera / sensor 402) captured using a full field of view. In this example, the multiple high-resolution image frames can be processed by a high-power image signal processor 406. The processed high-resolution image frames can then be encoded (e.g., compressed) by the encoder 502 and stored in the buffer 404. This provides the advantage of saving buffer memory and memory bandwidth over storing the high-resolution image frames prior to compression. This example instead compresses and stores compressed versions of the high-resolution image frames rather than the original uncompressed image frames.

[0089] Additionally, the sensor 402 can acquire multiple low-resolution image frames (simultaneously) that are captured using the full field of view. The low-resolution image frames can be provided to a low-power image signal processor 408. The processor 408 processes the low-resolution image frames to generate a stream (e.g., a YUV-formatted image stream) of images 410 (e.g., one or more images), which are de-bayered, color-corrected, and shade-corrected images.

[0090] The low-power image signal processor 408 may perform low-power computations 412 to analyze the generated stream of images 410, thereby detecting objects and / or regions of interest 414 within the images 410. Detection may include hand detection, object detection, paragraph detection, etc. Once an object and / or ROI of interest is detected, and the detection of such a region or object meets a threshold condition, a bounding box for the object and / or region is identified. The threshold condition may represent a particular high level of quality for the object or region in the image frames of the stream of images. For example, the threshold condition may relate to any or all determinations that the object or region is focused, still has low blur, and / or has adequate exposure. The bounding box may represent, for example, cropping coordinates for each of the four corners of the box.

[0091] 5B is a flow diagram 500B illustrating a process for generating partial image content from, for example, one or more image frames 410. Similar to FIG. 5A, the sensor 402 can capture and output two streams of image content including at least one high-resolution image frame and at least one low-resolution image frame. In this example, the at least one high-resolution image frame can be processed by the high-power image signal processor 406. The processed high-resolution image frame can then be encoded (e.g., compressed) by the encoder 502 and stored in the buffer 404.

[0092] Additionally, the sensor 402 may (simultaneously) acquire at least one low-resolution image frame, which may be provided to a low-power image signal processor 408. The processor 408 processes the at least one low-resolution image frame to generate at least one image 410 (e.g., a YUV-formatted image), which is a de-bayered, color-corrected, and shade-corrected image.

[0093] The low-power image signal processor 408 may perform low-power computations 412 to analyze at least one image 410, thereby detecting, for example, an object and / or region of interest 414 within the at least one image 410. The detection may include hand detection, object detection, paragraph detection, etc. Once the object and / or region of interest is detected, and the detection of such region or object meets a threshold condition, a bounding box for the object and / or region is identified. The threshold condition may represent a particular high level of quality for the object or region in the image frames of the stream of images. For example, the threshold condition may relate to any or all determinations that the object or region is focused, still has low blur, and / or has adequate exposure. The bounding box may represent, for example, cropping coordinates for each of the four corners of the box.

[0094] Once the ROI or object of interest is determined and a threshold condition is met, cropping coordinates (e.g., cropping coordinates 238) can be transmitted to or captured by a cropper (e.g., cropper 236). Flowchart 500B can then retrieve at least one high-resolution image from buffer 404. The retrieved high-resolution image is then cropped 504 to the region of interest or object of interest according to the cropping coordinates determined for the low-resolution image. The high-resolution image typically has the same capture time as the low-resolution image in which the region of interest or object of interest was identified. The output of the crop can include a cropped portion 506 of the high-resolution image. Additional downstream processing can be performed on-board or off-board wearable computing device 100.

[0095] Unlike conventional image analysis that occurs on devices with higher computational and power resources, the systems described herein can avoid sequential processing of high-resolution images. Furthermore, because scaling is not performed by wearable computing device 100, the systems described herein can avoid computational resources for scaling. Furthermore, typical image analysis systems that utilize both low-resolution and high-resolution image streams use separate buffered storage for each stream. The systems described herein utilize a buffer for the high-resolution images and avoid the use of a buffer for the low-resolution images, which provides the advantage of less memory traffic and / or memory.

[0096] 6A-6C illustrate example flow diagrams 600A, 600B, and 600C for performing image processing tasks on wearable computing device 100, according to implementations described throughout this disclosure. In the example of FIGS. 600A-600C, sensor 402 can function in a single-stream mode, where sensor 402 processes a single-mode analysis of either an incoming (i.e., captured) low-resolution image stream or an incoming (i.e., captured) high-resolution image stream. Sensor 402 can then be triggered to function in a dual-stream mode, where it processes both the incoming (i.e., captured) low-resolution image stream and the incoming (i.e., captured) high-resolution image stream consecutively and simultaneously.

[0097] 6A is a flow diagram 600A illustrating the use of an image sensor 402 installed in, for example, the wearable computing device 100. The sensor 402 can be configured to switch between dual-mode and single-mode operation. For example, the sensor 402 can initially function in a mode that streams / captures a low-resolution image stream (e.g., multiple sequential image frames). Such a mode can conserve computational and battery power for the wearable computing device 100 because high-resolution image analysis in the initial capture and / or image retrieval stage is avoided. The buffer 404 and high-power image signal processor 406 are available for use, but are disabled in this example.

[0098] During operation, low-resolution image frames may be provided to a low-power image signal processor 408. The processor 408 processes the low-resolution image frames to generate a stream of images 410 (e.g., one or more images) (e.g., a YUV-formatted image stream) that are de-bayered, color-corrected, and shade-corrected images. The low-power image signal processor 408 may perform low-power computations 412 to analyze the generated stream of images 410, thereby detecting objects and / or regions of interest 414 within the images 410. Detection may include hand detection, object detection, paragraph detection, etc.

[0099] When an object and / or ROI of interest is detected and the detection of such region or object meets a threshold condition, a bounding box for the object and / or region is identified. The threshold condition may represent a certain high level of quality for the object or region in the image frames of the image stream. For example, the threshold condition may relate to any or all determinations that the object or region is focused, remains low-blur, and / or has adequate exposure. The bounding box may represent, for example, cropping coordinates for each of the four corners of the box. Identification of the region or object of interest and cropping coordinates may be a trigger to enable the sensor 402 to begin operation in dual-stream mode.

[0100] In some implementations, a trigger for initiating dual-stream mode may occur when some, but not all, threshold conditions are met. For example, wearable computing device 100 may recognize an object of interest that meets a threshold condition for camera focus and a threshold condition for camera exposure. However, because the object may move (e.g., have motion blur), wearable computing device 100 may trigger use of dual-stream mode to prepare for capture of a high-resolution image using processor 118 when the final condition is met. This may be advantageous if there is a delay of several frames to switch on dual-stream mode. That is, the time period during which dual-stream mode is running may be a transitional period that begins when substantially all threshold conditions are met and the final threshold condition for detecting and determining is predicted to occur in the upcoming image frame.

[0101] FIG. 6B is a flow diagram 600B illustrating the use of an image sensor 402 installed in a wearable computing device 100, with the sensor 402 switched to dual-stream mode. A trigger for switching to dual-stream mode can include the detection of one or more objects or regions of interest. The dual-stream mode triggers the sensor 402 to begin capturing (or otherwise acquiring) high-resolution image frames in addition to the low-resolution image frames already streaming. In some implementations, several image frames from the high-resolution image stream can be used to configure automatic gain measurements, exposure settings, and focus settings. Once configuration for the high-resolution image stream is complete, the streams can be synchronized, and the low-resolution image frames can continue via the same flow as in FIG. 6A, as indicated by arrow 602. However, the sensor 402 can begin buffering high-resolution image frames using the processor (using buffer 404). The buffer 404 can be smaller than the buffer used if the sensor were to operate in continuous dual-stream mode from the start of image streaming.

[0102] 6C is a flow diagram 600C illustrating the use of an image sensor 402 installed in the wearable computing device 100 and configured to capture or receive a stream of image data after switching from single-mode streaming to dual-mode streaming. During operation, the sensor 402 may acquire multiple high-resolution image frames (e.g., the maximum resolution of the camera / sensor 402) captured using a full field of view. The multiple high-resolution image frames may be stored in a buffer 404.

[0103] Low-resolution image frames may be continuously captured and / or acquired by the sensor 402 and provided to a low-power image signal processor 408. The processor 408 processes the low-resolution image frames to generate a stream of images 410 (e.g., one or more images) (e.g., a YUV-formatted image stream) that are de-bayered, color-corrected, and shade-corrected images. The low-power image signal processor 408 may perform low-power computation 412 to analyze the generated stream of images 410, thereby detecting objects and / or regions of interest 414 within the images 410. Once objects and / or ROIs of interest are detected, and the detection of such regions or objects meets a threshold condition, a bounding box for the object and / or region is identified.

[0104] The threshold condition may represent a particular high level of quality for an object or region in an image frame of the stream of images. For example, the threshold condition may relate to any or all determinations that the object or region is focused, remains low-blur, and / or has adequate exposure. The bounding box may represent cropping coordinates that represent the region or object of interest 414. Once the object or region of interest 414 has been determined and the threshold condition has been met, the high-power image signal processor 406 retrieves (or receives) a high-resolution image frame from the buffer 404 (e.g., a raw frame that coincides with the same capture time as the low-resolution image frame in which the object or region of interest was identified).

[0105] The cropping coordinates and the retrieved high-resolution image frame may be provided as input to a high-power image signal processor 406 (e.g., arrow 604). The high-power image signal processor 406 may output a processed (e.g., de-bayered, color corrected, shade corrected, etc.) image 606 that is cropped to the object and / or region of interest determined from the low-resolution image frame. This cropped full-resolution image 606 may be provided to additional on-board or off-board devices for further processing.

[0106] 7A-7C illustrate example flow diagrams 700A, 700B, and 700C for performing image processing tasks on wearable computing device 100, according to implementations described throughout this disclosure. In the examples of FIGS. 700A-700C, sensor 402 can function in a first mode and then be triggered to function in a second mode. For example, sensor 402 can be triggered to begin streaming and / or capturing content in a low-resolution streaming mode, and low-resolution images are captured and utilized. Upon detecting one or more objects and / or regions of interest, sensor 402 can switch to a high-resolution streaming mode, and high-resolution images are utilized.

[0107] 7A is a flow diagram 700A illustrating the use of an image sensor 402 installed in, for example, wearable computing device 100. In this example, sensor 402 may initially function in a low-resolution streaming mode, streaming and / or capturing (or otherwise acquiring) a low-resolution image stream (e.g., multiple sequential image frames).

[0108] During operation, low-resolution image frames may be provided to a low-power image signal processor 408. The processor 408 processes the low-resolution image frames to generate a stream of images 410 (e.g., one or more images) (e.g., a YUV-formatted image stream) that are de-bayered, color-corrected, and shade-corrected images. The low-power image signal processor 408 may perform low-power computations 412 to analyze the generated stream of images 410, thereby detecting objects and / or regions of interest 414 within the images 410. Detection may include hand detection, object detection, paragraph detection, etc.

[0109] When an object and / or ROI of interest is detected and the detection of such region or object meets a threshold condition, a bounding box for the object and / or region is identified. The threshold condition may represent a certain high level of quality for the object or region in the image frames of the stream of images. For example, the threshold condition may relate to any or all determinations that the object or region is focused, remains low-blur, and / or has adequate exposure. The bounding box may represent, for example, cropping coordinates for each of the four corners of the box. Identification of the region or object of interest and cropping coordinates may be a trigger for enabling the sensor 402 to switch from a low-resolution streaming mode to a high-resolution streaming mode.

[0110] 7B is a flow diagram 700B illustrating the use of an image sensor 402 installed in a wearable computing device 100, where the sensor 402 has been triggered to switch from a low-resolution streaming mode to a high-resolution streaming mode. The trigger for such a mode switch may include the detection of one or more objects or regions of interest. The switch from the low-resolution streaming mode to the high-resolution streaming mode triggers the sensor 402 to begin capturing (or otherwise acquiring) high-resolution image frames while stopping the capture (or otherwise acquiring) of low-resolution image frames.

[0111] In this example, at least one high-resolution image frame may be processed by the high-power image signal processor 406. One or more of the processed high-resolution image frames may be provided to a scaler 702 where they are scaled (e.g., resolution scaled down) to generate low-resolution images 704 (or an image stream). The high-power image signal processor 406 may perform low-power computations 412 on the images 704 (or image stream). The low-power computations may include analyzing the at least one image 704 to detect, for example, objects and / or regions of interest 706 within the at least one image 704. The detection may include hand detection, object detection, paragraph detection, etc. Once the object and / or ROI of interest is detected, and the detection of such region or object meets a threshold condition, a bounding box for the object and / or region is identified. The threshold condition may represent a particular high level of quality for the object or region in the image frames of the stream of images. For example, the threshold conditions may relate to any or all determinations of being focused on the object or region, still having low blur, and / or having adequate exposure. The bounding box may represent, for example, cropping coordinates for each of the four corners of the box. The high-power image signal processor 406 may also buffer the full high-resolution image 708 in a buffer 710. The image 708 may be used for operations that can benefit from a high-resolution image.

[0112] FIG. 7C is a flow diagram 700C illustrating the use of an image sensor 402 installed in the wearable computing device 100. The sensor 402 in this example is switched to a high-resolution streaming mode in response to determining cropping coordinates of an ROI 706. For example, once the ROI 706 (or object of interest) is determined and a threshold condition is met, the cropping coordinates (e.g., cropping coordinates 238) can be transmitted to or acquired by a cropper (e.g., cropper 236). A high-resolution image 708 can be retrieved from a buffer 710. The retrieved high-resolution image 708 is then cropped 712 to the region or object of interest according to the cropping coordinates (e.g., of the ROI 706) determined for the low-resolution image 704. The high-resolution image 708 typically has the same capture time as the low-resolution image 704 in which the region or object of interest 706 was identified. The output of the crop can include a cropped portion 714 of the high-resolution image. Additional downstream processing may be performed on-board or off-board the wearable computing device 100.

[0113] In some implementations, the sensor 402 can determine when to perform a fast dynamic switch between using a low-resolution streaming mode and using a high-resolution streaming mode. For example, the sensor 402 can determine when to provide sensor data (e.g., image data, pixels, light data, etc.) to the high-power image signal processor 406 or the low-power image signal processor 408. The determination can be based on detected events occurring with respect to the image stream, for example, to reduce and / or otherwise minimize power usage and system latency. Exemplary detected events can include the detection of any one or more of an object, an ROI, motion or cessation of motion, a change in lighting, a new user, another user, another object, etc.

[0114] In this example, the wearable computing device 100 can be triggered to operate with content having a low resolution and to switch to content having a high resolution. Because both the low and high resolution image streams are available, the wearable computing device 100 can select between the two streams without retrieving the content from buffered data. This can save computation switching costs, memory costs, and the cost of retrieving the content from memory.

[0115] 8A and 8B show example flow diagrams 800A and 800B for emulating a dual-resolution image sensor using a single image sensor to perform image processing tasks, according to implementations described throughout this disclosure. In this example, a dual-stream output sensor (e.g., sensor 402) can be configured to output a first image stream with a low resolution and a wide field of view to represent a scene camera 802. Similarly, the sensor can be configured to output a second stream with a high resolution and a narrow field of view to represent a detail camera 804.

[0116] 8A is a flow diagram 800A illustrating the use of image sensors (e.g., a scene camera 802 and a detail camera 804) installed in a wearable computing device and configured to capture or receive a stream of image data. The detail camera 804 can acquire (e.g., capture) and output a high-resolution image stream (e.g., multiple image frames). The multiple high-resolution image frames can be stored in a buffer 806. Typically, a high-power image signal processor 808 is optionally available and can perform analysis of the high-resolution image frames.

[0117] The scene camera 802 can acquire (e.g., capture) and output multiple low-resolution image frames (simultaneously). The low-resolution image frames can be provided to a low-power image signal processor 810. The processor 810 processes the low-resolution image frames to generate a stream (e.g., a YUV-formatted image stream) of images 812 (e.g., one or more images), which are de-bayered, color-corrected, and shade-corrected images.

[0118] The low-power image signal processor 810 can perform low-power computations 814 to analyze the generated stream of images 812, thereby detecting objects and / or regions of interest 816 within the images 812. Detection can include hand detection, object detection, paragraph detection, etc. When an object and / or ROI of interest is detected, and the detection of such a region or object meets a threshold condition, a bounding box for the object and / or region is identified. The bounding box can represent cropping coordinates for each of the box's four corners. The threshold condition can represent a trigger event in the scene camera view, e.g., switching to a detail camera for a single capture (or short capture burst) based on a detected trigger event. The threshold condition (e.g., a trigger event) can occur, for example, when a condition and / or item is recognized within the view of the image.

[0119] 8B is a flow diagram 800B illustrating the use of image sensors (e.g., a scene camera 802 and a detail camera 804) installed on a wearable computing device and configured to capture or receive a stream of image data. Similar to FIG. 8A, the scene camera 802 can capture and output one or more low-resolution images (e.g., image frames), process the images using a low-power image signal processor 810, and / or perform computations using a low-power compute 814 to generate the low-resolution image 812 and determine an object and / or region of interest 816.

[0120] Once the object and / or region of interest 816 has been identified and the threshold conditions have been determined to be met, a high-resolution image can be retrieved from the buffer 806. For example, the high-power image signal processor 808 can use the cropping coordinates of the low-resolution image 812 and the cropped ROI 816 to identify the same ROI 816 in a higher-resolution, narrower field-of-view image stored in the buffer 806 that was captured at the same time as the image 812.

[0121] For example, the cropping coordinates (e.g., arrow 818) and the retrieved high-resolution image frame can be provided as input to a high-power image signal processor 808. The high-power image signal processor 808 can output a processed (e.g., de-bayered, color corrected, shade corrected, etc.) image 820 that is cropped to the object and / or region of interest determined from the low-resolution image frame. This cropped full-resolution image 820 can be provided to additional on-board or off-board devices for further processing.

[0122] In some implementations, a single high-resolution image can be processed through a high-power image signal processor rather than multiple image frames, saving computational and / or device power. Because downstream processing can operate on the cropped image 820 rather than the entire high-resolution image, such additional downstream processing may still incur additional computational and / or device power.

[0123] 4A through 8B illustrate flow diagrams for providing dual sensor (e.g., scene and detail camera) functionality utilizing a single sensor. Because the sensor can output at least two streams with different resolutions and / or different fields of view, in some implementations, the output of a single sensor can be configured to function in dual mode without requiring separate sensors. For example, a first stream can be configured to output a wide field of view for a low-resolution stream of images, representing the scene camera. Similarly, a second stream can be configured to output a narrow field of view for a high-resolution stream of images, representing the detail camera.

[0124] 9 is a flowchart illustrating an example of a process 900 for performing image processing tasks on a computing device, according to implementations described throughout this disclosure. In some implementations, the computing device is a battery-powered, wearable computing device. In some implementations, the computing device is a battery-powered, non-wearable computing device.

[0125] Process 900 may utilize an image processing system on a computing device having at least one processing device, a speaker, optional display capabilities, and a memory that stores instructions that, when executed, cause the processing device to perform several of the operations and computer-implemented steps set forth in the claims. Generally, wearable computing device 100, systems 200, and / or 1000 may be used to describe and implement process 900. The combination of wearable computing device 100 and systems 200 and / or 1000 may represent a single system in some implementations. Typically, process 900 utilizes the systems and algorithms described herein to detect image data for identifying one or more regions of interest in real time (e.g., at the time of capture), which may be cropped to reduce the amount of data used to perform image processing on wearable computing device 100.

[0126] At block 902, process 900 may include waiting to receive image content (or a request to identify image content). For example, image sensor 216 on wearable computing device 100 may receive, capture, or detect image content. In response to receiving the content or such a request to have sensor 216 identify image content associated with light data captured by the sensor, sensor 216 and / or processors 208, 222, and / or 224 may detect a first sensor data stream having a first image resolution, where the first sensor data stream may be based on the light data, at block 904. For example, image sensor 216 may detect a low image resolution data stream based on the light data received at image sensor 216 and may provide the low image resolution data stream to low-resolution image signal processor 222. Further, at block 906, the sensor 216 and / or the processors 208, 222, and / or 224 may detect a second sensor data stream having a second image resolution, where the second sensor data stream is also based on the light data. For example, the image sensor 216 may detect a high image resolution data stream based on the light data received by the image sensor 216 and provide the high image resolution data stream to the high resolution image signal processor 224.

[0127] During operation, the first sensor data stream and the second data stream can be captured and / or received at the same time and / or with a correlation to the capture time that can be cross-referenced to obtain a low-resolution image from the first sensor data stream that matches the timestamp of the high-resolution image from the second sensor data stream.

[0128] At block 908, process 900 includes identifying, by processing circuitry of wearable computing device 100, at least one region of interest in the first sensor data stream. For example, one or more ROIs 232 may be identified within an image frame 226 acquired by image sensor 216. Image frame 226 may have been analyzed by low-resolution image signal processor 222 and / or ROI detector 230. The ROI may be defined by pixels, edge detection coordinates, object shape or detection, or other region or object identification metric. In some implementations, the ROI may be an object or other image content in the image stream.

[0129] In some implementations, the processing circuitry may include a first image processor (e.g., low-resolution processor 222) configured to perform image signal processing on a first sensor data stream (e.g., a low-resolution stream) and a second image processor (e.g., high-resolution image signal processor 224) configured to perform image signal processing on a second sensor data stream (e.g., a high-resolution stream). Typically, the first image resolution of the first sensor data stream is lower than the second image resolution of the second sensor data stream.

[0130] At block 910, the process 900 includes determining, by processing circuitry, cropping coordinates 238 that define a first plurality of pixels of at least one region of interest in the first sensor data stream. For example, the low-resolution image signal processor 222 can utilize the image frame 226, the analyzed resolution 228, and the ROI detector 230 to identify cropping coordinates for defining an object 234 or region of interest (ROI) 232.

[0131] At block 912, process 900 includes generating, by processing circuitry, a cropped image representing the at least one region of interest. Generating the cropped image may include triggering cropper 236 to identify a second plurality of pixels in the second sensor data stream using cropping coordinates 238 that define a first plurality of pixels of the at least one region of interest in the first sensor data stream at block 914, and may also include cropping the second sensor data stream to the second plurality of pixels at block 916. The generated cropped image may be a high-resolution version of the originally identified region or object of interest.

[0132] In some implementations, generating a cropped image representative of at least one region of interest (or object) is performed in response to detecting that the at least one region of interest meets a threshold condition 256. Generally, the threshold condition 256 may represent a particular high level of quality for an object or region in an image frame of the stream of images. In some implementations, the threshold condition 256 may include detecting low blur of a second plurality of pixels in the second sensor data stream. In some implementations, the threshold condition 256 may include detecting that the second plurality of pixels has a particular image exposure measurement. In some implementations, the threshold condition 256 may include detecting that the second plurality of pixels is focused.

[0133] For example, the threshold condition may relate to any or all determinations of being focused on the object or region, still having low blur, and / or having adequate exposure. The bounding box may represent cropping coordinates that represent the region or object of interest 414. Once the object or region of interest (e.g., a second plurality of pixels) has been determined and the threshold condition has been met, the high-power image signal processor may retrieve (or receive) a high-resolution image frame (e.g., a raw frame that coincides with the same capture time as the low-resolution image frame in which the object or region of interest was identified).

[0134] The cropped high-resolution image can be used for additional image analysis. For example, processing circuitry 224 may include performing optical character resolution (OCR) on the cropped image to generate a machine-readable version of the cropped image (e.g., output 132). Processing circuitry 224 may then perform a search query using the machine-readable version of the cropped image (e.g., image 418), thereby generating multiple search results (e.g., shown in output 132), which may be triggered for display on display 244 of wearable computing device 100. In some implementations, instead of performing a search query, OCR may be performed on the cropped image to generate output 132, which may be displayed to the user on wearable computing device 100, read to the user, and / or otherwise provided as output for consumption by the user. For example, an audio output of the performed OCR may be provided through a speaker of wearable computing device 100.

[0135] In some implementations, the second sensor data stream is stored in memory on the wearable computing device 100. For example, the high-resolution image stream can be stored in the buffer 404 for later access to image frames and metadata associated with the high-resolution image stream. In response to identifying at least one region of interest (or object) in the first sensor data stream, the processing circuitry 224 can search for the corresponding at least one region of interest (or object) in the second sensor data stored in memory. For example, the buffered data can be accessed to obtain a high-resolution version of the image content identified in the region of interest (or object) in the low-resolution image stream. Additionally, the wearable computing device 100 can trigger a restriction of access to the first sensor data stream (e.g., the low-resolution image stream) while continuing to detect and access the second sensor data stream (e.g., the high-resolution image stream). That is, once the object or region of interest has been determined, the wearable computing device 100 may no longer wish to use the power or resources associated with the low-resolution image stream and may limit access to the stream to avoid excessive resource and / or power usage associated with streaming, storing and / or accessing the low-resolution image stream.

[0136] In some implementations, process 900 may also include, in response to generating a cropped image representing the at least one region of interest, transmitting the generated cropped image to a mobile device in communication with the computing device. For example, once the identification and cropping of the high-resolution image is performed on wearable computing device 100, wearable computing device 100 may transmit the cropped image to mobile device 202 for further processing. For example, mobile device 202 may generate data or information related to the region or object of interest. Mobile device 202 may have higher power and / or computational resources than wearable computing device 100, and thus the generated data and / or information may be retrieved and / or calculated on mobile device 202 and transmitted back to wearable computing device 100. Wearable computing device 100 may receive information related to the at least one region of interest from the mobile device and display the information on display 244 of wearable computing device 100.

[0137] In some implementations, the first image resolution has a low image resolution and the second image resolution has a high image resolution. Identifying the at least one region of interest or object of interest may further include using a machine learning algorithm executing on the wearable computing device 100 (e.g., via NN 254). The machine learning algorithm may use the first sensor data as input to identify text or at least one object represented in the first sensor data. For example, the first data sensor stream may be a low-resolution data sensor stream, but the machine learning algorithm may determine that text or an object is visible in the image stream.

[0138] For example, the wearable computing device 100 can use a machine learning algorithm to determine that certain image content (e.g., an object, an ROI) is present (e.g., depicted or represented) in the low-resolution image. The cropper 236 can crop the determined unknown image content to cropping coordinates 238 and provide the cropped image to a trained machine learning algorithm to determine whether the certain image content represents an object and / or region of interest. For example, the machine learning algorithm for determining the image content to crop can include identifying one or more details in the captured image. The identified details can be analyzed to determine whether the details include a recognizable object, symbol, object part, etc.

[0139] For example, once the detail is identified as a recognizable object, the wearable computing device 100 can attempt to recognize specific features of the object. For example, if the object is a soda can, the wearable computing device 100 can evaluate which brand of soda it is. A machine learning algorithm can run a lightweight detector architecture on the device (e.g., on the wearable computing device 100) across the low-resolution image to identify ROI (or object of interest) coordinates for a particular pattern. For example, the machine learning algorithm can determine as output that the brand of soda can is a clustered region of text that can be identified by the pattern, and that the pattern can be cropped out and used for further analysis. For example, the pattern can be cropped based on the machine learning algorithm output indicating the cropping coordinates. These coordinates can then be used to select a portion of a high-resolution image of the same scene (e.g., the clustered region of text). That portion can then be further analyzed to determine the brand. Because the high-resolution portion is used based on input retrieved from the low-resolution image, the machine learning algorithm can provide the advantage of optimizing energy and latency on the wearable computing device 100 because analysis of the full high-resolution image is avoided.

[0140] Examples described throughout this disclosure may refer to a computer and / or computing system. As used herein, a computer (and / or computing) system includes, without limitation, any suitable combination of one or more devices configured using hardware, firmware, and software for implementing one or more of the computerized techniques described herein. As used herein, a computer (and / or computing) system may be a single computing device or multiple computing devices acting collectively, with data storage and function execution distributed among the various computing devices.

[0141] Examples described throughout this disclosure may refer to augmented reality (AR). As used herein, AR refers to a user experience in which a computer system facilitates a perceptual perception that includes at least one virtual aspect and at least one real aspect. The AR experience may be provided by any of several types of computer systems, including, but not limited to, a battery-powered wearable computing device or a battery-powered non-wearable computing device. In some implementations, the wearable computing device may include an AR headset, which may include, but is not limited to, AR glasses, another wearable AR device, a tablet, a watch, or a laptop computer.

[0142] In some types of AR experiences, a user can directly perceive aspects of reality with their own senses, without the intervention of a computer system. Some AR glasses, such as wearable computing device 100, are designed to direct an image (e.g., a perceived virtual aspect) toward the user's retina, while the eye can also overlay other light not generated by the AR glasses. As another example, an in-lens microdisplay can be embedded in a see-through lens, or a projected display can be overlaid on a see-through lens. In other types of AR experiences, a computer system can enhance, complement, alter, and / or enable the user's impression of reality (e.g., a perceived real aspect) in one or more ways. In some implementations, the AR experience is perceived on the screen of a display device of the computer system. For example, some AR headsets and / or AR glasses are designed with a camera feedthrough to present a camera image of the user's surrounding environment on a display device placed in front of the user's eyes.

[0143] 10 illustrates an example of a computing device 1000 and a mobile computing device 1050 that can be used with the techniques described herein (e.g., to implement a wearable computing device 100 (e.g., a client computing device), a server computing device 204, and / or a mobile device 202). The computing device 1000 includes a processor 1002, a memory 1004, a storage device 1006, a high-speed interface 1008 that connects to the memory 1004 and a high-speed expansion port 1010, a low-speed bus 1014, and a low-speed interface 1012 that connects to the storage device 1006. Each of these components 1002, 1004, 1006, 1008, 1010, and 1012 are interconnected using various buses and may be mounted on a common motherboard or in any other suitable manner. The processor 1002 can process instructions for execution within the computing device 1000, including instructions stored in memory 1004 or on storage device 1006, for displaying graphical information for a GUI on an external input / output device, such as a display 1016 coupled to a high-speed interface 1008. In other implementations, multiple processors and / or multiple buses can be used, along with multiple memories and types of memory, where appropriate. Also, multiple computing devices 1000 can be connected, with each device providing some of the required operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).

[0144] The memory 1004 stores information within the computing device 1000. In one implementation, the memory 1004 is one or more volatile memory units. In another implementation, the memory 1004 is one or more non-volatile memory units. The memory 1004 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0145] The storage device 1006 can provide mass storage for the computing device 1000. In one embodiment, the storage device 1006 can be or include a computer-readable medium such as a floppy disk device, a hard disk device, an optical disk device or tape device, a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other configuration of devices. A computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 1004, the storage device 1006, or memory on the processor 1002.

[0146] High-speed controller 1008 manages bandwidth-intensive operations for computing device 1000, while low-speed controller 1012 manages lower-bandwidth-intensive operations. This allocation of functionality is merely exemplary. In one embodiment, high-speed controller 1008 is coupled to memory 1004, display 1016 (e.g., via a graphics processor or accelerator), and high-speed expansion port 1010, which can accept various expansion cards (not shown). In an embodiment, low-speed controller 1012 is coupled to storage device 1006 and low-speed expansion port 1014. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or a networking device, such as a switch or router, for example, via a network adapter.

[0147] The computing device 1000 can be implemented in several different forms, as shown. For example, the computing device 1000 can be implemented as a standard server 1020 or multiples of a group of such servers. The computing device 1000 can also be implemented as part of a rack server system 1024. Furthermore, the computing device 1000 can be implemented in a personal computer, such as a laptop computer 1022. Alternatively, components from the computing device 1000 can be combined with other components in a mobile device (not shown), such as device 1050. Each such device can include one or more of the computing devices 1000, 1050, and an entire system can be made up of multiple computing devices 1000, 1050 in communication with each other.

[0148] Computing device 1050 includes, among other components, a processor 1052, memory 1064, an input / output device such as a display 1054, a communication interface 1066, and a transceiver 1068. Device 1050 may also include a storage device such as a microdrive or other device to provide additional storage. Each of these components 1050, 1052, 1064, 1054, 1066, and 1068 are interconnected using various buses, and some of these components may be mounted on a common motherboard or in other suitable manner.

[0149] The processor 1052 can execute instructions within the computing device 1050, including instructions stored in the memory 1064. The processor can be implemented as a chipset of chips including individual analog and digital processors. The processor can provide coordination of other components of the device 1050, such as the user interface, applications executed by the device 1050, and control of wireless communications by the device 1050.

[0150] The processor 1052 can communicate with a user via a control interface 1058 and a display interface 1056 coupled to a display 1054. The display 1054 can be, for example, a TFT LCD (thin film transistor liquid crystal display), an LED (light emitting diode) or an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 1056 can include appropriate circuitry for driving the display 1054 to present graphics and other information to the user. The control interface 1058 can receive commands from the user and translate them for execution by the processor 1052. Additionally, an external interface 1062 can be provided in communication with the processor 1052 to enable near-field communication by the device 1050 with other devices. The external interface 1062 can provide, for example, wired communication in some implementations or wireless communication in other implementations, and multiple interfaces can also be used.

[0151] Memory 1064 stores information within computing device 1050. Memory 1064 may be embodied as one or more computer-readable media, one or more volatile memory units, or one or more nonvolatile memory units. Expansion memory 1074 may also be provided and connected to device 1050 via expansion interface 1072, which may include, for example, a SIMM (single in-line memory module) card interface. Such expansion memory 1074 may provide extra storage space for device 1050 or may store applications or other information for device 1050. In particular, expansion memory 1074 may include instructions for implementing or supplementing the processes described above and may also include secure information. Thus, for example, expansion memory 1074 may be provided as a security module for device 1050 and may be programmed with instructions that permit secure use of device 1050. Additionally, secure applications can be provided via SIMM cards along with additional information, for example identifying information can be placed on the SIMM card in a non-hackable manner.

[0152] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 1064, expansion memory 1074, or memory on processor 1052, which may be received, for example, via transceiver 1068 or external interface 1062.

[0153] Device 1050 can communicate wirelessly via communication interface 1066, which may include digital signal processing circuitry as needed. Communication interface 1066 can provide communications under various modes or protocols, such as GSM voice calling, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communications can occur, for example, via radio frequency transceiver 1068. Additionally, short-range communications can occur using Bluetooth, Wi-Fi, or other such transceivers (not shown). Furthermore, GPS (Global Positioning System) receiver module 1070 can provide additional navigation- and location-related wireless data to device 1050, which applications executing on device 1050 can use, if appropriate.

[0154] Device 1050 can also communicate audibly using voice codec 1060, which can receive spoken information from a user and convert it into usable digital information. Voice codec 1060 can also generate audible sounds for the user, such as through a speaker in the handset of device 1050. Such sounds can include sounds from a voice telephone call, can include recorded sounds (e.g., voice messages, music files, etc.), and can also include sounds generated by applications running on device 1050.

[0155] The computing device 1050 may be implemented in several different forms as shown in the figure. For example, the computing device 1050 may be implemented as a cellular telephone 1080. The computing device 1050 may also be implemented as part of a smartphone 1082, a personal digital assistant, or other similar mobile device.

[0156] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs that can be executed on and / or interpreted by a programmable system that includes at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0157] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0158] To provide for user interaction, the systems and techniques described herein can be implemented on a computer having a display device (such as an LED (light emitting diode) or OLED (organic LED), or LCD (liquid crystal display) monitor / screen) for displaying information to the user, as well as a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0159] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.

[0160] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0161] In some implementations, the illustrated computing device may include sensors that interface with the AR headset / HMD device 1090 to generate an augmented environment for viewing inserted content in a physical space. For example, one or more sensors included on the illustrated computing device 1050 or other computing devices may provide input to the AR headset 1090 or, generally, to the AR space. The sensors may include, but are not limited to, a touchscreen, an accelerometer, a gyroscope, a pressure sensor, a biometric sensor, a temperature sensor, a humidity sensor, and an ambient light sensor. The computing device 1050 may use these sensors to determine the absolute position and / or detected rotation of the computing device in the AR space, which may then be used as input to the AR space. For example, the computing device 1050 may be incorporated into the AR space as a virtual object such as a controller, a laser pointer, a keyboard, a weapon, or the like. When incorporated into the AR space, a user's positioning of the computing device / virtual object may enable the user to position the computing device to view the virtual object in a particular way within the AR space.

[0162] In some implementations, one or more input devices included on or connected to the computing device 1050 can be used as input to the AR space. The input devices can include, but are not limited to, a touchscreen, a keyboard, one or more buttons, a trackpad, a touchpad, a pointing device, a mouse, a trackball, a joystick, a camera, a microphone, earphones or buds with input functionality, a gaming controller, or other connectable input devices. When the computing device is integrated into the AR space, a user can cause certain actions in the AR space by interacting with the input devices included on the computing device 1050.

[0163] In some implementations, the touchscreen of the computing device 1050 can be displayed as a touchpad in the AR space. A user can interact with the touchscreen of the computing device 1050. The interaction is displayed, for example, in the AR headset 1090, as movements on the displayed touchpad in the AR space. The displayed movements can control virtual objects in the AR space.

[0164] In some implementations, one or more output devices included on the computing device 1050 can provide output and / or feedback to a user of the AR headset 1090 in the AR space. The output and feedback may be visual, tactile, or audio. The output and / or feedback may include, but is not limited to, vibration, turning on and off or blinking and / or flashing one or more lights or strobes, emitting an alarm, playing a chime, singing a song, and playing an audio file. Output devices may include, but are not limited to, vibration motors, vibration coils, piezoelectric devices, electrostatic devices, light-emitting diodes (LEDs), strobes, and speakers.

[0165] In some implementations, the computing device 1050 may appear as another object in the computer-generated 3D environment. A user's interaction with the computing device 1050 (e.g., rotating, shaking, touching a touchscreen, sliding a finger across a touchscreen) may be interpreted as an interaction with an object in the AR space. In the example of a laser pointer in the AR space, the computing device 1050 appears in the computer-generated 3D environment as a virtual laser pointer. As the user manipulates the computing device 1050, the user in the AR space sees the movement of the laser pointer. The user receives feedback from their interaction with the computing device 1050 in the AR environment on the computing device 1050 or on the AR headset 1090. The user's interaction with the computing device may be translated into an interaction with a user interface generated in the AR environment for the controllable device.

[0166] In some implementations, the computing device 1050 may include a touchscreen. For example, a user may interact with the touchscreen to interact with a user interface for the controllable device. For example, the touchscreen may include user interface elements, such as sliders, that may control characteristics of the controllable device.

[0167] Computing device 1000 is intended to represent various forms of digital computers and devices, including, but not limited to, laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 1050 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown, their connections and relationships, and their functions are meant to be merely exemplary and are not meant to limit the implementation of the subject matter described and / or claimed herein.

[0168] Several embodiments have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the specification.

[0169] Furthermore, the logic flows depicted in the figures do not require the particular order or sequential order shown to achieve desired results. Moreover, other steps may be provided or steps may be removed from the described flows, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the following claims.

[0170] In addition to the above, a user may have control over both whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or current location), as well as whether content or communications are sent to the user from the server. Additionally, certain data may be processed in one or more ways prior to storage or use so that personally identifiable information is removed. For example, the user's identity may be processed so that personally identifiable information for the user cannot be determined, or the user's geographic location from which location information (such as city, zip code, or country level) is obtained may be generalized so that the user's specific location cannot be determined. Thus, a user may have control over what information is collected about them, how that information is used, and what information is provided to them.

[0171] While certain features of the described embodiments have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to encompass all such modifications and changes as fall within the scope of the embodiments. They are presented by way of example only and are not limiting in any way, and it is to be understood that various changes in form and detail may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination except in mutually exclusive combinations. The embodiments described herein may include various combinations and / or subcombinations of the functions, components, and / or features of the different described embodiments.

Claims

1. An image processing method, comprising: receiving, by a computing device, from an image sensor, a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; determining a location for a region of interest in the first sensor data stream, the region of interest including text represented in the first sensor data stream; and generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; performing optical character resolution on the cropped image to generate a machine-readable version of the cropped image; performing a search query using the machine-readable version of the cropped image to generate a plurality of search results; displaying the plurality of search results on a display of the computing device; The image processing method further comprises:

2. An image processing method, comprising: receiving, by a computing device, from an image sensor, a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution and stored in a memory on the computing device; determining a location for a region of interest in the first sensor data stream, the region of interest including text represented in the first sensor data stream; and generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; responsive to identifying the text in the first sensor data stream, searching for a corresponding region of interest in the second sensor data stream stored in a memory; restricting access to the first sensor data stream while continuing to detect and access the second sensor data stream; The image processing method further comprises:

3. An image processing method, comprising: receiving, by a computing device, from an image sensor, a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; determining a location for a region of interest in the first sensor data stream, the region of interest including text represented in the first sensor data stream; and generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; transmitting the cropped image to a mobile device in communication with the computing device; receiving information about the region of interest from the mobile device; displaying the information on a display of the computing device; The image processing method further comprises:

4. An image processing method, comprising: receiving, by a computing device, from an image sensor, a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; determining a location for a region of interest in the first sensor data stream, the region of interest including text represented in the first sensor data stream; and generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; 11. The image processing method of claim 10, wherein generating the cropped image representing the region of interest is performed in response to detecting that the region of interest meets a threshold condition, the threshold condition including detecting that the second sensor data stream is low in blur.

5. the first image resolution is lower than the second image resolution; identifying the region of interest is performed using a machine learning algorithm executing on the computing device; The image processing method according to any one of claims 1 to 4.

6. The computing device comprises at least: a first image processor configured to perform image signal processing on the first sensor data stream; a second image processor configured to perform image signal processing on the second sensor data stream, wherein the first image resolution of the first sensor data stream is lower than the second image resolution of the second sensor data stream; The image processing method according to any one of claims 1 to 4.

7. A computing device comprising: a processing device; an image sensor configured to capture light data and output, based on the light data, a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; a memory storing instructions that, when executed, cause the computing device to perform operations; and the operation comprises: determining a location for a region of interest in the first sensor data stream; generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; A computing device, wherein the image sensor is a dual stream image sensor configured to operate in a low image resolution mode until triggered to switch to operation in a high image resolution mode.

8. A computing device comprising: a processing device; an image sensor configured to capture light data and output, based on the light data, a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; a memory storing instructions that, when executed, cause the computing device to perform operations; and the operation comprises: determining a location for a region of interest in the first sensor data stream; generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; performing optical character resolution on the cropped image to generate a machine-readable version of the cropped image; performing a search query using the machine-readable version of the cropped image to generate a plurality of search results; outputting the optical character resolution audio through a speaker of the computing device; and a computing device, 9. A computing device comprising: a processing device; an image sensor configured to capture light data and output, based on the light data, a first sensor data stream having a first image resolution and a second sensor data stream stored in a memory on the computing device and having a second image resolution different from the first image resolution; a memory storing instructions that, when executed, cause the computing device to perform operations; and the operation comprises: determining a location for a region of interest in the first sensor data stream; generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; responsive to identifying the region of interest in the first sensor data stream, searching for a corresponding region of interest in the second sensor data stream stored in a memory; restricting access to the first sensor data stream while continuing to detect and access the second sensor data stream; a computing device, 10. A computing device comprising: a processing device; an image sensor configured to capture light data and output, based on the light data, a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; a memory storing instructions that, when executed, cause the computing device to perform operations; and the operation comprises: determining a location for a region of interest in the first sensor data stream; generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; transmitting the cropped image to a mobile device in communication with the computing device; receiving information about the region of interest from the mobile device; outputting the information at the computing device; and a computing device,

11. the first image resolution is lower than the second image resolution; identifying the region of interest further includes identifying text represented in the first sensor data stream by a machine learning algorithm running on the computing device and using the first sensor data stream as input. A computing device according to any one of claims 7 to 10.

12. the processing device a first image processor configured to perform image signal processing on the first sensor data stream; a second image processor configured to perform image signal processing on the second sensor data stream, wherein the first image resolution of the first sensor data stream is lower than the second image resolution of the second sensor data stream. A computing device according to any one of claims 7 to 10.

13. A method for executing a program on a computing device, comprising: receiving a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; and identifying a location for a region of interest in the first sensor data stream, the region of interest including text represented in the first sensor data stream, the instructions causing the computing device to: generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; performing optical character resolution on the cropped image to generate a machine-readable version of the cropped image; performing a search query using the machine-readable version of the cropped image to generate a plurality of search results for display by the computing device; A computer program that performs the following:

14. A method for executing a program comprising: receiving a first sensor data stream having a first image resolution and a second sensor data stream stored in a memory on the computing device and having a second image resolution different from the first image resolution; and identifying a location for a region of interest in the first sensor data stream, the region of interest including text represented in the first sensor data stream, the instructions causing the computing device to: generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; responsive to identifying the text represented in the first sensor data stream, searching for a corresponding region of interest in the second sensor data stream stored in a memory; restricting access to the first sensor data stream while continuing to detect and access the second sensor data stream; A computer program that performs the following:

15. A method for executing a program comprising: receiving a first sensor data stream having a first image resolution and a second sensor data stream having a second image resolution different from the first image resolution; and identifying a location for a region of interest in the first sensor data stream, the region of interest including text represented in the first sensor data stream, the instructions causing the computing device to: generating a cropped image representing the region of interest in the second sensor data stream based on the location for the region of interest in the first sensor data stream; transmitting the cropped image to a mobile device in communication with the computing device; receiving information about the region of interest from the mobile device; displaying the information on a display of the computing device; A computer program that performs the following:

16. the first image resolution is lower than the second image resolution; identifying the region of interest using a machine learning algorithm running on the computing device; The computer program according to any one of claims 13 to 15.

17. The processing circuitry comprises at least a first image processor configured to perform image signal processing on the first sensor data stream; a second image processor configured to perform image signal processing on the second sensor data stream, wherein the first image resolution of the first sensor data stream is lower than the second image resolution of the second sensor data stream; The computer program according to any one of claims 13 to 15.

Citation Information

Patent Citations

  • Character extraction method, character extraction device, and program

    JP2007156741A

  • Coordinate detection device and learnt model

    JP2019046007A

  • Display control system, display device, and display control method

    JP2020087190A

  • Augmented Reality Display System

    JP2021507277A

  • Work support system and work support method

    JP2023005660A