Distributed sensor data processing using multiple classifiers on multiple devices

By distributing data processing tasks between wearable devices and computing devices and utilizing distributed machine learning models, the problems of high power consumption and short battery life of wearable devices when processing energy-intensive sensor data are solved, enabling long-term use and efficient computing of the devices.

CN114641806BActive Publication Date: 2025-09-09GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080035869.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-13
Publication Date
2025-09-09
Estimated Expiration
2040-10-13

AI Technical Summary

Technical Problem

Computing devices such as wearables face the challenges of high power consumption, short battery life, and unsuitability for prolonged wear when processing energy-intensive sensor data, especially when performing audio and image processing.

Method used

By offloading energy-intensive operations to a computing device or server computer, preliminary data processing is performed on the wearable device and transmitted via a wireless connection to the computing device for further processing, using distributed machine learning models for audio and image recognition.

Benefits of technology

It reduces the power consumption of wearable devices, extends battery life, reduces device weight and heat generation, and improves computing performance and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114641806B_ABST
    Figure CN114641806B_ABST
Patent Text Reader

Abstract

According to one aspect, a method for distributed sound / image recognition using a wearable device includes receiving sensor data via at least one sensor device and detecting, by a classifier on the wearable device, whether the sensor data includes an object of interest. The classifier is configured to execute a first machine learning (ML) model. The method includes, in response to detecting the object of interest within the sensor data, transmitting the sensor data to a computing device via a wireless connection, wherein the sensor data is configured to be used by a second ML model on the computing device or a server computer for further sound / image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to distributed sensor data processing using multiple classifiers on multiple devices. Background Art

[0002] Computing devices (e.g., wearable devices, smart glasses, smart speakers, sports cameras, etc.) are typically relatively small devices and, in some examples, can be worn on or around the human body for long periods of time. However, the computer processing requirements for processing sensor data (e.g., image data, audio data) may be relatively high, especially for devices that include display and perception capabilities. For example, the device may perform energy-intensive operations (e.g., audio and / or image processing, computer vision, etc.) that require multiple circuit components, which may lead to several challenges. For example, the device may generate a relatively large amount of heat, making it uncomfortable to wear the device close to the skin for a long period of time. In addition, the number of circuit components (including batteries) increases the weight of the device, thereby increasing the discomfort of wearing the device for a long period of time. In addition, energy-intensive operations (combined with battery capacity limitations) may result in relatively short battery life. Therefore, some conventional devices can only be used for a short duration in a day. Summary of the Invention

[0003] The present disclosure relates to low-power devices (e.g., smart glasses, wearable watches, portable action cameras, security cameras, smart speakers, etc.) connected to a computing device (e.g., a smartphone, laptop, tablet computer, etc.) via a wireless connection, wherein energy-intensive operations are offloaded to the computing device (or a server computer connected to the computing device), which results in improvements in device performance (e.g., power, bandwidth, latency, computing power, machine learning accuracy, etc.) and user experience. In some examples, the wireless connection is a short-range wireless connection, such as a Bluetooth connection or a near-field communication (NFC) connection. In some examples, the low-power device includes a head-mounted display device, such as smart glasses. However, the techniques discussed herein can be applied to other types of low-power devices, such as portable action cameras, security cameras, smart doorbells, smart watches, etc.

[0004] According to one aspect, a method for distributed sound recognition using a wearable device includes: receiving audio data via a microphone of the wearable device; detecting, by a sound classifier of the wearable device, whether the audio data includes a sound of interest, wherein the sound classifier executes a first machine learning (ML) model; and in response to detecting the sound of interest within the audio data, transmitting the audio data to a computing device via a wireless connection, wherein the audio data is configured to be used by a second ML model for further sound classification.

[0005] According to one aspect, a non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, causes the at least one processor to: receive audio data via a microphone of a wearable device; detect, by a sound classifier of the wearable device, whether the audio data includes a sound of interest, wherein the sound classifier is configured to execute a first machine learning (ML) model; and in response to detecting the sound of interest within the audio data, transmit the audio data to a computing device via a wireless connection, wherein the audio data is configured to be used by a second ML model on the computing device for further sound classification.

[0006] According to one aspect, a wearable device for distributed sound recognition includes: a microphone configured to capture audio data; a sound classifier configured to detect whether the audio data includes a sound of interest, the sound classifier including a first machine learning (ML) model; and a radio frequency (RF) transceiver configured to, in response to detecting the sound of interest within the audio data, transmit the audio data to a computing device via a wireless connection, wherein the audio data is configured to be used by a second ML model to convert the sound of interest into textual data.

[0007] According to one aspect, a computing device for sound recognition includes at least one processor; and a non-transitory computer-readable medium storing executable instructions, which, when executed by the at least one processor, causes the at least one processor to: receive audio data from a wearable device via a wireless connection, the audio data having a sound of interest detected by a sound classifier that executes a first machine learning (ML) model; determine, using a sound recognition engine on the computing device, whether to convert the sound of interest into text data; in response to the determination using the sound recognition engine on the computing device, convert the sound of interest into text data by the sound recognition engine, the sound recognition engine being configured to execute a second ML model; and send the text data to the wearable device via the wireless connection.

[0008] According to one aspect, a method for distributed image recognition using a wearable device includes receiving image data via at least one imaging sensor of the wearable device, detecting, by an image classifier of the wearable device, whether an object of interest is included within the image data, the image classifier executing a first machine learning (ML) model, and transmitting the image data to a computing device via a wireless connection, the image data being configured to be used by a second ML model on the computing device for further image classification.

[0009] According to one aspect, a non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, cause the at least one processor to receive image data from an imaging sensor on a wearable device, detect, by an image classifier of the wearable device, whether an object of interest is included in the image data, the image classifier being configured to execute a first machine learning (ML) model, and transmit the image data to a computing device via a wireless connection, the image data being configured to be used by a second ML model on the computing device to calculate object location data, the object location data identifying a location of an object of interest in the image data.

[0010] According to one aspect, a wearable device for distributed image recognition includes: at least one imaging sensor configured to capture image data, an image classifier configured to detect whether an object of interest is included within the image data, the image classifier configured to execute a first machine learning (ML) model, and a radio frequency (RF) transceiver configured to transmit the image data to a computing device via a wireless connection, the image data configured to be used by a second ML model on the computing device to calculate object location data, the object location data identifying a location of the object of interest in the image data.

[0011] According to one aspect, a computing device for distributed image recognition includes at least one processor and a non-transitory computer-readable medium storing executable instructions that, when executed by the at least one processor, cause the at least one processor to receive image data from a wearable device via a wireless connection, the image data having an object of interest detected by an image classifier executing a first machine learning (ML) model, calculate object location data based on the image data using a second ML model, the object location data identifying a location of the object of interest in the image data, and transmit the object location data to the wearable device via the wireless connection. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 Illustrated is a system for distributing image and / or audio processing across multiple devices, including wearable devices and computing devices, according to one aspect.

[0013] Figure 2 Illustrated is a system for distributing image and / or audio processing across a wearable device and a server computer according to one aspect.

[0014] Figure 3 Illustrated is a system for distributing image and / or audio processing across a wearable device, a computing device, and a server computer according to one aspect.

[0015] Figure 4 An example of a head mounted display device according to an aspect is illustrated.

[0016] Figure 5 An example of electronic components on a head mounted display device according to one aspect is illustrated.

[0017] Figure 6 Illustrated is a printed circuit board substrate for electronic components on a head mounted display device according to one aspect.

[0018] Figure 7A Illustrated is a system for distributing audio processing between a wearable device and a computing device according to one aspect.

[0019] Figure 7B A sound classifier according to an aspect is illustrated.

[0020] Figure 8 A system for distributing audio processing between a wearable device and a server computer is illustrated according to one aspect.

[0021] Figure 9 Illustrated is a system for audio processing using a wearable device and a computing device according to one aspect.

[0022] Figure 10 Illustrated is a system for audio processing using a wearable device, a computing device, and a server computer according to one aspect.

[0023] Figure 11 Illustrated is a flow diagram of audio processing using a wearable device according to one aspect.

[0024] Figure 12 A flow chart illustrating audio processing using a wearable device according to another aspect is illustrated.

[0025] Figure 13A Illustrated is a system for image processing using a wearable device and a computing device according to one aspect.

[0026] Figure 13B An image classifier according to an aspect is illustrated.

[0027] Figure 13C An example of a bounding box dataset according to an aspect is illustrated.

[0028] Figure 14 A flow chart illustrating image processing using a wearable device according to one aspect is illustrated.

[0029] Figure 15 Illustrated is a system for image processing using a wearable device and a computing device according to one aspect.

[0030] Figure 16 A flow chart for image processing using a wearable device is illustrated according to one aspect.

[0031] Figure 17 Illustrated is a system for image processing using a wearable device and a computing device according to one aspect.

[0032] Figure 18 Illustrated is a system for audio and / or image processing using a wearable device and a computing device according to one aspect.

[0033] Figure 19 A flow chart for image processing using a wearable device is illustrated according to one aspect.

[0034] Figure 20 Illustrated is a system for audio processing using a wearable device and a computing device according to one aspect.

[0035] Figure 21 Illustrated is a flow diagram for audio processing using a wearable device according to one aspect. DETAILED DESCRIPTION

[0036] For sensor data captured by one or more sensors on a wearable device, the wearable device performs a portion of audio and / or image processing (e.g., less energy-intensive operations), and the computing device (and / or a server computer and / or other devices) performs other portions of the audio and / or image processing (e.g., more energy-intensive operations). For example, the wearable device can use a relatively small machine learning (ML) model to intelligently detect the presence of sensor data (e.g., whether the audio data includes sounds of interest, such as speech, music, alarms, hot words for voice commands, etc., or whether the image data includes objects of interest, such as objects, text, barcodes, facial features, etc.), and if so, the sensor data can be streamed to the computing device via a wireless connection to perform more complex audio and / or image processing using the relatively larger ML model. The results of the more complex audio and / or image processing can be provided back to the wearable device via the wireless connection, which can cause the wearable device to perform an action (including additional image / audio processing) and / or cause the wearable device to render the results on a display of the wearable device.

[0037] In some examples, the hybrid architecture can enable a small form factor with fewer circuit components in a wearable device such as a head-mounted display device (e.g., smart glasses). For example, because the system offloads more energy-intensive operations to a connected computing device (and / or server computer), the wearable device can include less powerful / complex circuitry. In some examples, the architecture of the wearable device can enable a relatively small printed circuit board within the frame of the glasses, where the printed circuit board includes circuitry that is relatively low in power while still capable of performing wearable applications based on image processing and / or computer vision (such as object classification, optical character recognition (OCR), and / or barcode decoding). As a result, battery life can be increased, allowing a user to use the wearable device for an extended period of time.

[0038] In some examples, sound recognition operations are distributed between a wearable device and a computing device (and potentially a server computer or other computing device). For example, the wearable device includes a sound classifier (e.g., a small ML model) that is configured to detect whether a sound of interest (e.g., speech, music, an alarm, etc.) is included in the audio data captured by a microphone on the wearable device. If not, the sound classifier continues to monitor the audio data to determine whether a sound of interest is detected. If so, the wearable device can stream the audio data (e.g., original sound, compressed sound, sound fragments, extracted features and / or audio parameters, etc.) to the computing device via a wireless connection. The sound classifier can save power and latency through its relatively small ML model. The computing device includes a more powerful sound recognition engine (e.g., a more powerful classifier) ​​that executes the larger ML model to translate (or convert) the audio data into text data (or other forms of data), where the computing device transmits the text data back to the wearable device via a wireless connection for display on the wearable device's display and / or audibly read back to the user. In some examples, a computing device is connected to a server computer via a network (e.g., the internet), and the computing device transmits the audio data to the server computer, where the server computer executes a larger ML model to translate the audio data into text data (e.g., in the case of a different language). The text data is then routed back to the computing device and then back to the wearable device for display.

[0039] In some examples, image recognition operations are distributed between a wearable device and a computing device. In some examples, the image recognition operations include facial detection and tracking. However, the image recognition operations may include operations to detect (and track) other regions of interest in the image data, such as objects, barcodes, and / or text. The wearable device includes an image classifier (e.g., a small ML model) configured to detect whether an object of interest (e.g., facial features, text, OCR codes, etc.) is included in image data captured by one or more imaging sensors on the wearable device. If so, the wearable device may transmit the image frame (including the object of interest) to the computing device via a wireless connection. The computing device includes a more powerful object detector (e.g., a more powerful classifier) ​​that executes the larger ML model to calculate object location data (e.g., a bounding box dataset) identifying the location of the detected object of interest, where the computing device transmits the object location data back to the wearable device. The wearable device uses one or more low-complexity tracking mechanisms (e.g., inertial measurement unit (IMU)-based warping, blob detection, optical flow, etc.) to propagate the object location data for subsequent image frames captured on the wearable device. The wearable device can compress the cropped region and send it to the computing device, where an object detector on the computing device can perform object detection on the cropped region and send updated object location data back to the wearable device.

[0040] In some examples, sensing operations with multiple resolutions are distributed between the wearable device and the computing device. Sensing operations can include always-on sensing and sensing voice input requests (e.g., hot word detection). For example, the wearable device can include a low power / low resolution (LPLR) camera and a high power / high resolution (HPHR) camera. In some examples, the wearable device can include an image classifier that executes a small ML model to detect objects of interest (e.g., faces, text, barcodes, buildings, etc.) from image data captured by the LPLR camera. If an object of interest is detected, the HPHR camera can be triggered to capture one or more image frames with higher quality (e.g., higher resolution, less noise, etc.). For some applications, higher quality images may be required.

[0041] The image frames from the HPHR camera can then be transmitted via a wireless connection to a computing device, where the computing device executes a larger ML model to perform more complex image recognition operations on the image frames with higher quality. In some examples, the operation can be similar to the object detection example described above, where object location data (e.g., a bounding box dataset) is calculated and sent to the wearable device, and the wearable device uses one or more tracking mechanisms to propagate the object location data to subsequent frames, and then the wearable device crops and compresses the image area to be sent back to the computing device for further processing. In some examples, the image stream of a product can be used to capture label text or barcodes and look up associated product information (e.g., price, shopping suggestions, comparable products, etc.). This information can be displayed on a display surface present on the wearable device or read back to the user audibly.

[0042] In terms of sensing a voice input request, the wearable device may include a voice command detector that executes a small ML model (e.g., a gatekeeper model) to continuously (e.g., periodically) process microphone samples for an initial portion of a hot word (e.g., "okG" or "okD"). If the voice command detector detects the initial portion, the voice command detector may cause a buffer to capture subsequent audio data. In addition, the wearable device may transmit a portion of the buffer (e.g., 1-2 seconds of audio from the head of the buffer) to a computing device via a wireless connection, wherein the computing device includes a hot word recognition engine with a larger ML model to perform full hot word recognition. If the utterance is a false positive, the computing device may transmit a dismiss command to the wearable device, which discards the contents of the buffer. If the utterance is a true positive, the remainder of the audio buffer is transmitted to the computing device for automatic speech recognition and user-bound response generation.

[0043] The systems and techniques described herein can reduce the power consumption of wearable devices, increase battery life, reduce the heat generated by wearable devices, and / or reduce the amount of circuit components within wearable devices (which can result in weight reduction), which can lead to the use of wearable devices for extended periods of time. In some examples, in terms of power, the systems and techniques described herein can extend the battery life of wearable devices to extended periods of time (e.g., 5 to 15 hours, or more than 15 hours). In contrast, some conventional smart glasses and other image / audio processing products may only be used for a few hours.

[0044] In some examples, in terms of bandwidth, the systems and techniques described herein can distribute computational operations (e.g., inference operations) across wireless connections using gatekeeper models (e.g., small classifiers, binary classifiers, etc.) to limit unnecessary transmissions, which can reduce latency and reduce power usage. In some examples, in terms of latency, the systems and techniques described herein can enable the use of inference both near sensors of wearable devices and across components of computing devices (and potentially server computers), which can provide flexibility in tuning performance to meet the requirements of various applications. ML decisions can occur dynamically as application usage and power (e.g., remaining battery life) or computational requirements change during use. In some examples, in terms of computing power, the systems and techniques described herein can provide flexible use of computing resources to meet application requirements.

[0045] Figure 1 A system 100 is illustrated for distributing image and / or audio processing of sensor data 128 across multiple devices, including devices 102, computing devices 152, and / or server computers 160. In some examples, the sensor data 128 is real-time sensor data or near-real-time sensor data (e.g., data collected in real-time or near-real-time from one or more sensors 138). In some examples, the image and / or audio processing of the sensor data 128 can be distributed between the devices 102 and computing devices 152. In some examples, the image and / or audio processing of the sensor data 128 can be distributed between any two or more of the devices 102, computing devices 152, or server computers 160 (or any combination thereof). In some examples, the system 100 includes multiple devices 102 and / or multiple computing devices 152, each of which executes a classifier that makes decisions about whether and what data to relay to the next classifier, which can be on the same device or a different device.

[0046] Device 102 is configured to connect to computing device 152 via wireless connection 148. In some examples, wireless connection 148 is a short-range communication link, such as a near field communication (NFC) connection or a Bluetooth connection. Device 102 and computing device 152 can exchange information via wireless connection 148. In some examples, wireless connection 148 defines an application layer protocol that is implemented using protocol buffers with message types for drawing graphics primitives, configuring sensors 138 and peripherals, and changing device modes. In some examples, the application layer protocol defines another set of message types that can transmit sensor data 128 and remote procedure call (RPC) return values ​​back to computing device 152.

[0047] Computing device 152 can be coupled to server computer 160 via network 150. Server computer 160 can be a computing device in the form of multiple different devices, such as a standard server, a group of such servers, or a rack server system. In some examples, server computer 160 is a single system that shares components such as a processor and memory. Network 150 can include the Internet and / or other types of data networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, a satellite network, or other types of data networks. Network 150 can also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within network 150. In some examples, device 102 is also configured to connect to server computer 160 via network 150.

[0048] With respect to audio and / or image processing of sensor data 128 captured in real time or near real time by one or more sensors 138 on device 102, a portion of the audio and / or image processing (e.g., less energy-intensive operations) is performed at device 102, and another portion of the audio and / or image processing (e.g., more energy-intensive operations) is performed at computing device 152 (and / or server computer 160). In some examples, another portion of the audio and / or image processing is performed at another device. In some examples, another portion of the audio and / or image processing is performed at yet another device, and so on. In some examples, sensor data 128 includes audio data 131. In some examples, sensor data 128 includes image data 129. In some examples, sensor data 128 includes audio data 131 and image data 129.

[0049] Device 102 can intelligently detect the presence of certain types of data within sensor data 128 captured by sensor 138. In some examples, device 102 can detect whether audio data 131 captured by microphone 140 includes sounds of interest, such as speech, music, an alarm, or at least a portion of a hotword for command detection. In some examples, device 102 can detect whether image data 129 includes objects of interest (e.g., objects, text, barcodes, facial features, etc.). If device 102 detects relevant data within sensor data 128, device 102 can stream sensor data 128 to computing device 152 via wireless connection 148 to perform more complex audio and / or image processing. In some examples, device 102 can stream image data 129 to computing device 152. In some examples, device 102 can stream audio data 131 to computing device 152. In some examples, device 102 can stream both audio data 131 and image data 129 to computing device 152.

[0050] In some examples, device 102 compresses audio data 131 and / or image data 129 before transmitting to computing device 152. In some examples, device 102 extracts features from sensor data 128 and sends the extracted features to computing device 152. In some examples, device 102 extracts features from sensor data 128 and sends the extracted features to computing device 152. For example, the extracted features may include sound intensity, a calculated angle of arrival (e.g., the direction from which the sound is coming), and / or the type of sound (e.g., speech, music, alarm, etc.). In some examples, the extracted features may include compression encoding, which can save transmission bandwidth for certain types of sounds. The results of the more complex audio and / or image processing performed at computing device 152 can be provided back to device 102 via wireless connection 148 to cause device 102 to perform an action (including further audio and / or image processing), cause device 102 to render the results on device 102's display 116, and / or cause device 102 to audibly provide the results.

[0051] In some examples, device 102 is a display device that can be worn on or near a person's skin. In some examples, device 102 is a wearable device. In some examples, device 102 is a head-mounted display (HMD) device, such as an optical head-mounted display (HMD) device, a transparent head-up display (HUD) device, an augmented reality (AR) device, or other device, such as Google Glass or headphones, that has sensors, a display, and computing capabilities. In some examples, device 102 is smart glasses. Smart glasses are optical head-mounted displays designed in the shape of a pair of glasses. For example, smart glasses are glasses that add information (e.g., project a display 116) next to what the wearer is viewing through the glasses. In some examples, superimposing information (e.g., digital images) onto the field of view can be achieved through smart optics. Smart glasses are effectively wearable computers that can run self-contained mobile apps (e.g., application 112). In some examples, smart glasses can be hands-free and can communicate with the internet via natural language voice commands, while other smart glasses use touch buttons. In some examples, device 102 can include any type of low-power device. In some examples, device 102 includes a security camera. In some examples, device 102 includes an action camera. In some examples, device 102 includes a smartwatch. In some examples, device 102 includes a smart doorbell. As described above, system 100 can include multiple devices 102 (e.g., smartwatches, smart glasses, etc.), where each device 102 is configured to execute a classifier that can perform image / audio processing and then route the data to the next classifier in the classifier network.

[0052] Device 102 may include one or more processors 104, which may be formed in a substrate and configured to execute one or more machine-executable instructions or software, firmware, or a combination thereof. In some examples, processor 104 is included as part of a system-on-chip (SOC). Processor 104 may be semiconductor-based, i.e., the processor may include semiconductor material that can execute digital logic. Processor 104 includes a microcontroller 106. In some examples, microcontroller 106 is a subsystem within the SOC and may include processing, memory, and input / output peripherals. In some examples, microcontroller 106 is a dedicated hardware processor that executes a classifier. Device 102 may include a power management unit (PMU) 108. In some examples, PMU 108 is integrated with or included within the SOC. Microcontroller 106 is configured to execute a machine learning (ML) model 126 to use sensor data 128 to perform inference operations 124-1 related to audio and / or image processing. As discussed further below, the relatively small size of ML model 126 can save power and latency. In some examples, device 102 includes multiple microcontrollers 106 and multiple ML models 126 that perform multiple inference operations 124 - 1 , which can communicate with each other and / or other devices (e.g., computing device 152 and / or server computer 160 ).

[0053] Device 102 includes one or more memory devices 110. In some examples, memory device 110 includes flash memory. In some examples, memory device 110 may include main memory that stores information in a format that can be read and / or executed by processor 104, including microcontroller 106. Memory device 110 may store weights 109 (e.g., inference weights or model weights) for ML model 126 executed by microcontroller 106. In some examples, memory device 110 may store other assets such as fonts and images.

[0054] In some examples, device 102 includes one or more applications 112, which may be stored in memory device 110 and, when executed by processor 104, perform certain operations. Applications 112 may vary widely depending on the use case, but may include browser applications that search for web content, voice recognition applications such as speech-to-text applications, image recognition applications (including object and / or facial detection (and tracking) applications, barcode decoding applications, text OCR applications, etc.), and / or other applications that may enable device 102 to perform certain functions (e.g., capture images, record video, get directions, send messages, etc.). In some examples, applications 112 include an email application, a calendar application, a storage application, a voice calling application, and / or a messaging application.

[0055] Device 102 includes a display 116, which is a user interface for displaying information. In some examples, display 116 is projected onto the user's field of view. In some examples, display 116 is a built-in lens display. Display 116 may include a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting display (OLED), an electrophoretic display (EPD), or a micro-projection display using an LED light source. In some examples, display 116 may provide a transparent or translucent display so that a user wearing glasses can see the image provided by display 116, but can also see information in the field of view of the smart glasses behind the projected image. In some examples, device 102 includes a touchpad 117 that allows the user to control device 102 (e.g., it may allow swiping of an interface displayed on display 116). The device 102 includes a battery 120 configured to power circuit components, one or more radio frequency (RF) transceivers 114 that enable communication with a computing device 152 via a wireless connection 148 and / or with a server computer 160 via a network 150, a battery charger 122 configured to control charging of the battery 120, and one or more display regulators 118 that control information displayed by the display 116.

[0056] Device 102 includes multiple sensors 138, such as a microphone 140 configured to capture audio data 131, one or more imaging sensors 142 configured to capture image data, a light condition sensor 144 configured to obtain lighting condition information, and / or a motion sensor 146 configured to obtain motion information. Microphone 140 is a transducer device that converts sound into an electrical signal represented by audio data 131. Light condition sensor 144 can detect light exposure. In some examples, light condition sensor 144 includes an ambient light sensor that detects the amount of ambient light present, which can be used to ensure that image data 129 is captured with a desired signal-to-noise ratio (SNR). However, light condition sensor 144 can include other types of photometric (or colorimetric) sensors. Motion sensor 146 can obtain motion information, which can include blur estimation information. Motion sensor 146 is used to monitor device movement (such as tilt, shake, rotation, and / or sway) and / or to determine blur estimation.

[0057] Imaging sensor 142 is a sensor (e.g., a camera) that detects and communicates information used to create the image represented by image data 129. Imaging sensor 142 can take pictures and record video. In some examples, device 102 includes a single imaging sensor 142. In some examples, device 102 includes multiple imaging sensors 142. In some examples, imaging sensor 142 includes imaging sensor 142a and imaging sensor 142b. Imaging sensor 142a can be considered a low-power, low-resolution (LPLR) image sensor. Imaging sensor 142b can be considered a high-power, high-resolution (HPHR) image sensor. Images captured by imaging sensor 142b have higher quality (e.g., higher resolution, lower noise) than images captured by imaging sensor 142a. In some examples, device 102 includes two or more imaging sensors 142.

[0058] In some examples, imaging sensor 142a is configured to obtain image data 129 when device 102 is activated (e.g., continuously or periodically capture image data 129 when device 102 is activated). In some examples, imaging sensor 142a is configured to operate as a normally-on sensor. In some examples, imaging sensor 142b is activated (e.g., for a short duration) in response to detecting an object of interest, as discussed further below.

[0059] Computing device 152 can be any type of computing device capable of wirelessly connecting to device 102. In some examples, computing device 152 is a mobile computing device. In some examples, computing device 152 is a smartphone, tablet, or laptop. In some examples, computing device 152 is a wearable device. Computing device 152 can include one or more processors 154 formed in a substrate that are configured to execute one or more machine-executable instructions or software, firmware, or a combination thereof. Processor 154 can be semiconductor-based, i.e., the processor can include semiconductor material that can execute digital logic.

[0060] The computing device 152 may include one or more memory devices 156. The memory device 156 may include main memory that stores information in a format that can be read and / or executed by the processor 154. The operating system 155 is system software that manages computer hardware and software resources and provides common services to computer programs. Although not described in Figure 11 , but computing device 152 may include a display (e.g., a touch screen display, an LED display, etc.) that can display a user interface for an application 158 being executed by computing device 152. Application 158 may include any type of computer program executable by operating system 155. Application 158 may include a mobile application, such as a software program developed for a mobile platform or mobile device.

[0061] In some examples, the audio and / or image processing performed on the sensor data 128 obtained by the sensor 138 is referred to as an inference operation (or ML inference operation). An inference operation (e.g., inference operation 124-1 or inference operation 124-2) can refer to an audio and / or image processing operation, step, or sub-step involving an ML model that makes (or results in) one or more predictions. Certain types of audio and / or image processing use ML models to make predictions. For example, machine learning can use statistical algorithms that learn from data in existing data in order to make decisions about new data, a process known as inference. In other words, inference refers to the process of taking an already trained model and using the trained model to make predictions. Some examples of inference can include sound recognition (e.g., speech to text recognition), image recognition (e.g., facial recognition and tracking, etc.), and / or perception (e.g., always-on sensing, voice input request sensing, etc.).

[0062] In some examples, the ML model includes one or more neural networks. The neural network transforms the input received by the input layer, transforms it through a series of hidden layers, and produces an output via the output layer. Each layer consists of a subset of a node set. The nodes in the hidden layer are fully connected to all nodes in the previous layer and provide their output to all nodes in the next layer. The nodes in a single layer act independently of each other (i.e., do not share connections). The nodes in the output provide the transformed input to the requesting process. In some examples, the neural network is a convolutional neural network, which is a neural network that is not fully connected. Therefore, a convolutional neural network has a lower complexity than a fully connected neural network. Convolutional neural networks can also use pooling or maximum pooling to reduce the dimensionality (and therefore the complexity) of the data flowing through the neural network, thereby reducing the required level of calculation. This makes the calculation of the output in the convolutional neural network faster than the calculation of the output in the neural network.

[0063] With respect to certain inference types, device 102 may perform one or more portions of inference to intelligently detect the presence of sensor data 128 (e.g., whether audio data 131 includes a sound of interest, such as at least a portion of a voice, an alarm, or a hotword, and / or whether image data 129 includes an object of interest (e.g., a facial feature, text, an object, a barcode, etc.)), and if so, send sensor data 128 to computing device 152 via wireless connection 148, where computing device 152 performs one or more other portions of ML inference (e.g., more complex portions of audio and / or image processing) using sensor data 128. In other words, inference operations may be distributed between device 102 and computing device 152 (and potentially server computer 160), such that energy-intensive operations are performed at a more powerful computing device (e.g., computing device 152 or server computer 160) than at a relatively small computing device (e.g., device 102).

[0064] In some examples, the system 100 can include other devices (e.g., in addition to the device 102, the computing device 152, and the server computer 160), wherein one or more of these other devices can execute one or more classifiers (wherein each classifier executes an ML model related to object / sound recognition). For example, the system 100 can have one or more classifiers on the device 102, one or more wearable devices (e.g., one or more devices 102), and / or one or more classifiers on the computing device 152. In addition, the data can be sent to the server computer 160 for server-side processing, which can have an additional classification step. Thus, in some examples, the system 100 can include a classifier network that analyzes the audio / camera stream and decides whether and what to relay to the next node (or classifier).

[0065] In some examples, the microcontroller 106 of the device 102 can use sensor data 128 (e.g., audio data 131 from the microphone 140 and / or image data 129 from one or more of the imaging sensors 142) and the ML model 126 stored on the device 102 to perform the inference operation 124-1. In some examples, the ML model 126 can receive the sensor data 128 as input and detect whether the sensor data 128 has a classification that the ML model 126 is trained to classify (e.g., whether the audio data 131 includes a sound of interest or whether the image data 129 includes an object of interest). In some examples, the ML model 126 is a sound classifier that can evaluate the incoming sound based on specific criteria (e.g., frequency, amplitude, feature detection, etc.). In some examples, the analyzed criteria determine whether the audio data (e.g., original sound, compressed sound, sound clip, audio parameters, etc.) should be sent to other devices (including computing device 152, server computer 160, etc.) for further classification.

[0066] In some examples, ML model 126 is a speech classifier (e.g., a binary speech classifier) ​​that detects whether audio data 131 includes speech or does not include speech. In some examples, ML model 126 is an image object classifier (detector) that detects whether image data 129 includes an object of interest or does not include an object of interest. In some examples, ML model 126 is an object classifier that detects whether image data 129 includes facial features or does not include facial features. In some examples, ML model 126 is a classifier that determines whether audio data 131 includes at least a portion of a hotword for a voice command.

[0067] If the output of ML model 126 indicates that a classification has been detected, RF transceiver 114 of device 102 can transmit sensor data 128 to computing device 152 via wireless connection 148. In some examples, device 102 can compress sensor data 128 and then send the compressed sensor data 128 to computing device 152. Computing device 152 is then configured to perform inference operation 124-2 using sensor data 128 (received from device 102) and ML model 127 stored on computing device 152. In some examples, in terms of sound recognition (e.g., speech-to-text processing), ML model 127 is used to convert audio data 131 into text, where the results are transmitted back to device 102. In some examples, in terms of hotword command recognition, ML model 127 is used to perform all-hotword command recognition on audio data 131 received from device 102. In some examples, in terms of image processing, ML model 127 is used to calculate object location data (identifying the location of objects of interest in the image data), where the results are transmitted back to device 102 for further image processing, which is further described later in the specification.

[0068] However, in general, inference operation 124-2 may refer to an audio and / or image processing operation involving a different ML model than inference operation 124-1. In some examples, the inference operation includes a sound recognition operation, wherein inference operation 124-1 refers to a first sound recognition operation performed using ML model 126, and inference operation 124-2 refers to a second sound recognition operation performed using ML model 127. In some examples, the inference operation includes an image recognition operation, wherein inference operation 124-1 refers to a first image recognition operation performed using ML model 126, and inference operation 124-2 refers to a second image recognition operation performed using ML model 127. In some examples, the inference operation includes a perceptual sensing operation (e.g., always-on sensing, voice command sensing (e.g., hotword recognition), etc.), wherein inference operation 124-1 refers to a first perceptual sensing operation performed using ML model 126, and inference operation 124-2 refers to a second perceptual sensing operation performed using ML model 127.

[0069] The size of ML model 126 can be smaller (e.g., substantially smaller) than the size of ML model 127. In some examples, ML model 126 may be required to perform fewer computational operations to make predictions than ML model 127. In some examples, the size of a particular ML model can be represented by the number of parameters required by the model to make predictions. Parameters are configuration variables within the ML model, and their values ​​can be estimated from given data. ML model 126 can include parameters 111. For example, ML model 126 can define a number of parameters 111 required for ML model 126 to make predictions. ML model 127 can include parameters 113. For example, ML model 127 can define a number of parameters 113 required for ML model 127 to make predictions. The number of parameters 111 can be smaller (e.g., substantially smaller) than the number of parameters 113. In some examples, the number of parameters 113 is at least ten times greater than the number of parameters 111. In some examples, the number of parameters 113 is at least one hundred times greater than the number of parameters 111. In some examples, the number of parameters 113 is at least one thousand times greater than the number of parameters 111. In some examples, the number of parameters 113 is at least one million times greater than the number of parameters 111. In some examples, the number of parameters 111 is in a range between 10k and 100k. In some examples, the number of parameters 111 is less than 10k. In some examples, the number of parameters 113 is in a range between 1M and 10M. In some examples, the number of parameters 113 is greater than 10M.

[0070] In some examples, sound recognition operations (e.g., speech, alarms, or generally any type of sound) are distributed between device 102 and computing device 152. In some examples, sound recognition operations are distributed between device 102 and computing device 152. For example, microcontroller 106 is configured to perform inference operation 124-1 by invoking ML model 126 to detect whether a sound of interest is included in audio data 131 captured by microphone 140 on device 102. ML model 126 can be a classifier that classifies audio data 131 as containing a sound of interest or not containing a sound of interest. For example, ML model 126 receives audio data 131 from microphone 140 and calculates a prediction as to whether audio data 131 includes a sound of interest. If ML model 126 does not detect a sound of interest in audio data 131, ML model 126 continues to receive audio data 131 from microphone 140 as input to calculate a prediction as to whether a sound of interest is detected in audio data 131. If ML model 126 detects a sound of interest within audio data 131, device 102 streams audio data 131 (e.g., original sound, compressed sound, sound fragments, and / or audio parameters, etc.) to computing device 152 via wireless connection 148. In some examples, device 102 compresses audio data 131 and then sends the compressed audio data 131 to computing device 152 via wireless connection 148.

[0071] Computing device 152 receives audio data 131 from device 102 via wireless connection 148 and performs inference operation 124-2 by invoking ML model 127. ML model 127 can save power and latency due to its relatively small size. Computing device 152 includes a more powerful sound recognition engine (e.g., another type of classifier) ​​that executes ML model 127 (e.g., a larger ML model) to convert audio data 131 (potentially into text data), where computing device 152 transmits the text data back to device 102 via wireless connection 148 for display on the device's display. In some examples, computing device 152 is connected to server computer 160 via network 150 (e.g., the Internet), and computing device 152 transmits audio data 131 to server computer 160, where server computer 160 executes the larger ML model to convert audio data 131 into text data (e.g., in the case of translation into a different language). The text data is then routed back to computing device 152 and then to device 102 for display.

[0072] In some examples, image recognition operations are distributed between device 102 and computing device 152. In some examples, the image recognition operations include facial detection and tracking. However, the image recognition operations may include operations to detect (and track) other areas of interest in the image data, such as objects, text, and barcodes. The microcontroller 106 is configured to perform an inference operation 124-1 by calling the ML model 126 to detect whether an object of interest is included within the image data 129 captured by one or more imaging sensors 142 on the device 102. If so, the device 102 may transmit the image frame (including the object of interest) to the computing device 152 via the wireless connection 148. In some examples, the device 102 compresses the image frame and then transmits the compressed image frame to the computing device 152 via the wireless connection 148.

[0073] Computing device 152 is configured to perform inference operation 124-2 by invoking ML model 127 to perform more complex image processing operations using image data 129, such as computing object location data (e.g., a bounding box dataset) that identifies the location of an object of interest, wherein computing device 152 transmits the object location data back to device 102. Device 102 uses one or more low-complexity tracking mechanisms (e.g., IMU-based warping, blob detection, optical flow, etc.) to propagate the object location data for subsequent image frames captured on device 102. Device 102 may compress the cropped region and send it to computing device 152, wherein computing device 152 may perform image classification on the cropped region and send updated object location data back to device 102.

[0074] In some examples, sensing operations with multiple resolutions are distributed between the device 102 and the computing device 152. The sensing operations may include always-on sensing and sensing voice input requests (e.g., hot word detection). In some examples, when the user wears the device 102, the imaging sensor 142a (e.g., an LPLR camera) is activated to capture image data 129 at a relatively low resolution to search for an area of ​​interest. For example, the microcontroller 106 is configured to perform an inference operation 124-1 by calling the ML model 126 (using the image data 129 as input to the ML model 126) to detect an object of interest (e.g., a face, text, a barcode, a building, etc.). If an object of interest is detected, the imaging sensor 142b may be activated to capture one or more image frames with a higher resolution.

[0075] The higher-resolution image data 129 can then be transmitted to a computing device 152 via a wireless connection 148. In some examples, device 102 compresses the higher-resolution image data 129 and transmits the compressed image data 129 via the wireless connection 148. Computing device 152 is configured to perform inference operation 124-2 by invoking ML model 127 (which is fed with the higher-resolution image data 129) to perform image recognition. In some examples, the operation can be similar to the facial detection example described above, where object location data (e.g., a bounding box dataset) is calculated by computing device 152 and sent to device 102, and device 102 uses one or more tracking mechanisms to propagate the object location data to subsequent frames. Device 102 then crops and compresses the image region to be sent back to computing device 152 for further image classification. In some examples, the image stream of a product can be used to capture label text or barcodes and look up associated product information (e.g., price, shopping recommendations, comparable products, etc.). This information can be displayed on display 116 on device 102 or read back to the user audibly.

[0076] In terms of sensing a voice input request, microcontroller 106 is configured to perform inference operation 124-1 by invoking ML model 126 to continuously (e.g., periodically) process microphone samples (e.g., audio data 131) of an initial portion of a hot word (e.g., "ok G" or "ok D"). If ML model 126 detects the initial portion, microcontroller 106 may cause a buffer to capture subsequent audio data 131. In addition, device 102 may transmit a portion of the buffer (e.g., 1-2 seconds of audio starting from the beginning of the buffer) to computing device 152 via wireless connection 148. In some examples, the portion of the buffer is compressed before being transmitted to computing device 152. Computing device 152 is configured to perform inference operation 124-2 by invoking ML model 127 to perform full-hot word recognition using audio data 131. If the utterance is a false positive, computing device 152 may send a dismiss command to device 102, which discards the contents of the buffer. If the utterance is a true positive, the remainder of the audio buffer is compressed and transmitted to computing device 152 for automatic speech recognition and user-bound response generation.

[0077] In some examples, to improve transmission efficiency, device 102 may buffer multiple data packets 134 and transmit data packets 134 as a single transmission event 132 to computing device 152 via wireless connection 148. For example, each transmission event 132 may be associated with power consumption that causes power to be dissipated from battery 120. In some examples, device 102 determines the type of information to be transmitted to computing device 152. In some examples, if the type of information to be transmitted to computing device 152 involves latency-dependent information (e.g., audio streaming), device 102 may not buffer audio data 131, but instead stream audio data 131 without delay. In some examples, if the information to be transmitted is not latency-dependent information, device 102 may store the information as one or more data packets 134 in buffer 130 and transmit the information to computing device 152 at a later time. Buffer 130 may be part of memory device 110. In some examples, other non-latency related information may be combined with existing data in buffer 130 , and the information contained in buffer 130 may be transmitted to computing device 152 as a single transmission event 132 .

[0078] For example, buffer 130 may include data packet 136a and data packet 136b. Data packet 136a may include information obtained at a first time instance, and data packet 136b may include information obtained at a second time instance, where the second time instance is subsequent to the first time instance. However, rather than transmitting data packet 136a and data packet 136b as separate transmission events 132, device 102 may store data packet 136a and data packet 136b in buffer 130 and transmit data packet 136a and data packet 136b as a single transmission event 132. In this manner, the number of transmission events 132 may be reduced, which may improve the energy efficiency of delivering information to computing device 152.

[0079] Figure 2 A system 200 is illustrated for distributing image and / or audio processing across multiple devices including a device 202, a computing device 252, and a server computer 260. The system 200 may be Figure 1 2 is an example of a system 100 and may include any details disclosed with reference to those figures. Device 202 is connected to computing device 252 via wireless connection 248. In some examples, device 202 is a head-mounted display device, such as smart glasses. However, device 202 may be other types of low-power devices as discussed herein. Computing device 252 is connected to server computer 260 via network 250. Figure 2, the device 202 obtains sensor data 228 from one or more sensors 238 on the device 202. The sensor data 228 may include at least one of image data or audio data. Figure 1 The microcontroller 106 of the device 202 can perform an inference operation 224-1 by calling the ML model 226 to perform image and / or audio processing on the sensor data 228 to detect whether the sensor data 228 includes the type of data for which the ML model 226 was trained. In some examples, the device 202 can include multiple classifiers (e.g., multiple microcontrollers 106), where each classifier can make a decision to send the sensor data 228 (or a result of the decision) to another classifier, which can be on the device 202 or another device such as the computing device 252.

[0080] If the type of data for training ML model 226 is detected, device 202 can transmit sensor data 228 to computing device 252 via wireless connection 248. Computing device 252 can then transmit sensor data 228 to server computer 260 via network 250. In some examples, computing device 252 can include one or more classifiers that process the audio / image data captured by sensor 238 to make a decision about whether to invoke another classifier on computing device 252, device 202, or server computer 260. Server computer 260 includes one or more processors 262, which can be formed in a substrate configured to execute one or more machine-executable instructions or software, firmware, or a combination thereof. Processor 262 can be semiconductor-based, i.e., the processor can include semiconductor material that can execute digital logic. Server computer 260 includes one or more memory devices 264. Memory device 264 can include main memory that stores information in a format that can be read and / or executed by processor 262.

[0081] Server computer 260 is configured to perform inference operation 224-2 using sensor data 228 and ML model 229 stored on server computer 260. Inference operation 224-1 and inference operation 224-2 involve different audio and / or image processing operations. In some examples, inference operation 224-1 and inference operation 224-2 involve different audio processing operations. In some examples, inference operation 224-1 and inference operation 224-2 involve different image recognition operations. In some examples, inference operation 224-1 and inference operation 224-2 involve different perception operations.

[0082] The size of ML model 226 can be smaller (e.g., substantially smaller) than the size of ML model 229. ML model 226 can define a number of parameters 211 required for ML model 226 to make predictions. ML model 229 can define a number of parameters 215 required for ML model 229 to make predictions. The number of parameters 211 is smaller (e.g., substantially smaller) than the number of parameters 215. In some examples, the number of parameters 215 is at least one thousand times greater than the number of parameters 211. In some examples, the number of parameters 215 is at least one million times greater than the number of parameters 211. In some examples, the number of parameters 211 is in a range between 10k and 100k. In some examples, the number of parameters 211 is less than 10k. In some examples, the number of parameters 215 is in a range between 10M and 100M. In some examples, the number of parameters 215 is greater than 100M.

[0083] Figure 3 A system 300 is illustrated for distributing image and / or audio processing across multiple devices including a device 302, a computing device 352, and a server computer 360. The system 300 may be Figure 1 system 100 and / or Figure 2 302 is an example of a system 200 and may include any details disclosed with reference to those figures. Device 302 is connected to a computing device 352 via a wireless connection 348. In some examples, device 302 is a head-mounted display device, such as smart glasses. However, device 302 may be other types of low-power devices as discussed herein. Computing device 352 is connected to a server computer 360 via a network 350. Figure 3 , device 302 obtains sensor data 328 from one or more sensors 338 on device 302. Sensor data 328 may include at least one of image data or audio data. Figure 1 The microcontroller 106 may perform an inference operation 324 - 1 by calling the ML model 336 to perform image and / or audio processing on the sensor data 328 to detect whether the sensor data 328 includes the type of data for which the ML model 336 was trained.

[0084] If the type of data for training ML model 326 is detected, device 302 may transmit sensor data 328 to computing device 352 via wireless connection 348. Computing device 352 is configured to perform inference operation 324-2 using sensor data 328 and ML model 327 stored on computing device 352. Computing device 352 may then send the results of inference operation 324-2 and / or sensor data 328 to server computer 360 via network 350.

[0085] Server computer 360 is configured to perform inference operation 324-3 using the results of inference operation 324-2 and / or sensor data 328 and ML model 329 stored on server computer 360. Inference operation 324-1, inference operation 324-2, and inference operation 324-3 involve different audio and / or image processing operations. In some examples, inference operation 324-1, inference operation 324-2, and inference operation 324-3 involve different audio processing operations. In some examples, inference operation 324-1, inference operation 324-2, and inference operation 324-3 involve different image recognition operations. In some examples, inference operation 324-1, inference operation 324-2, and inference operation 324-3 involve different perception operations.

[0086] The size of ML model 326 can be smaller (e.g., substantially smaller) than the size of ML model 327. The size of ML model 327 can be smaller (e.g., substantially smaller) than the size of ML model 329. ML model 326 can define a plurality of parameters 311 required for ML model 326 to make predictions. ML model 327 can define a plurality of parameters 313 required for ML model 327 to make predictions. ML model 329 can define a plurality of parameters 315 required for ML model 329 to make predictions. The number of parameters 311 is smaller (e.g., substantially smaller) than the number of parameters 313. The number of parameters 313 is smaller (e.g., substantially smaller) than the number of parameters 315. In some examples, the number of parameters 311 is in a range between 10k and 100k. In some examples, the number of parameters 311 is less than 10k. In some examples, the number of parameters 313 is in a range between 100K and 1M. In some examples, the number of parameters 313 is greater than 1M. In some examples, the number of parameters 315 is in a range between 10M and 100M. In some examples, the number of parameters 315 is greater than 100M.

[0087] Figure 4 An example of a head mounted display device 402 according to an aspect is illustrated. The head mounted display device 402 may be Figure 1 Equipment 102, Figure 2 Device 202, and / or Figure 3An example of a device 302. The head mounted display device 402 includes smart glasses 469. Smart glasses 469 are glasses that add information (e.g., project a display 416) next to what the wearer is viewing through the glasses. In some examples, in addition to projecting information, the display 416 is an in-lens microdisplay. Smart glasses 469 (e.g., glasses or goggles) are vision aids that include lenses 472 (e.g., glass or hard plastic lenses) mounted in a frame 471 that typically holds them in front of a person's eyes with a bridge 473 that rests on the nose and legs 474 (e.g., temples or temple pieces) that rest on the ears. The smart glasses 469 include an electronic assembly 470 that includes circuitry for the smart glasses 469. In some examples, the electronic assembly 470 includes a housing that surrounds Figure 1 Equipment 102, Figure 2 Device 202 and / or Figure 3 The housing of the components of the device 302. In some examples, the electronic component 470 is included or integrated into one (or both) of the legs 474 of the smart glasses 469.

[0088] Figure 5 An example of an electronic component 570 of a pair of smart glasses according to an example is shown. The electronic component 570 may be Figure 4 5. The electronic components 570 of the smart glasses may include a display controller 518, a display 516, a flash memory 510, an RF transceiver 514, a universal serial bus (USB) interface 521, a power management unit (PMU) 508, a system on a chip (SOC) 504, a battery charger 522, a battery 520, a plurality of user controls 581, and a user light emitting diode (LED) 585. The display controller 518, the display 516, the RF transceiver 514, the battery charger 522, and the battery 520 may be Figure 1 The SOC 504 may include an example of a display regulator 118, a display 116, an RF transceiver 114, a battery charger 122, and a battery 120. Figure 1 The processor 104 (including the microcontroller 106) of the flash memory 510 can be Figure 1 The flash memory 510 may store weights for any ML models that may be executed by the SOC 504.

[0089] The SOC 504 can provide data and control information to a display 516 projected in the user's field of view. In some examples, the PMU 508 is included within or integrated with the SOC 504. A display regulator 518 is connected to the PMU 508. The display regulator 518 may include a first converter 576 (e.g., a VDDD DC-DC converter), a second converter 579 (e.g., a VDDA DC-DC converter), and an LED driver 580. The first converter 576 is configured to be activated in response to an enable signal, and the second converter 579 is configured to be activated in response to an enable signal. The LED driver 580 is configured to be driven according to a pulse width modulation (PWM) control signal. The plurality of user controls 581 may include a reset button 582, a power button 583, a first user button 584-1, and a second user button 584-2.

[0090] Figure 6 A printed circuit board (PCB) substrate 668 for smart glasses is shown according to one aspect. The PCB substrate 668 may be Figure 4 Electronic components 470 and / or Figure 5 Examples of electronic components 570 are provided and / or included therein. PCB substrate 668 includes multiple circuit components. In some examples, the circuit components are coupled to one side of PCB substrate 668. In some examples, the circuit components are coupled to both sides of PCB substrate 668. PCB substrate 668 may include battery charger 622, SOC 604, display flex 669, display regulator 618, and flash memory 610. PCB substrate 668 may be relatively small. For example, PCB substrate 668 may define a length (L) and a width (W). In some examples, the length (L) ranges from 40 mm to 80 mm. In some examples, the length (L) ranges from 50 mm to 70 mm. In some examples, the length (L) is 60 mm. In some examples, the width (W) ranges from 8 mm to 25 mm. In some examples, the width (W) ranges from 10 mm to 20 mm. In some examples, the width (W) is 14.5 mm.

[0091] Figure 7A and 7B A system 700 is shown for distributing voice recognition operations between a device 702 and a computing device 752. The system 700 may be Figure 1 System 100, Figure 2 System 200 and / or Figure 3 In some examples, device 702 may be an example of system 300 and may include any details discussed with reference to those figures. Figure 4In some examples, components of the device 702 may include: Figure 5 Electronic components 570 and / or Figure 6 electronic components 670.

[0092] like Figure 7A As shown, voice recognition operations are distributed between device 702 and computing device 752. Device 702 is connected to computing device 752 via wireless connection 748, such as a short-term wireless connection (such as a Bluetooth or NFC connection). In some examples, wireless connection 748 is a Bluetooth connection. In some examples, device 702 includes a voice recognition application that enables audio data 731 to be captured by microphone 740 on device 702 and text data 707 to be displayed on display 716 of device 702.

[0093] Device 702 includes a microcontroller 706 that executes a sound classifier 703 to detect whether a sound of interest (e.g., speech, alarm, etc.) is included in audio data 731 captured by a microphone 740 on device 702. Sound classifier 703 may include or be defined by an ML model 726. ML model 726 may define a number of parameters 711 required for ML model 726 to make a prediction (e.g., whether the sound of interest is included in audio data 731). ML model 726 may be relatively small because the actual conversion is offloaded to computing device 752. For example, the number of parameters 711 may be between 10k and 100k. Sound classifier 703 may save power and latency due to its relatively small ML model 726.

[0094] refer to Figure 7B In operation 721, the sound classifier 703 may receive audio data 731 from the microphone 740 on the device 702. In operation 723, the sound classifier 703 may determine whether a sound of interest is detected in the audio data 731. If no sound of interest is detected (No), the sound classifier 703 continues to monitor the audio data 731 received via the microphone 740 to determine whether a sound of interest is detected. If a sound of interest is detected (Yes), in operation 725, the device 702 streams the audio data 731 to the computing device 752 via the wireless connection 748. For example, the RF transceiver 714 on the device 702 may send the audio data 731 via the wireless connection 748. In some examples, the device 702 compresses the audio data 731 and then sends the compressed audio data 731 to the computing device 752.

[0095] refer to Figure 7A, computing device 752 includes a sound recognition engine 709 (e.g., another classifier) ​​that executes an ML model 727 (e.g., a larger ML model) to convert the sounds of audio data 731 into text data 707. ML model 727 may define a plurality of parameters 713 required for ML model 727 to make predictions. In some examples, the number of parameters 713 is at least ten times greater than the number of parameters 711. In some examples, the number of parameters 713 is at least one hundred times greater than the number of parameters 711. In some examples, the number of parameters 713 is at least one thousand times greater than the number of parameters 711. In some examples, the number of parameters 713 is at least one million times greater than the number of parameters 711. In some examples, the number of parameters 713 is between 1M and 10M. In some examples, the number of parameters 713 is greater than 10M. Computing device 752 transmits text data 707 to device 702 via wireless connection 748. Device 702 displays text data 707 on the device's display 716.

[0096] Figure 8 A system 800 is shown for distributing voice recognition operations between a device 802 and a server computer 860. The system 800 may be Figure 1 System 100, Figure 2 System 200, Figure 3 System 300 and / or Figure 7A and 7B In some examples, device 802 may be an example of system 700 and may include any details discussed with reference to those figures. Figure 4 In some examples, components of the device 802 may include: Figure 5 Electronic components 570 and / or Figure 6 electronic components 670.

[0097] like Figure 8 As shown, the voice recognition operation is distributed between the device 802 and the server computer 860, wherein the audio data 831 can be provided to the server computer 860 via the computing device 852. The device 802 is connected to the computing device 852 via a wireless connection 848, such as a short-term wireless connection (such as a Bluetooth or NFC connection). In some examples, the wireless connection 848 is a Bluetooth connection. The computing device 852 is connected to the server computer 860 via a network 850 (e.g., the Internet such as Wi-Fi or a mobile connection). In some examples, the device 802 includes a voice recognition application that enables the audio data 831 to be captured by the microphone 840 on the device 802 and the text data 807 to be displayed on the display 816 of the device 802.

[0098] Device 802 includes a microcontroller 806 that executes a sound classifier 803 to detect whether a sound of interest is included in audio data 831 captured by a microphone 840 on device 802. Sound classifier 803 may include or be defined by an ML model 826. ML model 826 may define a plurality of parameters 811 required for ML model 826 to make a prediction (e.g., whether a sound of interest is included in audio data 831). ML model 826 may be relatively small because the actual conversion is offloaded to server computer 860. For example, the number of parameters 811 may be in the range of between 10k and 100k. Sound classifier 803 may save power and latency through its relatively small ML model 826.

[0099] If no sound of interest is detected, the sound classifier 803 continues to monitor the audio data 831 received via the microphone 840 to determine whether a sound of interest is detected. If a sound of interest is detected, the device 802 streams the audio data 831 to the computing device 852 via the wireless connection 848. For example, the RF transceiver 814 on the device 802 can transmit the audio data 831 via the wireless connection 848. In some examples, the device 802 compresses the audio data 831 and then sends the compressed audio data 831 to the computing device 852.

[0100] In some examples, computing device 852 can transmit audio data 831 to server computer 860 via network 850. In some examples, computing device 852 determines whether computing device 852 has the ability to convert sound into text data 807. If not, computing device 852 can transmit audio data 831 to server computer 860. If so, computing device 852 can perform sound conversion, as described in reference to Figure 7A and 7B The system 700 is discussed.

[0101] In some examples, computing device 852 determines whether the voice conversion includes translation into another language. For example, audio data 831 may include speech in the English language, but parameters of the voice recognition application indicate that text data 807 is provided in another language, such as German. In some examples, if the conversation includes translation into another language, computing device 852 may transmit audio data 831 to server computer 860. In some examples, after receiving audio data 831 from device 802, computing device 852 may automatically transmit audio data 831 to server computer 860. In some examples, device 802 transmits audio data 831 directly to server computer 860 via network 850 (e.g., without using computing device 852), and device 802 receives text data 807 from server computer 860 via network 850 (e.g., without using computing device 852).

[0102] Server computer 860 includes a speech recognition engine 809 that executes an ML model 829 (e.g., a larger ML model) to convert the speech of audio data 831 into text data 807. In some examples, the speech-to-text data conversion 807 includes translation into a different language. ML model 829 can define a plurality of parameters 815 required for ML model 829 to make predictions (e.g., the speech-to-text data 807 conversion). In some examples, the number of parameters 815 is at least a thousand times greater than the number of parameters 811. In some examples, the number of parameters 815 is at least a million times greater than the number of parameters 811. In some examples, the number of parameters 815 is at least 100 million times greater than the number of parameters 811. In some examples, the number of parameters 815 is between 1M and 100M. In some examples, the number of parameters 815 is greater than 100M. Server computer 860 transmits text data 807 to computing device 852 via network 850. Computing device 852 transmits text data 807 to device 802 via wireless connection 848. Device 802 displays text data 807 on the device's display 816.

[0103] Figure 9 A system 900 is shown that performs voice recognition operations using a device 902. The system 900 may be Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 7A and 7B System 700 and / or Figure 8 In some examples, device 902 may be an example of system 800 and may include any details discussed with reference to those figures. Figure 4In some examples, components of the device 902 may include: Figure 5 Electronic components 570 and / or Figure 6 electronic components 670.

[0104] Device 902 is connected to a computing device 952 via a wireless connection 948, such as a short-term wireless connection (such as a Bluetooth or NFC connection). In some examples, wireless connection 948 is a Bluetooth connection. Computing device 952 may include a microphone 921 configured to capture audio data 931, and a voice recognition engine 909 configured to convert the sound of audio data 931 into text data 907. Voice recognition engine 909 may include or be defined by an ML model, as discussed with reference to the previous figures. After converting the sound into text data 907, computing device 952 may transmit text data 907 to device 902 via wireless connection 948, and device 902 receives text data 907 via RF transceiver 914 on device 902. Device 902 is configured to display text data 907 on a display 916 of device 902.

[0105] Figure 10 A system 1000 is shown for performing voice recognition operations using a device 1002. The system 1000 may be Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 7A and 7B System 700, Figure 8 System 800 and / or Figure 9 In some examples, device 1002 may be an example of system 900 and may include any details discussed with reference to those figures. Figure 4 In some examples, components of device 1002 may include Figure 5 Electronic components 570 and / or Figure 6 electronic components 670.

[0106] like Figure 10 As shown, voice recognition operations are distributed between computing device 1052 and server computer 1060, where text data 1007 is displayed via device 1002. Device 1002 is connected to computing device 1052 via wireless connection 1048, such as a short-term wireless connection (such as Bluetooth or NFC connection). In some examples, wireless connection 1048 is a Bluetooth connection. Computing device 1052 is connected to server computer 1060 via network 1050 (e.g., the Internet, such as Wi-Fi or a mobile connection).

[0107] The computing device 1052 includes a microphone 1021 configured to capture audio data 1031. In addition, the computing device 1052 includes a sound classifier 1003 (e.g., an ML model) to detect whether a sound of interest is included in the audio data 1031 captured by the microphone 1021 on the computing device 1052. If a sound of interest is not detected, the sound classifier 1003 continues to monitor the audio data 1031 received via the microphone 1021 to determine whether a sound of interest is detected. If a sound of interest is detected, the computing device 1052 streams the audio data 1031 to the server computer 1060 over the network 1050. In some examples, the computing device 1052 determines whether the computing device 1052 has the ability to convert the sound into text data 1007. If not, the computing device 1052 can transmit the audio data 1031 to the server computer 1060. If so, the computing device 1052 can perform the sound conversion, as described in reference to Figure 9 In some examples, computing device 1052 compresses audio data 1031 and sends the compressed audio data 1031 to server computer 1060.

[0108] In some examples, computing device 1052 determines whether the voice conversion includes translation into another language. For example, audio data 1031 may include speech in the English language, but parameters of the speech-to-text application indicate that text data 1007 is provided in a different language. In some examples, if the speech-to-text conversion includes translation into another language, computing device 1052 may transmit audio data 1031 to server computer 1060. In some examples, upon detecting speech within audio data 1031, computing device 1052 may automatically transmit audio data 1031 to server computer 1060.

[0109] Server computer 1060 includes a voice recognition engine 1009 that executes an ML model to convert the voice of audio data 1031 into text data 1007. In some examples, the conversion of voice into text data 1007 includes translation into a different language. Server computer 1060 transmits text data 1007 to computing device 1052 via network 1050. Computing device 1052 transmits text data 1007 to RF transceiver 1014 on device 1002 via wireless connection 1048. Device 1002 displays text data 1007 on the device's display 1016.

[0110] Figure 11 It is a depiction Figure 7A and Figure 7B 1100 is a flow chart of exemplary operations of the system 700. Although reference is made to Figure 7A and Figure 7B The system 700 illustrates Figure 11 1100, but the flowchart 1100 can be applied to any embodiment discussed herein, including Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 4 head-mounted display device 402, Figure 5 Electronic components 570, Figure 6 Electronic components 670, Figure 8 System 800, Figure 9 System 900 and / or Figure 10 System 1000. Although Figure 11 The flowchart 1100 shows the operations in a sequential order, but it should be appreciated that this is merely an example and additional or alternative operations may be included. Figure 11 The operations and related operations may be performed in a different order than shown, or in parallel or in an overlapping manner.

[0111] Operation 1102 includes receiving audio data 731 via microphone 740 of device 702. Operation 1104 includes detecting, by sound classifier 703, whether the audio data 731 includes a sound of interest (eg, speech), wherein sound classifier 703 executes a first ML model (eg, ML model 726).

[0112] Operation 1106 includes transmitting audio data 731 to computing device 752 via wireless connection 748, wherein audio data 731 is configured for use by computing device 752 to translate the sound of interest into text data 707 using a second ML model (e.g., ML model 727). Operation 1108 includes receiving text data 707 from computing device 752 via wireless connection 748. Operation 1110 includes displaying, by device 702, text data 707 on display 716 of device 702.

[0113] Figure 12 It is a depiction Figure 8 1200 is a flow chart of exemplary operations of the system 800. Although reference is made to Figure 8 The system 800 illustrates Figure 12 1200, but the flowchart 1200 can be applied to any embodiment discussed herein, including Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 4 head-mounted display device 402, Figure 5 Electronic components 570, Figure 6Electronic components 670, Figure 7A and 7B System 700, Figure 9 System 900 and / or Figure 10 System 1000. Although Figure 12 Flowchart 1200 shows the operations in a sequential order, but it should be appreciated that this is merely an example and additional or alternative operations may be included. Figure 12 The operations and related operations may be performed in a different order than shown, or in parallel or in an overlapping manner.

[0114] Operation 1202 includes receiving audio data 831 via microphone 840 of device 802. Operation 1204 includes detecting, by a sound classifier 803 of device 802, whether the audio data 831 includes a sound of interest (e.g., speech), wherein the sound classifier 803 executes a first ML model (e.g., ML model 826).

[0115] Operation 1206 includes transmitting, by device 802, audio data 831 to computing device 852 via wireless connection 848, wherein audio data 831 is further transmitted to server computer 860 over network 850 for translation of the sounds into text data 807 using a second ML model (e.g., ML model 829). Operation 1208 includes receiving, by device 802, text data 807 from computing device 852 via wireless connection 848. Operation 1210 includes displaying, by device 802, text data 807 on display 816 of device 802.

[0116] Figures 13A to 13C A system 1300 is shown for distributing image recognition operations between a device 1302 and a computing device 1352. The system 1300 may be Figure 1 System 100, Figure 2 System 200 and / or Figure 3 In some examples, device 1302 may be an example of system 300 and may include any details discussed with reference to those figures. Figure 4 In some examples, components of the device 1302 may include: Figure 5 Electronic components 570 and / or Figure 6 In some examples, the system 1300 also includes the capability of distributed voice recognition operations and may include reference Figure 7A and 7B System 700, Figure 8 System 800, Figure 9 System 900 and / or Figure 10 Any details discussed in detail for system 1000.

[0117] like Figure 13A As shown, image recognition operations are distributed between device 1302 and computing device 1352. In some examples, image recognition operations include facial detection and tracking. However, image recognition operations may include operations to detect (and track) other areas of interest in image data, such as objects, text, and barcodes. Device 1302 is connected to computing device 1352 via a wireless connection 1348 (such as a short-term wireless connection such as Bluetooth or NFC). In some examples, wireless connection 1348 is a Bluetooth connection. In some examples, device 1302 and / or computing device 1352 include an image recognition application that enables identification (and tracking) of objects via image data captured by one or more imaging sensors 1342.

[0118] Device 1302 includes a microcontroller 1306 that executes an image classifier 1303 to detect whether an object of interest 1333 is included in image data 1329 captured by an imaging sensor 1342 on device 1302. In some examples, object of interest 1333 includes facial features. In some examples, object of interest 1333 includes text data. In some examples, object of interest 1333 includes an OCR code. However, object of interest 1333 may be any type of object that can be detected in image data. Image classifier 1303 may include or be defined by an ML model 1326. ML model 1326 may define a number of parameters 1311 required for ML model 1326 to make a prediction (e.g., whether object of interest 1333 is included in image data 1329). ML model 1326 may be relatively small because some of the more intensive image recognition operations are offloaded to computing device 1352. For example, the number of parameters 1311 may range between 10k and 100k. The image classifier 1303 can save power and latency through its relatively small ML model 1326.

[0119] refer to Figure 13BIn operation 1321, image classifier 1303 may receive image data 1329 from imaging sensor 1342 on device 1302. In operation 1323, image classifier 1303 may be activated. In operation 1325, image classifier 1303 may determine whether object of interest 1333 is detected in image frame 1329a of image data 1329. If object of interest 1333 is not detected (No), image classifier 1303 (and / or imaging sensor 1342) may transition to a power saving state in operation 1328. In some examples, after a period of time, image classifier 1303 may be reactivated (e.g., the process returns to operation 1323) to determine whether object of interest 1333 is detected in image frame 1329a of image data 1329. If the object of interest is detected 1333 (YES), then in operation 1330, device 1302 transmits image frame 1329a to computing device 1352 via wireless connection 1348. For example, RF transceiver 1314 on device 1302 may transmit image frame 1329a via wireless connection 1348. In some examples, device 1302 compresses image frame 1329a and sends compressed image frame 1329a to computing device 1352.

[0120] refer to Figure 13A , computing device 1352 includes object detector 1309 that executes ML model 1327 (e.g., a larger ML model) to compute bounding box dataset 1341. In some examples, bounding box dataset 1341 is an example of object location data. Bounding box dataset 1341 can be data defining where an object of interest 1333 (e.g., a facial feature) is located within image frame 1329a. In some examples, reference Figure 13C , bounding box dataset 1341 defines coordinates of a bounding box 1381 that includes an object of interest 1333 within image frame 1329a. In some examples, the coordinates include a height coordinate 1383, a left coordinate 1385, a top coordinate 1387, and a width coordinate 1389. For example, height coordinate 1383 may be the height of bounding box 1381 as a ratio of the overall image height. Left coordinate 1385 may be the left coordinate of bounding box 1381 as a ratio of the overall image width. Top coordinate 1387 may be the top coordinate of bounding box 1381 as a ratio of the overall image height. Width coordinate 1389 may be the width of bounding box 1381 as a ratio of the overall image width.

[0121] ML model 1327 may define a number of parameters 1313 required for ML model 1327 to make predictions (e.g., calculation of bounding box dataset 1341). In some examples, the number of parameters 1313 is at least ten times greater than the number of parameters 1311. In some examples, the number of parameters 1313 is at least one hundred times greater than the number of parameters 1311. In some examples, the number of parameters 1313 is at least one thousand times greater than the number of parameters 1311. In some examples, the number of parameters 1313 is at least one million times greater than the number of parameters 1311. In some examples, the number of parameters 1313 is in a range between 1M and 10M. In some examples, the number of parameters 1313 is greater than 10M. Computing device 1352 transmits bounding box dataset 1341 to device 1302 via wireless connection 1348.

[0122] Device 1302 includes an object tracker 1335 configured to track an object of interest 1333 in one or more subsequent image frames 1329b using a bounding box dataset 1341. In some examples, object tracker 1335 is configured to perform a low-complexity tracking mechanism, such as warping, blob detection, or optical flow based on an inertial measurement unit (IMU). For example, object tracker 1335 may propagate bounding box dataset 1341 for subsequent image frames 1329b. Object tracker 1335 may include a cropper 1343 and a compressor 1345. Cropper 1343 may use bounding box dataset 1341 to identify an image region 1347 within image frame 1329b. Compressor 1345 may compress image region 1347. For example, image region 1347 may represent a region within image frame 1329b that has been cropped and compressed by object tracker 1335.

[0123] Device 1302 may then transmit image region 1347 to computing device 1352 via wireless connection 1348. For example, when object tracker 1335 is tracking object of interest 1333, computing device 1352 may receive image region stream 1347. At computing device 1352, object detector 1309 may perform image recognition on image region 1347 received from device 1302 via wireless connection 1348. In some examples, if object of interest 1333 is relatively close to the edge of image region 1347 (or not present at all), computing device 1352 may transmit a request to send a new complete frame (e.g., new image frame 1329a) again to recalculate bounding box dataset 1341. In some examples, if image frame 1329a does not contain object of interest 1333, computing device 1352 may transmit a request to enter a power save state to poll for objects of interest. In some examples, a visual indicator 1351 (eg, a visual box) may be provided on the display 1316 of the device 1302 , where the visual indicator 1351 identifies the object of interest 1333 (eg, a facial feature).

[0124] Figure 14 It is a depiction Figures 13A to 13C 1400 is a flow chart of exemplary operations of the system 1300. Although reference is made to 13A to 13C The system 1300 explains Figure 14 1400, but the flowchart 1400 can be applied to any embodiment discussed herein, including Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 4 head-mounted display device 402, Figure 5 Electronic components 570 and / or Figure 6 Electronic components 670, Figure 7A and 7B System 700. Although Figure 14 The flowchart 1400 shows the operations in a sequential order, but it should be appreciated that this is merely an example and additional or alternative operations may be included. Figure 14 The operations and related operations may be performed in a different order than shown, or in parallel or in an overlapping manner. In some examples, Figure 14 The operations of the flowchart 1400 may be Figure 11 Flowchart 1100 and / or Figure 12 The operations of flowchart 1200 are combined.

[0125] Operation 1402 includes receiving image data 1329 via at least one imaging sensor 1342 on the device 1302. Operation 1404 includes detecting, by the image classifier 1303 of the device 1302, whether the object of interest 1333 is included within the image data 1329, wherein the image classifier 1303 executes the ML model 1326.

[0126] Operation 1406 includes transmitting the image data 1329 (e.g., image frame 1329a) to the computing device 1352 via the wireless connection 1348, wherein the image frame 1329a includes the object of interest 1333. The image data 1329 is configured to be used by the computing device 1352 for image recognition using the ML model 1327.

[0127] Operation 1408 includes receiving bounding box dataset 1341 from computing device 1352 via wireless connection 1348. Operation 1410 includes using, by device 1302, bounding box dataset 1341 to identify image region 1347 in subsequent image data (e.g., image frame 1329b). Operation 1412 includes transmitting image region 1347 to computing device 1352 via wireless connection 1348, wherein image region 1347 is configured to be used by computing device 1352 for image recognition.

[0128] Figure 15 A system 1500 is illustrated for distributing image recognition operations between a device 1502 and a computing device 1552. The system 1500 may be Figure 1 System 100, Figure 2 System 200, Figure 3 System 300 and / or Figures 13A to 13C In some examples, device 1502 may be an example of system 1300 and may include any details discussed with reference to those figures. Figure 4 In some examples, components of device 1502 may include Figure 5 Electronic components 570 and / or Figure 6 In some examples, the system 1500 also includes the capability of distributed voice recognition operations and may include reference Figure 7A and 7B System 700, Figure 8 System 800, Figure 9 System 900 and / or Figure 10 Any details discussed in detail for system 1000.

[0129] like Figure 15As shown, image recognition operations are distributed between device 1502 and computing device 1552. In some examples, image recognition operations include facial detection and tracking. However, image recognition operations may include operations to detect (and track) other areas of interest in image data, such as objects, text, and barcodes. Device 1502 is connected to computing device 1552 via a wireless connection 1548 (such as a short-term wireless connection, such as a Bluetooth or NFC connection). In some examples, wireless connection 1548 is a Bluetooth connection. In some examples, device 1502 and / or computing device 1552 include an image recognition application that enables identification (and tracking) of objects via image data captured by imaging sensor 1542a and imaging sensor 1542b.

[0130] Imaging sensor 1542a can be considered a low-power, low-resolution (LPLR) image sensor. Imaging sensor 1542b can be considered a high-power, high-resolution (HPHR) image sensor. The resolution 1573b of image frame 1529b captured by imaging sensor 1542b is higher than the resolution 1573a of image frame 1529a captured by imaging sensor 1542a. In some examples, imaging sensor 1542a is configured to obtain image data (e.g., image frame 1529a) when device 1502 is activated and coupled to a user (e.g., continuously or periodically capture image frame 1529a when device 1502 is activated). In some examples, imaging sensor 1542a is configured to operate as a normally-on sensor. In some examples, imaging sensor 1542b is activated (e.g., for a short duration) in response to detecting an object of interest, as discussed further below.

[0131] Device 1502 includes a lighting condition sensor 1544 that is configured to estimate lighting conditions for capturing image data. In some examples, lighting condition sensor 1544 includes an ambient light sensor that detects the amount of ambient light present, which can be used to ensure that image frame 1529a is captured with a desired signal-to-noise ratio (SNR). However, lighting condition sensor 1544 may include other types of photometric (or colorimetric) sensors. Motion sensor 1546 can be used to monitor device movement, such as tilt, shake, rotation, and / or swing and / or for blur estimation. Sensor trigger 1571 can receive lighting condition information from lighting condition sensor 1544 and motion information from motion sensor 1546, and if the lighting condition information and motion information indicate that conditions are acceptable to obtain image frame 1529a, sensor trigger 1571 can activate imaging sensor 1542a to capture image frame 1529a.

[0132] The device 1502 includes a microcontroller 1506 that is configured to execute an image classifier 1503 that detects whether an object of interest is included in an image frame 1529a captured by an imaging sensor 1542a. Similar to other embodiments, the image classifier 1503 may include an ML model or be defined by an ML model. The ML model may define a plurality of parameters required for the ML model to make a prediction (e.g., whether an object of interest is included in the image frame 1529a). The ML model may be relatively small because some of the more intensive image recognition operations are offloaded to the computing device 1552. For example, the number of parameters may be in the range of between 10k and 100k. The image classifier 1503 may save power and latency due to its relatively small ML model.

[0133] If image classifier 1503 detects the presence of an object of interest within image frame 1529a, image classifier 1503 is configured to trigger imaging sensor 1542b to capture image frame 1529b. As described above, image frame 1529b has a higher resolution 1573b than resolution 1573a of image frame 1529a. Device 1502 transmits image frame 1529b to computing device 1552 via wireless connection 1548 for further processing. In some examples, device 1502 compresses image frame 1529b and then transmits the compressed image frame 1529b to computing device 1552. In some examples, motion information and / or lighting condition information is used to determine whether to transmit image frame 1529b. For example, if the motion information indicates motion above a threshold level (e.g., motion high), image frame 1529b may not be transmitted, and microcontroller 1506 may activate imaging sensor 1542b to capture another image frame. If the lighting condition information indicates that the lighting condition is below a threshold level, image frame 1529b may not be transmitted and microcontroller 1506 may activate imaging sensor 1542b to capture another image frame.

[0134] Computing device 1552 includes an object detector 1509 configured to perform image recognition operations (including computation of bounding box datasets) using image frames 1529b. Figures 13A to 13CIn an embodiment of the system 1300, the object detector 1509 executes a larger ML model to compute a bounding box dataset using a higher resolution image (e.g., image frame 1529b), which is transmitted back to the device 1502 via the wireless connection 1548. The device 1502 then uses the bounding box dataset to track the object of interest in one or more subsequent image frames. For example, the device 1502 can use a low-complexity tracking mechanism (such as warping, blob detection, or optical flow based on an inertial measurement unit (IMU)) to propagate the bounding box dataset for subsequent image frames. The device 1502 can use the bounding box dataset to identify an image region within the image frame 1529b, and the device 1502 can compress the image region and then transmit it back to the computing device 1552 for image recognition.

[0135] Figure 16 It is a depiction Figure 15 1600 is a flow chart of exemplary operations of the system 1500. Although reference is made to Figure 15 The System 1500 explains Figure 16 1600, but the flowchart 1600 can be applied to any embodiment discussed herein, including Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 4 head-mounted display device 402, Figure 5 Electronic components 570, Figure 6 Electronic components 670 and / or Figures 13A to 13C System 1300. Although Figure 16 Flowchart 1600 shows the operations in a sequential order, but it should be appreciated that this is merely an example and additional or alternative operations may be included. Figure 16 The operations and related operations may be performed in a different order than shown, or in parallel or in an overlapping manner. In some examples, Figure 16 The operations of the flowchart 1600 may be Figure 11 Flowchart 1100, Figure 12 Flowchart 1200 and / or Figure 14 The operations of flowchart 1400 are combined.

[0136] Operation 1602 includes receiving, by a first imaging sensor (eg, imaging sensor 1542a) of the device 1502, a first image frame 1529a. Operation 1604 includes detecting, by the image classifier 1503 of the device 1502, the presence of an object of interest in the first image frame 1529a.

[0137] Operation 1606 includes receiving, by a second imaging sensor of device 1502 (e.g., imaging sensor 1542b), a second image frame 1529b having a higher resolution 1573b than the resolution 1573a of first image frame 1529a, wherein second image frame 1529b is transmitted to computing device 1552 via wireless connection 1548, and second image frame 1529b is configured for use by object detector 1509 at computing device 1552.

[0138] Figure 17 A system 1700 is illustrated for distributing image recognition operations between a device 1702 and a computing device 1752. The system 1700 may be Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figures 13A to 13C System 1300 and / or Figure 15 In some examples, device 1702 may be an example of system 1500 and may include any details discussed with reference to those figures. Figure 4 In some examples, components of device 1702 may include Figure 5 Electronic components 570 and / or Figure 6 In some examples, the system 1700 also includes the capability of distributed voice recognition operations and may include reference Figure 7A and 7B System 700, Figure 8 System 800, Figure 9 System 900 and / or Figure 10 Any details discussed in detail for system 1000.

[0139] like Figure 17 As shown, image recognition operations are distributed between device 1702 and computing device 1752. In some examples, image recognition operations include facial detection and tracking. However, image recognition operations may include operations to detect (and track) other areas of interest in image data, such as objects, text, and barcodes. Device 1702 is connected to computing device 1752 via a wireless connection 1748 (such as a short-term wireless connection, such as a Bluetooth or NFC connection). In some examples, wireless connection 1748 is a Bluetooth connection. In some examples, device 1702 and / or computing device 1752 include an image recognition application that enables identification (and tracking) of objects via image data captured by imaging sensor 1742a and imaging sensor 1742b.

[0140] Imaging sensor 1742a can be considered a low-power, low-resolution (LPLR) image sensor. Imaging sensor 1742b can be considered a high-power, high-resolution (HPHR) image sensor. The resolution 1773b of image frame 1729b captured by imaging sensor 1742b is higher than the resolution 1773a of image frame 1729a captured by imaging sensor 1742a. ​​In some examples, imaging sensor 1742a is configured to obtain image data (e.g., image frame 1729a) when device 102 is activated and coupled to a user (e.g., continuously or periodically capture image frame 1729a while device 102 is activated). In some examples, imaging sensor 1742a is configured to operate as a normally-on sensor. In some examples, imaging sensor 1742b is activated (e.g., for a short duration) in response to detecting an object of interest, as discussed further below.

[0141] Device 1702 includes a light condition sensor 1744 that is configured to estimate the lighting conditions for capturing image data. In some examples, light condition sensor 1744 includes an ambient light sensor that detects the amount of ambient light present, which can be used to ensure that image frame 1729a is captured with a desired signal-to-noise ratio (SNR). However, light condition sensor 1744 may include other types of photometric (or colorimetric) sensors. Motion sensor 146 can be used to monitor device movement, such as tilt, shake, rotation, and / or swing, and / or for blur estimation. Sensor trigger 1771 can receive light condition information from light condition sensor 1744 and motion information from motion sensor 1746, and if the light condition information and motion information indicate that conditions are acceptable to obtain image frame 1729a, sensor trigger 1771 can activate imaging sensor 1742a to capture image frame 1729a.

[0142] Device 1702 includes a microcontroller 1706 configured to execute a classifier 1703 that detects whether a region of interest (ROI) 1789 is included within an image frame 1729a captured by an imaging sensor 1742a. ​​ROI 1789 is also referred to as an object of interest. Classifier 1703 may include or be defined by an ML model. The ML model may define a plurality of parameters required for the ML model to make a prediction (e.g., whether ROI 1789 is included within image frame 1729a). The ML model may be relatively small because some of the more intensive image recognition operations are offloaded to computing device 1752. For example, the number of parameters may be in the range of between 10k and 100k. Classifier 1703 may save power and latency due to its relatively small ML model.

[0143] If classifier 1703 detects the presence of ROI 1789 within image frame 1729a, classifier 1703 is configured to trigger imaging sensor 1742b to capture image frame 1729b. As described above, image frame 1729b has a higher resolution 1773b than resolution 1773a of image frame 1729a. Device 1702 transmits image frame 1729b to computing device 1752 via wireless connection 1748 for further processing. In some examples, device 1702 compresses image frame 1729b and transmits the compressed image frame 1729b to computing device 1752.

[0144] Computing device 1752 includes ROI classifier 1709, which executes an ML model (e.g., a larger ML model) to calculate ROI dataset 1741. In some examples, ROI dataset 1741 is an example of object location data and / or a bounding box dataset. ROI dataset 1741 can be data defining the location of ROI 1789 within image frame 1729b. Computing device 1752 can transmit ROI dataset 1741 to device 1702 via wireless connection 1748.

[0145] Device 1702 includes an ROI tracker 1735 configured to track an ROI 1789 in one or more subsequent image frames using an ROI dataset 1741. In some examples, ROI tracker 1735 is configured to implement a low-complexity tracking mechanism, such as warping based on an inertial measurement unit (IMU), blob detection, or optical flow. For example, ROI classifier 1709 can propagate ROI dataset 1741 for subsequent image frames. ROI tracker 1735 can include a cropper 1743 and a compressor 1745. Cropper 1743 can use ROI dataset 1741 to identify an image region 1747 within image frame 1729b. Compressor 1745 can compress image region 1747. For example, image region 1747 can represent a region within image frame 1729b that has been cropped and compressed by ROI tracker 1735, where image region 1747 includes ROI 1789.

[0146] Device 1702 may then transmit image region 1747 to computing device 1752 via wireless connection 1748. For example, while ROI tracker 1735 is tracking ROI 1789, computing device 1752 may receive a stream of image region 1747. At computing device 1752, ROI classifier 1709 may perform object detection on image region 1747 received from device 1702 via wireless connection 1748. In some examples, if ROI 1789 is relatively close to the edge of image region 1747 (or does not exist at all), computing device 1752 may transmit a request to send a new, complete frame (e.g., new image frame 1729b) to recalculate ROI dataset 1741. In some examples, if image frame 1729a does not contain ROI 1789, computing device 1752 may send a request to enter a power-saving state to poll for ROI 1789. In some examples, a visual indicator 1787 is provided on the display 1716 of the device 1702 , where the visual indicator 1787 identifies the ROI 1789 .

[0147] Figure 18 A system 1800 is illustrated for distributing image recognition operations between a device 1802 and a computing device 1852. The system 1800 may be Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figures 13A to 13C System 1300, Figure 15 System 1500 and Figure 17 In some examples, device 1802 may be an example of system 1700 and may include any details discussed with reference to those figures. Figure 4 In some examples, components of device 1802 may include Figure 5 Electronic components 570 and / or Figure 6 In some examples, the system 1800 also includes the capability of distributed voice recognition operations and may include reference Figure 7A and 7B System 700, Figure 8 System 800, Figure 9 System 900 and / or Figure 10 Any details discussed in detail for system 1000.

[0148] like Figure 18As shown, image recognition operations are distributed between device 1802 and computing device 1852. In some examples, image recognition operations include facial detection and tracking. However, image recognition operations may include operations to detect (and track) other areas of interest in image data (such as objects, text, and barcodes). Device 1802 is connected to computing device 1852 via a wireless connection (e.g., radio resources 1867) (such as a short-term wireless connection, such as a Bluetooth or NFC connection). In some examples, the wireless connection is a Bluetooth connection. In some examples, device 1802 and / or computing device 1852 include an image recognition application that enables identification (and tracking) of objects via image data captured by camera 1842a and camera 1842b.

[0149] Camera 1842a can be considered a low-power, low-resolution (LPLR) camera. Camera 1842b can be considered a high-power, high-resolution (HPHR) camera. The resolution of the image frames captured by camera 1842b is higher than the resolution of the image frames captured by camera 1842a. In some examples, camera 1842a is configured to obtain image data when device 1802 is activated and coupled to a user (e.g., continuously or periodically capture image frames when device 1802 is activated). In some examples, camera 1842a is configured to operate as a normally-on sensor. In some examples, camera 1842b is activated (e.g., for a short duration) in response to detecting an object of interest, as discussed further below.

[0150] Device 1802 includes a lighting condition sensor 1844 configured to estimate lighting conditions for capturing image data. In some examples, lighting condition sensor 1844 includes an ambient light sensor that detects the amount of ambient light present, which can be used to ensure that image frames are captured with a desired signal-to-noise ratio (SNR). However, lighting condition sensor 1844 may include other types of photometric (or colorimetric) sensors. Motion sensor 1846 can be used to monitor device movement, such as tilt, shake, rotation, and / or swing, and / or for blur estimation. Sensor trigger 1871 can receive lighting condition information from lighting condition sensor 1844 and motion information (e.g., blur estimation) from motion sensor 1846, and if the lighting condition information and motion information indicate acceptable conditions for obtaining an image frame, sensor trigger 1871 can activate camera 1842a to capture an image frame with a low resolution. In some examples, device 1802 includes microphone 1840 that provides audio data to classifier 1803.

[0151] Device 1802 includes a classifier 1803 that detects whether a region of interest is included within an image frame captured by camera 1842a. Classifier 1803 may include an ML model or be defined by an ML model. The ML model may define a plurality of parameters required for the ML model to make a prediction (e.g., whether a region of interest is included within an image frame). The ML model may be relatively small because some of the more intensive image recognition operations are offloaded to computing device 1852. For example, the number of parameters may be in the range of between 10k and 100k. Classifier 1803 may save power and latency due to its relatively small ML model.

[0152] If classifier 1803 detects the presence of a region of interest within an image frame captured by camera 1842a, classifier 1803 is configured to trigger camera 1842b to capture a higher resolution image. In some examples, device 1802 transmits the full image frame captured by camera 1842b via radio resource 1867.

[0153] The computing device 1852 includes a classifier 1809 that executes an ML model (e.g., a larger ML model) to calculate an ROI dataset (e.g., object box, x, y). The ROI dataset can be data that defines the location of an object of interest within an image frame. The computing device 1852 can transmit the ROI dataset to the device 1802. The classifier 1803 can provide the ROI dataset to the processor 1843, which crops subsequent image frames to identify image regions. The image regions are compressed by the compressor 1845 and transmitted to the computing device 1852 via the radio resource 1867. In some examples, the device 1802 includes an action manager 1865 that receives the ROI detection from the classifier 1809 and can provide a visual indicator or other action on the display 1816 of the device 1802.

[0154] Figure 19 It is a depiction Figure 17 1900 is a flow chart of exemplary operations of the system 1700. Although reference is made to Figure 17 The System 1700 explains Figure 19 1900, but the flowchart 1900 can be applied to any embodiment discussed herein, including Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 4 head-mounted display device 402, Figure 5 Electronic components 570, Figure 6 Electronic components 670, Figure 7A and 7B System 700, Figure 8 System 800, Figure 9 System 900, Figure 10 System 1000, 13A to 13C System 1300, Figure 15 System 1500 and / or Figure 18 System 1800. Although Figure 19 Flowchart 1900 shows the operations in a sequential order, but it should be appreciated that this is merely an example and additional or alternative operations may be included. Figure 19 The operations and related operations may be performed in a different order than shown, or in parallel or in an overlapping manner. In some examples, Figure 19 The operations of the flowchart 1900 may be Figure 11 Flowchart 1100, Figure 12 Flowchart 1200, Figure 14 Flowchart 1400 and / or Figure 16 The operations of flowchart 1600 are combined.

[0155] Operation 1902 includes activating a first imaging sensor 1742a of the device 1702 to capture first image data (e.g., image frame 1729a). Operation 1904 includes detecting, by a classifier 1703 of the device 1702, whether a region of interest (ROI) 1789 is included in the first image data, wherein the classifier 1703 executes the ML model.

[0156] Operation 1906 includes, in response to detecting the ROI 1789 within the first image data, activating the second imaging sensor 1742b of the device 1702 to capture second image data (e.g., image frame 1729b). The resolution 1773b of the second image data is higher than the resolution 1773a of the first image data. Operation 1908 includes transmitting the second imaging data to the computing device 1752 via the wireless connection 1748, wherein the second image data 1729b is used by the computing device 1752 for image processing using the ML model.

[0157] Figure 20 A system 2000 is illustrated for distributing image recognition operations between a device 2002 and a computing device 2052. The system 2000 may be Figure 1 System 100, Figure 2 System 200 and / or Figure 3 In some examples, device 2002 may be an example of system 300 and may include any details discussed with reference to those figures. Figure 4 In some examples, components of device 2002 may include Figure 5 Electronic components 570 and / or Figure 6In some examples, the system 2000 also includes the capability of distributed voice recognition operations and may include reference Figure 7A and 7B System 700, Figure 8 System 800, Figure 9 System 900 and / or Figure 10 In some examples, the system 2000 also includes the capability of distributed image recognition operations, which may include reference to Figures 13A to 13C System 1300, Figure 15 System 1500, Figure 17 System 1700 and Figure 18 Any details discussed in system 1800.

[0158] like Figure 20 As shown, the hotword recognition operation for the voice command is distributed between the device 2002 and the computing device 2052. The device 2002 may include a voice command detector 2093 that executes the ML model 2026 (e.g., a gatekeeper model) to continuously (e.g., periodically) process microphone samples (e.g., audio data 2031) from the microphone 2040 on the device 2002 for the initial portion of the hotword (e.g., "ok G" or "ok D") of the voice command 2090. If the voice command detector 2093 detects the initial portion, the voice command detector 2093 may cause the buffer 2091 to capture subsequent audio data 2031. In addition, the device 2002 may transmit the audio portion 2092 to the computing device 2052 via the wireless connection 2048. In some examples, the device 2002 compresses the audio portion 2092 and then transmits the compressed audio portion 2092. The audio portion 2092 may be part of the buffer. For example, the audio portion 2092 may be 1-2 seconds of audio data 2031 from the head of the buffer 2091 .

[0159] The computing device 2052 includes a hot word recognition engine 2094 that is configured to execute an ML model 2027 (e.g., a larger ML model) to perform full hot word recognition using the audio portion 2092. For example, the ML model 2027 receives the audio portion 2092 as input, and the ML model 2027 predicts whether the audio portion 2092 includes a hot word (e.g., "ok Google, ok device"). If the audio portion 2092 is a false positive 2094, the computing device 2052 can transmit a dismiss command 2096 to the device 2002, which discards the contents of the buffer 2091 (e.g., audio data 2031). If the audio portion 2092 is a true positive 2095, the remaining portion 2099 of the buffer 2091 is transmitted to the computing device 2052. In some examples, the device 2002 compresses the audio data 2031 within the buffer 2091 (or the remaining portion 2099 of the buffer 2091) and transmits the compressed audio data 2031 to the computing device 2052. The computing device 2052 includes a command generator 2097 that uses the audio data 2031 (e.g., the remaining portion 2099 of the buffer 2091 and the audio portion 2092) to determine an action command 2098 (e.g., compose an email, take a photo, etc.). The computing device 2052 can transmit the action command 2098 to the device 2002 via the wireless connection 2048.

[0160] Figure 21 It is a depiction Figure 20 2100 is a flow chart of exemplary operations of the system 2000. Although reference is made to Figure 20 System 2000 explains Figure 21 2100, but the flowchart 2100 can be applied to any embodiment discussed herein, including Figure 1 System 100, Figure 2 System 200, Figure 3 System 300, Figure 4 head-mounted display device 402, Figure 5 Electronic components 570, Figure 6 Electronic components 670, Figure 7A and 7B System 700, Figure 8 System 800, Figure 9 System 900, Figure 10 System 1000, Figures 13A to 13C System 1300, Figure 15 System 1500 and / or Figure 18 System 1800. Although Figure 21 Flowchart 2100 shows the operations in a sequential order, but it will be appreciated that this is merely an example and that additional or alternative operations may be included. Figure 21The operations and related operations may be performed in a different order than shown, or in parallel or in an overlapping manner. In some examples, Figure 21 The operations of the flowchart 2100 can be Figure 11 Flowchart 1100, Figure 12 Flowchart 1200, Figure 14 Flowchart 1400, Figure 16 Flowchart 1600 and / or Figure 19 The operations of flowchart 1900 are combined.

[0161] Operation 2102 includes receiving audio data 2031 via a microphone 2040 of the device 2002. Operation 2104 includes detecting, by a voice command detector 2093, the presence of a portion of a hot word from the audio data 2031, wherein the voice command detector 2093 executes the ML model.

[0162] Operation 2106 includes storing audio data 2031 received via microphone 2040 in response to detecting a portion of the hotword in a buffer 2091 of device 2002. Operation 2108 includes transmitting an audio portion 2092 of buffer 2091 to computing device 2052 via wireless connection 2048, wherein the audio portion 2092 of buffer 2091 is configured to be used by computing device 2052 to perform hotword recognition.

[0163] While the disclosed inventive concept includes the inventive concept defined in the appended claims, it should be understood that the inventive concept may also be defined in terms of the following embodiments:

[0164] Embodiment 1 is a method for distributed sound recognition using a wearable device, comprising: receiving audio data via a microphone of the wearable device; detecting, by a sound classifier of the wearable device, whether the audio data includes a sound of interest; and in response to detecting the sound of interest within the audio data, sending the audio data to a computing device via a wireless connection.

[0165] Embodiment 2 is a method according to embodiment 1, wherein the sound classifier executes a first machine learning (ML) model.

[0166] Embodiment 3 is the method of any one of embodiments 1 to 2, wherein the audio data is configured to be used by a computing device or server computer for further sound recognition using a second ML model.

[0167] Embodiment 4 is the method of any one of embodiments 1 to 3, wherein the audio data is configured to be used by the computing device for further sound recognition.

[0168] Embodiment 5 is a method according to any one of embodiments 1 to 4, wherein the audio data is configured to be used by a server computer for further sound recognition.

[0169] Embodiment 6 is a method according to any one of embodiments 1 to 5, wherein the server computer is connected to the computing device via a network.

[0170] Embodiment 7 is a method according to any one of embodiments 1 to 6, wherein the sound of interest comprises speech.

[0171] Embodiment 8 is the method of any one of embodiments 1 to 7, wherein the audio data is configured to be used by a computing device or a server computer to convert the speech to text data using a second ML model.

[0172] Embodiment 9 is a method according to any one of embodiments 1 to 8, wherein the method further comprises receiving text data from a computing device via a wireless connection.

[0173] Embodiment 10 is a method according to any one of embodiments 1 to 9, wherein the speech is in a first language and the text data is in a second language, the second language being different from the first language.

[0174] Embodiment 11 is a method according to any one of embodiments 1 to 10, further comprising displaying the text data on a display of the wearable device.

[0175] Embodiment 12 is a method according to any one of embodiments 1 to 11, further comprising compressing the audio data, wherein the compressed audio data is transmitted to the computing device via a wireless connection.

[0176] Embodiment 13 is a method according to any one of embodiments 1 to 12, further comprising extracting features from the audio data, wherein the extracted features are transmitted to the computing device via a wireless connection.

[0177] Embodiment 14 is a method according to any one of embodiments 1 to 13, wherein the wireless connection is a short-range wireless connection.

[0178] Embodiment 15 is a method according to any one of embodiments 1 to 14, wherein the wearable device includes smart glasses.

[0179] Embodiment 16 is a system comprising one or more computers and one or more storage devices storing instructions, wherein the instructions, when executed by the one or more computers, are operable to cause the one or more computers to perform the method according to any one of embodiments 1 to 15.

[0180] Embodiment 17 is a wearable device configured to perform any one of Embodiments 1 to 15.

[0181] Embodiment 18 is a computer storage medium encoded with a computer program, the program comprising instructions that, when executed by a data processing apparatus, are operable to cause the data processing apparatus to perform the method according to any one of embodiments 1 to 15.

[0182] Embodiment 19 is a non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, cause the at least one processor to receive audio data from a microphone of a wearable device, detect by a sound classifier of the wearable device whether the audio data includes a sound of interest, and, in response to detecting the sound of interest within the audio data, transmit the audio data to a computing device via a wireless connection.

[0183] Embodiment 20 is the non-transitory computer-readable medium of embodiment 19, wherein the sound classifier is configured to execute a first machine learning (ML) model.

[0184] Embodiment 21 is the non-transitory computer-readable medium of any one of embodiments 19 to 20, wherein the audio data is configured to be used by a computing device for further sound recognition using a second ML model.

[0185] Embodiment 22 is a non-transitory computer-readable medium according to any one of embodiments 19 to 21, wherein the executable instructions include, when executed by at least one processor, causing the at least one processor to continue, by the sound classifier, detecting whether the audio data includes a sound of interest in response to not detecting a sound of interest within the audio data.

[0186] Embodiment 23 is the non-transitory computer-readable medium of any one of embodiments 19 to 22, wherein the sound of interest comprises speech.

[0187] Embodiment 23 is the non-transitory computer-readable medium of any one of embodiments 19 to 22, wherein the audio data is configured to be used by a computing device to translate the speech into text data using a second ML model.

[0188] Embodiment 24 is a non-transitory computer-readable medium according to any one of embodiments 19 to 23, wherein the executable instructions include instructions that, when executed by the at least one processor, cause the at least one processor to receive text data from a computing device via a wireless connection.

[0189] Embodiment 25 is a non-transitory computer-readable medium according to any one of embodiments 19 to 24, wherein the executable instructions include instructions that, when executed by the at least one processor, cause the at least one processor to compress audio data, wherein the compressed audio data is transmitted to the computing device via a wireless connection.

[0190] Embodiment 26 is a non-transitory computer-readable medium according to any one of embodiments 19 to 25, wherein the executable instructions include instructions that, when executed by the at least one processor, cause the at least one processor to extract features from the audio data, wherein the extracted features are transmitted to the computing device via a wireless connection.

[0191] Embodiment 27 is the non-transitory computer-readable medium of any one of embodiments 19 to 26, wherein the wearable device comprises smart glasses.

[0192] Embodiment 28 is the non-transitory computer-readable medium of any one of embodiments 19 to 27, wherein the computing device comprises a smartphone.

[0193] Embodiment 29 is a method comprising the operations of the non-transitory computer-readable medium according to any one of embodiments 19 to 28.

[0194] Embodiment 30 is a wearable device comprising the features described in any one of embodiments 19 to 28.

[0195] Embodiment 31 is a wearable device for distributed sound recognition, the wearable device comprising a microphone configured to receive audio data, a sound classifier configured to detect whether the audio data includes a sound of interest, and a radio frequency (RF) transceiver configured to send the audio data to a computing device via a wireless connection in response to detecting a sound of interest within the audio data.

[0196] Embodiment 32 is a wearable device according to embodiment 31, wherein the sound classifier includes a first machine learning (ML) model.

[0197] Embodiment 33 is the wearable device of any one of embodiments 29 to 32, wherein the audio data is configured to be used by the computing device or server computer to translate the sounds of interest into textual data using the second ML model.

[0198] Embodiment 34 is the wearable device of any one of embodiments 29 to 33, wherein the RF transceiver is configured to receive text data from the computing device over a wireless connection.

[0199] Embodiment 35 is the wearable device of any one of embodiments 29 to 34, wherein the wearable device further comprises a display configured to display text data.

[0200] Embodiment 36 is a wearable device according to any one of embodiments 29 to 35, wherein the wearable device comprises smart glasses.

[0201] Embodiment 37 is a wearable device according to any one of embodiments 29 to 36, wherein the wireless connection is a Bluetooth connection.

[0202] Embodiment 38 is a computing device for sound recognition, comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions, wherein the executable instructions, when executed by the at least one processor, cause the at least one processor to receive audio data from a wearable device via a wireless connection, the audio data having a sound of interest detected by a sound classifier executing a first machine learning (ML) model, use a sound recognition engine on the computing device to determine whether to translate the sound of interest into text data, and in response to determining to use the sound recognition engine on the computing device, the sound recognition engine is configured to execute a second ML model, the sound of interest is translated into text by the sound recognition engine, and the text data is transmitted to the wearable device via the wireless connection.

[0203] Embodiment 39 is a computing device according to embodiment 38, wherein the executable instructions include instructions that, when executed by the at least one processor, cause the at least one processor to perform the following operations: in response to determining that the sound recognition engine on the computing device is not to be used, transmit audio data to a server computer over a network, and receive text data from the server computer over the network.

[0204] Embodiment 40 is the computing device of any one of embodiments 38 to 39, wherein the computing device comprises a smartphone.

[0205] Embodiment 41 is a method comprising the operation of a computing device according to any one of embodiments 38 to 39.

[0206] Embodiment 42 is a computer storage medium encoded with a computer program, the program comprising instructions that, when executed by a data processing device, are operable to cause the data processing device to perform the operations of a computing device according to any one of embodiments 38 to 39.

[0207] Embodiment 43 is a method for distributed image recognition using a wearable device, comprising: receiving image data via at least one imaging sensor of the wearable device, detecting, by an image classifier of the wearable device, whether an object of interest is included in the image data, and sending the image data to a computing device via a wireless connection.

[0208] Embodiment 44 is a method according to embodiment 43, wherein the image classifier executes a first machine learning (ML) model.

[0209] Embodiment 45 is a method according to any one of embodiments 43 to 44, wherein the image data is configured to be used by the computing device for further image recognition using a second ML model.

[0210] Embodiment 46 is a method according to any one of embodiments 43 to 45, further comprising receiving the bounding box dataset from a computing device via a wireless connection.

[0211] Embodiment 47 is a method according to any one of embodiments 43 to 46, further comprising using the bounding box dataset by an object tracker of the wearable device to identify image regions in subsequent image data captured by the at least one imaging sensor.

[0212] Embodiment 48 is a method according to any one of embodiments 43 to 47, further comprising transmitting the image region to a computing device via a wireless connection, the image region being configured to be used by the computing device for further image recognition.

[0213] Embodiment 49 is a method according to any one of embodiments 43 to 48, further comprising cropping, by the object tracker, an image region from subsequent image data.

[0214] Embodiment 50 is a method according to any one of embodiments 43 to 49, further comprising compressing the image region by the object tracker, wherein the compressed image region is transmitted to the computing device via the wireless network.

[0215] Embodiment 51 is a method according to any one of embodiments 43 to 50, wherein the object of interest includes facial features.

[0216] Embodiment 52 is a method according to any one of embodiments 43 to 51, further comprising activating a first imaging sensor of the wearable device to capture first image data.

[0217] Embodiment 53 is a method according to any one of embodiments 43 to 45, further comprising detecting, by an image classifier, whether the first image data includes the object of interest.

[0218] Embodiment 54 is a method according to any one of embodiments 43 to 53, further comprising activating a second imaging sensor to capture second image data.

[0219] Embodiment 55 is a method according to any one of embodiments 43 to 44, wherein the quality of the second image data is higher than the quality of the first image data.

[0220] Embodiment 56 is a method according to any one of embodiments 43 to 55, wherein the second image data is transmitted to the computing device via a wireless connection, and the second image data is configured to be used by the computing device for further image recognition.

[0221] Embodiment 57 is a method according to any one of embodiments 43 to 56, further comprising receiving light condition information via a light condition sensor of the wearable device.

[0222] Embodiment 58 is a method according to any one of embodiments 43 to 57, further comprising activating the first imaging sensor based on lighting condition information.

[0223] Embodiment 59 is a method according to any one of embodiments 43 to 58, further comprising receiving motion information via a motion sensor of the wearable device.

[0224] Embodiment 60 is a method according to any one of embodiments 43 to 59, further comprising activating the first imaging sensor based on the motion information.

[0225] Embodiment 61 is a method according to any one of embodiments 43 to 60, wherein the wireless connection is a short-range wireless connection.

[0226] Embodiment 62 is a method according to any one of embodiments 43 to 61, wherein the wearable device includes smart glasses.

[0227] Embodiment 63 is a method according to any one of embodiments 43 to 62, wherein the computing device comprises a smartphone.

[0228] Embodiment 64 is a system comprising one or more computers and one or more storage devices storing instructions, wherein the instructions, when executed by the one or more computers, are operable to cause the one or more computers to perform a method according to any one of embodiments 43 to 63.

[0229] Embodiment 65 is a wearable device configured to perform any one of embodiments 43 to 63.

[0230] Embodiment 66 is a computer storage medium encoded with a computer program, the program comprising instructions that, when executed by a data processing apparatus, are operable to cause the data processing apparatus to perform a method according to any one of embodiments 43 to 63.

[0231] Embodiment 67 is a non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, cause the at least one processor to receive image data from an imaging sensor on a wearable device, detect, by an image classifier of the wearable device, whether an object of interest is included in the image data, the image classifier being configured to execute a first machine learning (ML) model, and transmit the image data to a computing device via a wireless connection, the image data being configured to be used by the computing device to calculate a bounding box dataset using a second ML model.

[0232] Embodiment 68 is a non-transitory computer-readable medium according to embodiment 67, wherein the executable instructions include instructions that, when executed by at least one processor, cause the at least one processor to perform the following operations: receive a bounding box dataset from a computing device via a wireless connection, use the bounding box dataset by an object tracker of the wearable device to identify an image region in subsequent image data captured by at least one imaging sensor, and / or transmit the image region to the computing device via the wireless connection, the image region being configured to be used by the computing device for further image recognition.

[0233] Embodiment 69 is a non-transitory computer-readable medium according to any one of embodiments 67 to 68, wherein the executable instructions include instructions that, when executed by at least one processor, cause the at least one processor to perform the following operations: cropping an image area from subsequent image data by the object tracker and / or compressing an image area by the object tracker, wherein the compressed image area is transmitted to the computing device via a wireless network.

[0234] Embodiment 70 is the non-transitory computer-readable medium of any one of embodiments 67 to 69, wherein the object of interest comprises a barcode or text.

[0235] Embodiment 71 is a non-transitory computer-readable medium according to any one of embodiments 67 to 70, wherein the executable instructions include, when executed by at least one processor, causing the at least one processor to activate a first imaging sensor of the wearable device to capture first image data, detect by an image classifier whether the first image data includes an object of interest, and / or activate a second imaging sensor to capture second image data, the quality of the second image data being higher than the quality of the first image data, wherein the second image data is transmitted to the computing device via a wireless connection, and the second image data is configured to be used by the computing device for further image recognition.

[0236] Embodiment 72 is a non-transitory computer-readable medium according to any one of embodiments 67 to 71, wherein the executable instructions include instructions that, when executed by at least one processor, cause the at least one processor to compress second image data, wherein the compressed image data is transmitted to the computing device via a wireless connection.

[0237] Embodiment 73 is a non-transitory computer-readable medium according to any one of embodiments 67 to 72, wherein the executable instructions include instructions that, when executed by at least one processor, cause the at least one processor to receive lighting condition information from a lighting condition sensor of the wearable device and / or determine whether to transmit second image data based on the lighting condition information.

[0238] Embodiment 74 is a non-transitory computer-readable medium according to any one of embodiments 67 to 73, wherein the executable instructions include instructions that, when executed by at least one processor, cause the at least one processor to receive motion information from a motion sensor of the wearable device and determine whether to transmit second image data based on the motion information.

[0239] Embodiment 75 is a wearable device for distributed image recognition, the wearable device comprising at least one imaging sensor configured to capture image data, an image classifier configured to detect whether an object of interest is included within the image data, the image classifier configured to execute a first machine learning (ML) model, and a radio frequency (RF) transceiver configured to transmit the image data to a computing device via a wireless connection, the image data configured to be used by the computing device to compute a bounding box dataset using a second ML model.

[0240] Embodiment 76 is a wearable device according to embodiment 75, wherein the RF transceiver is configured to receive a bounding box dataset from a computing device via a wireless connection, the wearable device further comprising an object tracker configured to use the bounding box dataset to identify an image region in subsequent image data captured by at least one imaging sensor, wherein the RF transceiver is configured to transmit the image region to the computing device via the wireless connection, and the image region is configured to be used by the computing device for further image recognition.

[0241] Embodiment 77 is a wearable device according to any one of embodiments 75 to 76, wherein the wearable device further includes a sensor trigger configured to activate the first imaging sensor to capture first image data, the image classifier is configured to detect whether the first image data includes an object of interest, the sensor trigger is configured to activate the second imaging sensor to capture second image data in response to detecting the object of interest in the first image data, the quality of the second image data is higher than the quality of the first image data, and wherein the RF transceiver is configured to transmit the second image data to the computing device via a wireless connection.

[0242] Embodiment 78 is a computing device for distributed image recognition, the computing device comprising at least one processor and a non-transitory computer-readable medium storing executable instructions, the executable instructions, when executed by the at least one processor, causing the at least one processor to receive image data from a wearable device via a wireless connection, the image data having an object of interest detected by an image classifier executing a first machine learning (ML) model, calculate a bounding box dataset based on the image data using a second ML model, and transmit the bounding box dataset to the wearable device via the wireless connection.

[0243] Embodiment 79 is a computing device according to embodiment 78, wherein the executable instructions include instructions that, when executed by at least one processor, cause the at least one processor to receive the image region in subsequent image data via a wireless connection and / or perform object recognition on the image region by a second ML model.

[0244] Embodiment 80 is a method for distributed hotword recognition using a wearable device, comprising: receiving audio data via a microphone of the wearable device, detecting the presence of a portion of a hotword from the audio data by a voice command detector of the wearable device, the voice command detector executing a first machine learning (ML) model, storing the audio data received via the microphone in response to detecting the portion of the hotword in a buffer of the wearable device, and transmitting a portion of the audio data included in the buffer to a computing device via a wireless connection, the portion of the audio data being configured to be used by the computing device to perform hotword recognition using a second ML model.

[0245] Embodiment 81 is a method according to embodiment 80, further comprising transmitting the remaining portion of the audio data included in the buffer to the computing device via a wireless connection.

[0246] Embodiment 82 is a method according to any one of embodiments 80 to 81, further comprising receiving an action command from a computing device via a wireless connection, the action command causing the wearable device to perform an action.

[0247] Embodiment 83 is a method according to any one of embodiments 80 to 82, further comprising: receiving a release command from the computing device via a wireless connection and / or discarding the audio data included in the buffer in response to the release command.

[0248] Embodiment 84 is a system comprising one or more computers and one or more storage devices storing instructions, wherein the instructions, when executed by the one or more computers, are operable to cause the one or more computers to perform a method according to any one of embodiments 80 to 83.

[0249] Embodiment 85 is a wearable device configured to perform any one of embodiments 80 to 83.

[0250] Embodiment 86 is a computer storage medium encoded with a computer program, the program comprising instructions that, when executed by a data processing apparatus, are operable to cause the data processing apparatus to perform a method according to any one of embodiments 80 to 83.

[0251] Embodiment 87 is a method for sensing image data having multiple resolutions using a wearable device, the method comprising: activating a first imaging sensor of the wearable device to capture first image data, detecting, by a classifier of the wearable device, whether a region of interest (ROI) is included in the first image data, the classifier executing a first machine learning (ML) model, in response to detecting the ROI within the first image data, activating a second imaging sensor of the wearable device to capture second image data, the resolution of the second image data being higher than the resolution of the first image data, and sending the second image data to a computing device via a wireless connection, the second image data being configured to be used by the computing device for image processing using a second ML model.

[0252] In addition to the above description, controls may be provided to the user that allow the user to make choices about whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's social network, social actions or activities, occupation, the user's preferences, or the user's current location) and whether to send content or communications from the server to the user. Additionally, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information can be removed. For example, the user's identity may be processed so that personally identifiable information about the user cannot be determined, or the user's geographic location from which location information is obtained may be summarized (such as to a city, zip code, or state level) so that the user's specific location cannot be determined. Thus, the user can control what information is collected about the user, how that information is used, and what information is provided to the user.

[0253] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in the form of one or more computer programs executed and / or interpreted on a programmable system comprising at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to send data and instructions to the storage system, at least one input device, and at least one output device.

[0254] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms "machine-readable medium," "computer-readable medium," and "machine-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) that is used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that is used to provide machine instructions and / or data to a programmable processor.

[0255] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including sound, voice, or tactile input.

[0256] The systems and techniques described herein can be implemented as a computing system that includes a back-end component (e.g., as a data server) or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user interacts with implementations of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.

[0257] A computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0258] In this specification and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" do not exclude plural references. In addition, unless the context clearly dictates otherwise, conjunctions such as "and", "or", and "and / or" are inclusive. For example, "A and / or B" includes A alone, B alone, and A and B. In addition, the connecting lines or connectors shown in the various figures presented are intended to represent exemplary functional relationships and / or physical or logical couplings between the various elements. Many alternative or additional functional relationships, physical connections, or logical connections may exist in actual devices. In addition, unless an element is specifically described as "essential" or "critical", an item or component is not necessary to implement the embodiments disclosed herein.

[0259] Terms such as, but not limited to, approximately, substantially, generally, etc. are used herein to indicate that precise values ​​or ranges thereof are not required and need not be specified. As used herein, the terms discussed above will have their ready and immediate meaning to one of ordinary skill in the art.

[0260] Furthermore, terms such as upper, lower, top, bottom, side, end, front, back, etc. are used herein with reference to the orientation currently considered or illustrated. If considered relative to another orientation, it should be understood that these terms must be modified accordingly.

[0261] In addition, in this specification and the appended claims, the singular forms "a," "an," and "the" do not exclude plural reference unless the context clearly dictates otherwise. Furthermore, conjunctions such as "and," "or," and "and / or" are inclusive unless the context clearly dictates otherwise. For example, "A and / or B" includes A alone, B alone, and A and B.

[0262] Although certain exemplary methods, apparatus, and articles have been described herein, the scope of coverage of this patent is not limited thereto. It should be understood that the terminology used herein is for the purpose of describing particular aspects and is not intended to be limiting. On the contrary, this patent covers all methods, apparatus, and articles that fully fall within the scope of the claims of this patent.

Claims

1. A method for distributed image recognition using a wearable device, the method comprising: receiving a first image frame via at least one imaging sensor of the wearable device; detecting, by a first model, whether an object of interest is contained in the first image frame; in response to detecting the object of interest within the first image frame, transmitting the first image frame to a computing device communicatively coupled to the wearable device, the first image frame being configured to be used by a second model on the computing device for further image processing; receiving object location data from the computing device, the object location data identifying a location of the object of interest in the first image frame; identifying an image region in a second image frame using the object location data of the first image frame; as well as The image region is transmitted to the computing device.

2. The method according to claim 1, further comprising: cropping the image region from the second image frame; as well as A compressed image region is generated from the image region, wherein the compressed image region is transmitted to the computing device.

3. The method according to claim 1, wherein The object of interest includes facial features, wherein the method further comprises: Facial recognition is performed by the second model using at least one of the image regions of the first image frame or the second image frame.

4. The method of claim 1 , wherein the at least one imaging sensor comprises a first imaging sensor and a second imaging sensor, the second imaging sensor having a higher resolution than the first imaging sensor, wherein the method further comprises: activating the first imaging sensor of the wearable device to capture the first image frame; In response to detecting the object of interest within the first image frame, activating the second imaging sensor to capture a third image frame, the third image frame having a higher quality than the first image frame, The third image frame is transmitted to the computing device.

5. The method according to claim 4, further comprising: Receiving light condition information via a light condition sensor of the wearable device; as well as The first imaging sensor is activated based on the lighting condition information.

6. The method according to claim 4, further comprising: Receiving motion information via a motion sensor of the wearable device; as well as The first imaging sensor is activated based on the motion information.

7. The method according to any one of claims 1 to 6, wherein: The wearable device comprises an extended reality device configured to be worn on a user's head, and the computing device comprises a mobile user device.

8. A non-transitory computer-readable medium storing executable instructions, wherein the executable instructions, when executed by at least one processor, cause the at least one processor to perform the following operations, the operations comprising: receiving a first image frame from at least one imaging sensor on the wearable device; detecting, by a first model, whether an object of interest is contained in the first image frame; in response to detecting the object of interest within the first image frame, transmitting the first image frame to a computing device communicatively coupled to the wearable device, the first image frame being configured for use by a second model on the computing device; receiving object location data from the computing device, the object location data identifying a location of the object of interest in the first image frame; identifying an image region in a second image frame using the object location data of the first image frame; as well as The image region is transmitted to the computing device.

9. The non-transitory computer-readable medium of claim 8, wherein: The operations further include: selecting the image area from the second image frame; and A compressed image region is generated from the image region, wherein the compressed image region is transmitted to the computing device.

10. The non-transitory computer-readable medium of claim 8, wherein: The object of interest comprises a machine-readable representation of data, wherein the operations further comprise: The machine-readable representation is decoded from the image region by the second model.

11. The non-transitory computer-readable medium of claim 8, wherein: The at least one imaging sensor includes a first imaging sensor and a second imaging sensor, the second imaging sensor having a higher resolution than the first imaging sensor, wherein the operations further include: activating the first imaging sensor to capture the first image frame based on activation data received via one or more sensors on the wearable device; In response to detecting the object of interest within the first image frame, activating the second imaging sensor to capture a third image frame, the third image frame having a higher quality than the first image frame, The third image frame is transmitted to the computing device.

12. The non-transitory computer-readable medium of claim 11, wherein: The operations further include: A compressed third image frame is generated, wherein the compressed third image frame is transmitted to the computing device.

13. The non-transitory computer-readable medium of claim 11, wherein: The activation data includes lighting condition information, wherein the operations further include: receiving the light condition information from a light condition sensor of the wearable device, the light condition information indicating a level of ambient light; and In response to the level of ambient light reaching a threshold level, the first imaging sensor is activated.

14. The non-transitory computer-readable medium according to any one of claims 11 to 13, wherein: The activation data includes motion information, wherein the operations further include: receiving the motion information from a motion sensor of the wearable device; and The first imaging sensor is activated based on the motion information.

15. A wearable device for distributed image recognition, the wearable device comprising: at least one processor; as well as a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to: receiving a first image frame via a first imaging sensor of the wearable device; detecting whether an object of interest is contained in the first image frame; In response to detecting the object of interest within the first image frame, activating a second imaging sensor of the wearable device to capture a second image frame, the second image frame having a higher resolution than a resolution of the first image frame; transmitting the second image frame to a computing device communicatively coupled to the wearable device, the second image frame being configured to be used by a second model on the computing device for further image processing; receiving object location data from the computing device, the object location data identifying a location of the object of interest in the second image frame; identifying an image region in a third image frame using the object location data of the second image frame; as well as The image region is transmitted to the computing device.

16. The wearable device according to claim 15, wherein: The executable instructions include instructions to cause the at least one processor to: receiving activation data from one or more sensors of the wearable device, the activation data comprising at least one of lighting condition information or motion information; Based on the activation data, the first imaging sensor is activated to capture the first image frame.

17. The wearable device according to claim 15, wherein: The executable instructions include instructions to cause the at least one processor to: A compressed image region is generated from the image region, wherein the compressed image region is transmitted to the computing device.

18. The wearable device according to claim 15, wherein: The object of interest includes facial features, wherein the executable instructions include instructions to cause the at least one processor to: Facial recognition is performed by the second model using at least one of the image areas of the second image frame or the third image frame.

19. The wearable device according to claim 15, wherein: The wearable device comprises an extended reality device configured to be worn on a user's head, and the computing device comprises a mobile user device.

Citation Information

Patent Citations

  • Information processing device and image input device

    US20160080652A1