CASCADING APPROACH TO DETECTING AND INTERPRETING USER ACTIVITY
A resource-poor process detects human-relevant events and triggers resource-intensive processes only when needed, optimizing power and processing for efficient user activity detection and interpretation.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-04-02
AI Technical Summary
Existing techniques for detecting user activity are inadequate in terms of accuracy, power consumption, and the types of activities detected, particularly when involving combinations of movement, sounds, glances, etc.
A resource-intensive process is triggered and guided by a resource-poor process, utilizing sensors with low computational effort and power consumption to detect human-relevant events, which then activates more powerful processes only when necessary, such as cameras and large language models, to provide detailed interpretations.
This approach optimizes power consumption and processing efficiency by using lightweight models continuously while activating intensive models only when relevant events occur, enabling accurate and efficient detection and interpretation of user activities.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL AREA
[0001] The present disclosure relates generally to systems, methods and devices that detect and interpret user activity using a resource-intensive process that is triggered or guided by an analysis obtained from an initial resource-poor process. BACKGROUND
[0002] Existing techniques for detecting user activity can be improved in terms of accuracy, power consumption and / or the types of user activity detected (e.g., activities involving combinations of movement, sounds, glances, etc.). SUMMARY
[0003] The various implementations disclosed herein include devices, systems, and methods that detect and interpret user activity through a resource-intensive process triggered and / or guided by detections and decisions originally made by a resource-poor process. For example, a resource-poor process can operate continuously (or nearly continuously) using sensor data (e.g., hand position data, gaze data, audio data, inertial measurement unit (IMU) data) with fewer sensors and computing resources than a resource-intensive process that uses sensors and data such as cameras, camera settings, large language models (LLMs), etc., and requires more power and computing resources.
[0004] In some implementations, a resource-intensive process and a resource-light process can be configured to receive different multimodal inputs. For example, a resource-light process can detect and collect data associated with human-relevant events, such as sounds, actions related to interacting with an object, user movement actions, and so on. Similarly, a resource-light process can detect and classify current events and / or objects of interest with which interaction is taking place, in order to trigger and / or guide improved resource-intensive process decisions. In some implementations, upon detection of a triggering event, a low-level mediation (LLM) process (of the resource-intensive process) can be activated to further analyze the triggering event through powerful, multimodal processing that can receive inputs such as images, audio, and / or contextual data.Subsequent processing of the event may occur. This resource-intensive process can be configured to generate further output or a prediction, such as a text description of user activity. In some implementations, multimodal input may include, among other things, user voice input, user gaze input, user hand or finger input, body language input, and so on.
[0005] In some implementations, a first subset of resource-efficient processes (e.g., executed by audio sensors) can be configured to control the operation of a second subset of resource-efficient processes, allowing the second subset (e.g., executed by cameras) to be analyzed by an LLM after the LLM has analyzed the resulting audio signals. For example, the first subset of resource-efficient processes might include audio recognition (e.g., speech recognition), which triggers a limited capacity of the LLM (e.g., checking audio data without image data) for speech interpretation. Based on the interpreted speech, it can be determined whether a resource-intensive process is necessary. If it is determined that a resource-intensive process is necessary, the use of the second subset of resource-efficient processes (e.g., camera processing) can be avoided.Cameras) are triggered to capture images for resource-intensive processes such as LLM with a video encoder, image processing, etc., to analyze hand / gaze gestures, as described in the following example: . In some implementations, a user can specify a spoken command: "When was the Type A automobile manufactured?" In this case, an LLM is only needed in limited capacity for analyzing audio data, as there is no need to activate cameras to perform resource-intensive processes such as hand / gaze detection algorithms, since the spoken command does not include any words that would indicate the user is referring to an element in an actual physical environment.
[0006] Alternatively, if the user gives a spoken command such as "When was this car manufactured?", a limited-capacity LLM (requiring only audio data) can first be used to interpret the spoken command (e.g., the user is requesting information about a car in their vicinity). Based on the LLM's output, the process can determine that resource-intensive processes (e.g., image capture and object recognition) are necessary to identify the presence of a car in the physical environment. This allows cameras to be activated to capture images and more computationally intensive processes like image recognition to be performed. If only one car is detected, the process can infer that the user is referring to that specific vehicle.However, if there are two cars in the physical environment, the process is configured to activate another resource-intensive process, such as a hand / gaze detection process, to determine which car the user was referring to when they said the term "this".
[0007] In some implementations, a device includes a processor (e.g., one or more processors) that executes instructions stored on a non-transitory, computer-readable medium to perform a procedure. The procedure performs one or more steps or processes. In some implementations, the procedure performs a first process to generate output. The first process includes: detecting events based on an initial set of sensor data; identifying a subset of the events as human-relevant events corresponding to one or more predefined classes represented in the initial set of sensor data; and gathering information regarding the human-relevant events based on the initial set of sensor data. Based on the output of the first process, the procedure performs a second process to interpret user activity.The second process involves obtaining a second set of sensor data and interpreting user activity using the second set of sensor data.
[0008] According to some implementations, a device includes one or more processors, non-transitory memory, and one or more programs; the one or more programs being stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs including instructions to perform or cause to perform one of the procedures described herein. According to some implementations, instructions are stored in a non-transitory, computer-readable storage medium which, when executed by one or more processors of a device, cause the device to perform or cause to perform one of the procedures described herein.According to some implementations, a device includes: one or more processors, non-transitory memory, and means for performing or causing the performance of any of the procedures described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to enable the present disclosure to be understood by average professionals, a more detailed description is provided with reference to aspects of some illustrative implementations, some of which are shown in the accompanying drawings. Fig. Figure 1 illustrates an exemplary electronic device that operates in a physical environment according to some implementations. The Fig. 2A and Fig. Figure 2B illustrates a view of a system that includes a low-performance system that outputs information via a human relevance module to trigger a high-performance system according to some implementations. Fig. Figure 3 illustrates a view of a system that combines a low-level signal processing system with a powerful multimodal language modeling system to capture detailed behaviors and events in real time, according to some implementations. Fig. 4 illustrates a process for using the system from Fig. 2, to enable, for example, the location of a car in a garage, according to some implementations. Fig. Figure 5 illustrates a view of a system which, according to some implementations, allows system activation using low-level signals as triggers. Fig. Figure 6 is a flowchart representation of an exemplary procedure that detects and interprets user activities through a resource-intensive process triggered and / or guided by findings and decisions enabled by a resource-poor process according to some implementations. Fig. Figure 7 is a block diagram of an electronic device according to some implementations.
[0010] According to common practice, the various features illustrated in the drawings may not be drawn to scale. Accordingly, the dimensions of the various features may be enlarged or reduced for clarity. Furthermore, some of the drawings may not depict all components of a particular system, process, or device. Finally, the same reference numerals may be used to designate identical features consistently throughout the patent specification and figures. DESCRIPTION
[0011] Numerous details are described to provide a thorough understanding of the exemplary implementations shown in the drawings. However, the drawings only illustrate some exemplary aspects of the present disclosure and are therefore not to be considered limiting. The person skilled in the art will recognize that other effective aspects and / or variations do not include all of the specific details described herein. Furthermore, generally known systems, methods, components, devices, and circuits have not been described exhaustively in detail so as not to obscure more relevant aspects of the exemplary implementations described herein.
[0012] Fig. Figure 1 illustrates an exemplary electronic device 105 that operates in a physical environment 100. In the example of Fig. 1. The physical environment 100 is a room that includes a desk 120. The electronic device 105 may include one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about the physical environment 100 and the objects within it, as well as information about the user 102 of the electronic device 105. The information about the physical environment 100 and / or the user 102 can be used to provide visual and audio content and / or to identify the current position of the physical environment 100 and / or the position of the user within the physical environment 100.
[0013] In some implementations, views of an augmented reality (XR) environment can be provided to one or more participants (e.g., user 102 and / or other participants not shown) via the electronic device 105 (e.g., a wearable device such as an HMD). Such an XR environment may include views of a 3D environment generated based on camera images and / or depth camera images of the physical environment 100, as well as a representation of user 102 based on camera images and / or depth camera images of user 102. Such an XR environment may include virtual content positioned at 3D locations relative to a 3D coordinate system associated with the XR environment, which may correspond to a 3D coordinate system of the physical environment 100.
[0014] In some implementations, a system including the electronic device 105 may be configured to perform a first process, for example a resource-efficient process (e.g. using sensors with low computational effort and / or low power consumption, etc.), to produce an output that may be configured to trigger a second process or that may include information used to control the second process.
[0015] In some implementations, the first process may involve detecting events based on an initial set of sensor data. These events may include, for example, hand position events, gaze direction events, audio events, IMU (Inertial Measurement Unit) events, and so on.
[0016] In some implementations, the first process may further include identifying a subset of events as human-relevant events that correspond to one or more predefined classes represented in the initial set of sensor data. For example, identifying the subset of events as human-relevant events may involve using a Human Foundation Model (HFM) and a relevance decoder to identify events such as hand-gaze events, hand-object interaction events, events of visual attention to objects, text reading events, user-initiated speech- or sound-based events, human body movement events, and so on.
[0017] In some implementations, the first process may further include gathering information related to human-relevant events based on the initial set of sensor data. This information gathering can be achieved through the continuous collection of the initial set of sensor data.
[0018] In some implementations, a second process is executed or triggered based on an output from the first process. The second process is configured to interpret user activity by receiving a second set of sensor data (e.g., image sensor data, frame rate increase data, resolution increase data, etc.) and interpreting the user activity using this second set of sensor data. For example, the interpretation of user activity might include LLM processing.
[0019] Fig. 2A and Fig. Figure 2B illustrates a view of a system 200 that includes a low-power system 210, which outputs information via a human relevance module 202 to trigger a high-power system (multimodal language model) 225, which, according to some implementations, includes an LLM 228. In some implementations, the low-power system 210 may include event-driven low-power modules that act as initial detectors. Likewise, the high-power system 225 may only be activated when relevant signals are triggered (e.g., by the human relevance module 202). The system 200 ensures that computationally intensive models are used only when needed, thereby optimizing both the power consumption and processing efficiency of the system 200.
[0020] The low-power System 210 is configured to use multiple sensor modules (e.g., modules 211, 212, 215, and 216) to detect, track, and interpret various events or interactions (via the HFM 206 and the relevance decoder 204 of the human relevance module 202 for identifying events 205) in an environment without relying on a full-fledged language model. For example, the low-power System 210 can be configured to operate with minimal computing resources and focus on processing essential, low-dimensional signals from sensors (modules 211, 212, 215, and 216), such as cameras (for object, hand, and environment tracking), audio sensors (for sound classification, speech-to-text conversion, etc.), and additional sensor inputs (e.g., eye tracking, motion detection, etc.).
[0021] In some implementations, modules 211 and 212 and the associated sensors 217 and 218 are included in a first group 222 of low-power / low-computing modules / sensors. In some implementations, modules 215 and 216 and the associated sensors 221 and 220 are included in a second group 223 of low-power / low-computing modules / sensors, which have lower power consumption / computational overhead than the first group 222. In some implementations, the operation of the first group 222 of low-power / low-computing modules / sensors can be triggered by the high-performance system 225 when, for example, audio data (e.g., speech) is detected, thereby activating a limited capacity of an LLM 228 for interpreting the audio data.Based on the interpreted audio, the first group 222 of modules / sensors with low power consumption / low computing power (e.g. cameras) can be triggered to capture images for a resource-intensive process, for example to evaluate hand / gaze gestures.
[0022] Module 211 is a visual perception module configured to perform object detection activities. For example, Module 211 can be configured to use outward-facing cameras (OFC) 217 to detect objects, hands, people, and environmental features such as lighting, space, etc. Similarly, Module 211 can be configured to enable salience detection, such as identifying areas within a user's field of view that are most likely to be relevant to the user, such as objects with which they will interact.
[0023] Module 212 is configured to enable the gaze detection function. For example, Module 212 can be configured to activate inward-facing cameras (IFC) 218 to monitor user behavior, such as a user's gaze direction, focus or attention on objects or the environment, etc.
[0024] Module 215 is an audio perception module configured to use one or more sensors 221 (e.g., a microphone) to perform processes for sound classification, speech-to-text conversion, and behavior prediction. For example, a sound classification process can be configured to detect ambient noise or speech and classify activities such as chewing or eating. Similarly, a behavior prediction process is configured to perform audio-based behavior detection in combination with other sensory inputs, for example, to distinguish whether a person is speaking while walking or sitting (e.g., making a verbal utterance).
[0025] Module 216 is configured to detect user activities (e.g., via IMU sensors 220) such as walking, running, standing, sitting and / or transitions between these activities.
[0026] In some implementations, the low-power System 210 can operate continuously to detect events or changes in the environment or user behavior that may be relevant, such as a loud noise, a person pointing at, moving, or interacting with an object, etc. Once a relevant event is detected, the low-power System 210 can accordingly generate a low-power output (e.g., a trigger event) that signals the occurrence of a relevant event.
[0027] In some implementations, an audio signal from audio sensors (e.g., sensors 221) can be used to trigger limited use of the LLM 228. Based on the results of the analysis performed by the LLM 228 with respect to the audio signal, low-power sensors (e.g., sensors 217), such as cameras (e.g., cameras are more powerful sensors than the audio sensors), can be activated, thus enabling the LLM 228 to perform multimodal processing.
[0028] For example, the second group 223 of low-power, low-computation modules / sensors (associated with a subset of resource-efficient processes) can be configured to control the operation of the first group 222 of low-power, low-computation modules / sensors (another subset of resource-efficient processes). This allows processes executed by the first group 222 of low-power, low-computation modules / sensors (e.g., those performed by cameras) to be executed by the LLM 228 after the LLM 228 has analyzed the associated audio signals. Similarly, the second group 223 of low-power, low-computation modules / sensors can include audio recognition functionality (e.g., speech recognition) that utilizes the limited capacity of the LLM 228 (e.g., checking audio data without image data) to interpret speech.Based on the interpreted language, it can be determined whether a resource-intensive process should be implemented using the full capacity of the LLM 228. If it is determined that a resource-intensive process should be implemented, the use of the other sensors (e.g., the cameras of the first group 222 of modules / sensors with low power consumption / low processing power) can be triggered to capture images for resource-intensive processes such as using the LLM 228 with a video encoder of the image system 236, etc., as described in the following example: In some implementations, a user can specify a spoken command: "In what year was the Type A automobile manufactured?" In this case, the LLM 228, with its limited capacity, is capable of analyzing only audio data, as there is no need to activate cameras to perform resource-intensive processes such as hand / gaze detection algorithms, since the spoken command does not include any words indicating that the user is referring to an element in an actual physical environment.
[0029] Alternatively, if the user gives a spoken command such as "When was this car manufactured?", the LLM 228 can be initialized with limited capacity (to analyze audio data) to interpret the spoken command (e.g., the user is requesting information about a car in a current environment). Based on the LLM 228's output, the process can determine that resource-intensive processes (e.g., image capture and object recognition) are necessary to identify the presence of a car in the physical environment. In response, cameras can be activated to capture images, and higher-level computational processes such as image recognition can be performed. If only one car is detected, the process can infer that the user is referring to that vehicle. Similarly, the process can be configured to activate another resource-intensive process, such as...a hand / gaze recognition process to determine which car the user was referring to when they said the term "this" when there are two cars in the physical environment.
[0030] In some implementations, when System 200 detects a triggering event, it activates LLM 228 to further analyze the event using powerful, multimodal processing. This involves receiving inputs such as images (e.g., video images) from an Image System 236, audio from an Audio System 238, and contextual data from a Context System 234. For example, if the low-performance System 210 detects a hand-object interaction, the high-performance System 225 can be configured to process complete image sequences to generate a prediction, such as "the user picks up a coffee cup."
[0031] In some implementations, the powerful System 225 can be configured to process multiple input types, such as image sequences (from a camera) or audio clips (for speech or tone analysis in relation to a verbal utterance), to generate a more detailed understanding of the event. The various input types can be processed for input into the LLM 228 via modules 226, 227, 229, 230, and 232.
[0032] In some implementations, the powerful System 225 can be configured to use additional contextual data, such as location (from spatial detection), historical patterns (e.g., typical actions at a specific time), calendar data to refine analysis, and so on. This additional contextual data can enable event interpretations, such as determining that it is 8:00 AM, a user is in the kitchen, and the user typically drinks coffee at that time.
[0033] Following the processing of an event, the powerful System 225 can generate further output or prediction, including a text description such as: "A user picked up a coffee cup in the kitchen at 8:15 a.m." This further output or prediction can also include embeddings of higher-level features, such as: image or video analysis embeddings, audio-based analysis embeddings, behavioral embeddings for detecting temporal or action-based patterns, and so on. The aforementioned embeddings can be stored for later analysis, future queries, or integration with other systems.
[0034] Accordingly, System 200 implements a process that enables an always-active perception layer (low-power System 210) which conserves energy by using only lightweight perception models (e.g., Modules 211, 212, 215, and 216) to generate low-dimensional signals. More computationally intensive models are triggered only when a relevant event occurs. This ensures that powerful computing resources (e.g., full-screen processing or multimodal analysis) are used only when needed. System 200 can continuously develop an understanding of the environment and a user's activities and, when necessary, transition from initial low-power detection to more in-depth analysis.
[0035] Fig. Figure 3 illustrates a view of a system 300 that combines a low-level signal processing system 302 with a powerful multimodal language modeling system 315 to capture detailed behaviors and events in real time, according to some implementations.
[0036] The Low-Level Signal Processing System 302 can be configured to operate continuously to detect basic scene-level signals and patterns from sensors such as motion detectors, microphones, etc., and thus capture rough scene-level descriptions, such as that the user is having breakfast or going to work.
[0037] The powerful multimodal language modeling system 315 can only be activated when significant events are detected, thus providing detailed interpretations of the significant events.
[0038] For example, a multi-level process can be activated so that the low-level signal processing system 302 continuously monitors user behavior, while the powerful multimodal speech modeling system 315 is triggered only when specific events are detected. Accordingly, detailed, time-related behaviors such as "You took medication" or "You left a coffee cup on the dining table" can be captured.
[0039] The multi-layered process mentioned above results in a semantic index or log of daily activities, with each detected event being tagged and stored in real time. This allows for the tracking of detailed behaviors and relationships between events, which can be useful for various applications such as health monitoring, personal diaries, productivity analyses, and more.
[0040] Accordingly, the System 300 can provide detailed logging of behaviors in real time while balancing energy efficiency by using the low-level signal processing system 302 to trigger the more powerful, expensive, high-performance multimodal speech modeling system 315 when needed.
[0041] Fig. Figure 4 illustrates a process 400 for using the system 200 from Fig. 2. To enable, for example, the location of a car in a garage, according to some implementations. Process 400 is configured to help users remember events by intelligently capturing and processing important interactions. Process 400 enables a combination of real-time data collection and event-driven processing to balance power consumption with functionality.
[0042] In some implementations, Process 400 is configured to create a software assistant that logs interactions or events in real time for future queries. In other implementations, Process 400 is configured to collect data on user behavior, preferences, and the environment, logging only relevant events to conserve battery life and processing power. Accordingly, a selective data collection process can be implemented so that, instead of continuous video recording or monitoring every action or event, Process 400 triggers data collection only at specific moments likely to be important or useful to a user.
[0043] For example, if a user drives to the airport and glances at a sign 402 that reads: “Parking lot full. Go to level 3”, process 400 can detect that the user is reading the sign and can capture a snapshot 403 of this event. Similarly, when the user parks the car, process 400 can log a position 405 based on the detection that a wireless system (e.g., a Bluetooth system), GPS, or other sensor 407 is disconnected, and then collect contextual data 409, 411, and 412, such as notes 414 or 416 viewed by the user or buttons 418 activated by the user.
[0044] In some implementations, Process 400 can be configured to handle multiple data formats, including images, text, GPS signals, object interaction data, and more. These different data formats can be processed by detecting and identifying key moments (of the data), such as reading a parking sign, pressing an elevator button, exiting a car, and so on. In some implementations, the detected and identified key moments can be passed to a multimodal language model (e.g., the LLM in Figure 228) to provide answers to questions such as "Where did I park my car?" Subsequently, associated images, text, and interaction logs can be combined into a query that is sent to the multimodal language model to provide a correct answer to the question(s).
[0045] In some implementations, interactions associated with Process 400 can be stored in a log or database for later querying. For example, an interaction might be associated with moments when the user meaningfully interacted with the world, such as looking at a sign, pressing a button, and so on. If a user subsequently asks a question like, "Where did I park my car?", Process 400 can use a voice query to trigger a search in a log of important events and retrieve snapshots or contextual data points relevant to the question. Accordingly, Process 400 is configured to recognize when a user is engaged in meaningful activities (e.g., reading a sign, interacting with objects, etc.), thus minimizing power consumption while still collecting enough data to provide useful insights.
[0046] In some implementations, sensors such as eye tracking, wireless signal interruption, or object interaction tracking can serve as low-power, always-on components that signal when more detailed data should be collected. Instead of generating a continuous data feed (e.g., of the last few days), Process 400 can therefore filter and retrieve snapshots of the most relevant events.
[0047] Instead of reviewing, for example, a video from the last few hours, the 400 process can retrieve only the moments when a user interacted with a specific parking sign or parked a car. This streamlines the query process and allows useful information to be obtained more quickly and efficiently without unnecessary data overload.
[0048] In some implementations, the 400 process may involve machine learning (ML) to predict relevant user moments and to use those relevant user moments to allow the ML to improve over time in recognizing patterns and predicting when data collection might be useful.
[0049] In some implementations, the 400 process can enable complex queries, such as requesting summaries of a user's day or week by combining the above snapshots into a coherent timeline of important events.
[0050] Accordingly, Process 400 can intelligently balance the efficiency of data collection and processing by capturing key moments relevant to a user's activities, while simultaneously minimizing power consumption by using sensors to detect these crucial moments in real time. This selective, event-driven approach enables Process 400 to function as a highly effective personal assistant, providing timely assistance and callbacks based on past interactions.
[0051] In some implementations, Process 400 enables cascading sensor and processing operations, so that an initial subset of low-power sensors (e.g., audio sensors) is initialized to capture data, such as audio, for input into a low-level manifold (LLM) (operating with limited, single-modal capacity) configured to process the audio. Subsequently, additional, more powerful sensors, such as cameras, can be activated to provide input to the LLM (operating multimodally) and thus interpret environmental context as well as hand and gaze gestures. For example, the operation of the initial low-power / low-computation sensors can be triggered upon detection of audio data such as speech. This operation can be configured to utilize a limited capacity of the LLM for speech interpretation.Based on the interpreted language, the second set of different (and more powerful) sensors (e.g. cameras) can be triggered to capture images for evaluating hand and eye gestures, thus refining a result that indicates a user request or user command.
[0052] Fig. Figure 5 illustrates a view of a System 500 which, according to some implementations, enables system activation using low-level signals as triggers. In some implementations, low-level signals received from the Sensors 509 (e.g., gaze detection sensors, motion detection sensors, sensors for detecting the separation of a wireless signal (long and short range), etc.) can be used to trigger (via a human basic model 510 and an adapter 514) the activation of more powerful components, such as cameras or speech models like the LLM 522.
[0053] In some implementations, the System 500 can activate the LLM 522 upon detecting a trigger event to further analyze the event using powerful, multimodal processing. This involves receiving inputs such as images 511 (e.g., video images) for processing via a video encoder 512 and an adapter 517, as well as text 519 for processing via a text tokenizer 518.
[0054] In some implementations, the System 500 activates higher-level components (such as the LLM 522 or a camera) only when it detects a relevant, human-related event (for example, when a user asks, "What is this?" while pointing at an object such as a hat 506a). For example, when high-level components such as a camera or LLM 522 are triggered, low-level signals received from the sensors 509 (such as gaze focus, gestures, object interaction, etc.) are configured to provide additional context to guide the high-level system (System 500) and ensure it does not process unnecessary data.Instead of analyzing an entire scene of image 502, for example, system 500 is configured to process only a relevant part 504 of image 502 that is associated with an interaction such as a glance, a hand gesture, and / or an audible request such as "What is that?".
[0055] Accordingly, by selectively capturing information (e.g. part 504 of Figure 502), the amount of data to be processed can be reduced, thereby supporting the efficiency and latency of System 500.
[0056] In some implementations, the System 500 allows for cascading sensor and processing operations. A low-power sensor (e.g., an audio sensor from Sensors 509) is initialized to capture data, such as audio, for input into the LLM 522 (which operates with limited, single-modal capacity), configured to process the audio. Subsequently, a more powerful sensor (e.g., a camera from Sensors 509) can be activated to provide input to the LLM 522 (now operating multimodally) to interpret environmental context as well as hand and gaze gestures. For example, the operation of the first low-power / low-performance sensors can be triggered upon detection of audio data such as speech. This process can be configured to utilize a limited capacity of the LLM 522 to interpret the speech.Based on the interpreted language, the second (and more powerful) sensor (e.g., a camera) can be triggered to capture images for evaluating hand and eye gestures, thus refining a result that indicates a user request or command, such as answering a question like "What is this?".
[0057] Fig. Figure 6 is a flowchart representation of an exemplary procedure 600 that detects and interprets user activities through a resource-intensive process triggered and / or guided by sensing and decision-making enabled by a resource-light process according to some implementations. In some implementations, procedure 600 is performed by a device, such as a mobile device, desktop, laptop, HMD, or server device. In some implementations, the device includes a screen for displaying images and / or a screen for viewing stereoscopic images, such as a head-mounted display (HMD, such as device 105 from [reference missing]). Fig. 1) In some implementations, Procedure 600 is performed by processing logic, including hardware, firmware, software, or a combination thereof. In other implementations, Procedure 600 is performed by a processor executing code stored in a non-transitory, computer-readable medium (such as memory). Each block in Procedure 600 can be activated and executed in any order.
[0058] In block 602, procedure 600 executes a first process (e.g., a resource-poor process, which might use the low-power system 210). Fig. 2 is implemented) to generate an output to trigger a second process, for example a resource-intensive process, which might use the powerful System 225. Fig. 2 is implemented. In some implementations, the first process includes: 1. Recognition of events based on an initial set of sensor data such as hand positions, gaze data, audio data, IMU data, etc., obtained via sensors such as OFC 217, IFC 218, sensor(s) 221 and IMU sensors 220, etc., as described in relation to Fig. 2 described. 2. Identifying a subset of events as human-relevant events corresponding to one or more predefined classes represented in the first set of sensor data. This may involve, for example, using an HFM 206 and a Relevance Decoder 204 to identify events relating to: hands and gaze, hand-object interactions, visual attention to objects, reading text, user-initiated speech or sounds, human body movements, etc., as described in relation to Fig. 2 described. 3. Gathering information on events relevant to humans based on the first set of sensor data.
[0059] In some implementations, the first set of sensor data includes data selected from the group consisting of hand position data, gaze data, audio data, and IMU data.
[0060] In block 604, procedure 600 executes a second process based on the output of the first process (e.g., triggered by or using information from it) to interpret user activity. The second process might involve receiving a second set of sensor data (e.g., image sensors, increasing the frame rate, increasing the resolution, etc., as described in the following example). Fig. 1) and include the interpretation of user activity using the second set of sensor data. For example, the interpretation of user activity may involve computationally intensive processes, such as the processing by the LLM 228, as described in relation to Fig. 2 described.
[0061] In some implementations, the second process can be triggered by the first process based on the detection of a human-relevant event. In some implementations, the human-relevant event can include, but is not limited to, an audible sound, an interaction with an object, a user movement, etc.
[0062] In some implementations, the second process can use an output from the first process (e.g., information about events relevant to humans) to interpret user activity.
[0063] In some implementations, the interpretation of user activity may include the classification of current events from the subset of events.
[0064] In some implementations, the interpretation of user activity includes the interpretation of a verbal utterance in combination with the user's gaze, gesture, body movement, body language, or facial expression.
[0065] In some implementations, the second set of sensor data may include data such as image sensor data, frame rate data, video resolution data, etc.
[0066] In some implementations, the interpretation of user activity using the second sensor data set may involve processing by the LLM.
[0067] Fig. Figure 7 is a block diagram of an exemplary device 700. The device 700 illustrates an exemplary device configuration for the electronic device 105. Fig. 1. Although certain specific features are illustrated, those skilled in the art will recognize from the present disclosure that various other features have not been illustrated for the sake of brevity, so as not to obscure more relevant aspects of the implementations disclosed herein. For this purpose, in some implementations, the device 700 includes, as a non-limiting example, one or more processing units 702 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, and / or the like), one or more input / output devices (I / O devices) and sensors 706, one or more communication interfaces 708 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, Zigbee, SPI, I2C, and / or similar interface types), one or more programming interfaces (e.g., I / O interfaces) 710, output devices (e.g.,one or more displays) 712, one or more inward and / or outward-facing image sensor systems 714, a memory 720 and one or more communication buses 704 for connecting these and various other components.
[0068] In some implementations, the one or more communication buses 704 include switching logic that connects and controls communications between system components. In some implementations, the one or more I / O devices and sensors 706 include at least one inertial measurement unit (IMU), accelerometer, magnetometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more loudspeakers, a haptic engine, one or more depth sensors (e.g., a fringe projection sensor, a time-of-flight sensor, or the like), one or more cameras (e.g., inward-facing cameras and outward-facing cameras of a head-mounted display (HMD)), and one or more infrared sensors, one or more thermal imaging sensors, and / or the like.
[0069] In some implementations, one or more 712 displays are configured to present the user with a view of a physical environment, a graphical environment, an augmented reality environment, etc. In other implementations, one or more 712 displays are configured to present content to the user (determined based on a specific user / object position within the physical environment).In some implementations, the one or more display(s) 712 correspond to a holographic digital light processing (DLP) display, a liquid crystal display (LCD), a liquid crystal on silicon (LCoS) display, an organic light-emitting field-effect transistor (OLET), an organic light-emitting diode (OLED), a surface conduction electron emitter (SED) display, a field emission display (FED), a quantum dot light-emitting diode (QD-LED), a microelectromechanical system (MEMS), and / or similar display types. In some implementations, the one or more display(s) 712 correspond to diffraction, reflection, polarized, holographic, etc., waveguide displays. In one example, the device 700 includes a single display. In another example, the device 700 includes a display for each eye of the user.
[0070] In some implementations, the one or more image sensor systems 714 are configured to obtain image data corresponding to at least a part of the physical environment 100. For example, the one or more image sensor systems 714 include one or more RGB cameras (e.g., with a complementary metal-oxide-semiconductor image sensor (CMOS image sensor) or a charge-coupled device image sensor (CCD image sensor)), monochrome cameras, IR cameras, depth cameras, event-based cameras, and / or the like. In various implementations, the one or more image sensor systems 714 also include illumination sources that emit light, such as a flash. In various implementations, the one or more image sensor systems 714 also include an in-camera image signal processor (ISP) configured to perform a variety of processing operations on the image data.
[0071] In some implementations, the Device 700 includes an eye-tracking system for detecting the position and movements of the eyes (e.g., gaze detection). For example, an eye-tracking system may include one or more infrared light-emitting diodes (IR LEDs), an eye-tracking camera (e.g., a near-infrared camera (NIR camera)), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) toward the user's eyes. Furthermore, the illumination source of the Device 700 may emit NIR light to illuminate the user's eyes, and the NIR camera may capture images of the user's eyes. In some implementations, the images captured by the eye-tracking system may be analyzed to determine the position and movements of the user's eyes or to obtain other information about the eyes, such as pupil dilation or pupil diameter.Furthermore, the gaze point estimated from the eye-tracking images can enable gaze-based interaction with content shown on the near-eye display of the Device 700.
[0072] Memory 720 includes high-speed random-access memory such as DRAM, SRAM, DDR-RAM, or other solid-state random-access memory devices. In some implementations, Memory 720 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 720 optionally includes one or more storage devices located remotely from the one or more Processing Units 702. Memory 720 includes a non-transient, computer-readable storage medium.
[0073] In some implementations, the memory 720 or the non-transitory, computer-readable storage medium of the memory 720 stores an optional operating system 730 and one or more instruction sets 740. The operating system 730 includes procedures for handling various basic system services and performing hardware-dependent tasks. In some implementations, the one or more instruction sets 740 include executable software defined by binary information stored in the form of electrical charge. In some implementations, the one or more instruction sets 740 is / are software executable by the one or more processing units 702 to perform one or more of the techniques described herein.
[0074] The instruction set(s) 740 include an instruction set 742 for generating data elements, an instruction set 744 for executing the first process, and an instruction set 746 for executing the second process. The instruction set(s) 740 can be executed as a single executable software file or as multiple executable software files.
[0075] The first process, which executes instruction set 742, is configured with processor-executable instructions to (e.g., continuously) run a resource-efficient process with respect to initial sensor data in order to trigger a second process for a more detailed interpretation of user activity.
[0076] The second process, which [should be "carried out" here? Should it be in the figure? - note the difference between 742 and 744 in Fig. 7] executing instruction set 744 is configured with instructions that can be executed by a processor to run a resource-intensive process (triggered by the resource-poor process) to interpret user activity based on additional sensor data that differs from the original sensor data.
[0077] The utterance interpretation instruction set 746 [which I do not see in the figure] is configured with instructions that can be executed by a processor to interpret the utterance using a subset of the data elements based on the time attributes.
[0078] Although one or more instruction sets 740 are shown to reside on a single device, it is understood that in other implementations any combination of elements may reside on separate computing devices. Furthermore, Fig. Section 7 is intended more as a functional description of the various features present in a given implementation, as opposed to a structural scheme of the implementations described herein. As average professionals will recognize, elements shown separately could be combined, and some elements could be separated. The actual number of instruction sets and how features are assigned to them may vary from one implementation to another and may depend in part on the specific combination of hardware, software, and / or firmware chosen for a particular implementation.
[0079] Returning to Fig.1. A physical environment refers to a physical world that humans can perceive and / or interact with without the aid of electronic devices. The physical environment can include physical features, such as a physical surface or a physical object. For example, the physical environment corresponds to a physical park, which includes physical trees, physical buildings, and physical people. Humans can directly perceive and / or interact with the physical environment, such as through sight, touch, hearing, taste, and smell. In contrast, an augmented reality (XR) environment refers to a fully or partially simulated environment that humans perceive through an electronic system and / or interact with through an electronic device. For example, the XR environment can include augmented reality (AR), mixed reality (MR), virtual reality, and / or similar content.An XR system tracks a subset of a person's physical movements, or their representation, and in response, adjusts one or more properties of one or more virtual objects simulated in the XR environment to conform to at least one physical law. For example, an XR system can detect head movement and, in response, adjust the graphical content and acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. Another example is that the XR system can detect movement of the electronic device presenting the XR environment (e.g., a camera).a mobile phone, tablet, laptop, or similar device) and, in response, adjust the graphical content and acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some situations (e.g., for accessibility reasons), the XR system can adjust a characteristic of graphical content in the XR environment in response to representations of physical movements (e.g., voice commands).
[0080] There are many different types of electronic systems that allow a person to perceive and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays designed as lenses intended to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., body-worn or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accommodate an external opaque display (e.g., a smartphone).The head-worn system may include one or more imaging sensors to capture images or video of the physical environment, and / or one or more microphones to capture audio of the physical environment. Instead of an opaque display, a head-worn system may have a transparent or translucent display. The transparent or translucent display may have a medium through which light, representative of images, is directed toward a person's eyes. The display may utilize digital light projection, OLEDs, LEDs, uLEDs, liquid crystals on silicon, a laser scanning light source, or any combination of these technologies. The medium may be an optical fiber, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In some implementations, the transparent or translucent display may be configured to become selectively opaque.Projection-based systems can employ retinal projection technology, which projects graphical images onto a person's retina. Projection systems can also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface.
[0081] Experts will recognize that generally known systems, methods, components, devices, and circuits are not described in detail in order to avoid obscuring more relevant aspects of the exemplary implementations described herein. Furthermore, effective aspects and / or variations do not encompass all the specific details described herein. Thus, several details are described to provide a thorough understanding of the exemplary aspects shown in the drawings. Moreover, the drawings merely illustrate some exemplary embodiments of the present disclosure and are therefore not to be considered limiting.
[0082] While this patent specification contains many specific implementation details, these should not be interpreted as limitations on the scope of any invention or on what may be claimed, but rather as descriptions of features specific to certain embodiments of certain inventions. Certain features described in this patent specification in connection with separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in connection with a single embodiment may also be implemented separately in several embodiments or in any suitable subcombination.Furthermore, although features described above may be effective in certain combinations and may even be initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0083] Similarly, although operations are depicted in the drawings in a specific order, this should not be interpreted as meaning that such operations must be performed in the specific order shown or in sequential order, or that all depicted operations must be performed to achieve the desired results. Multitasking and parallel processing may be advantageous under certain circumstances. Furthermore, the separation of the various system components in the embodiments described above should not be interpreted as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged across multiple software products.
[0084] Thus, certain embodiments of the subject matter have been described. Other embodiments are within the scope of protection of the following claims. In some cases, the actions specified in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the accompanying figures do not necessarily require the specific order or sequential sequence to achieve the desired results. In certain implementations, multitasking and parallel processing may be advantageous.
[0085] Embodiments of the subject matter and operations described in this patent specification can be implemented in digital electronic circuit arrangements or in computer software, firmware, or hardware, including the structures disclosed in this patent specification and their structural equivalents, or in combinations of one or more thereof. Embodiments of the subject matter described in this patent specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a computer storage medium for execution by or control of the operation of a data processing device. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g.,A computer storage medium is a machine-generated electrical, optical, or electromagnetic signal produced to encode information for transmission to a suitable receiving device for execution by a data processing device. A computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a storage arrangement, or a random-accessible or serial-accessible storage device, or a combination of one or more of these. Furthermore, while a computer storage medium is not a propagated signal, it may be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium may also consist of one or more separate physical components or media (e.g.,multiple CDs, hard drives or other storage devices) or be included therein.
[0086] The term "data processing equipment" encompasses all types of equipment, devices, and machines for processing data, including, for example, a programmable processor, a computer, a system-on-a-chip, or several or combinations thereof. The equipment may include specialized logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to hardware, the equipment may also include code that creates an execution environment for the computer program in question, such as code representing processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of these.The setup and execution environment can implement various different data processing model infrastructures, such as web services, distributed data processing, and network data processing infrastructures. Unless specifically stated otherwise, discussions contained in this description, using terms such as "processing," "processing data," "computing," "determining," and "identifying," or the like, are understood to refer to actions or processes of a data processing device, such as one or more computers or similar electronic data processing devices, that manipulate or transform data represented as physical, electronic, or magnetic quantities in storage devices, registers, or other information storage devices, transmission devices, or display devices of the data processing platform.
[0087] The system or systems discussed here are not limited to any particular hardware architecture or configuration. A data processing device may include any suitable arrangement of components that provides a result based on one or more inputs. Suitable data processing devices include general-purpose microprocessor-based computer systems that access stored software, which programs or configures the data processing system from a general-purpose computing device to a specialized data processing device that implements one or more implementations of the subject matter here.Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in software to be used in programming or configuring a data processing device.
[0088] Implementations of the methods disclosed herein can be carried out during the operation of such data processing devices. For example, the order of the blocks shown in the examples above can be varied, blocks can be rearranged, combined, and / or decomposed into sub-blocks. Certain blocks or processes can be performed in parallel. The operations described in this patent can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.
[0089] The use of "adapted to" or "configured to" herein is intended as an open and inclusive formulation that does not exclude devices that are adapted or configured to perform additional tasks or steps. Similarly, the use of "based on" is intended to be open and inclusive in that a process, step, calculation, or other action that is "based" on one or more specified conditions or values may, in practice, be based on additional conditions or a value beyond those specified. Headings, lists, and numbering contained herein are provided for convenience only and are not intended to be restrictive.
[0090] It is also understood that, although the terms "first," "second," etc., may be used herein to describe different elements, these elements are not to be restricted by these terms. These terms are used only to distinguish one element from another. For example, a first node could be called a second node, and similarly, a second node could be called a first node without changing the meaning of the description, as long as every occurrence of "first node" is consistently renamed and every occurrence of "second node" is consistently renamed. The first node and the second node are both nodes, but they are not the same node.
[0091] The terminology used herein serves only to describe certain implementations and is not intended to limit the claims. As used in the description of the implementations and the accompanying claims, the singular forms "a," "an," "the," "a," and "a" are intended to include the plural forms unless explicitly stated otherwise in the context. It is also understood that the term "and / or," as used herein, refers to and includes any and all possible combinations of one or more of the related elements listed.It is further understood that the terms “include” and / or “comprehensive”, when used in this patent specification, indicate the presence of listed features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0092] As used herein, the term "if" can, depending on the context, be interpreted as "upon" or "in response to finding" or "according to a finding" or "in response to recognizing" that a stated antecedent condition is satisfied. Likewise, the phrase "when it is found [that a stated antecedent condition is satisfied]" or "when [a stated antecedent condition is satisfied]" can, depending on the context, be interpreted as "upon finding" or "in response to a finding that" or "according to a determination" or "upon recognizing" or "in response to recognizing" that the stated antecedent condition is satisfied.
Claims
[1] Procedure, encompassing: on a device with a processor and one or more sensors: Performing an initial process to generate an output, wherein the initial process includes: Detecting events based on an initial set of sensor data; Identifying a subset of events as human-relevant events that correspond to one or more predetermined classes represented in the first set of sensor data; and Gathering information regarding events relevant to humans based on the first set of sensor data; and Based on the output of the first process, a second process is executed to... to interpret a user activity, wherein the second process involves obtaining a second set of sensor data and interpreting the user activity using the second set of sensor data. [2] Method according to claim 1, wherein the second process is triggered by the first process based on the detection of an event relevant to humans. [3] Method according to one of claims 1-2, wherein the event relevant to humans comprises an event selected from the group consisting of an audible sound, an interaction with an object and a user movement. [4] Method according to any one of claims 1-3, wherein the second process uses the output of the first process to interpret the user activity, the output comprising the information relating to the events relevant to the human. [5] Method according to any one of claims 1-4, wherein the interpretation of user activity comprises classifying current events from the subset of events. [6] Method according to any one of claims 1-5, wherein the interpretation of user activity comprises the interpretation of a verbal utterance in combination with the gaze, gesture, body movement, body language or facial expression of a user. [7] Method according to any one of claims 1-6, wherein the first set of sensor data comprises data selected from the group consisting of hand position data, gaze data, audio data and IMU data. [8] Method according to any one of claims 1-7, wherein the second set of sensor data comprises data selected from the group consisting of image sensor data, frame rate data and video resolution data. [9] Method according to any one of claims 1-8, wherein the interpretation of user activity using the second set of sensor data comprises processing using a large language model (LLM). [10] System, encompassing: one or more sensors; a non-transient, computer-readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions which, when executed on the one or more processors, cause the system to perform operations comprising one of the methods according to claims 1-9. [11] Non-transitory, computer-readable storage medium that stores program instructions executable by one or more processors to perform operations comprising one of the methods of claims 1-9.