Verbal command timestamp recordation and utilization
The system accurately interprets verbal commands using time-stamped sensor data from non-verbal activities, addressing the inefficiencies in existing command detection systems by correlating user gestures and speech for seamless extended reality interactions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-04-02
AI Technical Summary
Existing systems fail to accurately detect input commands, particularly audible and physical-based commands, in devices like head-mounted displays (HMDs), leading to inefficiencies in interpreting complex verbal commands with non-verbal user activities.
A system that interprets verbal commands using time-stamped data elements from sensor inputs, including user gaze and hand gestures, through a multi-modal time series model and transformer models to correlate non-verbal activities with verbal commands, enabling seamless interaction in extended reality environments.
Enables accurate interpretation of complex verbal commands by correlating non-verbal user activities with timestamps, allowing for intuitive and natural user interfaces in extended reality environments.
Smart Images

Figure US2025047633_02042026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 097425-01493(P69319WO)VERBAL COMMAND TIMESTAMP RECORDATION AND UTILIZATIONTECHNICAL FIELD
[0001] The present disclosure generally relates to systems, methods, and devices that interpret verbal commands using temporal and sensor data corresponding to non-verbal user activities.BACKGROUND
[0002] It may be desirable to detect input commands associated with performing actions while a user is using a device, such as a head mounted device (HMD), etc. However, existing systems may not provide accurate detection of input commands, for example, with respect to audible and physical based input commands.SUMMARY
[0003] Various implementations disclosed herein include devices, systems, and methods that interpret complex verbal commands occurring over a period of time using time-stamped data elements associated with sensor data corresponding to non-verbal user activities. For example, a user wearing an HMD (presenting a real or an extended reality (XR) environment) or recite the phrase “move this box from here to there” while performing a non-verbal activity such as gazing or pointing a hand or finger at a location of the box and subsequently gazing or pointing a hand or finger at another location.
[0004] In some implementations, multi-modal input occurring over time may be enabled. For example, multi-modal input may include, inter alia, input that includes user voice input, user gaze input, user hand or finger input, body language input, etc.
[0005] In some implementations, input signals corresponding to non-verbal user activities may be separated into parts such as tokens, etc. In some implementations, tokens may be associated with timestamps and a resulting structure may be stored for useAttorney Docket No. 097425-01493(P69319WO) in interpreting verbal commands. For example, tokens may be associated with timestamps using a temporal encoding technique.[00061 In some implementations, a machine learning (ML) model (e.g., a transformer) that tracks relationships in sequential data may be configured to operate on a subset of tokens selected based on a time period during which an audible (e.g., verbal) command was uttered. For example, a subset of tokens may be selected based on a start time and end time during in which a verbal command was uttered. In some implementations, an ML model may use a time sequence of the tokens to interpret a verbal command.
[0007] In some implementations, a user may utter a verbal command while performing one or more user activities such as gazing at or pointing towards one or more virtual objects or locations to issue a command instructing a device (e.g., an HMD or optical see through (OST) device) to perform a specified function such as, for example, moving a virtual picture from a position on a table to placement on a wall. In response, the device may be configured to interpret the verbal command using timestamps such as a sequence of the tokens and / or associating particular tokens with particular words or phrases of the verbal command. For example, particular word or phrase may include a trigger word in a sentence.
[0008] In some implementations, a device has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method performs one or more steps or processes. In some implementations, the device obtains a sequence of portions of sensor data occurring over a period of time via the one or more sensors. Each portion of sensor data is associated with a data timestamp within the period of time. The method may further obtain an audio signal containing a sequence of audio segments occurring over the period of time via the one or more sensors. Each audio segment may be associated with an audio timestamp within the period of time. The method may further match a first data timestamp with a first audio timestamp and perform a function using a first portion of sensor data associated with the first data timestamp matched with the first audio timestamp.Attorney Docket No. 097425-01493(P69319WO)
[0009] In accordance with some implementations, a device includes one or more processors, a non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors and the one or more programs include instructions for performing or causing performance of any of the methods described herein. In accordance with some implementations, a non-transitory computer readable storage medium has stored therein instructions, which, when executed by one or more processors of a device, cause the device to perform or cause performance of any of the methods described herein. In accordance with some implementations, a device includes: one or more processors, a non-transitory memory, and means for performing or causing performance of any of the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] So that the present disclosure can be understood by those of ordinary skill in the art, a more detailed description may be had by reference to aspects of some illustrative implementations, some of which are shown in the accompanying drawings.
[0011] Figure 1 illustrates an exemplary electronic device operating in a physical environment, in accordance with some implementations.
[0012] Figure 2 illustrates a view of a system configured to accept multi-modal signals from a device and create and / or train a multi-modal time series model for interpreting a verbal command, in accordance with some implementations.
[0013] Figures 3A-3C illustrate views of a process implemented by the system of figure 2 for accepting and interpreting multi-modal signals, in accordance with some implementations.
[0014] Figure 4 illustrates a view of a system that includes a large language model (LLM) with an HFM configured to generate a multi-modal sequence, in accordance with some implementations.
[0015] Figure 5 illustrates a view of a system configured to implement a temporal encoding process for interpreting verbal commands using temporal and sensor data corresponding to non-verbal user activities, in accordance with some implementations.Attorney Docket No. 097425-01493(P69319WO)
[0016] Figure 6 is a flowchart representation of an exemplary method that interprets complex verbal commands occurring over a period of time using time-stamped data elements associated with sensor data corresponding to non-verbal user activities, in accordance with some implementations.
[0017] Figure 7 is a flowchart representation of an exemplary method that interprets audible commands occurring over a period of time using time-stamped tokens matched to audible timestamps, in accordance with some implementations.
[0018] Figure 8 is a block diagram of an electronic device of in accordance with some implementations.
[0019] In accordance with common practice the various features illustrated in the drawings may not be drawn to scale. Accordingly, the dimensions of the various features may be arbitrarily expanded or reduced for clarity. In addition, some of the drawings may not depict all of the components of a given system, method or device. Finally, like reference numerals may be used to denote like features throughout the specification and figures.DESCRIPTION
[0020] Numerous details are described in order to provide a thorough understanding of the example implementations shown in the drawings. However, the drawings merely show some example aspects of the present disclosure and are therefore not to be considered limiting. Those of ordinary skill in the art will appreciate that other effective aspects and / or variants do not include all of the specific details described herein. Moreover, well-known systems, methods, components, devices and circuits have not been described in exhaustive detail so as not to obscure more pertinent aspects of the example implementations described herein.
[0021] Figure 1 illustrates an exemplary electronic device 105 operating in a physical environment 100. In the example of Figure 1, the physical environment 100 is a room that includes a desk 120. The electronic device 105 may include one or more cameras, microphones, depth sensors, or other sensors that can be used to capture information about and evaluate the physical environment 100 and the objects within it, as well as informationAttorney Docket No. 097425-01493(P69319WO) about the user 102 of electronic device 105. The information about the physical environment 100 and / or user 102 may be used to provide visual and audio content and / or to identify the current location of the physical environment 100 and / or the location of the user within the physical environment 100.
[0022] In some implementations, views of an extended reality (XR) environment may be provided to one or more participants (e.g., user 102 and / or other participants not shown) via electronic device 105 (e.g., a wearable device such as an HMD, an optical see through (OST) device, artificial intelligence (A / I) glasses (e.g., integrated with generative Al assistance for real time translation, contextual answers and summarization), alternative reality (A / R) glasses, mixed reality glasses, connected glasses, etc.). Such an XR environment may include views of a 3D environment that is generated based on camera images and / or depth camera images of the physical environment 100 as well as a representation of user 102 based on camera images and / or depth camera images of the user 102. Such an XR environment may include virtual content that is positioned at 3D locations relative to a 3D coordinate system associated with the XR environment, which may correspond to a 3D coordinate system of the physical environment 100.
[0023] In some implementations, the electronic device 105 does not include a display and the user views the physical environment 100 directly, e.g., using optical see-through components and / or through transparent lenses on the electronic device 105. In some implementations, such lenses are not configured to display content. In some implementations, such lenses are configured to display content (e.g., presenting an extended reality (XR) environment by displaying augmentations or other virtual content (that augments the user’s view of the physical environment 100) using optical waveguides that transmit light to display content on the lenses). In some implementations, the electronic device 105 includes a display that presents views of the physical environment 100 that are based on images captured by outward- facing cameras on the electronic device 105, e.g., by providing passthrough video, with or without added virtual content. In some implementations, the display presents an XR environment by displaying augmentations or other virtual content (that augments displayed depictions of the physical environment 100).Attorney Docket No. 097425-01493(P69319WO)
[0024] In some implementations, a system including electronic device 105 may be configured to generate sequences of data elements. The sequences may be generated by separating data corresponding to input signals related to non-verbal user activities. In some implementations, the input signals may be obtained via the sensors and the data elements may be associated with timestamps. For example, data elements such as tokens corresponding to input signals may be generated. In some implementations the input signals may correspond to user eye / gaze direction, hand or body gestures, body language, etc.
[0025] In some implementations, timing attributes of an audio signal (e.g., an utterance of a user) occurring over a time period may be determined based on, for example, audio data obtained from the sensors. For example, identifying start time and stop time of a verbal command may be identified. Likewise, it may be determined when specified words in a verbal command occurred.
[0026] In some implementations, the utterance may be interpreted using a subset of the data elements with respect to the timing attributes.
[0027] In some implementations, a stream of verbal commands and additional multimodal inputs such as gestures, visual cues, and / or sensor data may be continuously monitored. The stream of verbal commands and additional multi-modal inputs may be subsequently time-stamped such that each piece of data is associated with a specific period in time. In some implementations, the system may scan for predefined trigger words or phrases in verbal commands occurring within the specific period in time. A trigger may additionally include a specific combination of multi-modal inputs (e.g., a spoken command paired with a gesture captured by a camera).
[0028] In some implementations when a trigger is detected, the system may reference a specific timestamp associated with that trigger event thereby enabling the system to correlate the verbal command with other multi-modal inputs captured at a same time. Based on the detected trigger and the associated multi-modal data, the system may perform a specific function such as, if a verbal command "turn on the lights" is detected and is accompanied by a physical gesture, the system may execute the command to turn on the lights. Likewise, the system may be configured to accept multiple commands such as a secondary triggering event occurring after a first event to enable a function to beAttorney Docket No. 097425-01493(P69319WO) executed. For example, this may involve a sequence of gestures, a combination of gestures and verbal commands, sequential verbal commands, etc.
[0029] Figure 2 illustrates a view of a system 200 configured to accept multi-modal signals 215 from a device 203 and create and / or train a multi-modal time series model 210 for interpreting a verbal command(s) 206 for performing actions with respect to virtual content 204, in accordance with some implementations. In some implementations, system 200 may be configured to create multi-modal time series model 210 (e.g., a human foundation model (HFM)) that may be configured to interpret and respond to user behavior (e g., from a user 202) in real-time. Multi-modal time series model 210 may be configured to accept inputs from various devices, such as device 203 (e.g., an HMD or optical see through (OST) device) and / or other sensors to interpret user 202 attention, intention, and / or references from speech (e.g., an utterance) and / or gestures (e.g., gaze, hand, etc.).
[0030] In some implementations a multi-modal time series model 210 may include a machine learning (ML) model trained on time-series data from multiple modalities (e.g., vision, speech, gestures, etc.) associated with multi-modal signals 215. In some implementations, multi-modal time series model 210 may be configured to interpret human behavior at a fundamental level such as, for example, detecting what user 202 is focusing on (e.g., attention) or is trying to achieve (e.g., intention).
[0031] In some implementations, multi-modal time series model 210 may include a rule-based deterministic model (e.g., a decision tree, a state machine, a rule-based filter, etc.) explicitly defining a set of rules or conditions that guide the decisions and outputs associated with time-series data from multiple modalities (e.g., vision, speech, gestures, etc.) associated with multi-modal signals 215.
[0032] In some implementations, multi-modal time series model 210 may include a heuristic model (e.g., a greedy algorithm, a genetic algorithm, a search algorithm, a Monte Carlo method, etc.) using practical, experience-based methods or rules of thumb to solve problems or make decisions with respect to time-series data from multiple modalities (e.g., vision, speech, gestures, etc.) associated with multi-modal signals 215.
[0033] In some implementations, system 200 is configured to enable multi-modal time series model 210 to monitor user 202 interactions with a virtual environment (e.g.,Attorney Docket No. 097425-01493(P69319WO) virtual content 204 viewed via device 203 such as an HMD, Al glasses, OST device, etc.) using natural behaviors such as, inter alia, looking at, speaking about, or gesturing toward objects to issue commands. For example, user 202 say “move that box from here to there” (i.e., via verbal command 206) and in response, system 200 may interpret verbal command 206 based on the user's 202 language, gaze, and / or gestures to subsequently execute an action (“move that box from here to there”) in the virtual environment. For example, based on the verbal command (s) 206, system 200 may be configured to determine which tokens to reference. Likewise, if system 200 determines that a user wants to move an object from one place to another, the system 200 may be configured to determine which words to reference back to.
[0034] In some implementations, multi-modal time series model 210 is configured to analyze sequences of inputs (e.g., multi-modal signals 215) from various modalities to process data from a behavior of user 202, an environmental context, and / or any verbal commands (e.g., verbal command(s) 206) to generate appropriate responses such as translating verbal command(s) 206 into a structured action such as moving a virtual object from a source position to a target position.
[0035] In some implementations, system 200 may use a human moderator to manually interpret and simulate responses from multi-modal time series model 210 based on user 202 behavior thereby gathering initial training data and fine-tuning multi-modal time series model 210 prior to deployment.
[0036] In some implementations, system 200 is configured to create an intuitive and natural user interface to allow seamless interactions with virtual environments using everyday human behaviors.
[0037] Figures 3A-3C illustrate views 300a-300c of a process implemented by system 200 of figure 2 for accepting and interpreting multi-modal signals to perform actions with respect to virtual content 302, in accordance with some implementations.
[0038] In Figure 3A, at a first instant in time corresponding to view 300a, a user (e.g., user 202 of FIG. 2) initializes a process for providing multi-modal input. In some embodiments, the user can issue a verbal command (e.g., “Hey Siri”).
[0039] Subsequently, at a second instant in time corresponding to view 300b, the user issues a verbal command (e.g., “move this from here to there”) simultaneously inAttomey Docket No. 097425-01493(P69319WO) combination with a gesture-based command that includes the user pointing (via hand 306) at a first position 308a (shown in Figure 3A) and at a second position 308b (illustrated in Figure 3B and subsequent to pointing at first position 308a illustrated in Figure 3A) for movement / placement of virtual object 310.
[0040] In Figure 3C at a third instant in time corresponding to view 300c, the multimodal input signal / commands of figure 3B (i.e., the verbal command and the gesturebased command of Figure 3B) are interpreted and in response the virtual object 310 is moved from first (initial) position 308a for placement at second position 308b (e.g., on a wall). For example, the verbal command and the gesture-based command are interpreted when the full command (e.g., “move this from here to there”) is received. The verbal command and the gesture-based command may be interpreted by reviewing tokens and timestamps to analyze and execute the full command. Accordingly, the user’s natural behavior is interpreted for performing an action such as moving virtual object 310 between positions 308a and 308b.
[0041] Figure 4 illustrates a view of a system 400 that includes a multi-modal temporal module (MTM) 401 a (e.g., large language model (LLM) 401) with an HFM 409 configured to generate a multi-modal sequence 402 with respect to multi-modal inputs 415, in accordance with some implementations. In some implementations, system 400 is a multi-modal system that integrates various types of data inputs (e.g., speech, gestures, etc.) into a unified framework that may be processed by MTM 401 for higher- order reasoning and decision-making.
[0042] In some implementations, HFM 409 may accept (as input) multi-modal inputs 415 to generate as output, HFM tokens 405. HFM tokens 405 are configured to represent core human interactions or data points that may be obtained from various sensors or input methods (e.g., cameras, microphones, etc.).
[0043] In some implementations, multi-modal inputs 415 (e.g., multi-modal signals) may include visual data to provide spatial and geometric information. Likewise, multimodal inputs 415 (e.g., multi-modal signals) may include audible inputs representing speech or sounds.
[0044] In some implementations, time series model / adapter 403 processes multimodal tokens 405 (i.e., an output from HFM 409) over time. For example, time seriesAttorney Docket No. 097425-01493(P69319WO) model / adapter 403 may be configured to capture continuous streams of behavior and convert them into a form that is understandable for further processing.
[0045] In some implementations, time series model / adapter 403 is configured to map multi-modal tokens 405 (produced by FM 409) to a language space. For example, time series model / adapter 403 may be configured to translate various forms of data (e.g., visual, spatial, auditory, etc.) into a common representational space such as a multi-modal sequence 402 to be input into a LLM 401 for processing.
[0046] In some implementations, LLM 401 uses multi-modal sequence 402 to interpret the sequence of multi-modal inputs 415 (e.g., gestures and speech) and generate a response or command such as an action. For example, multi-modal data that includes different data types (e.g., geometric information from user behavior and textual data from transcribed speech) may be obtained and combined into a single multi-modal sequence representing a context of user interactions. The single multi-modal sequence may be translated into tokens in a language space for input into an LLM to apply reasoning capabilities to generate an appropriate output such as, inter alia, a command, an action, etc. with respect to executing functions in an extended reality (XR) environment (e.g., moving virtual objects between locations in the XR environment).
[0047] In some implementations, system 400 may use a deterministic, rules- based / heuristic approach to generate a command, an action, etc. with respect to executing functions in an extended reality (XR) environment (e.g., moving virtual objects between locations in the XR environment). For example, multimodal sensor data (visual, audible, etc.) may be continuously collected and when a command is issued (e.g., spoken), system 400 may be configured to determine whether the command matches a list of predetermined words (e.g., “that”, “there,” “here,” etc.) and collect the multimodal sensor data at those instances of time where those words were spoken to understand the user’s command.
[0048] Figure 5 illustrates a view of a system 500 configured to implement a temporal encoding process for interpreting verbal commands using temporal and sensor data corresponding to non-verbal user activities, in accordance with some implementations. A temporal encoding process is configured to provide transformerAttorney Docket No. 097425-01493(P69319WO) models (e.g., transformer models 504a-504d and 507a-507d) with information related to positions of tokens within a sequence with respect to time.
[0049] In some implementations, a token may be defined as a basic unit of text that a model (e.g., multi-modal transformer encoder 514) may use to process language such as, for example, a word, a portion of a word, punctuation, etc. Likewise in a multi-modal setting, a token may represent a basic unit of input from various modalities, such as visual cues (e.g., sensor data 502a-502c including hand movements or gaze direction, etc.), auditory signals (e.g., speech), or other sensory inputs (e.g., sensor data 502d including UI entities).
[0050] In some implementations, a token may represent sensory inputs such as two- dimensional (2D) or three-dimensional (3D) regions (e.g., bounding boxes, bounding cuboids, segmentation polygons, etc.) corresponding to real-world objects detected or tracked by system 500 using visual or other sensor modalities such as, inter alia, LiDAR, sonar, radar, ultrasound, GPS, etc.
[0051] In some implementations, tokens may not just represent single words but may instead represent more abstract chunks of text with respect to a multi-modal context. For example, tokens may represent chunks of time series data that have been tokenized or transformed into a format for processing by multi-modal transformer encoder 514. The aforementioned chunking process may enable multi-modal transformer encoder 514 to learn abstract features that are temporally localized but may additionally capture information from across an entire input sequence due to the use of multiple layers of attention mechanisms. These attention layers may provide both local and global context thereby enabling multi-modal transformer encoder 514 to integrate information from different portions of a sequence to understand context and meaning.
[0052] In some implementations, system 500 executes a process that includes obtaining sensor data 502a-502d (e.g., hand pose data, hand / head location data, gaze data, UI entities data, voice data, tracked 3D objects, etc.) for input into transformers 504a- 504d (e.g., a hand pose transformer, a body pose transformer, a gaze neural network or transformer, a UI transformer, etc.), respectively. In response, transformers 504a-504d may generate as output, tokens 505a-505d for input into time series transformers 507a- 507d to subsequently generate associated time series tokens 508a-508n organized as timeAttorney Docket No. 097425-01493(P69319WO) and modality tokens 509 being input into multi-modal transformer encoder 510. Subsequently, tokens 512 are generated and input into interaction state recognition module 514 to generate an output 518 such as, inter alia, a command, an action, etc. with respect to executing functions in an XR environment.
[0053] Figure 6 is a flowchart representation of an exemplary method 600 that interprets complex verbal commands occurring over a period of time using time-stamped data elements associated with sensor data corresponding to non-verbal user activities, in accordance with some implementations. In some implementations, the method 600 is performed by a device, such as a mobile device, desktop, laptop, HMD, Al glasses, OST device or server device. In some implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images such as a head-mounted display (HMD such as e.g., device 105 of Figure 1). In some implementations, the method 600 is performed by processing logic, including hardware, firmware, software, or a combination thereof. In some implementations, the method 600 is performed by a processor executing code stored in a non-transitory computer- readable medium (e.g., a memory). Each of the blocks in the method 600 may be enabled and executed in any order.
[0054] At block 602, the method 600 generates one or more sequences of data elements by separating data corresponding to one or more input signals corresponding to one or more non-verbal user activities over a time period. In some implementations, the one or more input signals may be obtained via the one or more sensors and the data elements may be associated with timestamps. For example, a multi-modal time series model 210 may be configured to accept inputs from various devices such as a device 203 or other sensors as described with respect to figure 2. In some implementations, the data elements may be tokens such as tokens 405 as described with respect to figure 4.
[0055] In some implementations, timestamps associated with the data elements may provide common temporal position encoding for user activities corresponding to different input modalities such as hand gestures, gaze, body language, etc. as described with respect to figure 5.
[0056] At block 604, the method 600 determines timing attributes of an utterance of a user occurring over the time period. For example, a multi-modal time series model 210 may be configured to accept speech such as an utterance as described with respect toAttorney Docket No. 097425-01493(P69319WO) figure 2. The timing attributes may be determined based on audio data obtained from the one or more sensors.
[0057] At block 606, the method 600 interprets the utterance using a subset of the data elements based on the timing attributes. For example, a multi-modal time series model 210 may be configured to interpret a verbal command(s) 206 as described with respect to figure 2. For example, based on the verbal command (s) 206, the method 600 may determine what tokens to reference. For example, if method 600 determines that the user wants to move an object from one place to another, the method 600 is configured to determine which words to reference back to.
[0058] In some implementations, interpreting the utterance may include selecting the subset of the data elements based on the timing attributes. The timing attributes may include a start time and an end time of the time period during which the utterance occurred. In some implementations, logic for selecting the start and end time may include additional heuristics other than the exact time-period during which the utterance occurred. For example, if the duration of the utterance is shorter than a pre-defined minimum duration (or minimum number of elements), a start time may be adjusted such that the minimum duration (or minimum number of elements) is selected. For short utterances, this may result in the data elements including some time before the utterance began.
[0059] In some implementations, interpreting the utterance may include utilizing the timing attributes to interpret the sequencing of the one or more sequences of data elements.
[0060] In some implementations, interpreting the utterance may include identifying a data element corresponding to a word of the utterance based on the timing attributes.
[0061] In some implementations, interpreting the utterance may include using a machine learning model that tracks relationships in the sequences of one or more data elements. The machine learning model may be configured to operate on the subset of the data elements and the subset is identified based on the timing attributes.Attorney Docket No. 097425-01493(P69319WO)
[0062] In some implementations, the machine learning model comprises a transformer such as transformer models 504a-504d and 507a-507d as described with respect to figure 5.
[0063] At block 608, the method 600 performs an action within an extended reality (XR) environment based on interpreting the utterance.
[0064] Figure 7 is a flowchart representation of an exemplary method 700 that interprets audible commands occurring over a period of time using time-stamped tokens matched to audible timestamps, in accordance with some implementations. In some implementations, the method 700 is performed by a device, such as a mobile device, desktop, laptop, HMD, Al glasses, OST device or server device. In some implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images such as a head-mounted display (HMD such as e.g., device 105 of Figure 1). In some implementations, the method 700 is performed by processing logic, including hardware, firmware, software, or a combination thereof. In some implementations, the method 700 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory). Each of the blocks in the method 700 may be enabled and executed in any order.
[0065] At block 702, the method 700 obtains a sequence of tokens occurring over a period of time via the one or more sensors. For example, tokens corresponding to input signals such as eye / gaze direction, hand or body gestures, body language, etc. as described with respect to figure 1. Each token may be associated with a data timestamp within the period of time. For example, a time series model / adapter 403 may process multi-modal tokens 405 over time as described with respect to figure 4.
[0066] In some implementations, each data timestamp may provide common temporal position encoding for user activities corresponding to different input modalities such as hand gestures, gaze, body language, etc.
[0067] At block 704, the method 700 obtains an audio signal containing a sequence of audio segments occurring over the period of time via the one or more sensors. In some implementations, each audio segment may be associated with an audio timestamp within the period of time such as, for example, identifying a start time and a stop time of a verbal command as described with respect to figure 1.Attorney Docket No. 097425-01493(P69319WO)
[0068] In some implementations, the sequence of audio segments may include an utterance from a user.
[0069] At block 706, the method 700 matches a first data timestamp with a first audio timestamp as described with respect to figure 1.
[0070] In some implementations, a triggering audio element of the sequence of audio segments may be determined. In some implementations, triggering the audio element may enable identification of the first audio timestamp associated with the triggering audio element.
[0071] In some implementations, matching the first data timestamp with the first audio timestamp may include using a machine learning model that tracks relationships in the sequence of tokens. For example, a machine learning model that tracks relationships in the sequences of one or more data elements or tokens as described with respect to figure 5. In some implementations, a machine learning model may include a transformer (e.g., transformer models 504a-504d and 507a-507d as described with respect to figure 5).
[0072] In some implementations, matching the first data timestamp with the first audio timestamp may include using a deterministic, rules-based approach that tracks relationships in the sequence of tokens.
[0073] At block 708 the method 700 performs a function (e.g., an action with respect to virtual content 204 as described with respect to figure 2) using a first token associated with the first data timestamp matched with the first audio timestamp.
[0074] In some implementations, the function may be determined by analyzing the sequence of audio segments occurring over the period of time.
[0075] In some implementations, the function may be determined based on the utterance determined at block 704 and then based on the determined function, the method 700 may be configured to identify which words such as for example, “here” and / or “there” are pertinent to perform the determined function. Likewise, based on the aforementioned words, the method 700 may identify an associated user gesture / gaze input necessary to perform the determined function. The associated user gestures / gaze input may be identified based on the matching timestamp.Attorney Docket No. 097425-01493(P69319WO)
[0076] In some implementations, a second data timestamp may be matched with a second audio timestamp such that performing the function further uses a second token associated with the second data timestamp matched with the second audio timestamp.
[0077] Figure 8 is a block diagram of an example device 800. Device 800 illustrates an exemplary device configuration for electronic device 105 of Figure 1. While certain specific features are illustrated, those skilled in the art will appreciate from the present disclosure that various other features have not been illustrated for the sake of brevity, and so as not to obscure more pertinent aspects of the implementations disclosed herein. To that end, as a non-limiting example, in some implementations the device 800 includes one or more processing units 802 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, and / or the like), one or more input / output (I / O) devices and sensors 806, one or more communication interfaces 808 (e g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.1 lx, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, SPI, I2C, and / or the like type interface), one or more programming (e.g., I / O) interfaces 810, output devices (e.g., one or more displays) 812, one or more interior and / or exterior facing image sensor systems 814, a memory 820, and one or more communication buses 804 for interconnecting these and various other components.
[0078] In some implementations, the one or more communication buses 804 include circuitry that interconnects and controls communications between system components. In some implementations, the one or more I / O devices and sensors 806 include at least one of an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptics engine, one or more depth sensors (e.g., a structured light, a time-of-flight, or the like), one or more cameras (e.g., inward facing cameras and outward facing cameras of an HMD), one or more infrared sensors, one or more heat map sensors, and / or the like.
[0079] In some implementations, the one or more displays 812 are configured to present a view of a physical environment, a graphical environment, an extended reality environment, etc. to the user. In some implementations, the one or more displays 812 areAttorney Docket No. 097425-01493(P69319WO) configured to present content (determined based on a determined user / object location of the user within the physical environment) to the user. In some implementations, the one or more displays 812 correspond to holographic, digital light processing (DLP), liquidcrystal display (LCD), liquid-crystal on silicon (LCoS), organic light- emitting fieldeffect transitory (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum-dot lightemitting diode (QD-LED), micro-electromechanical system (MEMS), and / or the like display types. In some implementations, the one or more displays 812 correspond to diffractive, reflective, polarized, holographic, etc. waveguide displays. In one example, the device 800 includes a single display. In another example, the device 800 includes a display for each eye of the user.
[0080] In some implementations, the one or more image sensor systems 814 are configured to obtain image data that corresponds to at least a portion of the physical environment 100. For example, the one or more image sensor systems 814 include one or more RGB cameras (e.g., with a complimentary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, and / or the like. In various implementations, the one or more image sensor systems 814 further include illumination sources that emit light, such as a flash. In various implementations, the one or more image sensor systems 814 further include an on-camera image signal processor (ISP) configured to execute a plurality of processing operations on the image data.
[0081] In some implementations, the device 800 includes an eye tracking system for detecting eye position and eye movements (e.g., eye gaze detection). For example, an eye tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye tracking camera (e.g., near-IR (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) towards the eyes of the user. Moreover, the illumination source of the device 800 may emit NIR light to illuminate the eyes of the user and the NIR camera may capture images of the eyes of the user. In some implementations, images captured by the eye tracking system may be analyzed to detect position and movements of the eyes of the user, or to detect other information about the eyes such as pupil dilation or pupil diameter. Moreover, the point of gaze estimated fromAttorney Docket No. 097425-01493(P69319WO) the eye tracking images may enable gaze-based interaction with content shown on the near-eye display of the device 800.
[0082] The memory 820 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some implementations, the memory 820 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 820 optionally includes one or more storage devices remotely located from the one or more processing units 802. The memory 820 includes a non-transitory computer readable storage medium.
[0083] In some implementations, the memory 820 or the non-transitory computer readable storage medium of the memory 820 stores an optional operating system 830 and one or more instruction set(s) 840. The operating system 830 includes procedures for handling various basic system services and for performing hardware dependent tasks. In some implementations, the instruction set(s) 840 include executable software defined by binary information stored in the form of electrical charge. In some implementations, the instruction set(s) 840 are software that is executable by the one or more processing units 802 to carry out one or more of the techniques described herein.
[0084] The instruction set(s) 840 includes a data element generation instruction set 842, a timing attribute determination instruction set 844, and an utterance interpretation instruction set 846. The instruction set(s) 840 may be embodied as a single software executable or multiple software executables.
[0085] The data element generation instruction set 842 is configured with instructions executable by a processor to generate one or more sequences of data elements by separating data corresponding to one or more input signals corresponding to one or more non-verbal user activities.
[0086] The timing attribute determination instruction set 844 is configured with instructions executable by a processor to determine timing attributes of an utterance of a user occurring over a time period.Attorney Docket No. 097425-01493(P69319WO)
[0087] The uterance interpretation instruction set 846 is configured with instructions executable by a processor to interpret the utterance using a subset of the data elements based on the timing attributes.
[0088] Although the instruction set(s) 840 are shown as residing on a single device, it should be understood that in other implementations, any combination of the elements may be located in separate computing devices. Moreover, Figure 8 is intended more as functional description of the various features which are present in a particular implementation as opposed to a structural schematic of the implementations described herein. As recognized by those of ordinary skill in the art, items shown separately could be combined and some items could be separated. The actual number of instructions sets and how features are allocated among them may vary from one implementation to another and may depend in part on the particular combination of hardware, software, and / or firmware chosen for a particular implementation.
[0089] Returning to Figure 1, a physical environment refers to a physical world that people can sense and / or interact with without aid of electronic devices. The physical environment may include physical features such as a physical surface or a physical object. For example, the physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment such as through sight, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a wholly or partially simulated environment that people sense and / or interact with via an electronic device. For example, the XR environment may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, and / or the like. With an XR system, a subset of a person’s physical motions, or representations thereof, are tracked, and, in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that comports with at least one law of physics. As one example, the XR system may detect head movement and, in response, adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. As another example, the XR system may detect movement of the electronic device presenting the XR environment (e g., a mobile phone, a tablet, a laptop, or the like) and, in response,Attorney Docket No. 097425-01493(P69319WO) adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some situations (e.g., for accessibility reasons), the XR system may adjust characteristic(s) of graphical content in the XR environment in response to representations of physical motions (e.g., vocal commands).
[0090] There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head mountable systems, projection-based systems, heads-up displays (HUDs), vehicle windshields having integrated display capability, windows having integrated display capability, displays formed as lenses designed to be placed on a person’s eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head mountable system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head mountable system may be configured to accept an external opaque display (e.g., a smartphone). The head mountable system may incorporate one or more imaging sensors to capture images or video of the physical environment, and / or one or more microphones to capture audio of the physical environment. Rather than an opaque display, a head mountable system may have a transparent or translucent display. The transparent or translucent display may have a medium through which light representative of images is directed to a person’s eyes. The display may utilize digital light projection, OLEDs, LEDs, uLEDs, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a hologram medium, an optical combiner, an optical reflector, or any combination thereof. In some implementations, the transparent or translucent display may be configured to become opaque selectively. Projection-based systems may employ retinal projection technology that projects graphical images onto a person’s retina. Projection systems also may be configured to project virtual objects into the physical environment, for example, as a hologram or on a physical surface.
[0091] Those of ordinary skill in the art will appreciate that well-known systems, methods, components, devices, and circuits have not been described in exhaustive detail so as not to obscure more pertinent aspects of the example implementations describedAttorney Docket No. 097425-01493(P69319WO) herein. Moreover, other effective aspects and / or variants do not include all of the specific details described herein. Thus, several details are described in order to provide a thorough understanding of the example aspects as shown in the drawings. Moreover, the drawings merely show some example embodiments of the present disclosure and are therefore not to be considered limiting.
[0092] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombmation.
[0093] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0094] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.Attorney Docket No. 097425-01493(P69319WO)
[0095] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively, or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0096] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a crossplatform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures. Unless specifically stated otherwise, it is appreciated that throughout thisAttorney Docket No. 097425-01493(P69319WO) specification discussions utilizing the terms such as “processing,” “computing,” “calculating,” “determining,” and “identifying” or the like refer to actions or processes of a computing device, such as one or more computers or a similar electronic computing device or devices, that manipulate or transform data represented as physical electronic or magnetic quantities within memories, registers, or other information storage devices, transmission devices, or display devices of the computing platform.
[0097] The system or systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device can include any suitable arrangement of components that provides a result conditioned on one or more inputs. Suitable computing devices include multipurpose microprocessor-based computer systems accessing stored software that programs or configures the computing system from a general purpose computing apparatus to a specialized computing apparatus implementing one or more implementations of the present subject matter. Any suitable programming, scripting, or other type of language or combinations of languages may be used to implement the teachings contained herein in software to be used in programming or configuring a computing device.
[0098] Implementations of the methods disclosed herein may be performed in the operation of such computing devices. The order of the blocks presented in the examples above can be varied for example, blocks can be re-ordered, combined, and / or broken into sub-blocks. Certain blocks or processes can be performed in parallel. The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0099] The use of “adapted to” or “configured to” herein is meant as open and inclusive language that does not foreclose devices adapted to or configured to perform additional tasks or steps. Additionally, the use of “based on” is meant to be open and inclusive, in that a process, step, calculation, or other action “based on” one or more recited conditions or values may, in practice, be based on additional conditions or value beyond those recited. Headings, lists, and numbering included herein are for ease of explanation only and are not meant to be limiting.Attorney Docket No. 097425-01493(P69319WO)
[0100] It will also be understood that, although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first node could be termed a second node, and, similarly, a second node could be termed a first node, which changing the meaning of the description, so long as all occurrences of the “first node” are renamed consistently and all occurrences of the “second node” are renamed consistently. The first node and the second node are both nodes, but they are not the same node.
[0101] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the claims. As used in the description of the implementations and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0102] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in accordance with a determination” or “in response to detecting,” that a stated condition precedent is true, depending on the context. Similarly, the phrase “if it is determined [that a stated condition precedent is true]” or “if [a stated condition precedent is true]” or “when [a stated condition precedent is true]” may be construed to mean “upon determining” or “in response to determining” or “in accordance with a determination” or “upon detecting” or “in response to detecting” that the stated condition precedent is true, depending on the context.
Claims
Attorney Docket No. 097425-01493(P69319WO)What is claimed is:
1. A method comprising: at a device having a processor and one or more sensors: obtaining a sequence of portions of sensor data occurring over a period of time via the one or more sensors, each portion of the sensor data associated with a data timestamp within the period of time; obtaining an audio signal containing a sequence of audio segments occurring over the period of time via the one or more sensors, each audio segment associated with an audio timestamp within the period of time; matching a first data timestamp with a first audio timestamp; and performing a function using a first portion of the sensor data associated with the first data timestamp matched with the first audio timestamp.
2. The method of claim 1, wherein each data timestamp provides a common temporal position encoding for user activities corresponding to different input modalities.
3. The method of claim 1, further comprising: determining a triggering audio element of the sequence of audio segments.
4. The method of claim 3, wherein the triggering audio element enables the device to identify the first audio timestamp associated with the triggering audio element.
5. The method of claim 1, further comprising: determining the function by analyzing the sequence of audio segments occurring over the period of time.Attorney Docket No. 097425-01493(P69319WO)6. The method of claim 1, further comprising: matching a second data timestamp with a second audio timestamp, wherein said performing the function further uses a second portion of the sensor data associated with the second data timestamp matched with the second audio timestamp.
7. The method of claim 1, wherein the sequence of audio segments comprises an utterance from a user.
8. The method of claim 1, wherein the triggering audio element is based on an interpreted spoken command of the utterance.
9. The method of claim 1, wherein matching the first data timestamp with the first audio timestamp comprises using a machine learning model that tracks relationships in the sequence of tokens.
10. The method of claim 9, wherein the machine learning model comprises a transformer.Attorney Docket No. 097425-01493(P69319WO)11. An electronic device comprising: one or more sensors; a non-transitory computer-readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising: obtaining a sequence of portions of sensor data occurring over a period of time via the one or more sensors, each portion of the sensor data associated with a data timestamp within the period of time; obtaining an audio signal containing a sequence of audio segments occurring over the period of time via the one or more sensors, each audio segment associated with an audio timestamp within the period of time; matching a first data timestamp with a first audio timestamp; and performing a function using a first portion of the sensor data associated with the first data timestamp matched with the first audio timestamp.
12. The electronic device of claim 11, wherein each data timestamp provides a common temporal position encoding for user activities corresponding to different input modalities.
13. The electronic device of claim 11, wherein the program instructions, when executed on the one or more processors, further cause the electronic device to perform operations comprising: determining a triggering audio element of the sequence of audio segments.
14. The electronic device of claim 13, wherein the triggering audio element enables the device to identify the first audio timestamp associated with the triggering audio element.Attorney Docket No. 097425-01493(P69319WO)15. The electronic device of claim 11, wherein the program instructions, when executed on the one or more processors, further cause the electronic device to perform operations comprising: determining the function by analyzing the sequence of audio segments occurring over the period of time.
16. The electronic device of claim 11 , wherein the program instructions, when executed on the one or more processors, further cause the electronic device to perform operations comprising: matching a second data timestamp with a second audio timestamp, wherein said performing the function further uses a second portion of the sensor data associated with the second data timestamp matched with the second audio timestamp.
17. The electronic device of claim 11, wherein the sequence of audio segments comprises an utterance from a user.
18. The electronic device of claim 17, wherein the triggering audio element is based on an interpreted spoken command of the utterance.Attorney Docket No. 097425-01493(P69319WO)19. A non-transitory computer-readable storage medium storing program instructions executable via one or more processors to perform operations comprising: obtaining a sequence of portions of sensor data occurring over a period of time via the one or more sensors, each portion of the sensor data associated with a data timestamp within the period of time; obtaining an audio signal containing a sequence of audio segments occurring over the period of time via the one or more sensors, each audio segment associated with an audio timestamp within the period of time; matching a first data timestamp with a first audio timestamp; and performing a function using a first portion of the sensor data associated with the first data timestamp matched with the first audio timestamp.
20. The non-transitory computer-readable storage medium of claim 19, wherein each data timestamp provides a common temporal position encoding for user activities corresponding to different input modalities.
Citation Information
Patent Citations
System and method for continuous multimodal speech and gesture interaction
US20200150921A1
Time-based visual targeting for voice commands
US20200219501A1