Preserving engagement states based on context signals
By dynamically adjusting graphical user interface elements based on context signals, the system addresses user intent uncertainty in voice-enabled devices, enhancing interaction and reducing frustration by maintaining or expiring prompts accordingly.
Patent Information
- Application Number
- JP2024506978
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-06
- Filing Date
- 2022-07-18
- Publication Date
- 2025-09-01
- Estimated Expiration
- 2042-07-18
AI Technical Summary
Voice-enabled devices often fail to accurately determine user intent due to unreliable detection of user commands, leading to frustration when prompts or actions are missed or incorrectly timed out, especially in environments where users are engaged in multiple activities.
A system that analyzes context signals such as user proximity, presence, and attention to dynamically adjust the timeout duration of temporal graphical user interface elements, ensuring they remain visible when the user intends to interact and expiring them when disengaged.
Enhances user interaction by ensuring timely confirmation of actions and reducing frustration by accurately maintaining or expiring prompts based on user engagement, thereby improving the responsiveness and usability of voice-enabled devices.
Smart Images

Figure 0007732078000001 
Figure 0007732078000002 
Figure 0007732078000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to preserving engagement states based on context signals. [Background technology]
[0002] A voice-enabled environment (e.g., a home, a workplace, a school, an automobile, etc.) allows a user to speak queries or commands aloud to a computer-based system, which then fields and responds to the queries and / or performs functions based on the commands. A voice-enabled environment may be implemented using a network of connected microphone devices distributed throughout various rooms or areas of the environment. These devices may use hot words to help identify when a given utterance is directed to the system as opposed to speech directed to another individual present in the environment. Thus, a device may operate in a sleep or hibernation state and wake up only when detected speech includes the hot word. Once awake, the device can proceed to perform more expensive processing, such as fully on-device automated speech recognition (ASR) or server-based ASR. Summary of the Invention [Means for solving the problem]
[0003] One aspect of the present disclosure provides a computer-implemented method for dynamically changing graphical user interface elements. The computer-implemented method, when executed by data processing hardware, causes the data processing hardware to perform operations in response to detecting that a temporal user interface element is displayed on a user interface of a user device. The operations include receiving, at the user device, a context signal characterizing a user state. The operations further include determining, by the user device, that the context signal characterizing the user state indicates that the user intends to interact with the temporal user interface element. The operations also include modifying a state of each of the temporal user interface elements displayed on the user interface of the user device in response to determining that the context signal characterizing the user state indicates that the user intends to interact with the temporal user interface element.
[0004] Another aspect of the present disclosure provides a system for dynamically changing graphical user interface elements. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations are performed in response to detecting that temporal user interface elements are displayed on a user interface of a user device. The operations include receiving, at the user device, a context signal characterizing a user's state. The operations further include determining, by the user device, that the context signal characterizing the user's state indicates that the user intends to interact with the temporal user interface element. The operations also include modifying a state of each of the temporal user interface elements displayed on the user interface of the user device in response to determining that the context signal characterizing the user's state indicates that the user intends to interact with the temporal user interface element.
[0005] Implementations of any aspect of the present disclosure may include one or more of the following optional features: In some implementations, the user's state includes an engagement state indicating that the user is attempting to engage with or is engaged with a temporal user interface element displayed on a user interface of the user device. In some examples, modifying the respective states of the verbal user interface elements includes increasing or pausing a timeout duration of the temporal user interface element. In some configurations, the user's state includes a disengagement state indicating that the user has disengaged with a temporal user interface element displayed on a user interface of the user device. In these configurations, modifying the respective states of the temporal user interface elements in response to determining that the context signal characterizing the user's state includes a disengagement state may include removing the temporal user interface element before expiration of the timeout duration of the temporal user interface element. In these configurations, modifying the respective states of the temporal user interface elements in response to determining that the context signal characterizing the user's state includes a disengagement state may include reducing the timeout duration of the temporal user interface element. Some examples of context signals include a user proximity signal indicating a user's proximity to the user device, a presence detection signal indicating a user's presence within a field of view of a sensor associated with the user device, and an attention detection signal indicating a user's attention to the user device. The temporal user interface element may represent an action specified by a query detected in streaming audio captured by the user device.Optionally, the operations further include, in response to determining that the context signal characterizing the state of the user indicates that the user is attempting to interact with the temporal user interface element, determining that the state of each of the temporal user interface elements displayed on the user interface of the user device has previously failed to be modified a threshold number of times within the time period before modifying the state of each of the temporal user interface elements displayed on the user interface of the user device.
[0006] In some examples of any of the aspects of the present disclosure, the context signal includes a presence detection signal indicating a presence of a user within a field of view of a sensor associated with the user device, where determining that the context signal characterizing the state of the user indicates that the user is attempting to interact with the temporal user interface element includes determining that the presence detection signal indicates that the presence of the user within the field of view of the sensor has changed from absent to present, and modifying the state of each of the displayed temporal user interface elements on the user interface includes increasing a timeout duration or pausing a timeout duration of the temporal user interface element.
[0007] In some implementations of any of the aspects of the present disclosure, the context signal includes a user proximity signal indicating a proximity of a user to a user device, where determining that the context signal characterizing a state of the user indicates that the user is about to interact with the temporal user interface element includes determining that the user proximity signal indicates that the user's proximity to the user device has changed to be closer to the user device, and modifying the state of each of the displayed temporal user interface elements on the user interface includes increasing a timeout duration or pausing a timeout duration of the temporal user interface element.
[0008] In some configurations of any of the aspects of the present disclosure, the context signal includes an attention detection signal indicative of a user's attention to a user device, where determining that the context signal characterizing a state of the user indicates that the user is about to interact with a temporal user interface element includes determining that the attention detection signal indicates that the user's attention has changed to focus on the user device, and modifying the state of each of the displayed temporal user interface elements on the user interface includes increasing a timeout duration or pausing a timeout duration of the temporal user interface element.
[0009] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a schematic diagram of an exemplary voice-enabled environment. [Figure 2A] 2 is a schematic diagram of an exemplary repository for the voice-enabled environment of FIG. 1; [Figure 2B] 2 is a schematic diagram of an exemplary repository for the voice-enabled environment of FIG. 1; [Figure 3] 1 is a schematic diagram of an exemplary voice-enabled environment using a repository; [Figure 4] 1 is a flowchart of an exemplary arrangement of operations for a method of changing a state of a graphical user interface element. [Figure 5] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0011] Like reference symbols in the various drawings indicate like elements.
[0012] A voice-enabled device (e.g., a user device running a voice assistant) allows a user to speak queries or commands aloud, field responses to the queries, and / or perform functions based on the commands. Often, when a voice-enabled device responds to a query, the voice-enabled device generates a visual response representing the action requested by the query. For example, a user of the voice-enabled device may speak an utterance requesting that the voice-enabled device play a particular song on a music streaming application of the voice-enabled device. In response to this request by the user, the voice-enabled device displays a graphical element on a display associated with the voice-enabled device that indicates that the voice-enabled device is playing the requested song on a music streaming application associated with a music streaming service.
[0013] In some situations, the graphical element is temporary or temporal in nature, such that the graphical element is removed from the voice-enabled device's display after a certain amount of time. That is, the graphical element may have a configured timeout value, which refers to the duration the graphical element visually exists before being removed (i.e., times out). For example, the graphical element may be a temporary notification that visually indicates an action that the voice-enabled device is taking or can take in response to a query by the user.
[0014] In some examples, the voice-enabled device will not perform an action in response to a query until a user associated with the voice-enabled device somehow confirms that the user wants the action to be performed. Here, the confirmation may be a verbal confirmation, such as a voice input (e.g., "Yes"), a movement confirmation, such as a gesture, or a tactile confirmation, such as the user tapping the display of the voice-enabled device to affirm the performance of the action. When using a confirmation approach, the confirmation approach can prevent the voice-enabled device from inadvertently performing an action not desired by the user. Furthermore, this approach allows the voice-enabled device to suggest one or more actions that the voice-enabled device perceives as requested by the user and that the user should select, or to indicate whether any of the actions were actually requested by the user (i.e., any actions that the user intended the voice-enabled device to perform).
[0015] In some implementations, because the voice-enabled device lacks confidence that the user actually requested that an action be performed, the voice-enabled device suggests an action in response to a query by the user rather than automatically performing the action. For example, a user may be conversing with another user in proximity to the voice-enabled device, and the voice-enabled device may perceive that some aspect of the conversation was a query for the voice-enabled device to perform an action. Also, because this is a secondary voice conversation not directed at the voice-enabled device, the voice-enabled device cannot determine an action with a particular level of confidence for automatically performing the action. Here, due to the lower confidence in not automatically performing the action, the voice-enabled device instead displays a graphical element asking whether the user wants to perform the detected action. For example, the voice-enabled device displays a prompt stating, "Shall I play Michael Jackson's Thriller?" This prompt thus gives the user the ability to confirm that this is what the user intended (i.e., a selectable "yes" button in the prompt) or indicate that this was not what the user intended (i.e., a selectable "no" button in the prompt). Here, the prompt may be set up as a temporal graphical element that allows the user to ignore or fail to notice the prompt and not perform the action suggested by the prompt. In other words, if the voice-enabled device does not receive any input from the user regarding the prompt, the prompt, and therefore the suggested action, will time out and not be taken.
[0016] Unfortunately, situations arise in which a user requests that a particular action be performed by a voice-enabled device, but is unable to provide confirmation that the user wants the action to be performed. By way of example, the voice-enabled device may be in the kitchen of the user's home. As the user returns home from the grocery store, the user begins loading groceries into the kitchen in a series of trips from the car. Initially, while in the kitchen, the user requests that the voice-enabled device play a local radio station that the user was listening to on the way home from the grocery store. In this scenario, the user explicitly requested that the voice-enabled device play a local radio station, but the user may have been moving around, and as a result, the voice-enabled device did not reliably detect the command. Due to this lack of reliability, the voice-enabled device generates a prompt asking whether the user wants the voice-enabled device to play a local radio station. Even though this proposed action is indeed what the user requested, the user may miss the prompt and become frustrated that the voice-enabled device's voice capabilities are less than expected. Alternatively, the user may return from the trip to the car with arms full of groceries to find that the local radio station is not playing and that the prompt is merely present as a temporal graphical element on the voice-enabled device's display. Also, the prompt may disappear (i.e., time out) before the user has an opportunity to drop off the groceries or respond verbally. In this example, the user may have moved toward the voice-enabled device to tap yes on the prompt before the prompt / suggested action timed out.
[0017] In situations like the grocery example, the state of the temporal graphical element can benefit from being aware of the user's state. That is, in the grocery example, the user was moving toward the voice-enabled device when the prompt disappeared (i.e., timed out). That is, if some system associated with the voice-enabled device recognizes the user's state as attempting to interact with the voice-enabled device, that system may notify the voice-enabled device (or some system thereof) to extend the timeout duration of the temporal graphical element (e.g., a prompt asking if the user wanted to play a local radio station). Thus, if the timeout duration is extended due to the user's state, the user may have enough time to input confirmation of the action at the voice-enabled device, and as a result, the voice-enabled device will successfully play a local radio station to the user's enjoyment.
[0018] In contrast, in the grocery example, if the user walked into the kitchen with groceries and, while talking on their cell phone, said, "I just heard this great song playing on a local radio station," the temporal graphical element could still benefit from being aware of the user's state in order to accelerate its timeout or immediately time out (i.e., change the state of the temporal element). For example, in this alternative to the example, the user might gaze at the voice-enabled device, which notifies them that the voice-enabled device has generated a prompt to play a local radio station, and then walk away to continue bringing in groceries. Here, the combination of the user's visual recognition (i.e., gaze) and subsequent walking away can notify the voice-enabled device to immediately time out the prompt due to the user's implicit lack of interest in the proposed action.
[0019] In order for the temporal element to be aware of the user's state, one or more systems associated with the voice-enabled device may be configured to analyze the context signals to determine whether one or more of these signals indicate that the user's state should affect the state of the temporal element. In other words, if the context signals characterize the user's state as actively attempting to interact (e.g., attempting to engage) with the voice-enabled device (i.e., an engaged state), the state of the temporal element may be modified to accommodate the active interaction (e.g., a timeout may be extended or suspended entirely). On the other hand, if the context signals characterize the user's state as passively attempting to interact (e.g., attempting not to engage) with the voice-enabled device (i.e., a disengaged state), the state of the temporal element may be modified to accommodate the passive interaction (e.g., a timeout may be shortened or immediately executed). Using this approach, the temporal element may be preserved when the context signals (e.g., collected / received sensor data) warrant its preservation.
[0020] Referring to FIG. 1 , in some examples, audio environment 100 includes user 10 making utterance 20 within audible range of voice-enabled device 110 (also referred to as device 110 or user device 110) running digital assistant interface 120. Here, utterance 20 made by user 10 may be captured by device 110 in streaming audio 12 and may correspond to a query 22 for performing an action, or more specifically, a query 22 for digital assistant interface 120 to perform an action. User 10 may prefix query 22 with a hotword 24 (e.g., an invocation phrase) to trigger device 110 from a sleep or hibernation state when hotword 24 is detected in streaming audio 12 by a hotword detector (e.g., soft acceptor 200) operating on device 110 while it is in a sleep or hibernation state. User 10 may also endpoint query 22 with hotword 24. Other techniques may be used to trigger device 110 from a sleep or hibernation state other than speaking hotword 24. For example, an alert event may trigger device 110 from a sleep or hibernation state. Alert events may include, but are not limited to, a user making a gesture before or while speaking query 22, a user turning toward the device when speaking query 22, a user suddenly appearing in front of device 110 and speaking query 22, or a contextual cue indicating the likelihood that speech (e.g., query 22) is expected to be directed toward device 110. An action may also be referred to as a movement or a task. In this sense, user 10 may have a conversational interaction with digital assistant interface 120 running on voice-enabled device 110 to perform a computational activity or find an answer to a question.
[0021] Device 110 may correspond to any computing device associated with user 10 and capable of capturing audio from environment 100. In some examples, user device 110 includes, but is not limited to, mobile devices (e.g., mobile phones, tablets, laptops, e-book readers, etc.), computers, wearable devices (e.g., smart watches), music players, casting devices, smart appliances (e.g., smart televisions) and Internet of Things (IoT) devices, remote controls, smart speakers, etc. Device 110 includes data processing hardware 112 and memory hardware 114 in communication with data processing hardware 112 and storing instructions that, when executed by data processing hardware 112, cause data processing hardware 112 to perform one or more operations related to audio processing.
[0022] Device 110 further includes an audio subsystem 116 having an audio capture device 116, 116a (e.g., an array of one or more microphones) for capturing and converting audio in audio environment 100 into electronic signals (e.g., audio data 14). While device 110 implements audio capture device 116a (also commonly referred to as microphone 116a) in the illustrated example, audio capture device 116a may not be physically present on device 110 but may be in communication with audio subsystem 116 (e.g., peripheral to device 110). For example, device 110 may correspond to a vehicle infotainment system that utilizes an array of microphones located throughout the vehicle. In another example, audio capture device 116a may be present on a separate device in communication with user device 110, which is to perform an action. In addition, the audio subsystem 116 may include playback devices 116, 116b (e.g., one or more speakers 116b, etc.) for playing audio generated / output by the user device 110 (e.g., synthesized audio, synthesized speech, or audio associated with various types of media).
[0023] Device 110 may also include display 118 for displaying graphical user interface (GUI) elements (e.g., graphical user interface element 202) and / or graphical content. Some examples of GUI elements include windows, screens, icons, menus, etc. For example, device 110 may load or launch an application (local or remote) that generates GUI elements (e.g., GUI element 202) or other graphical content for display 118. Furthermore, the elements generated on display 118 may be selectable by user 10 and may also serve to provide some form of visual feedback for processing activities and / or operations occurring on device 110. For example, the elements represent actions that device 110 is performing or suggests performing in response to query 22 from user 10. Furthermore, because device 110 is a voice-enabled device 110, user 10 may interact with elements generated on display 118 using various voice commands as well as other types of commands (e.g., gesture commands or touch input commands). For example, display 118 may illustrate a menu of options for a particular application, and user 10 may use interface 120 to select an option by voice or other means of feedback (e.g., tactile input, movement / gesture input, etc.). When user 10 speaks to select a presented option, device 110 may be operating in a reduced speech recognition state or using a warm word model associated with device 110 to determine whether a particular phrase was spoken to select the presented option. As an example, the warm word model operates to detect a binary sound (e.g., "yes" or "no") during the time (e.g., during a timeout window) that options are being presented by device 110 (e.g., on display 118 of device 110).
[0024] A voice-enabled interface (e.g., a digital assistant interface) 120 may field queries 22 or commands conveyed in utterances 20 captured by device 110. Voice-enabled interface 120 (also referred to as interface 120 or assistant interface 120) generally facilitates receiving audio data 14 corresponding to utterances 20 and coordinating voice processing on audio data 14 or other activity resulting from utterances 20. Interface 120 may execute on data processing hardware 112 of device 110. Interface 120 may channel audio data 14, including utterances 20, to various systems related to voice processing or query fulfillment.
[0025] Additionally, device 110 is configured to communicate with remote system 140 via network 130. Remote system 140 may include scalable remote resources 142, such as remote data processing hardware 144 (e.g., a remote server or CPU) and / or remote memory hardware 146 (e.g., a remote database or other storage hardware). Device 110 may utilize remote resources 142 to perform various functions related to speech processing (e.g., by speech processing system 150) and / or GUI element state preservation (e.g., by saver 200). For example, device 110 is configured to perform speech recognition using speech recognition system 152 and / or speech interpretation using speech interpreter 154. In some examples, not shown, device 110 may additionally convert TTS during speech processing using a text-to-speech (TTS) system.
[0026] The device 110 is also configured to communicate with a speech processing system 150. The speech processing system 150 is generally capable of performing various functions related to speech processing, such as speech recognition and speech interpretation (also known as query interpretation). For example, the speech processing system 150 of FIG. 1 is shown to include a speech recognizer 152 that performs automated speech recognition (ASR), a speech interpreter 154 that determines the meaning of the recognized speech (i.e., to understand the speech), and a search engine 156 that retrieves any search results in response to a query specified in the recognized speech. When a hotword detector (e.g., associated with the assistant interface 120) detects a hotword event, the hotword detector passes the audio data 14 to the speech processing system 150. The hotword event indicates that the hotword detector accepts a portion of the audio data 14 (e.g., a first audio segment) as a hotword 24. When a portion of audio data 14 is identified as a hot word 24, the hot word detector and / or assistant interface 120 communicates the audio data 14 as a hot word event so that speech processing system 150 can perform speech processing on the audio data 14. By performing speech processing on the audio data 14, speech recognizer 152, in combination with speech interpreter 154, can determine whether a second audio segment of audio data 14 (e.g., designated as query 22) indicates a query-type utterance.
[0027] The speech recognizer 152 receives as input audio data 14 corresponding to hot word events and transcribes the audio data 14 into a transcription as output, referred to as a speech recognition result R. Generally, by converting the audio data 14 into a transcription, the speech recognizer 152 enables the device 110 to recognize when an utterance 20 made by the user 10 corresponds to a query 22 (or command) or some other form of audio communication. The transcription refers to a string of text that the device 110 (e.g., the assistant interface 120 or the voice processing system 150) may then use to generate a response to the query or command. The speech recognizer 152 and / or the interface 120 may provide the speech recognition result R to a speech interpreter 154 (e.g., a natural language understand (NLU) module) to perform a semantic interpretation on the speech recognition result R to determine whether the audio data 14 includes a query 22 that requests a particular action 158 to be performed. In other words, the speech interpreter 154 generates an interpretation I of the result R to identify a query 22 or command in the audio data 14 and to enable the speech processing system 150 to respond to the query 22 with a corresponding action 158 invoked by the query 22. For example, if the query 22 is a command to play music, the corresponding action 158 invoked by the query 22 is to play the music (e.g., by executing an application capable of playing music). In some examples, the speech processing system 150 uses the search engine 156 to retrieve search results that enable the speech processing system 150 to respond to the query 22 (i.e., fulfill the query 22).
[0028] The user device 110 also includes or is associated with a sensor system 160 configured with sensors 162 to capture sensor data 164 within the environment of the user device 110. The user device 110 may receive the sensor data 164 captured by the sensor system 160 continuously or at least during periodic intervals to determine a current state 16 of the user 10 of the user device 110. Some examples of the sensor data 164 include motion data, image data, connectivity data, noise data, audio data, or other data indicative of the state 16 of the user 10 / user device 110 or the state of the environment in the user device 110's vicinity. The motion data may include accelerometer data that characterizes the movement of the user 10 by movement of the user device 110. For example, when the user 10 is holding the device 110 and the user 10 moves their thumb to tap the display 118 of the device 110 (e.g., to interact with a GUI element), the motion data indicates that the state 16 of the user 10 is engaged by the device 110. Motion data may also be received at device 110 from another device associated with the user, such as a smartphone in the user's pocket or a smartwatch worn by the user. The image data may be used to detect the user's 10's attention, the user's 10's proximity to user device 110 (e.g., the distance between user 10 and user device 110), the user's 10's presence (e.g., whether the user 10 is present within the field of view of one or more sensors 162), and / or features of the user's 10's environment. For example, the image data detects the user's 10's attention by capturing features of the user 10 (e.g., to characterize the user's 10's gestures by body features, to characterize the user's 10's gaze by facial features, or to characterize the user's 10's posture / orientation). The connectivity / communication data may be used to determine whether the user device 110 is connected to or communicating with other electronic devices or devices (e.g., a smartwatch or mobile phone).For example, the connection data may be short-range connection / communication data, Bluetooth connection / communication data, Wi-Fi connection / communication data, or some other wireless band connection / communication data (e.g., ultra-wideband (UWB) data). Acoustic data, such as noise data or voice data, may be captured by the sensor system 160 (e.g., the microphone 116a of the device 110) and used to determine the environment of the user device 110 (e.g., a characteristic or property of the environment having a particular acoustic signature) or to identify whether the user 10 or another party is speaking. In some configurations, the sensor system 160 captures ultrasound data to detect the location of the user 10 or other objects within the environment of the device 110. For example, the device 110 utilizes a combination of its speaker 116b and its microphone 116a to capture ultrasound data about the environment of the device 110. The sensors 162 of sensor system 160 may be embedded or hosted on-device (e.g., a camera that captures image data or a microphone 116a that captures acoustic data), reside off-device but in communication with device 110, or some combination thereof. While FIG. 1 illustrates a camera included on device 110 as an exemplary sensor 162, other peripheral devices of device 110 may also function as sensors 162, such as microphone 116a and speaker 116b of audio subsystem 116.
[0029] Systems 150, 160, 200 may reside on device 110 (referred to as on-device systems) or may reside remotely (e.g., reside on remote system 140) but in communication with device 110. In some examples, some of these systems 150, 160, 200 reside locally or on-device, while others reside remotely. In other words, any of these systems 150, 160, 200 may be local or remote in any combination. For example, when the size or processing requirements of a system 150, 160, 200 are significant, the system 150, 160, 200 may reside on remote system 140. Also, when device 110 can support the size or processing requirements of one or more systems 150, 160, 200, one or more systems 150, 160, 200 may reside on device 110 using data processing hardware 112 and / or memory hardware 114. Optionally, one or more of systems 150, 160, 200 may exist both locally / on-device and remotely. For example, one or more of systems 150, 160, 200 may default to running on remote system 140 when a connection to network 130 between device 110 and remote system 140 is available, but when the connection is lost or network 130 is unavailable, systems 150, 160, 200 instead run locally on device 110.
[0030] The preserver 200 generally functions as a system for dynamically adapting the state S of GUI elements (also referred to as user interface (UI) elements) displayed on the user device 110 (e.g., a user interface such as the display 118 of the device 110). The state S of a GUI element broadly refers to the properties of the GUI element. That is, the state S of a GUI element may refer to the location of the GUI element, the amount of time the GUI element is presented (i.e., a timeout value), and / or the characteristics of the graphics associated with the GUI element (e.g., color, typeface, font, style, size, etc.). In this regard, altering the state S of a GUI element includes, for example, changing the size of the GUI element, changing the location of the GUI element within the display, changing the time the GUI element is presented, changing the font of the GUI element, changing the content of the GUI element (e.g., changing the text or media content of the GUI element), etc. While the preserver 200 is capable of adapting the state S of any GUI element, the examples herein more specifically illustrate the preserver 200 altering the state S of a temporal GUI element 202. A temporal GUI element 202 is a graphical element that is transient in nature. For example, the temporal GUI element 202 includes a timeout value that specifies the amount of time T that the temporal GUI element exists (eg, is displayed) before being removed or automatically closed.
[0031] Although the examples herein refer to temporal GUI elements displayed on the display 118, implementations herein are equally applicable to presenting temporal elements through non-graphical interfaces, such as flashing lights and / or audible output of sounds (e.g., beeps / chimes). A user may interact with this type of non-graphical interface element 202 by pressing physical buttons on the device, providing voice input, or performing gestures.
[0032] To identify whether and when the preserver 200 should change (i.e., dynamically adapt) the state S of the temporal GUI element 202 displayed on the user device 110, the preserver 200 is configured to receive or monitor context signals 204 accessible to the device 110. In some examples, sensor data 164 captured by the sensor system 160 serves as one or more context signals 204 characterizing aspects of the device's environment. The preserver 200 may use the sensor data 164 as the context signals 204 without any further processing of the sensor data 164, or may perform further processing on the sensor data 164 to generate context signals 204 characterizing aspects of the device's environment. The device 110 may utilize the preserver 200 to determine whether the context signals 204 indicate that the user 10 intends to interact with the temporal GUI element 202 before it disappears. That is, the context signal 204 may indicate a state 16 of the user 10, where the state 16 indicates an engaged state, in which the user 10 is attempting to interact with the temporal GUI element 202, or a disengaged state, in which the user 10 is not interacting with or intentionally disengaging from the temporal GUI element 202. Depending on the state 16 of the user 10 characterized by the context signal 204, the preserver 200 may either preserve the temporal GUI element 202 for a longer period of time than originally specified, hasten the disappearance (i.e., deletion) of the temporal GUI element 202, or maintain a time-based property (e.g., a timeout value) of the temporal GUI element 202 (i.e., not change any state S of the temporal GUI element 202).
[0033] Returning to the first grocery example, as user 10 moves toward device 110, preserver 200 may receive one or more context signals 204 that characterize state 16 of user 10 as an engagement state 16, 16e. For example, context signal 204 indicates that user 10's proximity to user device 110 is changing in a manner that brings user 10 closer to user device 110 (i.e., the distance between user 10 and user device 110 is decreasing). Because context signal 204 indicates that state 16 of user 10 is changing, this suggests that user 10 is about to interact with the prompt displayed on user device 110. Here, preserver 200 modifies state S of the prompt, "Would you like to play a local radio station?", which is temporal GUI element 202. In this example, preserver 200 would modify state S of the prompt by extending or pausing the prompt's timeout duration to allow user 10 to successfully interact with the prompt before it expires. The context signal 204 may indicate any action performed by the user that may characterize the state 16 of the user 10 as being in an engagement state 16e. As another example, the user 10 picking up a remote control may characterize the state of the user 10 as being in an engagement state 16 to interact with the temporal GUI element 202 displayed on the television.
[0034] In contrast, when user 10 is talking on the phone and does not instruct device 110 to play a local radio station, preserver 200 may receive one or more context signals 204 that characterize state 16 of user 10 as disengaged state 16d. For example, as described in this version of the example, user 10 gazes at device 110, which displays the prompt "Would you like to play a local radio station?", then turns around and moves on to continuing to bring in groceries. When these are user 10's actions, preserver 200 may receive context signals 204 that characterize user 10's attention as engaged with or gazing at device 110. These context signals 204 are then followed by context signals 204 that characterize disengagement with device 110 (i.e., looking away from device 110) without any action by user 10 indicating an attempt to engage with temporal GUI element 202. In practice, in that case, the context signal 204 received by the preserver 200 would indicate that the user 10 is decreasing proximity to the device 110, which indicates a disengagement state 16d. These changing context signals 204, which collectively indicate an engagement state 16e followed by a disengagement state 16d, may cause the preserver 200 to either maintain the state S of the temporal GUI element 202 (i.e., allow the prompt to expire), advance the state S of the temporal GUI element 202 (i.e., shorten the timeout duration or remaining timeout duration), or change the state S of the temporal GUI element 202 and immediately remove the temporal GUI element 202. In this example, the preserver 200 may interpret the context signal 204 as characterizing intentional disengagement due to a change from the engagement state 16e (e.g., a state 16 recognizing the temporal GUI element 202) to the disengagement state 16d, and may immediately remove the temporal GUI element 202 upon this interpretation.
[0035] 2A and 2B , the storer 200 includes a state determiner 210 and a modifier 220. The state determiner 210 is configured to receive one or more context signals 204 and determine a state 16 of the user 10 relative to the device 110. In some examples, the state determiner 210 receives sensor data 164 (e.g., raw sensor data) from the sensor system 160 and converts or processes the sensor data 164 into a context signal 204 that characterizes the state 16 of the user 10. For example, the state determiner 210 may receive the sensor data 164 and characterize a physical state of the user 10 relative to the device 110 (e.g., a position or orientation of the user 10) based on the sensor data 164 to form the context signal 204. Additionally or alternatively, the state determiner 210 may receive the sensor data 164 and characterize a non-physical state 16 of the user 10. For example, the state determiner 210 receives acoustic data 164 that characterizes the state 16 of the user 10. To illustrate, the state determiner 210 receives acoustic data 164 indicating that the user 10 was speaking and, in response to the device 110 displaying the temporal GUI element 202, paused or slowed down. In other words, the user 10 may have been engaged in a secondary voice conversation with another user and noticed that the device 110 displayed a prompt as the temporal GUI element 202. The user 10's notice of the prompt may have caused the user 10 to temporarily slow down in the conversation (i.e., pause for a moment to read or look at the prompt) and then ignored the prompt and continued the conversation. In this situation, a change in prosody in the acoustic data 164 may be a context signal 204 indicating the state 16 of the user 10. Here, in this example, the state determiner 210 may determine that the user 10 is in an engaged state 16e because the user 10 changed the prosody (e.g., rhythm) of their speech. Then, when the user 10 reverts to their original prosody, the state determiner 210 may determine that the user 10 is in a disengaged state 16d.
[0036] In some examples, the context signal 204 refers to a perceived change in the state 16 of the user 10. That is, the state determiner 210 identifies sensor data 164 from a first time instance, compares the sensor data 164 from the first time instance with the sensor data 164 from a second time instance, and determines whether the comparison of the sensor data 164 indicates a state change (e.g., a change in location or orientation) of the user 10. The perceived change in the state 16 of the user 10 may cause the state determiner 210 to determine whether the change in the state 16 indicates that the user 10 is attempting to interact with the device 110 in some way (i.e., engaging with the device 110 by interacting with the temporal GUI element 202) or is refraining from interacting with the device 110. When the change in the state 16 of the user 10 indicates that the user 10 is attempting to interact with the device 110, the user 10 is in an engaged state 16e. On the other hand, when the change in the state 16 of the user 10 indicates that the user 10 is refraining from interacting with the device 110, the user 10 is in a disengaged state 16d. In some configurations, the state determiner 210 may classify the state 16 of the user 10 at a higher granularity than engagement or disengagement. For example, the state determiner 210 may be configured to classify a type of user engagement (e.g., approaching, making a positive gesture, speaking a confirmation, etc.) or a type of user disengagement (e.g., moving away, making a negative gesture, speaking negatively, etc.).
[0037] Once the state determiner 210 identifies the state 16 of the user 10, the state determiner 210 passes the state 16 to the modifier 220. The modifier 220 is configured to modify the state S of the temporal GUI element 202 displayed on the device 110 based on the state 16 of the user 10. That is, the modifier 220 can enable the state S of the temporal GUI element 202 to adapt to the state 16 of the user 10 while the temporal GUI element is displayed on the device 110. Referring to FIG. 2A , the modifier 220 may change the state S of the temporal GUI element 202 from a first state S, S1 to a second state S, S2, or may determine that the state S of the temporal GUI element 202 should not change (e.g., remain in the first state S1) based on the state 16 of the user 10. The modifier 220 may be configured to modify the state S of the temporal GUI element 202 in different ways. In some examples, the modifier 220 modifies or changes the state S of the temporal GUI element 202 by increasing the time (or remaining time) that the temporal GUI element 202 is displayed. FIG. 2B illustrates that when the user state 16 is the engaged state 16e, the modifier 220 may increase the timeout value from a time T of 7 seconds to a time T of 15 seconds. In other examples, when the user state 16 is the engaged state 16e, the modifier 220 modifies the state S of the temporal GUI element 202 by pausing the timeout function of the temporal GUI element 202. That is, while the preserver 200 perceives that the user 10 is engaged with the device 110, the modifier 220 allows the temporal GUI element 202 to exist for an indefinite duration (or until the user state 16 changes to the unengaged state 16d).
[0038] 2B , the modifier 220 may modify the state S of the temporal GUI element 202 when the user state 16 is the disengagement state 16d. For example, when the user 10 is in the disengagement state 16d, the modifier 220 accelerates or quickens the remaining time until the temporal GUI element 202 times out. Here, FIG. 2B illustrates two situations in which this can occur. In a first scenario, when the user state 16 is in the disengagement state 16d, the modifier 220 responds by expiring the temporal GUI element 202 based on the user 10's lack of interest in engaging with the temporal GUI element 202. This can occur when the user state 16 changes from the engagement state 16e to the disengagement state 16d while the temporal GUI element 202 is displayed, much like the second grocery example in which the user 10 looks at the device 110 and then walks away. 2B illustrates this first scenario by illustrating a timeout period T of 7 seconds in a first state S1 changing to a timeout period T of 0 seconds in a second state S2. In a second scenario, when the user state 16 is the disengaged state 16d, the modifier 220 responds by advancing the time T until the temporal GUI element 202 expires. For example, the modifier 220 changes the timeout period T from 7 seconds to 3 seconds.
[0039] 2A , the storer 200 further includes a state monitor 230. The state monitor 230 may be configured to monitor and / or adjust the number of times that the state 16 of the user 10 affects the state S of the temporal GUI element 202. To illustrate why this may be advantageous, the user 10 in the grocery example may finish bringing the groceries in from the car and begin putting them away. During this time, the user 10 may be moving around in the kitchen, and as a result, the state determiner 210 perceives the user 10 as engaging and disengaging with the device 110 while the device 110 is displaying the temporal GUI element 202. This engagement and disengagement may cause the modifier 220 to change the state S of the temporal GUI element 202. For example, the modifier 220 may increase and then decrease the timeout duration in some iterative manner that mirrors the user's actions. Unfortunately, if this happens several times, it likely indicates that the user 10 has already selected or interacted with the temporal GUI element 202 and that the temporal GUI element 202 should not have been dynamically changed, but rather expired or left expired. To prevent state changes from going back and forth or too many state changes from occurring, the monitor 230 may be configured with a state change threshold 232. In some examples, when the number of state changes of the temporal GUI element 202 meets the state change threshold 232, the monitor 230 deactivates the modifier 220 for the temporal GUI element 202, or in some cases, allows / forces the temporal GUI element 202 to expire (e.g., time out). When the monitor 230 is operating, if the number of state changes of the temporal GUI element 202 does not meet the state change threshold 232, the monitor 230 does not prevent the modifier 220 from modifying the temporal GUI element 202 (e.g., the modifier 220 continues to operate).
[0040] In some configurations, the monitor 230 may adjust one or more thresholds associated with the state determiner 210. For example, the state determiner 210 may identify a state change of the user 10 based on the user 10 moving from a far location to a near location relative to the device 110. Here, the far field and near field may be defined by a threshold (e.g., shown in FIG. 3 as threshold 212) that establishes a proximity boundary between the near field and the far field. In other words, if the user 10 moves toward the device 110 and crosses the threshold, the user is entering the near field (e.g., leaving the far field), but if the user 10 moves away from the device 110 and crosses the threshold, the user is entering the far field (e.g., leaving the near field). Also, there are situations in which the user 10 may be moving around this boundary. To illustrate, the user 10 may be putting groceries away from a kitchen island into a refrigerator, and the boundary may be located between the kitchen island and the refrigerator. Because user 10 may be constantly moving back and forth on this boundary, monitor 230 may recognize this activity in the threshold and act as a kind of debouncer. That is, user 10 may only move a few feet, and instead of allowing state determiner 210 to identify this movement as changing between engaged state 16e and disengaged state 16d, monitor 230 adjusts the threshold to stabilize the state change sensitivity of state determiner 210. In other words, user 10 would simply have to move further toward or away from device 110 than the island to trigger a state 16 change.
[0041] 3 illustrates a scenario illustrating several types of context signals 204 that can characterize a state 16 of a user 10 (e.g., the user's physical state 16) at a first time instance T0 and a second time instance T1 that occurs after the first time instance T0. In this scenario, the user 10 walks into a room with the device 110 while talking on his or her mobile phone and says, "I just heard this great song playing on local radio station 94.5." In response to this utterance 20 by the user 10, the device 110 generates a prompt as a temporal GUI element 202 asking whether the user 10 wants to play local radio station 94.5. Here, the storer 200 receives sensor data 164 corresponding to image data at the first time instance T0 that illustrates the location of the user 10 relative to the user device 110. Using the image data, the state determiner 210 estimates distances D, D1 between the user 10 and the device 110 at the first time instance T0 based on the received image data to form a context signal 204 characterizing a state 16 of the user 10 with respect to the device 110. Here, the context signal 204 indicating the user's proximity to the device 110 is considered a user proximity signal. At the first time instance T0, the state determiner 210 determines that the estimated distances D, D1 mean that the user 10 is in a near field close to the device 110 within a near-far boundary 212. For example, the state determiner 210 determines that the estimated distance D1 has a distance D to the device 110 that is less than the distance from the device 110 to the boundary 212. Based on the user's proximity to the device 110, the state determiner 210 determines that the user 10 is in an engaged state 16e at the first time instance T0.Because the user proximity signal indicates that the user 10's proximity to the device 110 has changed to be closer to the device 110 (e.g., the user 10 walks toward the device 110 in FIG. 3 to be at a first estimated distance D1), the modifier 220 may then modify the state S of the temporal GUI element 202, for example, by increasing the timeout duration or pausing the timeout duration of the temporal GUI element 202.
[0042] In particular, while the prompt of temporal GUI element 202 is displayed, device 110 may begin fulfilling the perceived command to play local radio station 94.5 by initiating a connection with and streaming audio from station 94.5 (or a music streaming service capable of streaming station 94.5). However, device 110 may not begin audibly outputting the streaming audio until user 10 affirmatively provides a user input indication indicating selection of temporal GUI element 202 to stream the audio. Thus, in this example, because user 10 never directed their voice to device 110 to stream audio from station 94.5, device 110 would terminate the streaming connection in response to temporal GUI element 202 timing out (or the user affirmatively providing an input indication selecting “No”) without ever audibly outputting the streaming audio.
[0043] Additionally, while the examples herein are directed to the user device 110 generating the prompt as a temporal GUI element 202, the user device 110 may similarly unobtrusively output visual (e.g., a flashing light) and / or audible (e.g., a beep) prompts that allow the user to confirm or deny implementation of a recognized voice command within a time timeout period. For example, the user 10 may simply say "yes" or "no." In this scenario, the user device 110 may activate a warm word model that listens for binary terms (e.g., "yes" and "no") in the audio.
[0044] In addition to enabling the state determiner 210 to generate a user proximity signal for the context signal 204, the sensor data 164 in this scenario also enables the state determiner 210 to generate a context signal 204 referred to as an attention detection signal. The attention detection signal refers to a signal that characterizes whether the user 10 is paying attention to the device 110. The attention detection signal may indicate whether the user 10 is in an engaged state 16e or a disengaged state 16d. In other words, the attention detection signal (e.g., based on the sensor data 164) may indicate that the user 10 has shifted attention either toward or away from the device 110. When the user 10's attention shifts toward the device 110 while the device 110 is displaying the temporal GUI element 202, the state determiner 210 may determine that the user 10 is in the engaged state 16e. Some examples of the user 10 paying attention to the device 110 include the user's 10 gaze directed toward the device 110, the user's 10 gesture toward the device 110, or the user's 10 posture / orientation toward the device 110 (i.e., facing toward the device 110). Referring to FIG. 3 , image data captures the user's 10 gaze (e.g., shown as a dotted line of sight from the user's 10 face) directed toward the device 110. When this attention detection signal 204 indicates that the user 10 is focusing their attention on the device 110, the state determiner 210 determines that the attention detection signal 204 characterizes the user 10 as being in an engaged state 16e at the first time instance T0. Because the attention detection signal alone indicates that the user's 10's attention is focused on the device 110, the modifier 220 may modify the state S of the temporal GUI element 202 by, for example, increasing the timeout duration or pausing the timeout duration of the temporal GUI element 202.In some examples, such as this scenario, when multiple context signals 204 are available, the preserver 200 may utilize any or all of the context signals 204 to determine whether the state 16 of the user 10 should affect (e.g., modify) the state S of the temporal GUI element 202.
[0045] 3 also shows that the user 10 moves from a first location at a first distance D1 from the device 110 to a second location at a second distance D2 from the device 110 at a second time instance T1. At the second time instance T1, the context signal 204 indicates that the user 10 is in a disengaged state 16d in which the user 10 is not attempting to interact with the temporal GUI element 202. For example, at the second time instance T1, the user proximity signal indicates that the user 10 relative to the device 110 has changed to move away from the device 110 and is now in the far field beyond the boundary 212. Furthermore, at the second time instance T1, the attention detection signal indicates that the user 10 has shifted their attention away from the device 110. Because both context signals 204 indicate that user 10 is in disengaged state 16d, modifier 220 may modify state S of temporal GUI element 202 (e.g., by advancing or expiring temporal GUI element 202), or may maintain state S of temporal GUI element 202 and expire temporal GUI element 202 accordingly.
[0046] Although not shown, it may be the case that the user 10 is not in the far field at the second time instance T1, but has left the room entirely. In this situation, the state determiner 210 may generate a context signal 204, referred to as a presence detection signal. A presence detection signal is a signal that characterizes whether the user 10 is present within a particular field of view of one or more sensors 162. In other words, the presence detection signal may function as a binary decision of whether image data or some other form of sensor data 164 indicates that the user 10 is present within the field of view of the sensors 162 of the device 110. The presence detection signal may indicate a user state, because a change in the presence detection signal may indicate whether the user 10 is performing an action that characterizes the user 10 being in an engaged state 16e or a disengaged state 16d. For example, if the user 10 is not present in the field of view of the device 110 (e.g., of the sensor 162 associated with the device 110) and then comes into the field of view of the device 110, and this change occurs while the temporal GUI element 202 is displayed on the display 118 of the user device 110, the state determiner 210 can interpret the change in the user's presence as indicating that the user is about to interact with the temporal GUI element 202. That is, the user 10 has become present to engage with the temporal GUI element 202, and as a result, the state 16 of the user 10 is an engaged state. In the opposite situation, if the user 10 is first present within the field of view of the device 110 (e.g., of the sensor 162 associated with the device 110) and then is no longer present within the field of view of the device 110, and this change occurs while the temporal GUI element 202 is being displayed on the display 118 of the user device 110, the state determiner 210 can interpret the change in the user's presence as indicating that the user is not interested in interacting with the temporal GUI element 202 (e.g., is actively disengaged), and as a result, the state 16 of the user 10 is disengaged state 16d.When the presence detection signal indicates that the presence of user 10 within the field of view of sensor 162 has changed from absent to present, modifier 220 may modify the state S of temporal GUI element 202, for example, by increasing the timeout duration or pausing the timeout duration of temporal GUI element 202. In contrast, when the presence detection signal indicates that the presence of user 10 within the field of view of sensor 162 has changed from present to absent, modifier 220 may modify the state S of temporal GUI element 202, for example, by decreasing the timeout duration, causing temporal GUI element 202 to expire immediately, or maintaining the timeout duration of temporal GUI element 202 and allowing temporal GUI element 202 to expire accordingly.
[0047] 4 is a flowchart of an example arrangement of operations for a method 400 of changing the state of a graphical user interface element 202. The method 400 performs operations 402-406 in response to detecting that a temporal GUI element 202 is displayed on the user interface of the user device 110. At operation 402, the method 400 receives, at the user device 110, a context signal 204 characterizing a state 16 of the user 10. At operation 404, the method 400 determines, by the user device 110, that the context signal 204 characterizing the state 16 of the user 10 indicates that the user intends to interact with the temporal GUI element 202. In response to determining that the context signal 204 characterizing the state 16 of the user 10 indicates that the user intends to interact with the temporal GUI element 202, at operation 406, the method 400 modifies the state S of each of the temporal GUI elements 202 displayed on the user interface of the user device 110.
[0048] 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems (e.g., systems 150, 160, 200) and methods (e.g., method 400) described herein. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functionality are intended to be exemplary only and are not intended to limit the implementation of the examples described and / or claimed herein.
[0049] Computing device 500 includes a processor 510 (e.g., data processing hardware 112, 144), a memory 520 (e.g., memory hardware 114, 146), a storage device 530, a high-speed interface / controller 540 that connects to memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 that connects to a low-speed bus 570 and storage device 530. Each of components 510, 520, 530, 540, 550, and 560 are interconnected using various buses and may be mounted on a common motherboard or in other manners as appropriate. Processor 510 can process instructions for execution within computing device 500, including instructions stored in memory 520 or on storage device 530, for displaying graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 580 coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as appropriate. Also, multiple computing devices 500 may be connected, each providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).
[0050] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0051] The storage device 530 is capable of providing mass storage for the computing device 500. In some implementations, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 may be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or devices in a storage area network or other configuration. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 520, the storage device 530, or memory on the processor 510.
[0052] The high-speed controller 540 manages bandwidth-intensive operations for the computing device 500, while the low-speed controller 560 manages less bandwidth-intensive operations. Such allocation of duties is merely exemplary. In some implementations, the high-speed controller 540 is coupled to the memory 520, the display 580 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 550 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is coupled to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled, for example, through a network adapter, to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router.
[0053] Computing device 500, as shown, may be implemented in a number of different forms. For example, computing device 500 may be implemented as a standard server 500a, or multiple times in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0054] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuitry, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0055] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0056] The processes and logic flows described herein may be implemented by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows may also be implemented by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, e.g., magnetic disks, magneto-optical disks, or optical disks, for storing data, or is operably coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0057] To enable interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to enable interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0058] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0059] 10 users 12. Streaming Audio 14 Audio Data 16 State, engagement state, user state, physical state 16d Non-participation 16e Involvement Status 20 utterances 22 queries 24 Hot Words 100 Audio Environment, Environment, System 110 Voice-enabled devices, devices, user devices 112 Data Processing Hardware 114 Memory Hardware 116, 116a Audio capture device 116, 116b playback devices 116 Audio Subsystem 116a Audio capture device, microphone 116b Speaker 118 Display 120 Digital Assistant Interface, Interface, Voice-Enabled Interface, Assistant Interface 130 Network 140 Remote Systems 142 Remote Resources 144 Data Processing Hardware 146 Memory Hardware 150 Voice processing system, system 152 Voice Recognition System 154 Voice Interpreter 156 search engines 158 Action 160 Sensor Systems, Systems 162 sensors 164 sensor data, acoustic data 200 Soft acceptor, storage, system 202 Graphical User Interface Elements, GUI Elements, Non-Graphical Interface Elements, Temporal GUI Elements 204 Context Signals, Attention Detection Signals 210 State Decider 212 Threshold, Near-Far Boundary, Boundary 220 Modifier 230 Status monitor, monitor 232 State Change Threshold 400 Method, Computer-Implemented Method 500 computing devices 500a Standard Server 500b laptop computer 500c Rack Server System 510 Processors, Components, and Data Processing Hardware 520 Memory, Non-Transient Memory, Components, Memory Hardware 530 Storage devices, components 540 High-Speed Interface / Controller, High-Speed Interface, High-Speed Controller, Components 550 High-Speed Expansion Port, Components 560 Low-Speed Interface / Controller, Low-Speed Controller, Component 570 Slow Bus 580 Display 590 Low-Speed Expansion Port D distance D1 First distance D2 Second distance R Speech recognition results, results S state S1 First state S2 Second state T time T0 First time instance T1 Second time instance
Claims
1. A computer-implemented method (400) that, when executed by data processing hardware (510), causes the data processing hardware (510) to perform an operation, the operation comprising: displaying, for a timeout duration, a temporal user interface element on a user interface of a user device, the temporal user interface element including a prompt urging a user to initiate performance of a perceived command detected in streaming audio captured by the user device by providing a user input indication indicating selection of the temporal user interface element while the temporal user interface element is displayed for the timeout duration; In response to detecting that the temporal user interface element (202) is displayed on the user interface of the user device (110), receiving, at the user device (110), a context signal (204) characterizing a state (16) of a user (10); determining, by the user device (110), that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) intends to interact with the temporal user interface element (202); modifying the timeout duration of the temporal user interface element (202) displayed on the user interface of the user device (110) in response to determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) intends to interact with the temporal user interface element (202); and The method (400) includes:
2. 2. The method of claim 1, wherein the state of the user includes an engagement state indicating that the user is attempting to engage with or is engaged with the temporal user interface element displayed on the user interface of the user device.
3. 3. The method of claim 2, wherein modifying the timeout duration of the temporal user interface element comprises increasing the timeout duration of the temporal user interface element.
4. 3. The method of claim 2, wherein modifying the timeout duration of the temporal user interface element comprises pausing the timeout duration of the temporal user interface element.
5. 5. The method of claim 1, wherein the state of the user includes a disengagement state indicating that the user is no longer engaged with the temporal user interface element displayed on the user interface of the user device.
6. 6. The method of claim 5, wherein modifying the timeout duration of the temporal user interface element in response to determining that the context signal characterizing the state of the user includes the disengagement state comprises removing the temporal user interface element before expiration of the timeout duration of the temporal user interface element.
7. 6. The method of claim 5, wherein modifying the timeout duration of the temporal user interface element in response to determining that the context signal characterizing the state of the user includes the disengagement state comprises decreasing the timeout duration of the temporal user interface element.
8. The method of claim 1, wherein the context signal comprises a user proximity signal indicative of the proximity of the user to the user device.
9. The method of claim 1 , wherein the context signal comprises a presence detection signal indicative of the presence of the user within a field of view of a sensor associated with the user device.
10. The method of claim 1 , wherein the context signal comprises an attention detection signal indicative of the user's attention to the user device.
11. the context signal (204) comprises a presence detection signal indicative of the presence of the user (10) within a field of view of a sensor (162) associated with the user device (110); determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) intends to interact with the temporal user interface element (202) comprises determining that the presence detection signal (162) indicates that the presence of the user (10) within the field of view of the sensor (162) has changed from absence to presence; modifying the timeout duration of the temporal user interface element (202) displayed on the user interface includes increasing the timeout duration or pausing the timeout duration of the temporal user interface element (202).
10. The method (400) of claim 1.
12. the context signal (204) includes a user proximity signal indicative of the proximity of the user (10) to the user device (110); determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) is about to interact with the temporal user interface element (202) includes determining that the user proximity signal (204) indicates that the proximity of the user (10) to the user device (110) has changed to be closer to the user device (110); modifying the timeout duration of the temporal user interface element (202) displayed on the user interface includes increasing the timeout duration or pausing the timeout duration of the temporal user interface element (202).
10. The method (400) of claim 1.
13. the context signal (204) includes an attention detection signal indicative of the user's attention to the user device (110); determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) is about to interact with the temporal user interface element (202) includes determining that the attention detection signal indicates that the attention of the user (10) has changed to focus on the user device (110); modifying the timeout duration of the temporal user interface element (202) displayed on the user interface includes increasing the timeout duration or pausing the timeout duration of the temporal user interface element (202).
10. The method (400) of claim 1.
14. A computer-implemented method (400) that, when executed by data processing hardware (510), causes the data processing hardware (510) to perform an operation, the operation comprising: In response to detecting that the temporal user interface element (202) is displayed on a user interface of the user device (110), receiving, at the user device (110), a context signal (204) characterizing a state (16) of a user (10); determining, by the user device (110), that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) intends to interact with the temporal user interface element (202); determining that the state of each of the temporal user interface elements has previously failed to be modified a threshold number of times within a time period in response to determining that the context signal characterizing the state of the user indicates that the user is attempting to interact with the temporal user interface elements; modifying the respective states of the temporal user interface elements (202) displayed on the user interface of the user device (110); The method (400) includes:
15. 15. The method (400) of claim 14, wherein the temporal user interface element (202) represents an action specified by a query (22) detected in streaming audio captured by the user device (110).
16. A system (100), comprising: data processing hardware (510); and memory hardware (520) in communication with the data processing hardware (510), the memory hardware (520) storing instructions that, when executed on the data processing hardware (510), cause the data processing hardware (510) to perform operations, the operations comprising: displaying, for a timeout duration, a temporal user interface element on a user interface of a user device, the temporal user interface element including a prompt urging a user to initiate performance of a perceived command detected in streaming audio captured by the user device by providing a user input indication indicating selection of the temporal user interface element while the temporal user interface element is displayed for the timeout duration; In response to detecting that the temporal user interface element (202) is displayed on the user interface of the user device (110), receiving, at the user device (110), a context signal (204) characterizing a state (16) of a user (10); determining, by the user device (110), that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) intends to interact with the temporal user interface element (202); modifying the timeout duration of the temporal user interface element (202) displayed on the user interface of the user device (110) in response to determining that the context signal (204) characterizing the state (16) of the user (10) indicates an attempt by the user (10) to interact with the temporal user interface element (202); A system (100) comprising:
17. 17. The system of claim 16, wherein the state of the user includes an engagement state indicating that the user is attempting to engage with or is engaged with the temporal user interface element displayed on the user interface of the user device.
18. 20. The system of claim 17, wherein modifying the timeout duration of the temporal user interface element comprises increasing the timeout duration of the temporal user interface element.
19. 20. The system of claim 17, wherein modifying the timeout duration of the temporal user interface element comprises pausing the timeout duration of the temporal user interface element.
20. 20. The system (100) of any one of claims 16 to 19, wherein the state (16) of the user (10) includes a disengagement state (16d) indicating that the user (10) is no longer engaged with the temporal user interface element (202) displayed on the user interface of the user device (110).
21. 21. The system of claim 20, wherein modifying the timeout duration of the temporal user interface element in response to determining that the context signal characterizing the state of the user includes the disengagement state comprises removing the temporal user interface element before expiration of the timeout duration of the temporal user interface element.
22. 21. The system of claim 20, wherein modifying the timeout duration of the temporal user interface element in response to determining that the context signal characterizing the state of the user includes the disengagement state comprises decreasing the timeout duration of the temporal user interface element.
23. 17. The system (100) of claim 16, wherein the context signal (204) comprises a user proximity signal indicative of the proximity of the user (10) to the user device (110).
24. 17. The system of claim 16, wherein the context signal comprises a presence detection signal indicative of the presence of the user within a field of view of a sensor associated with the user device.
25. 17. The system (100) of claim 16, wherein the context signal (204) comprises an attention detection signal indicative of the user's (10) attention to the user device (110).
26. the context signal (204) comprises a presence detection signal indicative of the presence of the user (10) within a field of view of a sensor (162) associated with the user device (110); determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) intends to interact with the temporal user interface element (202) comprises determining that the presence detection signal (162) indicates that the presence of the user (10) within the field of view of the sensor (162) has changed from absence to presence; modifying the timeout duration of the temporal user interface element (202) displayed on the user interface includes increasing the timeout duration or pausing the timeout duration of the temporal user interface element (202).
17. The system (100) of claim 16.
27. the context signal (204) includes a user proximity signal indicative of the proximity of the user (10) to the user device (110); determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) is about to interact with the temporal user interface element (202) includes determining that the user proximity signal (204) indicates that the proximity of the user (10) to the user device (110) has changed to be closer to the user device (110); modifying the timeout duration of the temporal user interface element (202) displayed on the user interface includes increasing the timeout duration or pausing the timeout duration of the temporal user interface element (202).
17. The system (100) of claim 16.
28. the context signal (204) includes an attention detection signal indicative of the user's attention to the user device (110); determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) is about to interact with the temporal user interface element (202) includes determining that the attention detection signal indicates that the attention of the user (10) has changed to focus on the user device (110); modifying the timeout duration of the temporal user interface element (202) displayed on the user interface includes increasing the timeout duration or pausing the timeout duration of the temporal user interface element (202).
17. The system (100) of claim 16.
29. A system (100), comprising: data processing hardware (510); and memory hardware (520) in communication with the data processing hardware (510), the memory hardware (520) storing instructions that, when executed on the data processing hardware (510), cause the data processing hardware (510) to perform operations, the operations comprising: In response to detecting that the temporal user interface element (202) is displayed on a user interface of the user device (110), receiving, at the user device (110), a context signal (204) characterizing a state (16) of a user (10); determining, by the user device (110), that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) is attempting to interact with the temporal user interface elements (202); and, in response to determining that the context signal (204) characterizing the state (16) of the user (10) indicates that the user (10) is attempting to interact with the temporal user interface elements (202), determining that the state of each of the temporal user interface elements (202) has previously failed to be modified a threshold number of times within a time period; modifying the respective states of the temporal user interface elements (202) displayed on the user interface of the user device (110); A system (100) comprising:
30. 30. The system of claim 29, wherein the temporal user interface element represents an action specified by a query detected in streaming audio captured by the user device.
Citation Information
Patent Citations
Device and method for multimodal interface
JP1999249773A
Digital Assistant Experience based on Presence Detection
US20170289766A1
System and method for adjusting an idle time of a hardware device based on a pattern of user activity that indicates a period of time that the user is not in a predetermined area
US8601301B1