Contextual triggering of assistive features
The method addresses safety concerns by dynamically switching presentation modes based on user state and environment, ensuring visual awareness and safety through sensor-driven mode transitions.
Patent Information
- Application Number
- JP2024504847
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-26
- Filing Date
- 2022-07-06
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2042-07-06
AI Technical Summary
Users interacting with computing devices through digital assistants may not be visually aware of their surroundings due to media content occupying their vision, posing potential safety risks, especially in environments requiring visual awareness.
A computer-implemented method that utilizes sensor data to determine a user's current state and provides selectable options to switch between visual and audio presentation modes, allowing users to maintain awareness of their surroundings while consuming content.
Enhances user safety by enabling seamless transitions between presentation modes based on the user's activity and environment, ensuring visual awareness and reducing potential hazards.
Smart Images

Figure 0007733805000001 
Figure 0007733805000002 
Figure 0007733805000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to contextual triggering of auxiliary functions. [Background technology]
[0002] Users often interact with computing devices, such as smartphones, smartwatches, and smart speakers, through digital assistant interfaces. These digital assistant interfaces allow users to consume media content on various applications accessible by the computing device. When a user of a computing device consumes media content, the media content often occupies certain aspects of the user's senses. For example, when a user is reading a news article, the act of reading the news article occupies the user's vision. As a result, when the computing device occupies the user's vision, the user may not be visually aware of other activity occurring around the user. This can be problematic in situations where activity around the user requires visual awareness, for example, to prevent potential harm to the user and / or the computing device. For example, when a user is reading a news article while walking, the user may not be aware of an upcoming collision with another person approaching the user. As assistant interfaces become increasingly integrated with these various applications and operating systems running on the computing device, digital assistants can be leveraged to influence how media content is presented to the user to aid in the user's awareness of the computing device. Summary of the Invention [Means for solving the problem]
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware of a user device, causes the data processing hardware to perform operations for triggering an assistance function on the user device, including obtaining a current state of a user of the user device while the user device is using a first presentation mode to present content to the user of the user device. The operations also include providing as output from a user interface of the user device, based on the user's current state, a user-selectable option that, when selected, causes the user device to use a second presentation mode to present content to the user. The operations further include, in response to receiving a user input indication indicating selection of the user-selectable option, starting presentation of the content using the second presentation mode.
[0004] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, the operations further include receiving sensor data captured by the user device, and obtaining the current state of the user is based on the sensor data. In these implementations, the sensor data may include at least one of global positioning data, image data, noise data, accelerometer data, connection data indicating that the user device is connected to another device, or noise / voice data.
[0005] In some examples, the user's current state indicates one or more current activities the user is performing, where the user's current activity may include at least one of walking, driving, commuting, talking, or reading. In some implementations, providing user-selectable options as output from the user interface is further based on a current location of the user device. Additionally or alternatively, providing user-selectable options as output from the user interface is based on a type of content and / or a software application running on the user device that is providing the content.
[0006] In some examples, the first presentation mode includes one of a visual-based presentation mode or an audio-based presentation mode, and the second presentation mode includes the other of the visual-based presentation mode or the audio-based presentation mode. In some implementations, the operations further include, after initiating presentation of the content using the second presentation mode, concomitantly presenting the content using the second presentation mode while ending presentation of the content using the first presentation mode. Alternatively, the operations further include, after initiating presentation of the content using the second presentation mode, presenting the content using both the first presentation mode and the second presentation mode in parallel.
[0007] In some implementations, providing the user-selectable option as output from the user interface includes displaying the user-selectable option as a graphical element on a screen of the user device via the user interface, where the graphical element informs the user that a second presentation mode is available for presenting the content. In these implementations, receiving a user input indication includes one of receiving touch input on the screen that selects the displayed graphical element, receiving stylus input on the screen that selects the displayed graphical element, receiving gesture input indicating selection of the displayed graphical element, receiving gaze input indicating selection of the displayed graphical element, or receiving voice input indicating selection of the displayed graphical element.
[0008] In some examples, providing the user-selectable option as an output from the user interface includes providing the user-selectable option as an audible output from a speaker in communication with the user device via the user interface, where the audible output informs the user that a second presentation mode is available for presenting the content. In some implementations, receiving a user input indication indicating a selection of the user-selectable option includes receiving a voice input from the user indicating a user command to select the user-selectable option. In these implementations, these operations may further include activating a microphone to capture the voice input from the user in response to providing the user-selectable option as an output from the user interface.
[0009] Another aspect of the present disclosure provides a system for triggering an assistance function on a user device. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a current state of a user of the user device while the user device is using a first presentation mode to present content to the user of the user device. The operations also include providing, based on the current state of the user, as output from a user interface of the user device, a user-selectable option that, when selected, causes the user device to use a second presentation mode to present the content to the user. The operations further include, in response to receiving a user input indication indicating selection of the user-selectable option, starting presentation of the content using the second presentation mode.
[0010] This aspect may include one or more of the following optional features: In some implementations, the operations further include receiving sensor data captured by the user device, and wherein obtaining the current state of the user is based on the sensor data. In these implementations, the sensor data may include at least one of global positioning data, image data, noise data, accelerometer data, connection data indicating that the user device is connected to another device, or noise / voice data.
[0011] In some examples, the user's current state indicates one or more current activities the user is performing, where the user's current activity may include at least one of walking, driving, commuting, talking, or reading. In some implementations, providing the user-selectable options as output from the user interface is further based on a current location of the user device. Additionally or alternatively, providing the user-selectable options as output from the user interface is based on a type of content and / or a software application running on the user device that is providing the content.
[0012] In some examples, the first presentation mode includes one of a visual-based presentation mode or an audio-based presentation mode, and the second presentation mode includes the other of the visual-based presentation mode or the audio-based presentation mode. In some implementations, the operations further include, after initiating presentation of the content using the second presentation mode, concomitantly presenting the content using the second presentation mode while ending presentation of the content using the first presentation mode. Alternatively, the operations further include, after initiating presentation of the content using the second presentation mode, presenting the content using both the first presentation mode and the second presentation mode in parallel.
[0013] In some implementations, providing the user-selectable option as output from the user interface includes displaying the user-selectable option as a graphical element on a screen of the user device via the user interface, where the graphical element informs the user that a second presentation mode is available for presenting the content. In these implementations, receiving a user input indication includes one of receiving touch input on the screen that selects the displayed graphical element, receiving stylus input on the screen that selects the displayed graphical element, receiving gesture input indicating selection of the displayed graphical element, receiving gaze input indicating selection of the displayed graphical element, or receiving voice input indicating selection of the displayed graphical element.
[0014] In some examples, providing the user-selectable option as output from the user interface includes providing the user-selectable option as an audible output from a speaker in communication with the user device via the user interface, where the audible output informs the user that a second presentation mode is available for presenting the content. In some implementations, receiving a user input indication indicating a selection of the user-selectable option includes receiving a voice input from the user indicating a user command to select the user-selectable option. In these implementations, these operations may further include activating a microphone to capture the voice input from the user in response to providing the user-selectable option as output from the user interface.
[0015] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the detailed description below. Other aspects, features, and advantages will be apparent from the detailed description and drawings, and from the claims. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a schematic diagram of an exemplary environment including a user using context trigger assistance. [Figure 2] FIG. 1 is a schematic diagram of an example context-triggered assistance process. [Figure 3A] 3 is a schematic diagram 300a of an exemplary user device with context trigger assistance enabled in an environment of user devices. [Figure 3B] 3 is a schematic diagram 300b of an exemplary user device in an environment of user devices with context trigger assistance enabled. [Figure 3C] 3 is a schematic diagram 300c of an exemplary user device in an environment of user devices with context trigger assistance enabled. [Figure 4A] 4 is a schematic diagram of an exemplary display 400a rendered on a screen of a user device for displaying an assistive function. [Figure 4B] FIG. 4 is a schematic diagram of an exemplary display 400b rendered on a screen of a user device for displaying assistive features. [Figure 4C] FIG. 4 is a schematic diagram of an exemplary display 400c rendered on a screen of a user device for displaying assistive features. [Figure 5] 1 is a flowchart of an exemplary arrangement of operations for a method for triggering an assistance function of a user device. [Figure 6] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0017] Like reference symbols in the various drawings indicate similar elements.
[0018] 1 is an example system 100 for triggering assistance functions based on a current state of a user 10 while using the user device 110. Briefly, as described in more detail below, an assistant application 140 running on the user device 110 obtains the current state 212 (FIG. 2) of the user 10 while the user 10 is using the user device 110 in a first presentation mode 234, 234a. Based on the current state 212 of the user 10, the assistant application 140 provides a selection of a presentation mode option 232 that, when selected, initiates the presentation of a second presentation mode 234, 234b on the user device 110.
[0019] The system 100 includes a user device 110 running an assistant application 140 with which a user 10 can interact. Here, the user device 110 corresponds to a smartphone. However, the user device 110 may be any computing device, such as, without limitation, a tablet, a smart display, a desk / laptop, a smart watch, a smart appliance, a smart speaker, headphones, or a vehicle infotainment device. The user device 110 includes data processing hardware 112 and memory hardware 114 that stores instructions that, when executed on the data processing hardware 112, cause the data processing hardware 112 to perform one or more operations (e.g., related to contextual assistance features). The user device 110 includes an array of one or more microphones 116 configured to capture acoustic sounds, such as voices, or other audible noises directed to the user device 110. The user device 110 may also include or be in communication with an audio output device (e.g., a speaker) 118 that may output notifications 404 and / or audio, such as synthesized speech (e.g., from the assistant application 140). The user device 110 may include an automatic speech recognition (ARS) system 142 that includes an audio subsystem configured to receive voice input from the user 10 via one or more microphones 116 of the user device 110 and process the voice input (e.g., to perform various voice-related functions).
[0020] The user device 110 may be configured to communicate with the remote system 130 via the network 120. The remote system 130 may include remote resources, such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). In some examples, some functions of the assistant application 140 reside locally or on the device, while other functions reside remotely. In other words, any of the functions of the assistant application 140 may be local or remote in any combination. For example, when the assistant application 140 performs automatic speech recognition (ASR), which involves significant processing requirements, the remote system 130 may perform that processing. Furthermore, when the user device 110 can support the processing requirements, for example, when the user device 110 is performing hotword detection or running end-to-end ASR (e.g., with processing requirements supported by the device), the data processing hardware 112 and / or memory hardware 114 may perform that processing. Optionally, assistant application 140 functionality can exist both locally / on-device and remotely (e.g., as a hybrid of local and remote).
[0021] The user device 110 includes a sensor system 150 configured to capture sensor data 152 within the environment of the user device 110. The user device 110 may receive the sensor data 152 captured by the sensor system 150 continuously or at least at periodic intervals to determine a current state 212 of the user 10 of the user device 110. Some examples of sensor data 152 include global positioning data, motion data, image data, connectivity data, noise data, audio data, or other data indicative of the state of the user device 110 or the state of the environment in the user device 110's vicinity. In the case of global positioning data, a system associated with the user device 110 may detect the location and / or orientation of the user 10. The motion data may include accelerometer data that characterizes the movement of the user 10 through movement of the user device 110. The image data may be used to detect features of the user 10 (e.g., gestures by the user 10 or facial features that characterize the user's 10 gaze) and / or features of the user's 10 environment. Connection data may be used to determine whether user device 110 is connected to other electronic devices or devices (e.g., docked with a vehicle infotainment system or headphones). Acoustic data, such as noise data or voice data, may be captured by sensor system 150 and used to determine the environment of user device 110 (e.g., characteristics or properties of an environment having a particular acoustic signature) or to identify whether user 10 or another party is speaking. In some implementations, sensor data 152 includes wireless communication signals (i.e., signal data), such as Bluetooth or Ultrasonic, that represent other computing devices (e.g., other user devices) in the vicinity of user device 110.
[0022] In some implementations, the user device 110 executes an assistant application 140 that implements a state determiner process 200 (FIG. 2) and a presenter 202 that manages which presentation mode options 232 are made available (i.e., presented) to the user 110. That is, the assistant application 140 determines the current state 212 of the user 10 and controls, as options 232, which presentation modes 234 are available to the user 10 based on the current state 212 of the user 10. In this sense, the assistant application 140 is configured (e.g., via a graphical user interface (GUI) 400) to present content (e.g., audio / visual content) on the user device 110 in different formats, called presentation modes 234. Furthermore, the assistant application 140 can facilitate which presentation modes 234 are available at a particular time depending on the content being communicated and / or the perceived current state 212 of the user 10. Advantageously, assistant application 140 allows user 10 to select presentation mode 234 from presentation mode options 232 using an interface 400, such as a graphical user interface (GUI) 400. As used herein, GUI 400 may receive user input instructions via any one of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylus) to interact with assistant application 140.
[0023] The assistant application 140 running on the user device 110 may render, for display on the GUI 400, presentation mode options 232 from which the user 10 can select to present content on the user device 110. The presentation mode options 232 rendered on the GUI 400 may include, for each presentation mode 234, a respective graphic 402 (FIGS. 4A-4C) that identifies the presentation mode 234 available to the current state 212 of the user 10. In other words, by displaying a graphical element 402 for each presentation mode option 232 on the GUI 400, the assistant application 140 can inform the user 10 which presentation mode options 232 are available for presenting content. For example, when the current state 212 of the user 10 is driving (i.e., a visually engaged activity), the presentation mode options 232 rendered for display on the GUI 400 may include graphical elements 402 for an audio-based presentation mode 234 that allows the user 10 of the user device 110 to listen to content using speakers 118 (e.g., headphones, vehicle infotainment, etc.) connected to the user device 110. On the other hand, when the current state 212 of the user 10 is commuting (e.g., walking or using public transportation), the presentation mode options 232 rendered for display on the GUI 400 may include graphics 402 for a visual-based presentation mode 234a or an audio-based presentation mode 234b.
[0024] 1, user 10 is walking through urban environment 12 while using user device 110 in visual-based presentation mode 234a. For example, visual-based presentation mode 234a may correspond to user 10 reading a news article displayed on GUI 400 of user device 110 (e.g., reading from a web browser or a news-specific application). In this example, vehicle 16 passes user 10 and honks its horn 18. In urban environment 12, it may be advantageous for user 10 to be more alert rather than looking at user device 110. Thus, audio-based presentation mode 234 may enable user 10 to pay closer attention to features of urban environment 12 while walking.
[0025] While the user device 110 is in the visual-based presentation mode 234a, the sensor system 150 of the user device 110 detects noises from the vehicle 16 (e.g., a horn honking 18) as sensor data 152. The sensor data 152 may further include geographic coordinate data indicating the geographic location of the user 10. The sensor data 152 is input to the assistant application 140, including the presenter 202, which determines the current state 212 of the user 10. Because the sensor data 152 indicates that the environment 12 is noisy within the vicinity of the user device 110 and / or indicates that the user 10 is currently located in a congested urban area near an intersection, the presenter 202 may determine that the user 10 may want to switch from the visual-based presentation mode 234a to the audio-based presentation mode 234b to enable the user 10 to have a visual awareness of their surroundings. In response, the presenter 202 provides the user 10 with the audio-based presentation mode 234b as a presentation mode option 232 (e.g., in addition to other presentation modes 234, such as the visual-based presentation mode 234a) as a graphical element 402 selectable by the user 10. The presentation mode options 232 and the corresponding graphical element 402 may be rendered unobtrusively on the GUI 400 as a "peek" to inform the user that, based on the current state 212, another presentation mode 234 may be a more suitable option for the user 10. The user 10 then provides a user input indication 14 indicating a selection of the option 232 representing the audio-based presentation mode 234b. For example, the user input indication 14 indicating the selection may cause the user device 110 to switch to the audio-based presentation mode 234b so that the assistant application 140 presents a news article (i.e., by outputting synthesized playback audio).
[0026] Referring to FIG. 2 illustrating an example state determiner process 200, the presenter 202 may include a state determiner 210 and a mode suggester 230. The state determiner 210 may be configured to identify a current state 212 of the user 10 based on sensor data 152 collected by the sensor system 150. In other words, the presenter 202 may use the sensor data 152 to derive / confirm the current state 212 of the user 10. For example, the current sensor data 152 (or recent sensor data 152) represents the current state 212 of the user 10. Given the current state 212, the mode suggester 230 may select a corresponding set of one or more presentation mode options 232, 232a-n for presenting content on the user device 110. In some examples, the mode suggester 230 accesses a data store 240 that stores all presentation modes 234, 234a-n that the user device 110 has in place to present to the user 10 as presentation mode options 232. In some examples, presentation mode 234 is associated with or dependent on the application hosting the content being displayed on user device 110 of user 10. For example, a news application may have a reading mode with a set or customizable text size and an audio mode in which text from articles is read aloud as synthesized speech (i.e., output from a speaker associated with user device 110). Accordingly, current state 212 may also indicate the current application hosting the content being presented to user 10.
[0027] In some implementations, the state determiner 210 maintains a record of a previous state 220 of the user 10. Here, the previous state 220 may refer to a state of the user 10 characterized by sensor data 152 that is not the most recent (i.e., most recent) sensor data 152 from the sensor system 150. For example, the previous state 220 of the user 10 may be walking in an environment that is apparently free of distractions to the user 10. In this example, after receiving the sensor data 152, the state determiner 210 may determine that the current state 212 of the user 10 is walking in a noisy and / or busy environment (i.e., urban environment 12). This change between the previous state 220 and the current state 212 of the user 10 triggers the mode suggester 230 to provide presentation mode options 232 to the user 10. However, if the state determiner 210 determines that the previous state 220 and the current state 212 are the same, the state determiner 210 may not send the current state 212 to the mode suggester 230, and the presenter will not present any presentation mode options 232 to the user 10.
[0028] In some examples, when there is a detected difference (e.g., a difference in the sensor data 152) between the previous state 220 and the current state 212, the state determiner 210 simply outputs the current state 212 to the mode suggester 230 (thereby triggering the mode suggester 230). For example, the state determiner 210 may be configured with a state change threshold, and when the detected difference between the previous state 220 and the current state 212 meets (e.g., exceeds) the state change threshold, the state determiner 210 outputs the current state 212 to the mode suggester 230. The threshold may be zero, in which case a slight difference between the previous state 220 and the current state 212 detected by the state determiner 210 may trigger the mode suggester 230 of the presenter 202 to provide presentation mode options 232 to the user 10. Conversely, as a type of user-interruption sensitivity mechanism, the threshold may be higher than zero to prevent unnecessary triggering of the mode suggester 230.
[0029] The current state 212 of the user 10 may indicate one or more current activities the user 10 is performing. For example, the current activity of the user 10 may include at least one of walking, driving, commuting, talking, or reading. Additionally, the current state 212 may characterize the environment the user 10 is in, such as a noisy / busy environment or a quiet / remote environment. Furthermore, the current state 212 of the user 10 may include the current location of the user device 110. For example, the sensor data 152 includes global positioning data that defines the current location of the user device 110. By way of illustration, the user 10 may be near a dangerous location, such as an intersection or a railroad crossing, and a change in presentation mode 234 may be advantageous to the sensory perception awareness of the user 10 at or near the current location. In other words, including the current location as part of the current state 212 may be relevant to the presenter 202 determining when and / or which options 232 to present to the user 10. Furthermore, the current situation 212 indicating that the user 10 is still reading content rendered for display on the GUI 400 in the visual-based presentation mode 234a can provide additional confidence that the user's 10 sensitivity-perceptual awareness needs are important, thereby warranting presenting the presentation option 232 to switch to the audio-based presentation mode 234b (if available).
[0030] The mode suggester 230 may receive the current state 212 of the user 10 as input and select a presentation mode option 232 associated with the current state 212 from a list of available presentation modes 234 (e.g., from the presentation mode data store 240). In these examples, the mode suggester 230 may discard presentation modes 234 that are not associated with the user's current state 212. In some examples, the mode suggester 230 simply retrieves the presentation mode 234 associated with the current state 212 from the presentation mode data store 240. For example, when the current state 212 of the user 10 is talking, the presentation modes 234 associated with the current state 212 may exclude audio-based presentation modes 234. When the current state 212 of the user 10 is driving, the presentation modes 234 associated with the current state 212 may exclude video-based presentation modes 234. In other words, each current state 212 of the user 10 is associated with one or more presentation modes 234 from which the mode suggester 230 makes its decision.
[0031] The current state 212 may also communicate to auxiliary components connected to the user device 110. For example, in the above example where the user is walking through a crowded urban environment 12 while actively reading a news article presented by the GUI 400 via a visual-based presentation mode 234a, a current state 212 further indicating that headphones are paired with the user device 110 may provide the mode suggester 230 with additional confidence to determine the need to present a presentation mode option 232 to switch to an audio-based presentation mode 234b so that the news article is dictated as synthesized speech for audible output through the headphones. In another example, while presenting content in the visual-based presentation mode 234a, the current state 212 may further communicate that the orientation and proximity of the user device 110 relative to the face of the user 10 are approximately in a state that indicates that the user 10 is having difficulty reading the content. Here, the mode suggester 230 may present a presentation mode option 232 to increase the text size of content presented in a visual-based presentation mode 232a in addition to, or instead of, a presentation mode option 232 to present content in an audio-based presentation mode 234b.
[0032] In some implementations, the mode suggester 230 determines the presentation mode options 232 to output to the user 10 by considering content 250 currently presented on the user device 110. For example, the content 250 may indicate that the user 10 is currently using a web browser that has the capability for dictating a news article that the user 10 is reading. Accordingly, the mode suggester 230 includes an audio-based presentation mode 234 among the presentation mode options 232 provided to the user 10. Additionally or alternatively, the content 250 may indicate that the user 10 is currently using an application that has subtitle and video capabilities, but no dictation capabilities. In these examples, the mode suggester 230 includes the video-based presentation mode 234 and the subtitle presentation mode 234, but excludes (i.e., discards / ignores) the audio-based presentation mode 234 from the presentation mode options 232 provided to the user 10.
[0033] 3A-3C show schematic diagrams 300a-c of a user 10 with a user device 110 running an assistant application 140 as the user 10 moves around an environment (e.g., an urban environment 12). FIGS. 4A-4C show exemplary GUIs 400a-c rendered on the screen of the user device 110 to display respective sets of presentation mode options 232 that complement the current state 212 of the user 10 determined in FIGS. 3A-3C. As discussed above, each presentation mode option 232 may be rendered in the GUI 400 as a respective graphic 402 representing a corresponding presentation mode 234. As apparent in FIGS. 3A-4C, the presentation mode options 232 rendered in each of the GUIs 400a-c change based on the current state 212 of the user 10 as the user 10 moves around the environment 12.
[0034] 3A and 4A, user 10 is performing a current activity of walking along a street within environment 12. Additionally, user 10 is reading a news article rendered in GUI 400a. Thus, user device 110 may be described as presenting the news article to user 10 in a visual-based presentation mode 234a. For example, user 10 previously stopped to read the news article, such that user 10's previous state 220 may be characterized as a static state. As discussed with reference to FIGS. 1 and 2, user device 110 may continuously (or at periodic intervals) acquire sensor data 152 captured by sensor system 150 to determine user 10's current state 212, whereby current state 212 is associated with presentation mode options 232 available to user 10. Referring to FIGS. 3A and 4A, sensor data 152 may indicate that external speakers 118 (i.e., Bluetooth headphones) are currently connected to user device 110. Thus, the presentation modes 234 presented to the user 10 as presentation mode options 232 take this headphone connectivity into account.
[0035] The sensor system 150 of the user device 110 can pass sensor data 152 to the state determiner 210 of the presenter 202, causing the state determiner 210 to determine that the current state 212 of the user 10 is walking. The state determiner 210 may form this determination based on the user 10's changing location within the environment 12 (e.g., as indicated by the location / movement sensor data 152). After determining the current state 212 of the user 10, the mode suggester 230 may determine a presentation mode option 232 associated with the current state 212 of the user 10 from presentation modes 234 available on the user device 110. As noted above, the mode suggester 230 may ignore presentation modes 234 that are not associated with the current state 212 of the user 10 when determining presentation mode options 232 to present to the user 10.
[0036] After determining the presentation mode options 232 associated with the user's 10's current state 212, the user device 110 generates (e.g., using the assistant application 140) the presentation mode options 232 for display on the GUI 400a of FIG. 4A. As shown in FIG. 4A, the user device 110 renders / displays, on the GUI 400a, a first graphical element 402, 402a for the first presentation mode option 232, 232a and a second graphical element 402, 402b for the second presentation mode option 232, 232b at the bottom of the screen. The two graphical elements 402a-b may inform the user 10 that two different presentation modes 234 are available for presenting the content (i.e., a news article).
[0037] In the illustrated example, the assistant application 140 of the user device 110 may further render / display a graphical element 404a representing text asking the user 10, "Would you like me to read this aloud?" The presentation mode options 232 associated with the current walking state 212 may include a first audio-based presentation option 232a corresponding to a first audio-based presentation mode 234b in which the news article is dictated using a connected external speaker 118 (e.g., headphones), and a second audio-based presentation option 232b corresponding to another second audio-based presentation mode 234c in which the news article is dictated using the internal speaker 118 of the user device 110. Here, the user 10 may provide a user input indication 14a indicating selection of the second audio-based presentation option 232b to use the internal speaker 118 of the user device 110 (e.g., by touching a graphical button in the GUI 400a that generally represents "Speaker"). This selection then causes the user device 110 to begin presenting the news article in a dictated second audio-based presentation mode 234c.
[0038] 3B and 4B, in response to user 10 providing user input instruction 14a indicating selection of internal speaker 118 of user device 110 displayed in GUI 400a of FIG. 4A to cause assistant application 140 to begin presenting a second audio-based presentation mode 234c, assistant application 140 of user device 110 dictates a news article to user 10 while current status 212 indicates that user 10 is walking. As shown in FIG. 4B, GUI 400b displays / renders second audio-based presentation mode 234c in parallel with visual-based presentation mode 234a displayed / rendered in GUI 400a. In particular, GUI 400b displays a waveform graphic to indicate that content is currently being audibly output from an audio output device in second audio-based presentation mode 234c. The graphic may further provide playback options such as pause / play, as well as options for scanning forward / backward through the content. In some implementations, the initiation of a presentation mode 234 terminates the presentation of content in the previous presentation mode 234. For example, the presentation of a news article in an audio-based presentation mode 234 occurs simultaneously with the termination of the presentation of a visual-based presentation mode 234 that was rendered while the user 10 was in a static state.
[0039] 3B , a vehicle 16 passes the user 10 and honks its horn 18. In an urban environment 12, this sudden loud noise may make it difficult for the user to hear the audio-based presentation mode 234 from the speaker 118 of the user device 110. The sensor system 150 may detect the honking sound 18 as sensor data 152 and provide the sensor data 152 to the presenter 202. Based on the sensor data 152, the state determiner 210 of the presenter 202 may determine that the current state 212 of the user 10 indicates that the user 10 is walking in a noisy environment. The state determiner 210 may make its determination based on a single environmental factor captured as sensor data 152 (e.g., the honking sound 18) or an aggregation of environmental factors captured as sensor data 152 (e.g., the honking sound 18 of the vehicle 16 plus a geographic location in proximity to a busy street). Additionally, when the vehicle 16 honks, the current sensor data 152 still indicates that an external speaker 118 (ie, Bluetooth headphones) is connected to the user device 110 .
[0040] After determining the current state 212 of the user 10, the mode suggester 230 may determine presentation mode options 232 associated with the current state 212 of the user 10 from the presentation modes 234 available on the user device 110. In particular, the presentation mode 234 of the internal speaker 118 of the user device 110 may be excluded from the presentation mode options 232 because it is the current presentation mode 234 rendered / displayed on the GUI 400b. As mentioned above, the mode suggester 230 may also ignore presentation modes 234 that are not associated with the current state 212 of the user 10 when determining presentation mode options 232 to present to the user 10.
[0041] After determining the presentation mode option 232 associated with the current state 212 of the user 10, the user device 110 generates (i.e., using the assistant application 140) a presentation mode option 232a for display on the GUI 400b of FIG. 4B. As shown in FIG. 4B, the user device 110 renders / displays a graphical element 402a of the presentation mode option 232a at the bottom of the screen on the GUI 400b. The graphical element 402a informs the user 10 that the presentation mode option 232a is available as an audio-based presentation mode 234b for presenting the content (i.e., a news article).
[0042] In the illustrated example, the assistant application 140 of the user device 110 further renders / displays a graphical element 404b representing text asking the user 10, "Would you like to switch to Bluetooth?" The presentation mode options 232 associated with the current state 212 of walking in a noisy environment may include a first audio-based presentation mode option 232a having a corresponding first presentation mode 234b of dictating a news article using a connected external speaker 118 (e.g., headphones). Here, the user 10 may provide a user input indication 14b indicating a selection of the connected external speaker 118 of the user device 110 (e.g., by touching a graphical button in the GUI 400b that generally represents "headphones") to initiate presentation of the news article in the first audio-based presentation mode 234b of dictation to the external speaker 118 connected to the user device 110.
[0043] 3C and 4C, the assistant application 140 of the user device 110 is dictating a news article to the user 10 through the connected external speaker 118 while the current situation 212 of the user 10 is walking in a busy environment in response to the user 10 providing a user input instruction 14b indicating a selection of the connected external speaker 118 of the user device 110, displayed in GUI 400b of FIG. 4B, to cause the assistant application 140 to begin presenting the first audio-based presentation mode 234b. As shown in FIG. 4C, the GUI 400b displays / renders the first audio-based presentation mode 234b in parallel with the visual-based presentation mode 234a displayed / rendered in GUI 400a.
[0044] As shown in this example, user 10 is currently walking toward a crosswalk. In an urban environment 12, this crosswalk is potentially dangerous to user 10 if user 10 is not paying attention to the environment 12. Sensor system 150 may detect from sensor data 152 that user 10 is approaching the crosswalk and provide sensor data 152 to presenter 202. A state determiner 210 of presenter 202 may determine, based on sensor data 152, that a current state 212 of user 10 indicates that a potential danger is imminent. State determiner 210 may make that determination based on sensor data 152 that indicates environmental factors, such as a busy street, in addition to the crosswalk that user 10 is approaching. Additionally, current sensor data 152 still indicates that external speakers 118 (i.e., Bluetooth headphones) are connected to user device 110.
[0045] After determining the current state 212 of the user 10, the mode suggester 230 may determine presentation mode options 232 associated with the current state 212 of the user 10 from the presentation modes 234 available on the user device 110. In particular, the first audio-based presentation mode 234b of the connected external speaker 118 of the user device 110 may be excluded from the presentation mode options 232 because it is the current presentation mode 234 rendered / displayed on the GUI 400c. As mentioned above, the mode suggester 230 may also ignore presentation modes 234 that are not associated with the current state 212 of the user 10 when determining presentation mode options 232 to present to the user 10.
[0046] After determining the presentation mode options 232 associated with the current state 212 of the user 10, the user device 110 generates (i.e., using the assistant application 140) the presentation mode options 232 for display on the GUI 400c of FIG. 4C. As shown in FIG. 4C, the user device 110 renders / displays a graphical element 402 of the presentation mode options 232 at the bottom of the screen on the GUI 400c. The graphical element 402 informs the user 10 that the presentation mode options 232 are available for presenting the content (i.e., the news article).
[0047] In the illustrated examples, the assistant application 140 of the user device 110 further renders / displays a graphical element 404c representing a notification or warning to the user that reads, "Warning: You are approaching a crosswalk. Do you wish to pause?" The presentation mode options 232 associated with the current state 212 of approaching a crosswalk may include a graphical element 402 for pausing the audio-based presentation mode 234 and a graphical element 402 for switching to the visual-based presentation mode 234 for viewing the content later. Additionally or alternatively, because the sensor data 152 indicates that an external speaker 118 is connected to the user device 110, the assistant application 140 may output a warning to the user 10 as synthesized speech 122. In these examples, the user 10 may provide a user input indication 14 indicating a selection of the presentation mode option 232 of the user device 110 by providing voice input to the user device 110. For example, the voice input is an utterance spoken by the user 10 that is a user command to begin presentation of a news article in a particular presentation mode 234. In other words, the user 10 can speak a command to select a graphical element 402 to cause the user device 110 to begin presenting the news article in the presentation mode 234 associated with the selected graphical element 402.
[0048] In some implementations, in response to providing the synthesized speech 122 to the user 10, the user device 110 (e.g., via the assistant application 140) may activate the microphone 116 of the user device 110 to capture voice input from the user 10. In these implementations, the assistant application 140 of the user device 110 may be trained to detect, but not recognize, specific warm words associated with the presentation mode options 232 (e.g., "yes," "no," "video-based presentation mode," etc.) via the microphone 116, without performing full speech recognition of spoken utterances containing these specific warm words. This would maintain the privacy of the user 10, such that no unintended audio is recorded while the microphone 116 is active, while still reducing the power / computation required to detect the associated warm words.
[0049] 5 includes a flowchart of an example configuration of operations of a method 500 for triggering an assistance function of a user device 110. At operation 502, the method 500 includes obtaining a current state 212 of a user 10 of the user device 110 while the user device 110 is using a first presentation mode to present content to the user 10 of the user device 110. At operation 504, the method 500 further includes providing, based on the current state 212 of the user 10, as an output from a user interface 400 of the user device 110, a user-selectable option 402 that, when selected, causes the user device 110 to use a second presentation mode to present content to the user 10. At operation 506, the method 500 further includes, in response to receiving a user input indication 14 indicating a selection of the user-selectable option 402, starting the presentation of the content using the second presentation mode.
[0050] 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described herein. Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functionality are meant to be exemplary only and are not meant to limit the implementation of the invention(s) described and / or claimed herein.
[0051] Computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connecting to memory 620 and a high-speed expansion port 650, and a low-speed interface / controller 660 connecting to a low-speed bus 670 and storage device 630. Each of components 610, 620, 630, 640, 650, and 660 may be interconnected using various buses and mounted on a common motherboard or otherwise as appropriate. Processor 610 may process instructions for execution within computing device 600, including instructions stored in memory 620 or on storage device 630 to display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 680 coupled to high-speed interface 640. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memories. Additionally, multiple computing devices 600 may be connected, with each device providing a portion of the required operations (eg, as a server bank, a group of blade servers, or a multi-processor system).
[0052] The memory 620 stores information non-transiently within the computing device 600. The memory 620 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 620 may be a physical device used to temporarily or permanently store programs (e.g., a sequence of instructions) or data (e.g., program state information) for use by the computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), as well as disk or tape.
[0053] The storage device 630 is capable of providing mass storage for the computing device 600. In some implementations, the storage device 630 is a computer-readable medium. In various different implementations, the storage device 630 can be an array of devices, including a floppy disk drive, a hard disk drive, an optical disk drive, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 620, the storage device 630, or memory on the processor 610.
[0054] The high-speed controller 640 manages bandwidth-intensive operations for the computing device 600, and the low-speed controller 660 manages less bandwidth-intensive operations. Such an allocation of duties is merely exemplary. In some implementations, the high-speed controller 640 is coupled to the memory 620, the display 680 (e.g., through a graphics processor or accelerometer), and the high-speed expansion port 650, which can accept various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to the storage device 630 and the low-speed expansion port 690. The low-speed expansion port 690, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or to a networking device, such as a switch or router, for example, through a network adapter.
[0055] Computing device 600 may be implemented in several different forms, as shown in the figure. For example, computing device 600 may be implemented as a standard server 600a, or may be implemented multiple times within a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.
[0056] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0057] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," an "app," or a "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0058] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0059] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data, or is operably coupled to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0060] To provide for user interaction, one or more aspects of the present disclosure may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen for displaying information to a user, and, optionally, a keyboard and pointing device, e.g., a mouse or trackball, by which the user may provide input to the computer. Other types of devices may likewise be used to provide for user interaction; for example, feedback provided to the user may be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic input, voice input, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from devices used by the user, e.g., by sending web pages to a web browser on the user's client device in response to a request received from that web browser.
[0061] Although several implementations have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0062] 10 users 12 Urban environment 14 User Input Instructions 14a User Input Instructions 14b User Input Instructions 16 vehicles 18 Honking the horn, the sound of honking the horn 100 systems 110 User Devices 112 Data Processing Hardware 114 Memory Hardware 116 microphones 118 Audio output devices, speakers, external speakers, internal speakers 120 Network 130 Remote Systems 132 Remote Data Processing Hardware 134 Remote Memory Hardware 140 Assistant Applications 142 Automatic Speech Recognition (ARS) System 150 Sensor System 152 Sensor Data 200 State determiner process 202 Presenter 210 State Determinants 212 Current situation 220 Previous state 230 Mode Suggestor 232 Presentation Mode Options, Options 232a~n Presentation Mode Options 232a visual-based presentation mode, first presentation mode option, first audio-based presentation option, presentation mode option 232b Second Presentation Mode Option, Second Audio-Based Presentation Option 234 first presentation mode, second presentation mode, presentation mode, audio-based presentation mode, video-based presentation mode, subtitle presentation mode, visual-based presentation mode 234a~n Presentation Mode 234a First Presentation Mode, Visually Based Presentation Mode 234b second presentation mode, audio-based presentation mode, first audio-based presentation mode, first presentation mode 234c Second Audio-Based Presentation Mode 240 data stores 250 Contents 400 Graphical User Interface (GUI), Interface 400a~400c Display, GUI 402 graphic, graphical element, first graphical element, second graphical element, user selectable option 402a First Graphical Element 402b Secondary Graphical Element 404 notification 500 ways 600 computing devices 610 Processor, component 620 Memory, Components, Non-Temporary Memory 630 Storage Devices, Components 640 High-Speed Interface / Controller, Components 650 High-Speed Expansion Port, Components 660 Low-Speed Interface / Controller, Components 670 Slow Bus 680 display
Claims
1. 1. A computer-implemented method (500) that, when executed on data processing hardware (112) of a user device (110), causes the data processing hardware (112) to perform operations for triggering an assistance function on the user device (110), the operations comprising: While the user device (110) is using a first presentation mode (234) to present content to a user (10) of the user device (110), receiving sensor data captured by the user device, the sensor data including an orientation and proximity of the user device to the user's face; obtaining a current state (212) of the user (10) of the user device (110) based on the sensor data; providing, based on the current state (212) of the user (10), as output from a user interface (400) of the user device (110), a user-selectable option (402) that, when selected, causes the user device (110) to use a second presentation mode (234) to present the content to the user (10); initiating presentation of the content using the second presentation mode (234) in response to receiving a user input indication (14) indicating selection of the user-selectable option (402); A computer-implemented method (500) comprising:
2. 10. The computer-implemented method of claim 1, wherein the sensor data further comprises at least one of global positioning data, image data, noise data, accelerometer data, connection data indicating that the user device is connected to another device, or noise / voice data.
3. 10. The computer-implemented method of claim 1, wherein the current state of the user indicates one or more current activities being performed by the user.
4. 4. The computer-implemented method of claim 3, wherein the current activity of the user includes at least one of walking, driving, commuting, talking, or reading.
5. 10. The computer-implemented method of claim 1, wherein providing the user-selectable options as output from the user interface is further based on a current location of the user device.
6. 10. The computer-implemented method of claim 1, wherein providing the user-selectable options as output from the user interface is further based on the type of content and / or a software application running on the user device that is providing the content.
7. the first presentation mode (234) comprises one of a visual-based presentation mode (234a) or an audio-based presentation mode (234b); the second presentation mode (234) includes the other of the visual-based presentation mode (234a) or the audio-based presentation mode (234b).
10. The computer-implemented method (500) of claim 1.
8. 10. The computer-implemented method of claim 1, wherein the operations further include, after starting presentation of the content using the second presentation mode, simultaneously presenting the content using the second presentation mode while ending presentation of the content using the first presentation mode.
9. 2. The computer-implemented method of claim 1, wherein the operations further include, after initiating presentation of the content using the second presentation mode, presenting the content using both the first presentation mode and the second presentation mode in parallel.
10. 2. The computer-implemented method of claim 1, wherein providing the user-selectable options as output from the user interface includes displaying the user-selectable options as graphical elements on a screen of the user device via the user interface, the graphical elements informing the user that the second presentation mode is available for presenting the content.
11. The step of receiving a user input instruction (14) comprises: receiving touch input on the screen that selects the displayed graphical element; receiving a stylus input on the screen that selects the displayed graphical element; receiving a gesture input indicating a selection of the displayed graphical element; receiving a gaze input indicating a selection of the displayed graphical element; or receiving a voice input indicating a selection of the displayed graphical element.
11. The computer-implemented method (500) of claim 10, comprising one of:
12. 2. The computer-implemented method of claim 1, wherein providing the user-selectable options as an output from the user interface includes providing the user-selectable options as an audible output from a speaker in communication with the user device via the user interface, the audible output informing the user that the second presentation mode is available for presenting the content.
13. 13. The computer-implemented method (500) of any one of claims 1 to 12, wherein receiving the user input indication (14) indicating a selection of the user-selectable option (402) comprises receiving a voice input from the user (10) indicating a user command to select the user-selectable option (402).
14. 14. The computer-implemented method (500) of claim 13, wherein the operations further include activating a microphone (116) to capture the audio input from the user (10) in response to providing the user-selectable option (402) as output from the user interface (400).
15. A system (100), data processing hardware (112); memory hardware (114) in communication with said data processing hardware (112); When the memory hardware (114) is executed on the data processing hardware (112), the memory hardware (114) causes the data processing hardware (112) to While the user device (110) is using a first presentation mode (234) to present content to a user (10) of said user device (110), receiving sensor data captured by the user device, the sensor data including an orientation and proximity of the user device to the user's face; obtaining a current state (212) of the user (10) of the user device (110) based on the sensor data; providing, based on the current state (212) of the user (10), as output from a user interface (400) of the user device (110), a user-selectable option (402) that, when selected, causes the user device (110) to use a second presentation mode (234) to present the content to the user (10); initiating presentation of the content using the second presentation mode (234) in response to receiving a user input indication (14) indicating selection of the user-selectable option (402); A system (100) storing instructions for performing operations including:
16. 16. The system of claim 15, wherein the sensor data further comprises at least one of global positioning data, image data, noise data, accelerometer data, connection data indicating that the user device is connected to another device, or noise / voice data.
17. 16. The system of claim 15, wherein the current state of the user indicates one or more current activities that the user is performing.
18. 20. The system (100) of claim 17, wherein the current activity of the user (10) comprises at least one of walking, driving, commuting, talking, or reading.
19. 16. The system (100) of claim 15, wherein providing the user-selectable options (402) as output from the user interface (400) is further based on a current location of the user device (110).
20. 16. The system (100) of claim 15, wherein providing the user-selectable options (402) as output from the user interface (400) is further based on the type of content and / or a software application (140) running on the user device (110) that is providing the content.
21. the first presentation mode (234) comprises one of a visual-based presentation mode (234a) or an audio-based presentation mode (234b); the second presentation mode (234) includes the other of the visual-based presentation mode (234a) or the audio-based presentation mode (234b).
16. The system (100) of claim 15.
22. 16. The system (100) of claim 15, wherein the operations further include, after starting to present the content using the second presentation mode (234), ending to present the content using the first presentation mode (234) while simultaneously presenting the content using the second presentation mode (234).
23. 16. The system (100) of claim 15, wherein the operations further include, after initiating presentation of the content using the second presentation mode (234), presenting the content using both the first presentation mode (234) and the second presentation mode (234) in parallel.
24. 17. The system of claim 16, wherein providing the user-selectable options as output from the user interface includes displaying the user-selectable options as graphical elements on a screen of the user device via the user interface, the graphical elements informing the user that the second presentation mode is available for presenting the content.
25. receiving the user input instructions (14); receiving a touch input on the screen that selects the displayed graphical element; receiving a stylus input on the screen that selects the displayed graphical element; receiving a gesture input indicating a selection of the displayed graphical element; receiving a gaze input indicating a selection of the displayed graphical element; or receiving a voice input indicating a selection of the displayed graphical element; 25. The system (100) of claim 24, comprising one of:
26. 16. The system of claim 15, wherein providing the user-selectable options as an output from the user interface includes providing the user-selectable options as an audible output from a speaker in communication with the user device via the user interface, the audible output informing the user that the second presentation mode is available for presenting the content.
27. 27. The system (100) of any one of claims 15 to 26, wherein receiving the user input indication (14) indicating a selection of the user-selectable option (402) comprises receiving a voice input from the user (10) indicating a user command to select the user-selectable option (402).
28. 28. The system (100) of claim 27, wherein the operations further include activating a microphone (116) to capture the voice input from the user (10) in response to providing the user-selectable option (402) as output from the user interface (400).
Citation Information
Patent Citations
Information processing apparatus, information processing method, and program
JP2019128846A
Operation support method and operation support device
JP2020064492A
Assist layer with automated extraction
US20160350136A1