Contextual triggers for accessibility features
By obtaining the current state of the user's device and providing switching options, the problem of users not being aware of their surroundings when using computing devices is solved, achieving the effect of improving security while maintaining device usability.
Patent Information
- Application Number
- CN202280063542.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-26
- Filing Date
- 2022-07-06
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-07-06
AI Technical Summary
When users are visually preoccupied with using computing devices, they may be unable to notice activities in their surroundings, which could lead to potential dangers, such as not noticing oncoming vehicles while walking.
The user's current state is obtained through the sensor system of the user's device, and based on this state, optional options are provided, allowing the user to choose to switch to a more suitable presentation mode, such as switching from visual mode to audio mode, so that the user can be aware of the surrounding environment.
This allows users to continue using computing devices while remaining aware of their surroundings, improving user safety and attention allocation.
Smart Images

Figure CN118020059B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to context triggering of accessibility features. BACKGROUND
[0002] Users often interact with computing devices (e.g., smartphones, smartwatches, and smart speakers) through digital assistant interfaces. These digital assistant interfaces enable users to consume media content on various applications accessible to the computing devices. When a user of a computing device consumes media content, the media content often occupies some aspect of the user’s senses. For example, when a user is reading a news article, the act of reading the news article occupies the user’s vision. Thus, in situations where the computing device occupies the user’s vision, the user can not be visually aware of other activities occurring around the user. This can be problematic in situations where the activities around the user require visual awareness to, for example, prevent potential harm to the user and / or the computing device. For example, if the user is walking and reading a news article, the user can not be aware of a head-on collision with another person approaching the user. As assistant interfaces become more integrated with these various applications and operating systems running on computing devices, digital assistants can be used to influence how media content is presented to users of computing devices to help users be aware. SUMMARY
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware of a user device, causes the data processing hardware to perform operations for triggering accessibility features on the user device, the operations comprising: while the user device is presenting content to a user of the user device using a first presentation mode, obtaining a current state of the user of the user device. The operations further comprise providing, as output from a user interface of the user device, a user-selectable option that, when selected, causes the user device to present content to the user using a second presentation mode based on the current state of the user. The operations further comprise initiating presentation of the content using the second presentation mode in response to receiving an indication of user input indicative of selection of the user-selectable option.
[0004] Implementations of the present disclosure can include one or more of the following optional features. In some implementations, the operations further comprise receiving sensor data captured by the user device, wherein obtaining the current state of the user is based on the sensor data. In these implementations, the sensor data can include at least one of global positioning data, image data, noise data, accelerometer data, connection data indicative of the user device being connected to another device, or noise / speech data.
[0005] In some examples, the current status of the user indicates one or more current activities being performed by the user. Here, the current activities of the user can include at least one of walking, driving, commuting, talking, or reading. In some implementations, providing the user-selectable option as output from the user interface is further based on a current location of the user device. Additionally or alternatively, providing the user-selectable option as output from the user interface is based on a type of the content and / or a software application running on the user device that provided the content.
[0006] In some examples, the first presentation mode includes one of a visual-based presentation mode or an audio-based presentation mode, and the second presentation mode includes the other of the visual-based presentation mode or the audio-based presentation mode. In some implementations, the operations further include, after initiating presentation of the content using the second presentation mode, presenting the content using the second presentation mode while ceasing presentation of the content using the first presentation mode. Alternatively, the operations further include, after initiating presentation of the content using the second presentation mode, presenting the content using both the first presentation mode and the second presentation mode in parallel.
[0007] In some implementations, providing the user-selectable option as output from the user interface includes displaying the user-selectable option as a graphical element on a screen of the user device via the user interface. Here, the graphical element informs the user that the second presentation mode is available for presenting the content. In these implementations, receiving the user input indication includes receiving one of: a touch input on the screen that selects the displayed graphical element; receiving a stylus input on the screen that selects the displayed graphical element; receiving a gesture input that indicates selection of the displayed graphical element; receiving a gaze input that indicates selection of the displayed graphical element; or receiving a voice input that indicates selection of the displayed graphical element.
[0008] In some examples, providing the user-selectable option as output from the user interface includes providing, via the user interface, the user-selectable option as audible output from a speaker in communication with the user device. Here, the audible output notifies the user that the second presentation mode is available for presentation of the content. In some implementations, receiving the user input indication indicating selection of the user-selectable option includes receiving, from the user, speech input indicating a user command to select the user-selectable option. In these implementations, the operations can further include, in response to providing the user-selectable option as output from the user interface, activating a microphone to capture the speech input from the user.
[0009] Another aspect of the disclosure provides a system for triggering assistive functionality on a user device. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a current state of a user of the user device when the user device is presenting content to the user of the user device using a first presentation mode. The operations also include providing a user-selectable option as output from a user interface of the user device based on the current state of the user, the user-selectable option causing the user device to present the content to the user using a second presentation mode when the user-selectable option is selected. The operations further include initiating presentation of the content using the second presentation mode in response to receiving a user input indication indicating selection of the user-selectable option.
[0010] This aspect can include one or more of the following optional features. In some implementations, the operations further include receiving sensor data captured by the user device, wherein obtaining the current state of the user is based on the sensor data. In these implementations, the sensor data can include at least one of global positioning data, image data, noise data, accelerometer data, connection data indicating that the user device is connected to another device, or noise / speech data.
[0011] In some examples, the current status of the user indicates one or more current activities being performed by the user. Here, the current activities of the user can include at least one of walking, driving, commuting, talking, or reading. In some implementations, providing the user-selectable option as output from the user interface is further based on a current location of the user device. Additionally or alternatively, providing the user-selectable option as output from the user interface is based on a type of the content and / or a software application running on the user device that provided the content.
[0012] In some examples, the first presentation mode includes one of a visual-based presentation mode or an audio-based presentation mode, and the second presentation mode includes the other of the visual-based presentation mode or the audio-based presentation mode. In some implementations, the operations further include, after initiating presentation of the content using the second presentation mode, presenting the content using the second presentation mode while ceasing presentation of the content using the first presentation mode. Alternatively, the operations further include, after initiating presentation of the content using the second presentation mode, presenting the content using both the first presentation mode and the second presentation mode in parallel.
[0013] In some implementations, providing the user-selectable option as output from the user interface includes displaying the user-selectable option as a graphical element on a screen of the user device via the user interface. Here, the graphical element informs the user that the second presentation mode is available for presenting the content. In these implementations, receiving the user input indication includes receiving one of: a touch input on the screen that selects the displayed graphical element; receiving a stylus input on the screen that selects the displayed graphical element; receiving a gesture input that indicates selection of the displayed graphical element; receiving a gaze input that indicates selection of the displayed graphical element; or receiving a voice input that indicates selection of the displayed graphical element.
[0014] In some examples, providing the user-selectable option as output from the user interface includes providing, via the user interface, the user-selectable option as audible output from a speaker in communication with the user device. Here, the audible output informs the user that the second presentation mode is available for presentation of the content. In some implementations, receiving the user input indicative of selection of the user-selectable option includes receiving, from the user, speech input indicative of a user command to select the user-selectable option. In these implementations, the operations can further include, in response to providing the user-selectable option as output from the user interface, activating a microphone to capture the speech input from the user.
[0015] The details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a schematic diagram of an example environment including a user using context-triggered assistive functionality.
[0017] Figure 2 is a schematic diagram of an example context-triggered assistive process.
[0018] Figures 3A-3C is a schematic diagram 300a-c of an example user device that enables context-triggered assistive functionality in an environment of the user device.
[0019] Figure 4A -C is a schematic diagram of an example display 400a-c of display assistive functionality rendered on a screen of a user device.
[0020] Figure 5 is a flowchart of an example arrangement of operations for a method for triggering assistive functionality of a user device.
[0021] Figure 6 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0022] Like reference symbols in the various drawings indicate like elements. DETAILED DESCRIPTION
[0023] Figure 1 is an example system 100 for triggering assistive functionality based on a current state of a user 10 when using a user device 110. In brief, as described in more detail below, when the user 10 is using the user device 110 in a first presentation mode 234, 234a, an assistant application 140 executed on the user device 110 obtains a current state 212 of the user 10 (Figure 2 ). Based on the current state 212 of the user 10, the assistant application 140 provides a selection of a presentation mode option 232, which when selected initiates presentation of a second presentation mode 234, 234b on the user device 110.
[0024] The system 100 includes a user device 110 that executes an assistant application 140 with which a user 10 can interact. Here, the user device 110 corresponds to a smartphone. However, the user device 110 can be any computing device such as, but not limited to, a tablet, a smart display, a desktop / laptop computer, a smart watch, a smart appliance, a smart speaker, a headset, or an in-vehicle infotainment device. The user device 110 includes data processing hardware 112 and memory hardware 114 that stores instructions that, when executed on the data processing hardware 112, cause the data processing hardware 112 to perform one or more operations (e.g., related to contextual assistant functionality). The user device 110 includes an array of one or more microphones 116 configured to capture acoustic sound, such as speech directed to the user device 110 or other audible noise. The user device 110 can also include or be in communication with an audio output device (e.g., a speaker) 118 that can output audio, such as notifications 404 and / or synthesized speech (e.g., from the assistant application 140). The user device 110 can include an automatic speech recognition (ASR) system 142 that includes an audio subsystem configured to receive speech input from the user 10 via the one or more microphones 116 of the user device 110 and process the speech input (e.g., perform various speech-related functionality).
[0025] The user device 110 can be configured to communicate with a remote system 130 via a network 120. The remote system 130 can include remote resources, such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). In some examples, some functionality of the assistant application 140 resides locally or on-device, while other functionality resides remotely. In other words, any functionality of the assistant application 140 can be any combination of local or remote. For example, when the assistant application 140 performs automatic speech recognition (ASR) that includes a heavy processing requirement, the remote system 130 can perform the processing. However, when the user device 110 can support the processing requirement, such as when the user device 110 is performing hotword detection or operating end-to-end ASR (e.g., with device-supported processing requirements), the data processing hardware 112 and / or the memory hardware 114 can perform the processing. Alternatively, the assistant application 140 functionality can reside locally / on-device and remotely (e.g., as a hybrid of local and remote).
[0026] User equipment 110 includes a sensor system 150 configured to capture sensor data 152 within the environment of user equipment 110. User equipment 110 may continuously or at least at periodic intervals receive the sensor data 152 captured by sensor system 150 to determine the current state 212 of user 10 of user equipment 110. Some examples of sensor data 152 include global positioning data, motion data, image data, connectivity data, noise data, voice data, or other data indicating the state of user equipment 110 or the state of the environment near user equipment 110. Using global positioning data, systems associated with user equipment 110 can detect the position and / or orientation of user 10. Motion data may include accelerometer data characterizing the motion of user 10 via the motion of user equipment 110. Image data may be used to detect features of user 10 (e.g., gestures of user 10 or facial features characterizing the gaze of user 10) and / or environmental features of user 10. Connectivity data can be used to determine whether user equipment 110 is connected to other electronic devices or equipment (e.g., docking with a vehicle infotainment system or headphones). Acoustic data (e.g., noise data or voice data) can be captured by sensor system 150 and used to determine the environment of user equipment 110 (e.g., characteristics or attributes of an environment with specific acoustic features) or to identify whether user 10 or another party is speaking. In some embodiments, sensor data 152 includes wireless communication signals (i.e., signal data) (e.g., Bluetooth or ultrasound) representing other computing devices (e.g., other user equipment) in the vicinity of user equipment 110.
[0027] In some implementations, user equipment 110 executes assistant application 140, which implements state determiner process 200. Figure 2The assistant application 140 manages which presentation mode options 232 are available to the user 110 (i.e., presentation). That is, the assistant application 140 determines the current state 212 of the user 10 and controls which presentation modes 234 are offered to the user 10 as options 232 based on the current state 212 of the user 10. In this sense, the assistant application 140 (e.g., via a graphical user interface (GUI) 400) is configured to present content (e.g., audio / visual content) on the user device 110 in different formats referred to as presentation modes 234. Furthermore, the assistant application 140 can facilitate which presentation modes 234 are available at a particular time based on the content being conveyed and / or the perceived current state 212 of the user 10. Advantageously, the assistant application 140 allows the user 10 to select a presentation mode 234 from the presentation mode options 232 using the interface 400 (e.g., a graphical user interface (GUI) 400). As used herein, GUI 400 may receive user input instructions via touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylus) to interact with assistant application 140.
[0028] The assistant application 140 running on user device 110 can render presentation mode options 232 for display on GUI 400, and user 10 can select presentation mode options 232 to present content on user device 110. The presentation mode options 232 rendered on GUI 400 may include a corresponding graphic 402 for each presentation mode 234. Figures 4A-4C The graphic 402 identifies the presentation mode 234 available for the user 10's current state 212. In other words, by displaying a graphic element 402 for each presentation mode option 232 on the GUI 400, the assistant application 140 can inform the user 10 which presentation mode options 232 are available for presenting content. For example, when the user 10's current state 212 is driving (i.e., visually engaging activity), the presentation mode options 232 rendered to be displayed on the GUI 400 may include a graphic element 402 of an audio-based presentation mode 234 that allows the user 10 of the user device 110 to listen to content using a speaker 118 (e.g., headphones, in-vehicle infotainment, etc.) connected to the user device 110. On the other hand, when the current state 212 of user 10 is commuting (e.g., walking or using public transportation), the rendering mode options 232 rendered to be displayed on GUI 400 may include a visual-based rendering mode 234a or an audio-based rendering mode 234b of graphics 402.
[0029] exist Figure 1In the example, the user 10 is walking in an urban environment 12 while using the user device 110 in a visual-based presentation mode 234a. For example, the visual-based presentation mode 234a can correspond to the user 10 reading a news article (e.g., from a web browser or a news-specific application) displayed on the GUI 400 of the user device 110. In this example, the vehicle 16 drives by the user 10 and honks 18. In the urban environment 12, it can be advantageous for the user 10 to be more alert than looking at the user device 110. Accordingly, the audio-based presentation mode 234 can allow the user 10 to more closely focus on aspects of the urban environment 12 while walking.
[0030] While the user device 110 is in the visual-based presentation mode 234a, the sensor system 150 of the user device 110 detects noise from the vehicle 16 (e.g., the sound of the vehicle honk 18) as sensor data 152. The sensor data 152 can further include geographic coordinate data indicative of the geographic location of the user 10. The sensor data 152 is input to the assistant application 140 including the presenter 202, which determines the current state 212 of the user 10. Because the sensor data 152 indicates that the environment 12 near the user device 110 is noisy and / or indicates that the user 10 is currently located in a crowded urban area near an intersection, the presenter 202 can determine that the user 10 can wish to switch from the visual-based presentation mode 234a to the audio-based presentation mode 234b to enable the user 10 to have an auditory perception of his / her surroundings. Accordingly, the presenter 202 provides the audio-based presentation mode 234b as a presentation mode option 232 (e.g., in addition to other presentation modes 234 such as the visual-based presentation mode 234a) to the user 10 as a graphical element 402 that can be selected by the user 10. The presentation mode option 232 and the corresponding graphical element 402 can be rendered on the GUI 400 as a "peek" in an unobtrusive manner to notify the user that another presentation mode 234 can be a more suitable option for the user 10 based on the current state 212. The user 10 subsequently provides a user input indication 14 indicating a selection of the option 232 representing the audio-based presentation mode 234b. For example, the user input indication 14 indicating the selection can cause the user device 110 to switch to the audio-based presentation mode 234b such that the assistant application 140 reads the news article (i.e., outputs synthesized playback audio).
[0031] Reference Figure 2which depicts an example state determiner process 200, the presenter 202 can include a state determiner 210 and a mode recommender 230. The state determiner 210 can be configured to identify a current state 212 of the user 10 based on sensor data 152 collected by the sensor system 150. In other words, the presenter 202 uses the sensor data 152 to derive / ascertain the current state 212 of the user 10. For example, the current sensor data 152 (or the most recent sensor data 152) is representative of the current state 212 of the user 10. With the current state 212, the mode recommender 230 can select a corresponding one or more presentation mode options 232, set 232a-n, for presenting content on the user device 110. In some examples, the mode recommender 230 accesses a data store 240 that stores all presentation modes 234, 234a-n, which the user device 110 is equipped with as presentation mode options 232 to present to the user 10. In some examples, the presentation modes 234 are associated with or dependent on an application that hosts content displayed on the user device 110 of the user 10. For example, a news application can have a reading mode with set or customizable text size and an audio mode where text from an article is read aloud (i.e., output from a speaker associated with the user device 110) as synthesized speech. Thus, the current state 212 can also indicate a current application that hosts content being presented to the user 10.
[0032] In some implementations, the state determiner 210 maintains a record of a previous state 220 of the user 10. Here, the previous state 220 can refer to a state of the user 10 characterized by sensor data 152 that is not the most recent (i.e., up-to-date) sensor data 152 from the sensor system 150. For example, the previous state 220 of the user 10 can be walking in an environment where the user 10 is not significantly distracted. In this example, after receiving the sensor data 152, the state determiner 210 can determine that the current state 212 of the user 10 is walking in a noisy and / or busy environment (i.e., the urban environment 12). This change between the previous state 220 and the current state 212 of the user 10 triggers the mode recommender 230 to provide the presentation mode options 232 to the user 10. However, if the state determiner 210 determines that the previous state 220 and the current state 212 are the same, the state determiner 210 can not send the current state 212 to the mode recommender 230, and the presenter does not present any presentation mode options 232 to the user 10.
[0033] In some examples, the state determiner 210 outputs the current state 212 to the mode recommender 230 (and thus triggers the mode recommender 230) only when a difference between the previous state 220 and the current state 212 is detected (e.g., a difference in the sensor data 152). For example, the state determiner 210 can be configured with a state change threshold, and the state determiner 210 outputs the current state 212 to the mode recommender 230 when the detected difference between the previous state 220 and the current state 212 satisfies the state change threshold (e.g., exceeds the threshold). The threshold can be zero, where the slightest difference between the previous state 220 and the current state 212 detected by the state determiner 210 can trigger the mode recommender 230 of the presenter 202 to provide the presentation mode options 232 to the user 10. Conversely, the threshold can be higher than zero to prevent unnecessary triggering of the mode recommender 230 as a kind of user interruption sensitive mechanism.
[0034] The current state 212 of the user 10 can indicate one or more current activities being performed by the user 10. For example, the current activities of the user 10 can include at least one of walking, driving, commuting, talking, or reading. Further, the current state 212 can characterize an environment in which the user 10 is located, such as a noisy / busy environment or a quiet / remote environment. Further still, the current state 212 of the user 10 can include a current location of the user device 110. For example, the sensor data 152 includes global positioning data defining the current location of the user device 110. To illustrate, the user 10 can be near a dangerous location, such as a crosswalk or a train track crossing, and a change in the presentation mode 234 can be beneficial to the sensory perception of the user 10 at or near the current location. In other words, including the current location as part of the current state 212 can be relevant to the decision of the presenter 202 as to when to present the options 232 to the user 10 and / or which options 232 to present. Further, the current state 212 indicating that the user 10 is also reading content that is rendered for display on the GUI 400 in the visual-based presentation mode 234a can provide additional confidence that the sensory perception of the user 10 is critical and thus it is warranted to present the presentation option 232 for switching to the audio-based presentation mode 234a (if available).
[0035] The mode recommender 230 receives the current state 212 of the user 10 as input and can select a presentation mode option 232 associated with the current state 212 from a list of available presentation modes 234 (e.g., from the presentation mode data store 240). In these examples, the mode recommender 230 can discard presentation modes 234 that are not associated with the current state 212 of the user. In some examples, the mode recommender 230 only retrieves from the presentation mode data store 240 the presentation modes 234 associated with the current state 212. For example, when the current state 212 of the user 10 is that the user is speaking, the presentation modes 234 associated with the current state 212 can exclude audio-based presentation modes 234. When the current state 212 of the user 10 is that the user is driving, the presentation modes 234 associated with the current state 212 can exclude video-based presentation modes 234. In other words, each current state 212 of the user 10 is associated with one or more presentation modes 234, from which the mode recommender 230 makes its determinations.
[0036] The current state 212 can also convey auxiliary components connected to the user device 110. For example, in the above example where the user is walking in a crowded urban environment 12 while actively reading a news article being presented by the GUI 400 via a vision-based presentation mode 234a, a current state 212 further indicating that earphones are paired with the user device 110 can provide the mode recommender 230 with additional confidence to determine a need to present a presentation mode option 232 to switch to an audio-based presentation mode 234b so that the news article is read aloud as synthesized speech that is outputtable through the earphones. In other examples, the current state 212 can further convey that the orientation and proximity of the user device 110 relative to the face of the user 10 is close when presenting content in a vision-based presentation mode 234a to indicate that the user 10 is having difficulty reading the content. Here, in addition to or instead of a presentation mode option 232 to present the content in an audio-based presentation mode 234b, the mode recommender 230 can present a presentation mode option 232 to increase the text size of the content being presented in the vision-based presentation mode 232a.
[0037] In some implementations, the mode recommender 230 determines the presentation mode options 232 to output to the user 10 by taking into account the content 250 that is currently being presented on the user device 110. For example, the content 250 can indicate that the user 10 is currently using a web browser that has the ability to speak a news article that the user 10 is reading. Accordingly, the mode recommender 230 includes an audio-based presentation mode 234 in the presentation mode options 232 provided to the user 10. Additionally or alternatively, the content 250 can indicate that the user 10 is currently using an application that has closed-captioning capabilities and video capabilities but not speaking capabilities. In these examples, the mode recommender 230 includes a video-based presentation mode 234 and a closed-captioning presentation mode 234 in the presentation mode options 232 provided to the user 10, but excludes (i.e., discards / ignores) the audio-based presentation mode 234.
[0038] Figures 3A-3C Figures 3a-c show a schematic 300a-c of the user 10 as the user 10 with the user device 110 executing the assistive application 140 moves about an environment (e.g., the urban environment 12). Figures 4A-4C Figures 4a-c show example GUIs 400a-c rendered on a screen of the user device 110 to display a respective set of presentation mode options 232 that supplement the current state 212 of the user 10 determined in Figures 3A-3C As described above, each presentation mode option 232 can be rendered in the GUI 400 as a respective graphic 402 that represents the corresponding presentation mode 234. As will become apparent in Figures 3A-4C As will become apparent in
[0039] Referring to Figure 3A and 4A the user 10 is currently engaged in a street-walking activity in the environment 12. Further, the user 10 is reading a news article that is rendered in the GUI 400a. As such, the user device 110 can be described as presenting the news article to the user 10 in a vision-based presentation mode 234a. For example, the user 10 was previously standing still to read the news article, such that the previous state 220 of the user 10 can have been characterized as stationary. As discussed with reference to Figure 1 and 2 the user device 110 can continuously (or at periodic intervals) obtain the sensor data 152 captured by the sensor system 150 to determine the current state 212 of the user 10, whereby the current state 212 is associated with the presentation mode options 232 available to the user 10. With reference toFigure 3A and 4A Sensor data 152 can indicate that an external speaker 118 (i.e., Bluetooth headset) is currently connected to user device 110. Therefore, presentation mode 234, which is presented to user 10 as presentation mode option 232, takes into account this headset connection.
[0040] The sensor system 150 of user equipment 110 can transmit sensor data 152 to the state determiner 210 of the presenter 202, whereby the state determiner 210 determines that the current state 212 of user 10 is walking. The state determiner 210 may form this determination based on changes in the user 10's position in the environment 12 (e.g., as indicated by the position / movement sensor data 152). After determining the current state 212 of user 10, the mode proposer 230 can determine a presentation mode option 232 associated with the current state 212 of user 10 from the presentation modes 234 available on user equipment 110. As described above, when determining the presentation mode option 232 presented to user 10, the mode proposer 230 may ignore presentation modes 234 that are not related to the current state 212 of user 10.
[0041] After determining the presentation mode option 232 associated with the current state 212 of user 10, user device 110 generates (i.e., using assistant application 140) the presentation mode option 232 for display. Figure 4A On the GUI 400a. (e.g.) Figure 4A As shown, on GUI 400a, user device 110 renders / displays first graphical elements 402, 402a for first presentation mode options 232, 232a and second graphical elements 402, 402b for second presentation mode options 232, 232b at the bottom of the screen. The two graphical elements 402a-b can inform user 10 that two different presentation modes 234 are available for presenting content (i.e., news articles).
[0042] In the illustrated example, the assistant application 140 of the user device 110 can further render / display a graphical element 404a representing text asking the user 10 "Would you like to have this article read aloud?". The presentation mode options 232 associated with the current state 212 of walking can include a first audio-based presentation option 232a corresponding to a first audio-based presentation mode 234b of having the news article read aloud using a connected external speaker 118 (e.g., headphones) and a second audio-based presentation option 232b corresponding to another second audio-based presentation mode 234c of having the news article read aloud using the internal speaker 118 of the user device 110. Here, the user 10 can provide a user input indication 14a indicating a selection of the second audio-based presentation option 232b to use the internal speaker 118 of the user device 110 (e.g., by touching a graphical button in the GUI 400a generally representing "Speaker"). This selection subsequently causes the user device 110 to initiate presentation of the news article in the second audio-based presentation mode 234c of being read aloud.
[0043] Referring to Figure 3B and 4B , in response to the user 10 providing the user input indication 14a indicating a selection of the internal speaker 118 of the user device 110 displayed in the GUI 400a of FIG. 4a, the assistant application 140 of the user device 110 is reading aloud the news article to the user 10 while the current state 212 indicates that the user 10 is walking. As shown in FIG. 4b, the GUI 400b displays / renders the second audio-based presentation mode 234c in parallel with the visual-based presentation mode 234a displayed / rendered in the GUI 400a. Notably, the GUI 400b is displaying a waveform graph to indicate that the content is currently being output audibly from the audio output device in the second audio-based presentation mode 234c. This graph can further provide playback options, such as pause / play, and options to scan the content forward / backward. In some implementations, initiating a presentation mode 234 stops presentation of content in a previous presentation mode 234. For example, presentation of the news article in the audio-based presentation mode 234 stops presentation of the visual-based presentation mode 234 rendered while the user 10 was stationary. Figure 4B
[0044] As Figure 3B As shown, the vehicle 16 drives by the user 10 and honks 18. In the urban environment 12, this sudden loud noise can make it difficult for the user to hear the audio-based presentation mode 234 from the speaker 118 of the user device 110. The sensor system 150 can detect the honk 18 as sensor data 152 and provide the sensor data 152 to the presenter 202. Based on the sensor data 152, the state determiner 210 of the presenter 202 can determine that the current state 212 of the user 10 indicates that the user 10 is walking in a noisy environment. The state determiner 210 can make its determination based on a single environmental factor captured as sensor data 152 (e.g., the honk 18) or an aggregation of environmental factors captured as sensor data 152 (e.g., the geographic location near a crowded street in addition to the honk 18 of the vehicle 16). Moreover, the current sensor data 152 still indicates that the external speaker 118 (i.e., the Bluetooth earbuds) is connected to the user device 110 when the vehicle 16 honks.
[0045] Upon determining the current state 212 of the user 10, the mode recommender 230 can determine the presentation mode options 232 associated with the current state 212 of the user 10 from the presentation modes 234 available on the user device 110. Notably, the presentation mode 234 of the internal speaker 118 of the user device 110 can be excluded from the presentation mode options 232 as this is the current presentation mode 234 rendered / displayed on the GUI 400b. As mentioned above, the mode recommender 230 can also disregard presentation modes 234 that are not relevant to the current state 212 of the user 10 when determining the presentation mode options 232 to present to the user 10.
[0046] Upon determining the presentation mode options 232 associated with the current state 212 of the user 10, the user device 110 generates (i.e., using the assistant application 140) the presentation mode options 232a for display on the GUI 400b of the user device 110. Figure 4B As shown, on the GUI 400b, the user device 110 renders / displays the graphical elements 402a of the presentation mode options 232a at the bottom of the screen. The graphical elements 402a inform the user 10 that the presentation mode options 232a are available as audio-based presentation modes 234b for presenting the content (i.e., the news article). Figure 4B
[0047] In the illustrated example, the assistant application 140 of the user device 110 further renders / displays a graphical element 404b representing text that asks the user 10, "Would you like to switch to Bluetooth?". The presentation mode options 232 associated with the current state 212 of the user 10 walking in a noisy environment can include a first audio-based presentation mode option 232a with a corresponding first presentation mode 234b that uses the connected external speaker 118 (e.g., headphones) to read the news article aloud. Here, the user 10 can provide a user input indication 14b that indicates a selection of the connected external speaker 118 of the user device 110 (e.g., by touching a graphical button in the GUI 400b generally representing "headphones") to cause the user device 110 to initiate presentation of the news article in the first audio-based presentation mode 234b that reads aloud to the connected external speaker 118.
[0048] Referring to Figure 3C and 4C in response to the user 10 providing the user input indication 14b that indicates a selection of the connected external speaker 118 of the user device 110 displayed in the GUI 400b of FIG. 4b to cause the assistant application 140 to initiate presentation of the first audio-based presentation mode 234b, the assistant application 140 of the user device 110 is reading the news article aloud to the user 10 via the connected external speaker 118 while the current state 212 of the user 10 is walking in a busy environment. As shown in FIG. 4b, the GUI 400b displays / renders the first audio-based presentation mode 234b in parallel with the visual-based presentation mode 234a displayed / rendered in the GUI 400a. Figure 4C
[0049] As shown in the example, the user 10 is walking up to a crosswalk. In the urban environment 12, the crosswalk is a potential danger to the user 10 if the user 10 is not paying attention to the environment 12. The sensor system 150 can detect from the sensor data 152 that the user 10 is approaching the crosswalk and provide the sensor data 152 to the presenter 202. Based on the sensor data 152, the state determiner 210 of the presenter 202 can determine that the current state 212 of the user 10 indicates that a potential danger is approaching. The state determiner 210 can base its determination on the sensor data 152 that indicates an environmental factor (e.g., a crowded street in addition to the crosswalk that the user 10 is approaching). Additionally, the current sensor data 152 still indicates that the external speaker 118 (i.e., Bluetooth headphones) is connected to the user device 110.
[0050] After determining the current state 212 of the user 10, the mode recommender 230 can determine the presentation mode options 232 associated with the current state 212 of the user 10 from the presentation modes 234 available on the user device 110. Notably, the first audio-based presentation mode 234b of the connected external speaker 118 of the user device 110 can be excluded from the presentation mode options 232 because this is the current presentation mode 234 rendered / displayed on the GUI 400c. As noted above, the mode recommender 230 can also disregard presentation modes 234 that are not relevant to the current state 212 of the user 10 when determining the presentation mode options 232 to present to the user 10.
[0051] After determining the presentation mode options 232 associated with the current state 212 of the user 10, the user device 110 generates (i.e., using the assistant application 140) the presentation mode options 232 for display on the GUI 400c of the news article. Figure 4C As shown in FIG. 4C, the user device 110 renders / displays the graphical elements 402 of the presentation mode options 232 at the bottom of the screen on the GUI 400c. The graphical elements 402 inform the user 10 that the presentation mode options 232 are available for presenting the content (i.e., the news article). Figure 4C
[0052] In the illustrated example, the assistant application 140 of the user device 110 further renders / displays a graphical element 404c that represents a notification or warning to the user, “Warning: you are approaching a crosswalk. Do you want to pause?” The presentation mode options 232 associated with the current state 212 of approaching a crosswalk can include a graphical element 402 that pauses the audio-based presentation mode 234, and a graphical element 402 that switches to the visual-based presentation mode 234 for later viewing of the content. Additionally or alternatively, because the sensor data 152 indicates that the external speaker 118 is connected to the user device 110, the assistant application 140 can output the warning to the user 10 as synthesized speech 122. In these examples, the user 10 can provide a user input indication 14 by providing a voice input to the user device 110 that indicates a selection of a presentation mode option 232 of the user device 110. For example, the voice input is a spoken utterance of the user 10 that is a user command to initiate presentation of the news article in the presentation mode 234 associated with the selected graphical element 402. In other words, the user 10 can speak the command to select the graphical element 402 to cause the user device 110 to initiate presentation of the news article in the presentation mode 234 associated with the selected graphical element 402.
[0053] In some implementations, in response to providing the synthesized speech 122 to the user 10, the user device 110 (e.g., via the assistant application 140) can activate the microphone 116 of the user device 110 to capture speech input from the user 10. In these implementations, the assistant application 140 of the user device 110 can be trained to detect (but not recognize) the particular warm word (e.g., “yes,” “no,” “video-based presentation mode,” etc.) associated with the presentation mode option 232 via the microphone 116 without full speech recognition of the spoken utterance that includes the particular warm word. This would protect the privacy of the user 10 such that all unintentional speech is not recorded while the microphone 116 is active, while also reducing the power / computations required to detect the relevant warm word.
[0054] Figure 5 A flowchart of an example operational arrangement of the method 500 including triggering an assistive function of the user device 110. At operation 502, the method 500 includes obtaining a current state 212 of the user 10 of the user device 110 when the user device 110 is presenting content to the user 10 of the user device 110 using a first presentation mode. The method 500 further includes, at operation 504, based on the current state 212 of the user 10, providing a user-selectable option 402 as output from the user interface 400 of the user device 110 that, when selected, causes the user device 110 to present the content to the user 10 using a second presentation mode. At operation 506, the method 500 further includes, in response to receiving an indication of user input 14 that indicates selection of the user-selectable option 402, initiating presentation of the content using the second presentation mode.
[0055] Figure 6 A schematic diagram of an example computing device 600 that can be used to implement the systems and methods described herein. The computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the applications described and / or claimed in this document.
[0056] Computing device 600 includes a processor 610, memory 620, a storage device 630, a high-speed interface / controller 640 connecting to memory 620 and high-speed expansion ports 650, and a low speed interface / controller 660 connecting to a low speed bus 670 and storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 are interconnected using various busses, and can be mounted on a common motherboard or in other ways as appropriate. Processor 610 can process instructions for execution within the computing device 600, including instructions stored in the memory 620 or on the storage device 630 to display graphical information for a graphical user interface (GUI) on an external input / output device, such as display 680 coupled to high speed interface 640. In other implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices 600 can be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
[0057] Memory 620 stores information non-transitorily within computing device 600. Memory 620 can be a computer-readable medium, a volatile memory unit(s), or non-volatile memory unit(s). The non-transitory memory 620 can be physical devices used to store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by or
[0058] Storage device 630 can provide mass storage for computing device 600. In some implementations, storage device 630 is a computer-readable medium. In various implementations, storage device 630 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 620, the storage device 630, or memory on processor 610.
[0059] The high-speed controller 640 manages bandwidth-intensive operations for the computing device 600, while the low-speed controller 660 manages lower bandwidth-intensive operations. Such allocation of functions is exemplary only. In some implementations, the high-speed controller 640 is coupled to memory 620, display 680 (e.g., through a graphics processor or accelerator), and to high-speed expansion ports 650, which can accept various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to storage device 630 and low-speed expansion port 690. The low-speed expansion port 690, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
[0060] As shown, the computing device 600 can be implemented using a variety of different forms. For example, it can be implemented as a standard server 600a or multiple servers 600a in a cluster 600b, as a laptop computer 600c, or as part of a rack server system 600d.
[0061] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system or systems. These various implementations can also include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system or systems.
[0062] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform a task. In some examples, a software application can be referred to as an “application,” an “app,” or a “program.” Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0063] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer readable medium, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0064] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0065] To provide for interaction with a user, one or more aspects of the disclosure can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen) for displaying information to the user and optionally a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser.
[0066] A number of implementations have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims.
Claims
1. A computer-implemented method, characterized in that, When executed on the data processing hardware of a user equipment, the computer-implemented method causes the data processing hardware to perform operations for triggering accessibility functions on the user equipment, the operations including: When the user device is using a first presentation mode that includes a vision-based presentation mode to present content to the user of the user device: Receive sensor data captured by the user equipment, the sensor data conveying the orientation and proximity of the user equipment relative to the user's face; The current state of the user on the user equipment is obtained based on the sensor data. Determining the user's current state indicates that the user is currently: The current activity of reading the content is being performed, and the content is being presented to the user using the first presentation mode; and Approaching places that pose a potential danger to the user; and Based on the determination of the user's current state indicating that the user is currently engaged in reading, the following is provided as output from the user interface of the user device: User-selectable options, when selected, cause the user device to present the content to the user using a second presentation mode including an audio-based presentation mode; and A warning to the user, the warning informing the user of the potential danger; and In response to receiving a user input instruction indicating the selection of the user's optional options, the content is continued to be presented to the user of the user device by initiating the presentation of the content using the second presentation mode and the first presentation mode in parallel.
2. The computer-implemented method according to claim 1, characterized in that, The sensor data includes at least one of the following: global positioning data, image data, noise data, accelerometer data, connection data indicating that the user equipment is connected to another device, or noise / voice data.
3. The computer-implemented method according to claim 1, characterized in that, The user's current state also indicates one or more other current activities that the user is currently engaged in.
4. The computer-implemented method according to claim 3, characterized in that, The one or more other current activities include at least one of walking, driving, commuting, or talking.
5. The computer-implemented method according to claim 1, characterized in that, Providing the user-selectable options as output from the user interface is further based on the current location of the user device.
6. The computer-implemented method according to claim 1, characterized in that, Providing the user-selectable options as output from the user interface is further based on the type of content and / or the software application running on the user device providing the content.
7. The computer-implemented method according to claim 1, characterized in that, Providing the user-selectable option as output from the user interface includes displaying the user-selectable option as a graphical element on the screen of the user device via the user interface, the graphical element informing the user that the second presentation mode is available for presenting the content.
8. The computer-implemented method according to claim 7, characterized in that, Receiving the user input instruction includes receiving one of the following: The touch input on the screen selects the displayed graphic element; Receive pen input on the screen, wherein the pen input selects the displayed graphic element; Receive gesture input, the gesture input indicating the selection of the displayed graphic element; Receive gaze input, the gaze input indicating a selection of the displayed graphical element; or Receive voice input, which indicates the selection of the displayed graphical element.
9. The computer-implemented method according to claim 1, characterized in that, Providing the user-selectable option as output from the user interface includes providing the user-selectable option as an audible output from a speaker (118) communicating with the user device, the audible output informing the user that the second presentation mode is available for presenting the content.
10. The computer-implemented method according to claim 1, characterized in that, The user input instruction that receives the instruction to select the user's optional option includes receiving voice input from the user, the voice input indicating a user command to select the user's optional option.
11. The computer-implemented method according to claim 10, characterized in that, The operation further includes activating a microphone to capture the voice input from the user in response to providing the user-selectable option as an output from the user interface.
12. A system, characterized in that, include: Data processing hardware; and Memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to operate, the operations including: When a user device is using a first presentation mode that includes a vision-based presentation mode to present content to the user of the user device: Receive sensor data captured by the user equipment, the sensor data conveying the orientation and proximity of the user equipment relative to the user's face; The current state of the user on the user equipment is obtained based on the sensor data. Determining the user's current state indicates that the user is currently: The current activity of reading the content is being performed, and the content is being presented to the user using the first presentation mode; and Approaching places that pose a potential danger to the user; and Based on the determination of the user's current state indicating that the user is currently engaged in reading, the following is provided as output from the user interface of the user device: User-selectable options, when selected, cause the user device to present the content to the user using a second presentation mode including an audio-based presentation mode; and A warning to the user, the warning informing the user of the potential danger; and In response to receiving a user input instruction indicating the selection of the user's optional options, the content is presented to the user of the user device by initiating the presentation of the content using the second presentation mode and the first presentation mode in parallel.
13. The system according to claim 12, characterized in that, The sensor data includes at least one of the following: global positioning data, image data, noise data, accelerometer data, connection data indicating that the user equipment is connected to another device, or noise / voice data.
14. The system according to claim 12, characterized in that, The user's current state indicates one or more other current activities that the user is currently engaged in.
15. The system according to claim 14, characterized in that, The one or more other current activities include at least one of walking, driving, commuting, or talking.
16. The system according to claim 12, characterized in that, Providing the user-selectable options as output from the user interface is further based on the current location of the user device.
17. The system according to claim 12, characterized in that, Providing the user-selectable options as output from the user interface is further based on the type of content and / or the software application running on the user device providing the content.
18. The system according to claim 12, characterized in that, Providing the user-selectable option as output from the user interface includes displaying the user-selectable option as a graphical element on the screen of the user device via the user interface, the graphical element informing the user that the second presentation mode is available for presenting the content.
19. The system according to claim 18, characterized in that, Receiving the user input instruction includes receiving one of the following: The touch input on the screen selects the displayed graphic element; Receive pen input on the screen, wherein the pen input selects the displayed graphic element; Receive gesture input, the gesture input indicating the selection of the displayed graphic element; Receive gaze input, the gaze input indicating a selection of the displayed graphical element; or Receive voice input, which indicates the selection of the displayed graphical element.
20. The system according to claim 12, characterized in that, Providing the user-selectable option as output from the user interface includes providing the user-selectable option as an audible output from a speaker communicating with the user device, the audible output informing the user that the second presentation mode is available for presenting the content.
21. The system according to claim 12, characterized in that, The user input instruction that receives the instruction to select the user's optional option includes receiving voice input from the user, the voice input indicating a user command to select the user's optional option.
22. The system according to claim 21, characterized in that, The operation further includes activating a microphone to capture the voice input from the user in response to providing the user-selectable option as an output from the user interface.
Citation Information
Patent Citations
Tailoring user interface presentations based on user state
CN109154894A
Pedestrian information system
US20160093207A1
Changing information output modalities
US20180358012A1