Selectable control for automatic voice response system
The computing device enhances IVR system usability by providing selectable controls through text descriptions of options during calls, addressing navigation challenges for users with disabilities.
Patent Information
- Application Number
- JP2025103652
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-22
AI Technical Summary
IVR systems present multiple selectable options that are difficult for users with communication disabilities, speech impairments, or short-term memory issues to navigate and understand, leading to challenges in interacting with these systems.
A computing device provides selectable controls by determining text descriptions of IVR options during a voice or video call and displaying them on the screen, allowing users to select options through text-based interactions.
Enhances user experience by enabling easier navigation and understanding of IVR systems for users with disabilities, improving usability and interaction capabilities.
Smart Images

Figure 2025160161000001_ABST
Abstract
Description
[Background technology]
[0001] background An interactive voice response (IVR) system, or phone tree, allows callers to interact with a computer-operated telephone system through voice input or a numeric keypad. For example, telephone systems may use IVR for mobile phone purchases, bank payments, services, retail orders, utility services, travel information, and weather forecasts. IVR systems typically use a series of voice menus to identify and segment callers. These menus contain multiple options that may be difficult for callers to understand, navigate, or remember. Summary of the Invention
[0002] overview Described herein are systems and techniques for providing selectable controls for IVR systems. The described systems and techniques may determine whether audio data associated with a voice or video call between a user of a computing device and a third party includes multiple selectable options. The third party audibly provides the selectable options during the call. In response to determining that the audio data includes selectable options, the computing device may determine a text description of the multiple selectable options. The described systems and techniques may then display two or more selectable controls on a display. A user may select a selectable control to indicate a selected option of the multiple selectable options. In this manner, the described systems and techniques may improve the user experience of voice and video calls by making IVR systems easier to navigate and understand.
[0003] The described systems and techniques can improve usability for users, such as users with certain communication disabilities, when interacting with an IVR system. As an example, the described systems and techniques can enable users who are deaf and who might otherwise find it difficult or impossible to interact with an IVR system to provide responses to an IVR system. Similarly, the described systems and techniques can enable users who have speech disabilities and who might otherwise find it difficult or impossible to interact with an IVR system to provide responses to an IVR system. The described systems and techniques can also assist users with short-term memory impairments who are unable to remember the list of options provided by an IVR system in providing responses to an IVR system. The described systems and techniques can also improve usability for users who have difficulty understanding options provided in a voice or video call, for example, when the voice is distorted or they are distracted by background noise not resulting from the voice or video call.
[0004] For example, a computing device obtains audio data output from a communication application running on the computing device. The audio data includes an audible portion of a voice or video call between a user of the computing device and a third party. The computing device uses the audible portion of the voice or video call to determine whether the audio data includes two or more selectable options. The third party audibly provides two or more selectable options during the voice or video call. The audio data includes two or more selectable options. In response to determining that the display includes a text description of the two or more selectable options, the text description provides a transcription of at least a portion of the two or more selectable options. The computing device then displays two or more selectable controls. The two or more selectable controls may be selectable to indicate a selected option of the two or more selectable options to a third party. Each of the two or more selectable controls provides a text description of a respective selectable option.
[0005] Other methods, configurations, and systems for providing selectable controls for IVR systems are also described herein.
[0006] This Summary is provided to introduce simplified concepts for providing selectable controls for an IVR system, which is further described in the Detailed Description and Drawings. This Summary is not intended to identify essential features of the claimed subject matter, nor is it intended for use in determining the scope of the claimed subject matter.
[0007] The details of one or more aspects of a visual user interface for providing selectable controls for an IVR system are described herein with reference to the following drawings, in which like numbers are used to reference like features and components throughout the several drawings. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 illustrates an example environment including a computing device that can provide selectable controls for an IVR system. [Figure 2] FIG. 1 illustrates an example of a computing device capable of providing a visual user interface for an interactive voice response system. [Figure 3]FIG. 1 illustrates an example of a machine learning model of a computing device capable of providing text descriptions for selectable controls in response to an IVR system. [Figure 4] 1 is a flowchart illustrating an example operation of a computing device capable of providing selectable controls and user data related to voice and video calls. [Figure 5] FIG. 1 illustrates an operational example of providing selectable controls for an IVR system. [Figure 6A] 1A-1C illustrate example user interfaces of computing devices that assist users in voice and video calls. [Figure 6B] 1A-1C illustrate example user interfaces of computing devices that assist users in voice and video calls. [Figure 6C] 1A-1C illustrate example user interfaces of computing devices that assist users in voice and video calls. [Figure 6D] 1A-1C illustrate example user interfaces of computing devices that assist users in voice and video calls. [Figure 7A] 10A-10C illustrate other examples of user interfaces for computing devices that assist users with voice and video calls. [Figure 7B] 10A-10C illustrate other examples of user interfaces for computing devices that assist users with voice and video calls. [Figure 7C] 10A-10C illustrate other examples of user interfaces for computing devices that assist users with voice and video calls. [Figure 8A] 10A-10C illustrate other examples of user interfaces for computing devices that assist users with voice and video calls. [Figure 8B] 10A-10C illustrate other examples of user interfaces for computing devices that assist users with voice and video calls. [Figure 8C] 10A-10C illustrate other examples of user interfaces for computing devices that assist users with voice and video calls. [Figure 8D] FIG. 1 illustrates another example of a computing device that assists users with voice and video calls. DETAILED DESCRIPTION OF THE INVENTION
[0009] Detailed Description Overview This specification describes techniques and systems for providing selectable controls on a computing device for an IVR system. As noted above, an IVR system can use voice input or Dual-Tone Multimedia (DTMF) generated by a numeric keypad. IVR systems allow callers to interact with telephone systems through voice commands (Voice Commands, Multi-Frequency Tone). IVR systems can present a series of menus, each containing multiple selectable options. Voice menus can be confusing and difficult for callers to navigate. For example, some IVR systems have many options in each menu, or detailed options that are difficult to recall. Deaf users may find it difficult or impossible to hear the options and therefore may not be able to provide responses that typically select an option. Users with speech impairments may not be able to respond aloud to options. Users with short-term memory impairments may not be able to remember the options presented by an IVR system when providing responses.
[0010] Consider a smartphone with a communications application that allows a user to make voice or video calls. For example, the user can use the communications application to call a doctor's office. The doctor's office can use an IVR system to direct the caller to the appropriate information, personnel, or department. An initial voice menu can ask the user to select an appropriate language. After selecting a language, either audibly or by pressing the number associated with the desired language, the IVR system can present another menu of options. For example, the IVR system can direct the caller to additional menus related to billing, scheduling, medical questions, service providers, and personnel questions.
[0011] Communication applications generally do not assist users in navigating IVR systems. Instead, communication applications and computing devices typically require users to invoke menu options or navigate voice menus using voice input or a numeric keypad.
[0012] The described techniques and systems can assist users in navigating IVR systems by providing selectable controls associated with selectable options. In particular, the described techniques and systems can obtain audio data from a voice or video call and determine whether the conversation includes two or more selectable options. In response to determining that the conversation includes selectable options, the described techniques and systems can determine text descriptions associated with the selectable options.
[0013] Consider the doctor's office scenario described above. A smartphone can listen to the voice call and determine whether the doctor's office audibly provides an IVR menu of selectable options. In response to determining that the doctor's office audibly provides selectable options, the described systems and techniques can determine text descriptions of the selectable options and display selectable controls on the smartphone display. Each of the selectable controls provides a text description of the respective selectable option. By selecting one of the selectable controls, the user can cause the smartphone to display the selected option. In this manner, the described techniques and systems can provide a user-friendly experience that allows smartphone users to easily navigate IVR systems, enabling users who would not normally be able to interact with IVR systems to interact with such systems. The described techniques and systems are compatible with a variety of different existing IVR systems.
[0014] As a non-limiting example, a computing device may obtain audio data output from a communications application. The audio data includes an audible portion of a voice or video call between a user of the computing device and a third party. The computing device uses the audible portion to determine whether the audio data includes two or more selectable options audibly provided by the third party during the voice or video call. In response to determining that the audio data includes two or more selectable options, the computing device determines a text description of the two or more selectable options. The text description includes a transcription of at least a portion of the two or more selectable options. The computing device then displays two or more selectable controls on a display of the computing device. The two or more selectable controls provide a text description of each selectable option. A user may select a selectable control to indicate to the third party a selected option from among the two or more selectable options.
[0015] A computing device may use information from audio data only after the computing device receives explicit permission from a user of the computing device. For example, in the situations described above where a computing device may collect audio data from voice and video calls, individual users may be provided with an opportunity to provide input to control whether programs or features of the computing device collect and use the information. Additionally, individual users may be provided with an opportunity to control what the programs or features can or cannot do with the information.
[0016] This example is merely one way of illustrating how the above-described selectable controls for an IVR system can enhance the user experience on a computing device and enable users with communication disabilities to interact with the IVR system. Other examples and implementations are described throughout this specification. This specification next describes additional example configurations, components, and methods for providing selectable controls for an IVR system on a computing device.
[0017] Example environment 1 illustrates an example environment 100 that includes an example computing device 102 that can provide selectable controls for an IVR system. In addition to the computing device 102, the environment 100 includes a computing system 104 and a caller system 106. The computing device 102, the computing system 104, and the caller system 106 are communicatively coupled to a network 108.
[0018] Although the operations of computing device 102 are described as being performed locally, in some examples, the operations may be performed by multiple computing devices and systems (e.g., computing system 104), including additional computing devices and systems beyond those shown in Figure 1. For example, the operations may be performed by computing system 104, caller system 106, or a network. Other devices or systems communicatively coupled to computing device 108 may perform some or all of the functionality of computing device 102, and vice versa.
[0019] Computing system 104 represents any combination of one or more computers, mainframes, servers, cloud computing systems, or other types of remote computing systems that can exchange information with computing device 102 over network 108. Computing system 104 may store or provide access to additional processors, stored data, or other computing resources required by computing device 102 to implement the described systems and techniques for providing selectable controls for an IVR system on computing device 102.
[0020] The caller-side system 106 can execute the IVR system 110 to send and receive telephony data to and from the computing device 102 over the network 108. For example, the caller-side system 106 can be a mobile phone, a landline phone, a laptop computer, a telephone call center workstation, or other computing device configured to present the IVR system 110 to a caller. The caller-side system 106 can also represent any combination of computers, computing devices, mainframes, servers, cloud computing systems, or other types of remote computing systems that can communicate information over the network 108 to conduct a voice or video call between the caller-side system 106 and the computing device 102.
[0021] The network 108 represents any public or private communication network for transmitting data (e.g., voice communications, video communications, data packages) between computing systems, servers, and computing devices. For example, the network 108 may include a public switched telephone network (PSTN), a wireless network (e.g., a cellular network, a wireless local area network (WLAN)), a wired network (e.g., a local area network (LAN), a wide area network (WAN)), an Internet Protocol (IP) telephony network (e.g., a voice-over-IP (VoIP) network), or any combination thereof. The network 108 may include a network hub, a network switch, a network router, or other network equipment operably interconnected. The computing device 102, the computing system 104, and the caller system 106 may transmit and receive data across the network 108 using any suitable communication technology. The computing device 102, the computing system 104, and the caller system 106 may be operatively coupled to a network 108 using respective network links.
[0022] The computing device 102 represents any suitable computing device capable of providing selectable controls for an IVR system. For example, the computing device 102 may be a smartphone that provides input for a user to make or accept a voice or video call with a calling entity (e.g., the calling system 106).
[0023] The computing device 102 includes one or more communication units 112. The communication units 112 enable the computing device 102 to communicate with wireless networks including the network 108. The communication unit 112 enables communication over a network or a wired network. For example, the communication unit 112 may include a transceiver for cellular communication or network data communication. The computing device 102 can tune the communication unit 112 and supporting circuitry (e.g., antenna, front-end module, amplifier) to one or more frequency bands defined by various communication standards.
[0024] Computing device 102 includes user interface components 114, including audio components 116, display components 118, and input components 120. Computing device 102 also includes an operating system 122 and communication applications 124. These and other components (not shown) of computing device 102 are operably coupled in various ways, including by wired and wireless buses and links. Computing device 102 may include additional components and interfaces that are omitted from FIG. 1 for clarity.
[0025] The user interface component 114 manages input and output to a user interface 126 controlled by the operating system 122 or applications running on the computing device 102. For example, a communication application 124 may cause the user interface 126 to display various user interface elements, including input controls, navigation components, information components, or a combination thereof.
[0026] As described above, the user interface component 114 may include an audio component 116, a display component 118, and an input component 120. The audio component 116, the display component 118, and the input component 120 may be separate or integrated into a single component. The audio component 116 (e.g., a single speaker or multiple speakers) may receive an audio signal as input and convert the audio signal into audible sound. The display component 118 may display visual elements on the user interface 126. The display component 118 may include any suitable display technology, including light-emitting diode (LED), organic light-emitting diode (OLED), and liquid crystal display (LCD) technology. The input component 120 may be a microphone, a presence sensing device, a touchscreen, a mouse, a keyboard, or another type of component configured to receive user input.
[0027] The operating system 122 generally controls the computing device 102, including the communication unit 112, the user interface component 114, and other peripherals. For example, the operating system 122 may manage the hardware and software resources of the computing device 102 and provide common services to applications. As another example, the operating system 122 may control task scheduling. The operating system 122 and applications are generally executable by one or more processors (e.g., a system on chip (SoC), central processing unit (CPU)) to enable communication and user interaction with the computing device 102. The operating system 122 generally provides for user interaction via a user interface 126.
[0028] The operating system 122 also provides an execution environment for applications, such as a communications application 124. The communications application 124 enables the computing device 102 to make and receive voice and video calls with callers, including the caller system 106.
[0029] During an audio or video call, the communication application 124 may cause the user interface 126 to display a caller box 128, a numeric keypad icon 130, a speakerphone icon 132, selectable controls 134, and an end call icon 136. The caller box 128 may display the name and phone number of the caller (e.g., the caller system 106). The numeric keypad icon 130 is a selectable icon that, when selected, causes a numeric keypad to be displayed on the user interface 126. The speakerphone icon 132 is a selectable icon that, when selected, causes the computing device 102 to use the speakerphone feature for the audio or video call.
[0030] The selectable control 134 is selectable by a user of the computing device 102 to perform a particular action or function. In the illustrated example, the selectable control 134 is selectable by a user to indicate to the caller system 106 a selected option from the selectable options provided by the IVR system 110. The selectable control 134 may include a button, a toggle, selectable text, a slider, a checkbox, or an icon. The end call icon 136 allows the user of the computing device 102 to end the voice or video call.
[0031] The operating system 122 can associate the input detected at the input component 120 with an element of the user interface 126. In response to receiving the input (e.g., a tap) at the input component 120, the operating system 122 or the communication application 124 can receive information about the detected input from the user interface component 114. The operating system 122 or the communication application 124 can perform a function or action in response to the detected input. For example, the operating system 122 can determine that the input corresponds to a user selecting one of the selectable controls 134 and, in response, send an indication of the corresponding selected option to the caller system 106.
[0032] During operation, the operating system 122 or the communication application 124 can automatically generate selectable controls 134 that correspond to selectable options of the IVR system 110 provided by the caller system 106. The computing device 102 can obtain audio data from an audio mixer or sound engine of the operating system 122. The audio data generally includes the audible portion of a voice or video call, including the IVR options provided by the IVR system 110.
[0033] Configuration example This section describes example configurations of systems that provide selectable controls for IVR systems, some or all of which may occur separately or together. This section describes various example configurations and, for ease of reading, associates each example configuration with a drawing.
[0034] FIG. 2 illustrates an example of a computing device 202 that can provide selectable controls for an IVR system (e.g., IVR system 110). Computing device 202 is an example of computing device 102 with some additional details.
[0035] As shown in FIG. 2, computing device 202 may be a smartphone 202-1, a tablet device 202-2, a laptop computer 202-3, a desktop computer 202-4, a computerized wristwatch 202-5 or other wearable device, a voice assistant system 202-6, a smart display system, or a computing system installed in a vehicle.
[0036] In addition to the communication unit 112 and the user interface component 114 , the computing device 202 includes one or more processors 204 and a computer-readable storage medium (CRM) 206 .
[0037] The processor 204 may include any combination of one or more controllers, microcontrollers, processors, microprocessors, hardware processors, hardware processing units, digital signal processors, graphics processors, graphics processing units, etc. For example, the processor 204 may be an integrated processor and memory subsystem including, by way of non-limiting example, an SoC, a CPU, a graphics processing unit, or a tensor processing unit. An SoC typically integrates many of the components of the computing device 202, including a central processing unit, memory, and input / output ports, into a single device. The CPU typically executes commands and processes required by the computing device 202. The graphics processing unit performs operations to display graphics for the computing device 202 and may perform other specific computational tasks. The tensor processing unit typically performs symbolic matching operations in neural network machine learning applications. The processor 204 may include a single core or multiple cores.
[0038] The CRM 206 can provide persistent and non-persistent storage of executable instructions (e.g., firmware, recovery firmware, software, applications, modules, programs, functions) and data (e.g., user data, operational data) to support the execution of the executable instructions for the computing device 202. For example, the CRM 206 includes instructions that, when executed by the processor 204, execute the operating system 122 and communications application 124. Examples of the CRM 206 include volatile and non-volatile memory, fixed and removable media devices, and any suitable memory device or electronic data storage device that holds executable instructions and supporting data. The CRM 206 may include various implementations of random-access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), non-volatile RAM (NVRAM), read-only memory (ROM), flash memory, and other storage memory types in various memory device configurations. The CRM 206 excludes propagating signals. The CRM 206 may be a solid-state drive (SSD) or a hard disk drive (HDD).
[0039] The operating system 122 may also include or control an audio mixer 208 and a caption module 210. The audio mixer 208 and the caption module 210 may be dedicated hardware components, software components, or a combination thereof. 8 and the caption module 210 are separate from the operating system 122 (eg, as a system plug-in or additional add-on service installed locally on the computing device 202).
[0040] The audio mixer 208 can acquire and consolidate audio data generated by applications, including the communication application 124, running on the computing device 202. The audio mixer 208 acquires audio streams from applications, such as the communication application 124, and generates an audio output signal that, when consolidated and output from the audio component 116, reproduces the sound encoded in the audio stream. The audio mixer 208 may also adjust the audio signal in other ways, such as controlling focus, intent, and volume. The audio mixer provides an interface between the application source that generates the content and the audio component 116 that generates sound from the content. The audio mixer 208 can manage the raw audio data, analyze it, and direct the audio signal to be output by the audio component 116 or transmitted to another computing device (e.g., the caller system 106) via the communication unit 112.
[0041] The caption module 210 is configured to analyze audio data in raw form received by the audio mixer 208 (e.g., as a byte stream). For example, the caption module 210 may perform speech recognition on the audio data to determine whether the audio data includes selectable options for an IVR system, a request for user information, or communication information related to the call context. Rather than processing each audio signal, the caption module 210 may identify individual pre-mixed audio data streams suitable for captioning. For example, the caption module 210 may automatically caption spoken audio data but not notification or sonification audio data (e.g., system beeps, ringtones). The caption module 210 may apply filters to the byte stream received by the audio mixer 208 to identify audio data suitable for captioning. The caption module 210 may use machine learning models to determine descriptions of audio data from the audible portions of a voice or video call.
[0042] Rather than captioning all audio data, the operating system 122 can use metadata to focus captioning on specific portions of audio data. For example, the caption module 210 can focus on audio data related to selectable controls in an IVR system, user information responding to requests, or providing communication information related to call context. In other words, the operating system 122 can identify “captionable” audio data based on metadata and avoid captioning all audio data. Some examples of metadata include context indicators that specify the content of a voice or video call. An audio mixer can use the context indicators to control routing, focus, and captioning decisions regarding the audio data.
[0043] Some computing devices can transcribe voice or video calls. However, transcriptions typically provide a direct transcription of the audible portion of the call and cannot determine whether the conversation includes selectable options in an IVR system, requests for user information, or conveyed information relevant to the context of the call. Users often need to select desired menu options, requested user information, or other information. Users still need to read the transcript to determine the delivery. Thus, even if a computing device provides a transcription, users may still find it difficult to navigate the IVR system and select desired options. In contrast, the described systems and techniques assist users in navigating IVR systems, providing user information upon request, and managing delivery information from voice and video calls by displaying selectable controls and message elements along with associated information.
[0044] The computing device 202 also includes one or more sensors 214. The sensors 214 obtain contextual information indicative of the physical operating environment of the computing device 202 or characteristics of the computing device 202 while functioning in the physical operating environment. For example, the caption module 210 can use this contextual information as metadata to focus audio data processing. Examples of sensors 214 include motion sensors, temperature sensors, position sensors, proximity sensors, ambient light sensors, moisture sensors, and pressure sensors, among others.
[0045] During operation, the operating system 122 or the caption module 210 determines whether audio data is intended for captioning. For example, the caption module 210 may determine whether the audio data includes selectable options in an IVR system, a request for user information, or communication information related to the call context. In response to determining that the audio data is intended for captioning, the operating system 122 determines a description of the audio data. For example, the operating system 122 may execute a machine learning model (e.g., an end-to-end recurrent neural network transducer automatic speech recognition model) trained to generate a description of the audible portions of a voice or video call. The machine learning model may be any type of model suitable for learning descriptions of speech, including transcription of spoken speech. The machine learning model used by the operating system 122 may be smaller and less complex than other machine learning models because it only needs to be trained to identify the audible portions of voice and video calls. The machine learning model may avoid processing all of the audio data sent to the audio mixer 208. In this manner, the described systems and techniques can avoid the use of remote processing resources (e.g., machine learning models on remote computing devices) to avoid unnecessary privacy risks and potential processing latency.
[0046] By relying on the original audio data rather than the speech signal generated by the audio component 116, the machine learning model can generate descriptions that more accurately represent the audible portions of voice and video calls. By determining whether audio data is for captioning before using the machine learning model, the operating system 122 can avoid wasting resources over-analyzing all audio data output by the communication application 124. By determining captioning in this manner, the computing device 202 can implement more efficient, smaller, and less complex machine learning models. In this manner, the machine learning model can perform automatic speech recognition and classification techniques locally to maintain privacy.
[0047] The operating system 122 receives the description of the machine learning model and displays it using the display component 118. The display component 118 may also display other visual elements related to the description (e.g., selectable controls that allow a user to perform actions on the computing device 202). The rating system 122 can present visual elements (e.g., selectable controls 134) as part of the user interface 126. The descriptions can include transcriptions or summaries of the audible portions of voice and video calls (e.g., telephone conversations). The descriptions can also identify the context of the audible portions of the audio data. The details and operation of the machine learning model are described in more detail with respect to FIG. 3.
[0048] 3 is a diagram 300 illustrating an example of a machine learning model 302 of a computing device 202 that can provide text descriptions for selectable controls in response to an IVR system. In other implementations, the computing device 202 may be the computing device 102 of FIG. 1 or a similar computing device.
[0049] 3 , the machine learning model 302 may be part of the caption module 210. The machine learning model 302 may convert the audio data 304 into a text description 306 of the audible portion of a voice or video call (e.g., a text description of a selectable option provided by the IVR system 110) without converting the audio data 304 into sound. The audio data 304 may include different types, forms, or variations of data from the communication application 124. For example, the audio data 304 may include raw, pre-mixed voice byte stream data, or processed byte stream data. The machine learning model 302 may include multiple machine learning models combined into a single model that provides the text description 306 in response to the audio data 304.
[0050] Applications, including communication applications 124, can use machine learning models 302 to process audio data 304 into text descriptions 306. For example, communication applications 124 can communicate with machine learning models 302 through operating system 122 or caption module 210 using an application programming interface (API) (e.g., a public API across all applications). In some implementations, machine learning models 302 can process audio data 304 within a secure section or enclave of operating system 122 or CRM 206 to ensure user privacy and security.
[0051] The machine learning model 302 can perform inference. In particular, the machine learning model 302 can be trained to receive audio data 304 as input and provide a text description 306 of the audible portion of the speech as output data. By performing inference using the machine learning model 302, the caption module 210 can process the audio data 304 locally. The machine learning model 302 can also perform classification, regression, clustering, anomaly detection, recommendation generation, and other tasks.
[0052] An engineer can train the machine learning model 302 using supervised learning techniques. For example, an engineer can train the machine learning model 302 using training data 308 (e.g., truth data) that includes example statements inferred from examples of audio data 304 from a series of voice and video calls. The inferences can be applied manually by an engineer or other expert, generated through crowdsourcing, or provided by other techniques (e.g., complex speech recognition and content recognition algorithms). The training data 308 can include audio data from voice and video calls for the audio data 304. As an example, the audio data 304 may include voice calls with an IVR system used in a doctor's office. The training data 308 for the machine learning model 302 may include a wide range of voice calls with the IVR system. The training data 308 may include many audio data files from voice and video calls. As another example, the audio data 304 may include voice calls with a company's customer representatives. The training data 308 may include many audio data files from a wide range of similar voice and video calls. Engineers may also use unsupervised learning techniques to train the machine learning model 302.
[0053] The machine learning models 302 may be trained on a training computing system and then provided for storage and implementation on one or more computing devices 202. For example, the training computing system may include a model trainer. The training computing system may be included in or separate from the computing device 202 that implements the machine learning models 302.
[0054] An engineer can also train the machine learning model 302 online or offline. In offline training (e.g., batch learning), an engineer trains the machine learning model 302 on an entire static set of training data 308. In online learning, an engineer continuously trains the machine learning model 302 as new training data 308 becomes available (e.g., while the machine learning model 302 is being used on the computing device 202 to perform inference). For example, an engineer can initially train the machine learning model 302 to replicate descriptions applied to the audible portions of voice and video calls (e.g., captioned IVR systems, captioned telephone conversations). Once the machine learning model 302 infers text descriptions 306 from the audio data 304, the computing device 202 can feed the text descriptions 306 (and corresponding portions of the audio data 304) back to the machine learning model 302 as new training data 308. In this way, the machine learning model 302 can continuously improve the accuracy of the text descriptions 306. In some implementations, a user of the computing device 202 can provide input to the machine learning model 302 to flag a particular statement as having an error, which the computing device 202 can use to train the machine learning model 302 to improve future predictions.
[0055] An engineer or trainer may perform centralized training of multiple machine learning models 302 (e.g., based on a centrally stored dataset). In other implementations, a trainer or engineer may use distributed training techniques, including distributed training or federated learning, to train, update, or personalize machine-learned models 302. An engineer may use user information to personalize a machine learning model 302 only after receiving explicit permission from the user. For example, in situations where a computing device 202 may collect user information, individual users may be provided with an opportunity to provide input to control whether a program or function of a machine learning model 302 can collect and use user information. Additionally, individual users may be provided with an opportunity to control what the program or function can or cannot do with the user information.
[0056] The machine learning model 302 may be or include one or more artificial neural networks. In such implementations, the machine learning model 302 may include a group of connected or not fully connected nodes (e.g., neurons). Engineers may also organize the machine learning model 302 into one or more layers (e.g., a deep network). In a deep network implementation, the machine learning model 302 may include an input layer, an output layer, and one or more hidden layers disposed between the input layer and the output layer.
[0057] The machine learning model 302 may also include one or more recurrent neural networks. For example, the machine learning model 302 may be an end-to-end recurrent neural network-transducer automatic speech recognition model. Examples of recurrent neural networks include long short-term memory (LSTM) recurrent neural networks, gated recurrent units, bidirectional recurrent neural networks, continuous-time recurrent neural networks, neural history compressors, echo state networks, Elman networks, Jordan networks, recursive neural networks, Hopfield networks, fully recurrent networks, and sequence-to-sequence configurations.
[0058] At least some of the nodes of a recurrent neural network may form cycles. When configured as a recurrent neural network, the machine learning model 302 may be particularly useful for processing continuous input data (e.g., audio data 304). For example, the recurrent neural network may use recurrent or directed circular node connections to pass or store information from previous portions of the audio data 304 to subsequent portions of the audio data 304.
[0059] The audio data 304 may also include time-series data (e.g., audio data over time). As a recurrent neural network, the machine learning model 302 may analyze the audio data 304 over time to detect or predict spoken and associated non-speech sounds to generate a text description 306 of at least a portion of the audio data 304. For example, successive sounds from the audio data 304 may indicate spoken words in a sentence (e.g., natural language processing, speech detection, or processing).
[0060] The machine learning model 302 may also include one or more convolutional neural networks. A convolutional neural network may include multiple convolutional layers that perform convolutions on input data using trained filters or kernels. Engineers commonly use convolutional neural networks to diagnose vision problems in still or video images. Engineers may also apply convolutional neural networks to natural language processing of audio data 304 to generate text descriptions 306.
[0061] The operation of the caption module 210 and the machine learning model 302 is described in more detail herein with respect to FIG.
[0062] Example method 4 is a flowchart illustrating example operations 400 of a computing device that can provide selectable controls and user data related to voice and video calls. Operations 400 are described below in the context of computing device 202 of FIG. 2. In other implementations, computing device 202 may be computing device 102 of FIG. 1 or a similar computing device. Operations 400 may be performed in a different order than shown in FIG. 4 and may be performed with additional or fewer operations.
[0063] At 402, the computing device optionally retrieves content including user information for the computing device user. The computing device can use the user information to assist the user in retrieving requested information or storing communications related to voice and video calls. Before retrieving the user information or performing the options described below, the computing device 202 may request consent from the user to use the user information for voice and video calls. For example, the computing device 202 may use the user information only after receiving explicit consent. The computing device 202 may obtain the user information from user input into an application on the computing device 202 (e.g., entering contact information into a user profile, entering an account number via a third-party application) or may learn it from information received in an application (e.g., an account number included in an emailed statement, a saved calendar entry).
[0064] The computing device displays a graphical user interface of the communication application at 404. For example, the computing device 202 may direct the display component 118 to display the user interface 126 of the communication application 124 in response to the user placing or receiving a voice or video call.
[0065] At 406, the computing device obtains audio data output from a communication application executing on the computing device. The audio data includes the audible portion of a voice or video call. For example, the communication application 124 enables a user of the computing device 202 to make and receive voice and video calls. The audio mixer 208 obtains audio data 304 output from the communication application 124 during the voice and video calls. The audio data 304 includes the audible portion of the voice or video call between the user of the computing device 202 and a third party. The caption module 210 can extract the audio data 304 from the audio mixer 208 to provide the user with selectable controls and other information during the voice or video call.
[0066] At 408, the computing device uses the audible portion of the voice or video call to determine whether the audio data includes relevant information. The relevant information may be two or more selectable options of an IVR system (e.g., phone tree options), a request for user information (e.g., a request for a credit card number, address, or account number), or communication information (e.g., reservation details, contact information, account information). For example, the caption module 210 can use the machine learning model 302 to determine whether the audio data 304 includes relevant information. The relevant information may include two or more selectable options of an IVR system, a request for user information, or communication information. A user or a third party audibly provides the relevant information during the voice or video call. The caption module 210 or the machine learning model 302 may filter audio data 304 that does not require processing, such as notification sounds and background noise. Examples of the machine learning model 302 determining whether the audio data 304 includes two or more selectable options are shown in FIGS. 6A and 8A. Examples of the machine learning model 302 determining whether the audio data 304 includes a request for user information are shown in Figures 6B, 6C, 7A, and 8B. Examples of the machine learning model 302 determining whether the audio data 304 includes communication information are shown in Figures 6D, 7B, 7C, and 8C.
[0067] If the audio data does not include relevant information, the computing device displays a user interface of a communication application at 416. For example, in response to determining that the audio data 304 does not include relevant information, the computing device 202 displays the user interface 126 of the communication application 124.
[0068] If the audio data is determined to contain relevant information, then at 410 the computer The video device determines a text description of the associated information. The text description transcribes the associated information. For example, the caption module 210 can use a machine learning model 302 to perform speech recognition on the audio data 304 to determine a text description 306 of the associated information. The text description 306 provides a transcription of at least a portion of two or more selectable options, a request for user information, or a communication. Examples of the machine learning model 302 determining the text description 306 of two or more selectable options are shown in FIGS. 6A and 8A. Examples of the machine learning model 302 determining the text description 306 of a request for user information are shown in FIGS. 6B, 6C, 7A, and 8B. Examples of the machine learning model 302 determining the text description of a communication are shown in FIGS. 6D, 7B, 7C, and 8C.
[0069] The caption module 210 can improve the accuracy of the text description 306 in various ways, including biasing the machine learning model 302 based on the context of the computing device 202. For example, the caption module 210 may bias the machine learning model 302 based on the identity of a third party in an audio or video call. Suppose a user of the computing device 202 places an audio call to a doctor's office. The caption module 210 can bias the machine learning model 302 using common words from the doctor's office conversation. In this way, the computing device 202 can improve the text description 306 of the audio call. The caption module 210 can use other types of contextual information, including location information obtained from sensors 214 and information from other applications, to bias the machine learning model 302.
[0070] In some implementations, the computing device 202 may translate the text description 306 into another language before displaying it. For example, the caption module 210 may determine the user's preferred language from the operating system 122 and translate the text description 306 into the preferred language. In this way, a Japanese user can view the text description 306 in Japanese even if the audio data 304 is in a different language (e.g., Chinese or English).
[0071] At 412, the computing device optionally identifies user data in response to the request for user information. The computing device does not perform this operation if the audio data does not include a request for user information. For example, in response to determining that a third party has requested user information, the computing device 202 can identify user data in response to the request for user information. The computing device 202 can retrieve user data from the CRM 206, the communications application 124, another application on the computing device 202, or a remote computing device associated with the user or the computing device 202. Consider the clinic call scenario described above. The clinic receptionist can request that the user provide insurance information. In response, the computing device 202 can retrieve the health insurance company and user account number from an email previously received by the user and stored on the computing device 202. Examples of the computing device 202 identifying a user data response to a request for user information are shown in FIGS. 6B, 6C, 7A, and 8B.
[0072] A computing device may use information in response to a request for user information only after receiving explicit permission from a user of the computing device. For example, in the situations described above in which a computing device may collect user data, individual users may be provided with an opportunity to provide input to control whether programs or features of the computing device can collect and use the user data. Additionally, individual users may be provided with the opportunity to control what a program or feature can or cannot do with their user data.
[0073] At 414, the computing device displays user data or selectable controls. The selectable controls are user-selectable and include text descriptions. Consider the following scenario: The audio data includes a request for user information. In this scenario, the computing device can display the identified user data. The audio data includes two or more selectable options from an IVR system. In this scenario, the user can use the selectable control to indicate to a third party a selected option from the two or more selectable options. Consider the following scenario: The audio data includes a communication. In this scenario, the user can use the selectable control to save the communication to the computing device, a communication application, or another application. For example, the computing device 202 can cause the display component 118 to display the user data or selectable controls 134. The display component 118 can provide the user data as a text notification on the user interface 126. Consider the clinic call scenario described above: The display component 118 can display the health insurance company and user account information as text boxes on the user interface 126 during the voice call. The display component 118 can also provide the selectable controls 134. The display component 118 may provide the text description 306 or the requested information as part of a button on the user interface 126 of the communication application 124. Examples of the display component 118 displaying the selectable controls 134 are shown in Figures 6A and 8A. Examples of the display component 118 displaying user data are shown in Figures 6B, 6C, 7A, and 8B. Examples of the display component 118 displaying the selectable controls 134 and user data in response to a communication are shown in Figures 6D, 7B, 7C, and 8C.
[0074] Suppose a clinic uses IVR system 110 to direct a voice call to a receptionist. Display component 118 can display selectable controls 134. Selectable controls 134 provide text descriptions 318 of each of two or more selectable options provided by IVR system 110. A user can use selectable controls 134 to indicate to the clinic a selected option from the two or more selectable options.
[0075] Also consider that a user makes an appointment with a doctor's office. The display component 118 can display a selectable control 134. The selectable control 134 includes a text description of the appointment. The user can use the selectable control 134 to save the appointment details in a calendar application.
[0076] At 416, the computing device displays a user interface of the communication application. For example, the display component 118 can display the user interface 126 associated with the communication application 124. The user interface 126 can include user data and selectable controls 134.
[0077] 5 illustrates example operations 500 for providing selectable controls for an IVR system. The operations 500 are described in the context of computing device 202 of FIG. 2. The operations 500 may be performed in a different order or with additional or fewer operations.
[0078] At 502, a computing device obtains audio data output from a communication application executing on the computing device. The audio data includes an audible portion of a voice or video call between a user of the computing device and a third party. For example, the audio mixer 208 of the computing device 202 may obtain audio data 304 output from the communication application 124 executing on the computing device 202. The caption module 210 may receive the audio data 304 from the audio mixer 208. The audio data 304 includes an audible portion of a voice or video call between a user of the computing device 202 and a third party (e.g., a person, a computerized IVR system).
[0079] At 504, the computing device uses the audible portion to determine whether the audio data includes two or more selectable options. The third party audibly provides the two or more selectable options during the voice or video call. For example, the machine learning model 302 of the caption module 210 can use the audible portion of the audio data 304 to determine whether the audio data 304 includes two or more selectable options (e.g., numbered options in an IVR menu or phone tree). The third party audibly provides the two or more selectable options during the voice or video call.
[0080] At 506, in response to determining that the audio data includes two or more selectable options, the computing device determines a text description of the two or more selectable options. The text description provides a transcription of at least a portion of the two or more selectable options. For example, in response to determining that the audio data 304 includes two or more selectable options, the machine learning model 302 determines a text description 306 of the two or more selectable options. The text description 306 provides a transcription of at least a portion of the two or more selectable options. In some implementations, the text description 306 includes a word-by-word transcription of the two or more selectable options. In other implementations, the text description 306 provides a paraphrase of the two or more selectable options.
[0081] At 508, the computing device displays two or more selectable controls. The two or more selectable controls are selectable by a user to indicate to a third party a selected option of the two or more selectable options. Each of the two or more selectable controls provides a text description of the respective selectable option. For example, the display component 118 displays two or more selectable controls 134 on a display of the computing device 202. The display includes the user interface 126. The two or more selectable controls 134 are selectable by a user to provide to a third party an indication of a selected option of the two or more selectable options. Each of the two or more selectable controls provides a text description 306 of the respective selectable option.
[0082] Example of implementation This section describes implementations of the described systems and techniques, some or all of which may occur separately or together, that can assist users with voice and video calls. This section describes various implementations, and for ease of reading, each is outlined in relation to specific figures.
[0083] 6A-6D are diagrams illustrating a computing device that assists a user in voice and video calls. 6A-6D are sequentially described in the context of computing device 202 of FIG. 2. Computing device 202 may provide a different user interface with fewer or additional features than those shown in FIGS. 6A-6D.
[0084] 6A, computing device 202 causes display component 118 to display user interface 126. User interface 126 is associated with communication application 124. User interface 126 includes a caller box 128, a numeric keypad icon 130, a speakerphone icon 132, selectable controls 134, and an end call icon 136.
[0085] Suppose a user calls a new healthcare provider, a doctor's office. In this implementation, the user places the voice call using the communications application 124. In other implementations, the user can place a video call using the communications application 124 or another application on the computing device 202. The caller's box 128 displays the third-party business name (e.g., a doctor's office) and phone number (e.g., (111) 555-1234). The doctor's office uses the IVR system 110 to audibly present a menu of selectable options. The IVR system 110 can direct the caller to the appropriate personnel and staff at the doctor's office. When answering the voice call, the IVR system 110 provides the following dialog: "Thank you for calling the doctor's office. Please listen to the options below and select the option that best suits the purpose of your call today. For a prescription refill, press 1. For an appointment, press 2. For billing, press 3. If you would like to speak to a nurse, press 4."
[0086] When the IVR system 110 audibly provides the selectable options, the caption module 210 obtains audio data 304 output from the communication application 124. As described above, the audio mixer 208 can send the audio data 304 to the caption module 210. The caption module 210 then determines that the audio data 304 includes multiple selectable options. In response to this determination, the caption module 210 determines text descriptions 306 of the selectable options. For example, the machine learning model 302 can transcribe at least a portion of the selectable options. The transcription may be a word-for-word transcription or a paraphrase of each of the selectable options.
[0087] The caption module 210 then causes the display component 118 to display the selectable controls 134 on the user interface 126. The selectable controls 134 include a selectable control associated with each selectable option provided by the IVR system 110, namely, a first selectable control 134-1, a second selectable control 134-2, a third selectable control 134-3, and a fourth selectable control 134-4. The selectable controls 134 include a text description 306 associated with each selectable option. For example, the first selectable control 134-1 includes the text "1-Prescription Refill." The number "1" indicates that the first selectable control 134-1 is associated with the first selectable option provided by the IVR system 110. The second selectable control 134-2 provides the text "2-Book an Appointment." The third selectable control 134-3 displays the text "3-Charge." And the fourth selectable control 134-4 includes the text “4-Call Nurse.” In some implementations, the selectable controls 134 may omit the numbers associated with each selectable option.
[0088] As mentioned above, the selectable control 134 can be presented in a variety of forms on the user interface 126. For example, the selectable control 134 can be a button, a toggle, selectable text, a slider, a checkbox, or an icon. A user can select the selectable control 134 to cause the computing device 202 to indicate to the IVR system 110 the selected option of multiple selectable options.
[0089] In response to the IVR system 110 providing selectable options, the user can select the numeric keypad icon 130 to display a numeric keypad and select a number associated with the desired selectable option. For example, the user can select the number "2" on the numeric keypad to make a reservation. In response, the computing device 202 can transmit a DTMF tone to the IVR system 110. In other implementations, the IVR system 110 may allow the user to provide the selected option by audibly saying the number "2." The described systems and techniques also allow the user to select a selectable control 134 associated with the desired option. In this example, the user selects the second selectable control 134-2 to make a new reservation. In response to the user selecting the second selectable control 134-2, the input component 120 causes the computing device 202 to transmit a DTMF tone associated with the number "2" or an audible communication of the number "2" to the IVR system 110. In this manner, the described systems and techniques assist the user in navigating selectable IVR menu options and selecting a desired option.
[0090] In some implementations, the computing device 202 may provide a series of selectable controls 134 corresponding to different levels of the IVR menu. The computing device 202 may update the selectable controls 134 to correspond to the current selectable options. In other implementations, the computing device 202 may provide an option to display a previous menu of previous selectable options for the voice or video call.
[0091] FIG. 6B is an example of the user interface 126 responding to a request for user information. In response to the user selecting the second selectable control 134-2 in the previous scenario, the IVR system 110 directs the user to a receptionist at a medical clinic. Because the user is a new patient, the receptionist may ask a series of questions to set up an account or profile associated with the user. For example, the receptionist may request the user's medical insurance information. In such a situation, the audio data 304 may include the question, "Do you have medical insurance?" The machine learning model 302 can use the audible portion of the voice conversation with the medical clinic to determine whether the audio data 304 includes a request for user information. In this example, the machine learning model 302 can use the words "medical insurance," along with other portions of the conversation and the context that the third party is a medical clinic, to determine that the audio data 304 includes a request for user information.
[0092] The machine learning model 302 can responsively determine a text description 306 of the request for user information. In this example, the machine learning model 302 or the caption module 210 determines that the text description 306 includes "medical insurance." The caption module 210 or the computing device 202 can then identify user data in response to the request for medical insurance information in the CRM 206 and cause the display component 118 to display it on the user interface 126. In this example, the user data may include an insurance company, a policy number, or an account identifier. The computing device The computing device 202 may also retrieve health insurance information from profile information stored in an email or contacts application within an email application. In some implementations, the computing device 202 may store and retrieve sensitive user data from a secure enclave in the CRM 206 or other memory within the computing device 202.
[0093] The display component 118 can display user data (e.g., insurance company and policy number) in a message element 600 on the user interface 126. The message element 600 can be an icon, notification, message box, or similar user interface element for displaying text information. The message element 600 can also include a text description 306 of the request for user information to provide context. In this example, the message element 600 provides the text "Your Insurance Company: Apex Health Insurance Company" and "Your Policy Number: 123456789-0." In the implementation shown, the message element 600 provides both sets of user data in a single message element 600. In other implementations, the display component 118 can include user data in multiple message elements 604.
[0094] The display component 118 displays the message element 600 on the user interface 126 immediately after the receptionist asks the question. In some implementations, the computing device 202 can determine from the audio data 304 that the user is a new patient at a doctor's office. In response to this context, the machine learning model 302 or the caption module 210 can predict that the receptionist will ask for medical insurance information and retrieve this user data. In other implementations, the machine learning model 302 or the caption module 210 can predict that medical insurance information is likely to be requested when the user calls the doctor's office. In such a situation, the medical insurance information can be displayed in response to a request for this information.
[0095] The computing device 202 can use the sensors 214 to determine the context of the computing device 202. In response to determining that the user is not looking at the display, the computing device 202 can cause the audio component 116 to provide an audio signal or haptic feedback. The audio signal can alert the user that user data related to a user information request is being displayed. For example, if the computing device 202 determines (e.g., by using a proximity sensor, a gyroscope, or an accelerometer) that the user is holding the computing device 202 to their ear, the computing device 202 can cause the audio component 116 to provide an audio signal (e.g., a soft tone) that only the user can hear. In other implementations, the computing device 202 can provide haptic feedback to the user as an alert.
[0096] In response to retrieving the message element 600 containing medical insurance information, the user may audibly provide this information to the receptionist. In some situations, the user may be in a public place and may not wish to audibly provide user data. As a result, the user may select one of a number of selectable controls 134. The display component 118 displays a fifth selectable control 134-5 and a sixth selectable control 134-6. The fifth selectable control 134-5 includes the text "Read my insurance company." The sixth selectable control 134-6 reads the text "Read my insurance number." In response to the user selecting one of the selectable controls 134, the computing device 202 transmits the information to the audio mixer 208 without requiring the user to audibly provide this information. The receptionist then reads the user data audibly to the receptionist. In other implementations, the computing device 202 may provide the user with additional selectable controls 134 to email, text, or otherwise transmit user data (e.g., medical insurance information) to the receptionist. In this manner, the described techniques and systems provide a secure and private way to share sensitive user data with another person or entity during voice and video calls.
[0097] In FIG. 6C , the computing device 202 provides user data in response to a suggested appointment time. Consider a previous voice call to a doctor's office. After the user provides medical insurance information, the receptionist suggests an appointment for Tuesday at 11:00 AM. For example, the audio data 304 includes the receptionist's question, "Is next Tuesday at 11:00 AM okay?" In response to the suggested time, the computing device 202 can check the user's calendar information in a calendar application to identify potential conflicts. In this example, the user has a dental appointment scheduled for Tuesday at 11:15 AM. The computing device 202 causes the display component 118 to display this information in the message element 600. For example, the display component 118 can display the text, "Dentist appointment for 11:15 AM." In some implementations, the computing device 202 can also automatically suggest alternative times based on the user's calendar information. The display component 118 can display the text, "We have a conflict, so how about these times instead: Tuesday 9:30 AM [or] Wednesday 1:00 PM." In this way, the computing device 202 helps the user make a new appointment with the clinic. The user should not call up a previously made dentist appointment or open a calendar application on the computing device 202 while speaking with the receptionist. The user can also avoid calling the clinic again to make the appointment after remembering the conflict.
[0098] In FIG. 6D , the computing device 202 displays a communication related to a voice call. Consider a previous voice call to a doctor's office. The receptionist confirmed the appointment by saying, "Your appointment is scheduled for Wednesday, November 4th, at 1:00 PM," indicating that an appointment slot is available on Wednesday, November 4th, at 1:00 PM. In response, the computing device 202 can cause the display component 118 to display the appointment details in a message element 600. For example, the message element 600 can provide the communication, "Appointment at the doctor's office on Wednesday, November 4th, 2020, at 1:00 PM."
[0099] The computing device 202 may also provide the user with several selectable controls related to communication information, including a seventh selectable control 134-7 and an eighth selectable control 134-8. In this example, the seventh selectable control 134-7 displays the text "Save to Calendar." When selected, the seventh selectable control 134-7 causes the computing device 202 to save the appointment information to a calendar application. The eighth selectable control 134-8 displays the text "Send to Spouse." When selected, the eighth selectable control 134-8 causes the computing device 202 to send the appointment information to a spouse. The user may also cause the computing device 202 to save the appointment information to a calendar application via an audible command.
[0100] The computing device 202 can cause the display component 118 to leave the appointment-related message element 600 and selectable controls 134 on the user interface 126 until the voice call ends and for several minutes thereafter. In another implementation, the user can select the clinic conversation in the history menu of the communication application 124. 6A-6D , the computing device 202 can provide a more user-friendly experience in voice and video calls.
[0101] 7A-7C illustrate other examples of user interfaces of computing devices that assist users in voice and video calls. Figures 7A-7C are sequentially described in the context of computing device 202. Computing device 202 may provide a different user interface with fewer or additional features than those illustrated in Figures 7A-7C.
[0102] In Figure 7A, the computing device 202 causes the display component to display the user interface 126. Assume that a user uses the communication application 124 to place a voice call to their friend Amy. The caller box 128 provides Amy's name and phone number (e.g., (111) 555-6789). During the voice call, Amy asks the user for the user's new address. As shown in Figure 7A, the audio data 304 includes the phrase, "What's your new address?"
[0103] In response to determining that the audio data 304 includes a request for user information (e.g., a user address), the computing device 202 determines a description of the request. In this example, the caption module 210 determines that the text description 306 of the request includes the user's home address. The computing device 202 locates the home address in the CRM 206 and displays it on the user interface 126. For example, the display component 118 may cause a message element 700 to provide the text description 306 and the responsive user data. The message element 700 provides the information, "Your Address: 100 1st Street, San Francisco, California, Zip Code 94016." In most cases, the user will remember this user data, but may need help remembering certain details (such as the zip code).
[0104] The computing device 202 may also cause the display component 118 to display selectable controls 702. A user may audibly provide Amy with her home address. In some situations, a user may be in a public place and may not want to audibly provide their address. As a result, the user may select one of the selectable controls 702. In this example, the selectable controls 702 include a first selectable control 702-1, a second selectable control 702-2, and a third selectable control 702-3. The first selectable control 702-1 includes the text "Read my address." When selected, the first selectable control 702-1 causes the audio mixer 208 to audibly read Amy's home address without requiring the user to audibly provide this information. The second selectable control 702-2 includes the text "Text my address." When selected, the second selectable control 702-2 causes the communication application 124 or another application to send a text message with Amy's home address using the communication unit 116. The third selectable control 702-3 includes the text "Email Address." When selected, the third selectable control 702-3 causes an email application to send Amy an email with her home address. The computing device 202 can obtain Amy's email address from a contacts application. In this manner, the computing device 202 can contact Amy in a voice or video call. It provides users with a secure way to share sensitive user data without audibly broadcasting it to nearby people.
[0105] 7B, the computing device 202 displays a communication related to a voice call. Consider a previous voice call with Amy, in which Amy provides new contact information (e.g., her new work email address). In response, the computing device 202 provides the communication to the user. The caption module 210 determines that the audio data 304 includes Amy providing a new email address: "My email address is amy@email.com." The display component 118 then displays the new email address in a message element 702. The message element provides the text "Amy's email address: amy@email.com."
[0106] In some implementations, the computing device 202 may verify that the new email address is not stored on the computing device 202 (e.g., in a contacts application or an email application). If the new email address is stored, the computing device 202 may prevent the caption module 210 from displaying the communication. If the new email address is not stored, the computing device 202 may cause the caption module 210 to display the communication.
[0107] The computing device 202 may display a fourth selectable control 702-4 that includes the text "Save to Contacts." When selected, the fourth selectable control 702-4 causes the computing device 202 to save the email address in a contacts application.
[0108] In FIG. 7C , the computing device 202 provides additional selectable controls in response to a communication during a voice call. Consider that in a previous voice call with Amy, the user and Amy agreed to meet for lunch. The audio data 304 includes the phrase "Meet me at Mary's restaurant in 20 minutes," audibly spoken by the user. In response to this communication, the computing device 202 can display the address of Mary's restaurant in a message element 702. The message element 702 includes the text "Mary's Restaurant Address: 500 20th Street, San Francisco, California, Zip Code 94016." The computing device 202 can also display a fifth selectable control 702-5. The fifth selectable control 702-5 displays the text "Directions to Mary's Restaurant." When selected, the fifth selectable control 702-5 causes the computing device 202 to initiate navigation instructions from a navigation application.
[0109] In some implementations, the fifth selectable control 702-5 may be a navigation application slice window that provides a subset of the navigation application's functionality related to the communication. For example, the navigation application slice window may allow a user to select walking directions, driving directions, or public transportation directions to Mary's Restaurant.
[0110] 8A-8D illustrate other examples of user interfaces for a computing device that facilitates voice and video calling for a user. 8A-8D are sequentially described in the context of computing device 202 of FIG. 2. 202 may provide a different user interface with fewer or additional features than those shown in FIGS. 8A-8D.
[0111] 8A, in response to a selectable option from the IVR system 110, the computing device 202 causes the display component 118 to display a user interface 126 having a message element 800 and a selectable control 802. Suppose a user places a voice call to a new utility company. The caller's box 128 displays the business name (e.g., utility company) and phone number (e.g., (111) 555-2345) of the called party.
[0112] The IVR system 110 uses a voice response system that prompts the caller to provide voice responses to a series of questions and statements. Suppose the audio data 304 includes the sentence, "Thank you for contacting us about registering as a new customer. Please tell us what type of service you are interested in." The IVR system 110 can listen for phrases that match or closely match a list of services offered. For example, a utility company can listen for one of the following selectable options: home internet service, home phone, or television service. The computing device 202 can determine that the audio data 304 includes an implicit list of two or more selectable options. The display component 118 can display the following text in the message element 800: "Below is a list of common responses provided by new customers." In this example, the selectable controls 802 can include a first selectable control 802-1 (e.g., "Home Internet Service"), a second selectable control 802-2 (e.g., "Home Phone"), and a third selectable control 802-3 (e.g., "Television Service"). The selectable controls 802 may include additional or fewer suggestions. The user can select one of the selectable controls 802 to cause the audio mixer 208 to audibly provide the selected option to the IVR system 110.
[0113] The computing device 202 can determine potential offers based on the audio data 304 by decoding available services from the audible portion of the voice call. The computing device 202 can also determine selectable options based on data obtained from other computing devices that have received similar requests from the same utility or similar companies. In this way, the computing device 202 can help the user navigate open-ended IVR prompts and avoid ineffective responses or reboot the system.
[0114] FIG. 8B is an example of a user interface 126 responding to a request for user information (e.g., payment information). In response to the user selecting home Internet service, the IVR system 110 directs the user to an account specialist to set up a new account and initiate home Internet service. Because the user is a new account holder, the account specialist collects payment information, including a credit card number, to set up the account. For example, the audio data 304 may include a request from the specialist: "Please provide your preferred payment method for the new service." In response to determining that the audio data 304 includes a request for user information, the computing device 202 determines a text description 306 of the request. In this example, the caption module 210 determines that the text description 306 requests credit card information. The computing device 202 identifies the credit card information in the CRM 206 and displays the user data on the user interface 126. The response element 800 displays the following: "Your credit card information: ####-####-####-1234, [expiration date]" Contains the information "01 / 21, [PIN]789".
[0115] The computing device 202 may also determine whether the user data includes sensitive information. In response to determining that a portion of the user data is sensitive information, the computing device 202 may obscure the portion of the sensitive information (e.g., by replacing at least some digits of a credit card number with different symbols, including "#" or "*," or by omitting them). In this manner, the computing device 202 may keep the sensitive information private and less visible to others.
[0116] The display component 118 can display selectable controls 802 to maintain the confidentiality of user data. In this example, the display component 118 displays a fourth selectable control 802-4 that includes the text "Read my credit card information." When selected, the fourth selectable control 802-4 causes the computing device 202 to audibly read the entire credit card number, expiration date, and PIN to the account professional. In this manner, the computing device 202 provides a secure way for a user to share sensitive credit card information with an account professional.
[0117] In FIG. 8C , the computing device 202 displays communication information related to a voice call. Consider a previous voice call to a utility company. An account specialist provides account information (e.g., an account number and a personal identification number (PIN)) to the user. In this situation, the audio data 304 includes the sentence, "Your new account number is UTIL12345, and the PIN associated with your account is 6789." In response, the computing device 202 displays the account number and PIN in a message element 800. Specifically, the message element 802 displays, "Your account number: UTIL12345, Your PIN: 6789." The computing device 202 may provide the user with a fifth selectable control 802-5 and a sixth selectable control 802-6. The fifth selectable control 802-5 includes the text, "Save to Contact." When selected, the fifth selectable control 802-5 causes the computing device 202 to save the account number and PIN to a contacts application. The sixth selectable control 802-6 includes the text "Save to Secure Memory." When selected, the sixth selectable control 802-6 causes the computing device 202 to store the account number and PIN in secure memory that requires special privileges by an application or user to access.
[0118] In FIG. 8D , the computing device 202 displays a communication related to a previous voice call. Consider a previous voice call to a utility company. In this example, the user was unable to view the communication displayed on the user interface during or immediately after the voice call. The computing device 202 can store the message element 802, the fifth selectable control 802-5, the sixth selectable control 802-6, or a combination thereof, related to the voice call. In this manner, the user can later access the text description 306 of the communication.
[0119] The call history may provide a user interface 126 associated with each voice or video call. For example, a user interface 126 associated with a history of voice calls with a utility company may include a history element 804. The history element 804 may include history information about the voice call, including the text "Called on November 2nd."
[0120] In some situations, a user may end a voice call with a utility company and immediately place another voice call. A user may need to make a voice or video call or perform another function on the computing device 202. The computing device 202 may store the message elements 800 and selectable controls 802 associated with each voice or video call in memory associated with the communication application 124. The communication application 124 may include a call history. In this manner, the user may later retrieve the message elements 800 and selectable controls 802 associated with the voice or video call at a convenient time.
[0121] example The following section provides an example.
[0122] Example 1: A method, including: a computing device obtaining audio data output from a communications application executing on the computing device, the audio data including an audible portion of a voice or video call between a user of the computing device and a third party; the method further including: the computing device using the audible portion to determine whether the audio data includes two or more selectable options, the two or more selectable options being audibly provided by the third party during the voice or video call; the method further including: in response to determining that the audio data includes the two or more selectable options, the computing device determining a text description of the two or more selectable options, the text description providing a transcription of at least a portion of the two or more selectable options; and the method further including displaying two or more selectable controls on a display of the computing device, the two or more selectable controls configured to be selectable by a user to provide to the third party an indication of a selected option of the two or more selectable options, each of the two or more selectable controls providing the text description of a respective selectable option.
[0123] Example 2: The method of Example 1, wherein the method further includes receiving a selection of one selectable control of two or more selectable controls associated with the selected option, the selection being made by a user during the voice or video call, and the method further includes, in response to receiving the selection of the one selectable control, the computing device communicating the selected option to a third party.
[0124] Example 3: The method of Example 2, wherein communicating the selected option to the third party includes the computing device sending a voice response or dual-tone multi-frequency (DTMF) tones to the third party without audibly communicating the user-selected option.
[0125] Example 4: The method of Example 2 or 3, wherein the method further includes, in response to communicating the selected option to the third party, the computing device obtaining additional audio data output from the communication application, the additional audio data including two or more additional selectable options audibly provided by the third party during the voice or video call in response to the selected option.
[0126] Example 5: The method further includes the computing device using the audible portion to determine whether the audio data includes a request for user information, the request for user information being audibly provided by a third party during the voice or video call, the method further including the computing device using the audible portion to identify user data in response to the request for user information, and during the voice or video call, the computing device displaying the user data on a display or and wherein the user device provides the user data to a third party.
[0127] Example 6: The method of any one of the preceding examples, wherein the computing device further includes determining, using the audible portion, whether the audio data includes communication information, the communication information relating to the context of the voice or video call and provided audibly by a third party or the user during the voice or video call, and wherein in response to determining that the audio data includes communication information, the computing device further includes determining a text description of the communication information, the text description of the communication information providing a transcription of at least a portion of the communication information, and wherein the method further includes displaying another selectable control on the display, the other selectable control providing the text description of the communication information and configured to be selectable by a user to save the communication information to at least one of the computing device, the application, or another application on the computing device.
[0128] Example 7: The method of any one of the preceding examples, wherein determining the textual descriptions of the two or more selectable options includes the computing device executing a machine learning model to determine the textual descriptions of the two or more selectable options, the machine learning model being trained to determine the textual descriptions from audio data, the audio data being received from an audio mixer of the computing device.
[0129] Example 8: The method of example 7, wherein the machine learning model includes an end-to-end recurrent neural network transducer automatic speech recognition model.
[0130] Example 9: The method of any one of the preceding examples, wherein the two or more selectable options are a menu representing options for an interactive voice response (IVR) system or a voice response unit (VRU) system, and the IVR system or VRU system is configured to interact with the user and direct the user to at least one of another menu of the IVR system or VRU system, personnel associated with the third party, a department associated with the third party, a service associated with the third party, or information associated with the third party.
[0131] Example 10: The method of any one of the preceding examples, wherein the two or more selectable controls include at least one of a button, a toggle, selectable text, a slider, a checkbox, or an icon and are included in a user interface of a communication application.
[0132] Example 11: The method of any one of the preceding examples, wherein the text description includes a number associated with each of the two or more selectable options, and each of the selectable controls includes a visual representation of the number associated with each of the two or more selectable options.
[0133] Example 12: The method of any one of the preceding examples, wherein the display of the computing device includes a touch-sensitive screen, and the selectable controls are presented on the touch-sensitive screen.
[0134] Example 13: The method of any one of the preceding examples, wherein the computing device includes a smartphone, a computerized watch, a tablet device, a wearable device, or a laptop computer.
[0135] Example 14: A computing device comprising at least one processor configured to perform any one of the methods described in Examples 1-13.
[0136] Example 15: A computer-readable storage medium comprising instructions that, when executed, configure a processor of a computing device to perform any one of the methods described in Examples 1-13.
[0137] conclusion While various configurations and methods for providing selectable controls on a computing device for an IVR system have been described in feature- and / or method-specific language, it should be understood that the subject matter of the appended claims is not necessarily limited to the particular features or methods described. Rather, the particular features and methods are disclosed as non-limiting examples for providing selectable controls on a computing device for an IVR system. Furthermore, while various examples are described above, and each example has particular features, it should be understood that the particular features of one example need not be used exclusively with that example. Instead, any of the features described above and / or shown in the drawings can be combined with any of the examples in addition to, or in place of, any of the other features of those examples.
Claims
1. 1. A method comprising: and a computing device obtaining audio data output from a communication application executing on the computing device, the audio data including an audible portion of a voice or video call between a user of the computing device and a third party, the method further comprising: the computing device using the audible portion to determine whether the audio data includes two or more selectable options, the two or more selectable options being audibly provided by the third party during the voice call or the video call, the method further comprising: In response to determining that the audio data includes the two or more selectable options, the computing device determines a text description of the two or more selectable options, the text description providing a transcription of at least a portion of the two or more selectable options, the method further comprising:
1. A method comprising: displaying two or more selectable controls on a display of the computing device, the two or more selectable controls configured to be selectable by the user to provide to the third party an indication of a selected option of the two or more selectable options, each of the two or more selectable controls providing the text description of a respective selectable option.
2. The method further comprises: receiving a selection of one selectable control of the two or more selectable controls associated with the selected option, the selection being made by the user during the voice call or the video call, the method further comprising: The method of claim 1 , further comprising: in response to receiving a selection of the one selectable control, the computing device communicating the selected option to the third party.
3. 3. The method of claim 2, wherein communicating the selected option to the third party includes the computing device sending a voice response or dual-tone multi-frequency (DTMF) tones to the third party without the user audibly communicating the selected option.
4. The method further comprises:
4. The method of claim 2 or 3, further comprising: in response to communicating the selected option to the third party, the computing device obtaining additional audio data output from the communication application, the additional audio data including two or more additional selectable options audibly provided by the third party during the voice or video call in response to the selected option.
5. The method further comprises: the computing device using the audible portion to determine whether the audio data includes a request for user information, the request for user information being provided audibly by the third party during the voice call or the video call, the method further comprising: said computing device using said audible portion to identify user data in response to said request for user information; During the voice call or the video call, the computing device 10. The method of any one of the preceding claims, comprising displaying user data on the display or the computing device providing the user data to the third party.
6. The method further comprises: and the computing device using the audible portion to determine whether the audio data includes communication information related to a context of the voice or video call and audibly provided by the third party or the user during the voice or video call, the method further comprising: In response to determining that the audio data includes the communication, the computing device determines a text description of the communication, the text description of the communication providing a transcription of at least a portion of the communication, the method further comprising:
10. The method of any one of the preceding claims, comprising displaying another selectable control on the display, the another selectable control providing the text description of the communication and configured to be selectable by the user to save the communication to at least one of the computing device, the application, or another application on the computing device.
7. 10. The method of claim 1, wherein determining the textual descriptions of the two or more selectable options comprises the computing device executing a machine learning model to determine the textual descriptions of the two or more selectable options, the machine learning model being trained to determine the textual descriptions from the audio data, the audio data being received from an audio mixer of the computing device.
8. The method of claim 7 , wherein the machine learning model comprises an end-to-end recurrent neural network transducer automatic speech recognition model.
9. 10. The method of any one of the preceding claims, wherein the two or more selectable options are a menu representing options for an interactive voice response (IVR) system or a voice response unit (VRU) system, the IVR system or the VRU system being configured to interact with the user and direct the user to at least one of another menu of the IVR system or the VRU system, personnel associated with the third party, a department associated with the third party, a service associated with the third party, or information associated with the third party.
10. 10. The method of any one of the preceding claims, wherein the two or more selectable controls comprise at least one of a button, a toggle, selectable text, a slider, a checkbox, or an icon and are included in a user interface of the communication application.
11. 10. The method of any one of the preceding claims, wherein the textual description includes a number associated with each of the two or more selectable options, and each of the selectable controls includes a visual representation of the number associated with each of the two or more selectable options.
12. 10. The method of claim 1, wherein the display of the computing device includes a touch-sensitive screen, and the selectable controls are presented on the touch-sensitive screen.
13. 10. The method of any one of the preceding claims, wherein the computing device comprises a smartphone, a computerized watch, a tablet device, a wearable device, or a laptop computer.
14. A computing device comprising at least one processor configured to perform any one of the methods according to claims 1 to 13.
15. A computer-readable storage medium comprising instructions that, when executed, configure a processor of a computing device to perform any one of the methods of claims 1-13.