Selectable Controls for an Automated Speech Response System

The computing device transcribes IVR audio into text and displays selectable controls, improving navigation and interaction for users with impairments in IVR systems.

JP7701448B2Active Publication Date: 2025-07-01GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023534701
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-08
Publication Date
2025-07-01
Estimated Expiration
2040-12-08

AI Technical Summary

Technical Problem

IVR systems are difficult for users with hearing, speech, or memory impairments to navigate due to complex voice menus and the inability to easily understand or remember options.

Method used

A computing device provides selectable controls by transcribing audible IVR options into text and displaying them on a screen, allowing users to select options visually.

Benefits of technology

Enhances user experience by enabling easier navigation and interaction with IVR systems for individuals with impairments, facilitating communication through visual selection of options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701448000001
    Figure 0007701448000001
  • Figure 0007701448000002
    Figure 0007701448000002
  • Figure 0007701448000003
    Figure 0007701448000003
Patent Text Reader

Abstract

Described herein are systems and techniques that enable selectable controls for interactive voice response (IVR) systems. The described systems and techniques may determine whether audio data associated with a voice or video call between a user of a computing device and a third party includes multiple selectable options. The third party audibly provides the selectable options during the call. In response to determining that the audio data includes selectable options, the computing device may determine a text description of the multiple selectable options. The described systems and techniques may then display two or more selectable controls on a display. A user may select a selectable control to indicate a selected option of the multiple selectable options. In this manner, the described systems and techniques may improve the user experience of voice and video calls by making IVR systems easier to navigate and understand.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Background An interactive voice response (IVR) system, or phone tree, enables a caller to interact with a computerized phone system via voice input or a keypad. For example, phone systems can use IVR for purchases made with mobile phones, bank payments, services, retail orders, public services, travel information, and weather forecasts. IVR systems generally use a series of voice menus to identify and segment callers. These menus contain multiple options that can be difficult for callers to understand, navigate, or remember.

Summary of the Invention

[0002] Summary This document describes systems and techniques for providing selectable controls for an IVR system. The systems and techniques described can determine whether audio data associated with a voice call or video call between a user of a computing device and a third party includes multiple selectable options. The third party audibly provides the selectable options during the call. In response to determining that the audio data includes selectable options, the computing device can determine text descriptions of the multiple selectable options. Next, the systems and techniques described can display two or more selectable controls on a display. The user can select a selectable control to indicate a selected option among the multiple selectable options. In this way, the systems and techniques described can improve the user experience of voice calls and video calls by making the IVR system easier to navigate and understand.

[0003] The systems and techniques described can improve the ease of use when a user, such as a user with a particular communication impairment, interacts with an IVR system. As an example, the systems and techniques described can enable a user who is hard of hearing and who may otherwise find it difficult or impossible to interact with an IVR system to provide a response to the IVR system. Similarly, the systems and techniques described can enable a user who has a speech impairment and who may otherwise find it difficult or impossible to interact with an IVR system to provide a response to the IVR system. Also, the systems and techniques described can assist a user with short-term memory impairment who cannot remember the list of options provided by the IVR system to provide a response to the IVR system. Further, the systems and techniques described can improve the ease of use when a user has difficulty understanding the options provided in a voice call or video call, for example, when the voice is distorted or the user is distracted by ambient noise that is not from the voice call or video call, when the user interacts with the IVR system.

[0004] For example, a computing device obtains audio data output from a communication application running on the computing device. The audio data includes the audible portion of a voice call or video call between a user of the computing device and a third party. The computing device uses the audible portion of the voice call or video call to determine whether the audio data includes two or more selectable options. The third party audibly provides two or more selectable options during the voice call or video call. In response to determining that the audio data includes two or more selectable options, the computing device determines a text description of the two or more selectable options, where the text description provides a transcription of at least a portion of the two or more selectable options. Next, the computing device displays two or more selectable controls. The two or more selectable controls may be selectable to indicate the selected option of the two or more selectable options to the third party. Each of the two or more selectable controls provides a text description of a respective selectable option.

[0005] Other methods, configurations, and systems for providing selectable controls for an IVR system are also described herein.

[0006] This summary is provided to introduce a simplified concept for providing selectable controls for an IVR system that is further described in the detailed description and drawings. This summary is not intended to identify essential features of the claimed subject matter, nor is it intended to be used to determine the scope of the claimed subject matter.

[0007] Details of one or more aspects of a visual user interface for providing selectable controls for an IVR system are described herein with reference to the following drawings. The same numbers are used throughout the several drawings to refer to like features and components.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 7A

Figure 7B

Figure 7C

Figure 8A

Figure 8B

Figure 8C

Figure 8D

[0009] Detailed Description Overview In this specification, techniques and systems for providing selectable controls on a computing device for an IVR system are described. As described above, an IVR system enables a caller to interact with a telephone system via voice input or DTMF (Dual-Tone Multi-Frequency-Tone) generated by a keypad. The IVR system can provide a series of menus, each containing a plurality of selectable options. Voice menus can be difficult to understand and navigate for the caller. For example, some IVR systems may have many options in each menu or may have detailed options that are difficult to call. Users with hearing impairments may have difficulty or be unable to hear the options and thus may not be able to provide a response to select an option. Users with speech impairments may not be able to respond verbally to the options. Users with short-term memory impairments may not be able to remember the options provided by the IVR system when providing a response.

[0010] Consider a smartphone equipped with a communication application that enables a user to make voice or video calls. For example, a user can use the communication application to call a clinic. At the clinic, an IVR system can be used to direct the caller to the appropriate information, personnel, or department. In the initial voice menu, the user can be asked to select an appropriate language. When the language is selected, either audibly or by pressing a number associated with the desired language, the IVR system can present another option menu. For example, the IVR system can direct the caller to additional menus regarding billing, scheduling, medical questions, service providers, and questions about personnel.

[0011] Communication applications generally do not assist the user in navigating an IVR system. Instead, the communication application and the computing device typically require the user to use voice input or the keypad to invoke menu options or navigate the voice menu.

[0012] The techniques and systems described can assist a user in navigating an IVR system by providing selectable controls associated with selectable options. In particular, the techniques and systems described can obtain audio data from a voice or video call and determine whether the conversation includes two or more selectable options. In response to determining that the conversation includes selectable options, the techniques and systems described can determine a text description associated with the selectable options.

[0013] Consider the above clinic scenario. The smartphone can listen to the voice call and determine whether the clinic can audibly provide an IVR menu of selectable options. In response to determining that the clinic can audibly provide selectable options, the system and technology to be described can determine the text description of the selectable options and display selectable controls on the smartphone's display. Each of the selectable controls provides the text description of the respective selectable option. By selecting one of the selectable controls, the user can cause the smartphone to display the selected option. Thus, the technology and system to be described provide a user-friendly experience that enables smartphone users to easily navigate an IVR system and allows users who normally would not be able to interact with an IVR system to do so. The technology and system to be described are compatible with a variety of different existing IVR systems.

[0014] As a non-limiting example, a computing device can obtain audio data output from a communication application. The audio data includes the audible portion of a voice call or video call between a user of the computing device and a third party. The computing device uses the audible portion to determine whether the audio data includes two or more selectable options that are audibly provided by the third party during a voice call or video call. In response to determining that the audio data includes two or more selectable options, the computing device determines a text description of the two or more selectable options. The text description includes a transcription of at least a portion of the two or more selectable options. The computing device then displays two or more selectable controls on a display of the computing device. The two or more selectable controls provide a text description of each respective selectable option. The user can select a selectable control to indicate to the third party an option selected from the two or more selectable options.

[0015] The computing device may use information from the audio data only after the computing device has obtained explicit permission from the user of the computing device. For example, in the above-described situation where the computing device may collect audio data from voice calls and video calls, an individual user may be provided with the opportunity to provide input to control whether the program or function of the computing device can collect and utilize information. Further, an individual user may be provided with the opportunity to control what the program or function can or cannot do with the information.

[0016] This example is merely an example showing that the selectable controls for the IVR system described above improve the user experience on a computing device and enable users with communication impairments to interact with the IVR system. Other examples and implementations will be described throughout this specification. Next, this specification will describe additional examples of configurations, components, and methods for providing selectable controls for an IVR system on a computing device.

[0017] Example Environment FIG. 1 shows an example of an environment 100 that includes an example of a computing device 102 that can provide selectable controls for an IVR system. In addition to the computing device 102, the environment 100 includes a computing system 104 and a caller-side system 106. The computing device 102, the computing system 104, and the caller-side system 106 are communicatively coupled to a network 108.

[0018] The operation of the computing device 102 is described as being executed locally, but in some examples, the operation may be executed by a plurality of computing devices and systems, including additional computing devices and systems beyond those shown in FIG. 1 (e.g., the computing system 104). For example, the computing system 104, the caller-side system 106, or other devices or systems communicatively coupled to the network 108 may execute some or all of the functions of the computing device 102, and vice versa.

[0019] Computing system 104 represents any combination of one or more computers, mainframes, servers, cloud computing systems, or other types of remote computing systems that can exchange information with computing device 102 via network 108. Computing system 104 can store or provide access to additional processors, stored data, or other computing resources required by computing device 102 to implement the systems and techniques described for providing selectable controls for an IVR system on computing device 102.

[0020] Caller-side system 106 can execute an IVR system 110 and transmit and receive telephony data with computing device 102 via network 108. For example, caller-side system 106 can be a mobile phone, landline phone, laptop computer, workstation in a telephone call center, or other computing device configured to present IVR system 110 to the caller. Also, caller-side system 106 can represent any combination of a computer, computing device, mainframe, server, cloud computing system, or other type of remote computing system that can transmit information via network 108 to conduct a voice or video call between caller-side system 106 and computing device 102.

[0021] Network 108 represents any public or private communication network for transmitting data (e.g., voice communication, video communication, data packages) among computing systems, servers, and computing devices. For example, network 108 may include a public switched telephone network (PSTN), a wireless network (e.g., a cellular network, a wireless local area network (WLAN)), a wired network (e.g., a local area network (LAN), a wide area network (WAN)), an Internet Protocol (IP) telephony network (e.g., a voice-over-IP (VoIP) network), or any combination thereof. Network 108 may include network hubs, network switches, network routers, or other network devices operatively interconnected. Computing device 102, computing system 104, and originator system 106 may send and receive data over network 108 using any suitable communication technology. Computing device 102, computing system 104, and originator system 106 may be operatively coupled to network 108 using respective network links.

[0022] Computing device 102 represents any suitable computing device capable of providing selectable controls for an IVR system. For example, computing device 102 may be a smartphone that provides an input for a user to make or receive a voice or video call with an originator entity (e.g., originator system 106).

[0023] Computing device 102 includes one or more communication units 112. The communication units 112 enable the computing device 102 to communicate over a wireless network or a wired network including network 108. For example, the communication units 112 may include a transceiver for cellular phone communication or network data communication. The computing device 102 can tune the communication units 112 and support circuitry (e.g., antennas, front-end modules, amplifiers) to one or more frequency bands defined by various communication standards.

[0024] Computing device 102 includes a user interface component 114 that includes an audio component 116, a display component 118, and an input component 120. Computing device 102 also includes an operating system 122 and a communication application 124. These components of computing device 102 and other components (not shown) are operatively coupled in various ways including wired and wireless buses and links. Computing device 102 may include additional components and interfaces that are omitted from FIG. 1 for clarity.

[0025] User interface component 114 manages input / output to user interface 126 that is controlled by an operating system 122 or an application running on computing device 102. For example, communication application 124 can cause user interface 126 to display various user interface elements including input controls, navigation components, information components, or combinations thereof.

[0026] As described above, the user interface component 114 may include an audio component 116, a display component 118, and an input component 120. The audio component 116, the display component 118, and the input component 120 can be separate or integrated as a single component. The audio component 116 (e.g., a single speaker or multiple speakers) can receive an audio signal as input and convert the audio signal into an audible sound. The display component 118 can display visual elements on the user interface 126. The display component 118 can include any suitable display technology, including light-emitting diode (LED), organic light-emitting diode (OLED), and liquid crystal display (LCD) technology. The input component 120 can be a microphone, a presence sensing device, a touch screen, a mouse, a keyboard, or another type of component configured to receive user input.

[0027] The operating system 122 generally controls a computing device 102 that includes a communication unit 112, a user interface component 114, and other peripheral devices. For example, the operating system 122 can manage the hardware resources and software resources of the computing device 102 and provide services common to applications. As another example, the operating system 122 can control task scheduling. The operating system 122 and applications are generally executable by one or more processors (e.g., a system on chip (SoC), a central processing unit (CPU)) to enable communication with the computing device 102 and user interaction. The operating system 122 generally provides interaction with the user via a user interface 126.

[0028] The operating system 122 also provides an execution environment for applications such as, for example, a communication application 124. With the communication application 124, the computing device 102 can make and receive outgoing and incoming voice calls and video calls with a caller including a caller-side system 106.

[0029] During a voice call or a video call, the communication application 124 can cause the user interface 126 to display a caller-side box 128, a keypad icon 130, a speakerphone icon 132, selectable controls 134, and a call end icon 136. The caller-side box 128 can display the name and phone number of the caller (e.g., the caller-side system 106). The keypad icon 130 is a selectable icon that, when selected, causes a keypad to be displayed on the user interface 126. The speakerphone icon 132 is a selectable icon that, when selected, causes the computing device 102 to use the speakerphone function for a voice call or a video call.

[0030] The selectable control 134 can be selected by a user of the computing device 102 to perform a particular action or function. In the illustrated example, the selectable control 134 can be selected by the user to indicate an option selected from among the selectable options provided by the IVR system 110 to the originating system 106. The selectable control 134 can include buttons, toggles, selectable text, sliders, check boxes, or icons. By the call end icon 136, a user of the computing device 102 can end an audio call or a video call.

[0031] The operating system 122 can associate the input detected at the input component 120 with an element of the user interface 126. In response to receiving an input (e.g., a tap) at the input component 120, the operating system 122 or the communication application 124 can receive information regarding the detected input from the user interface component 114. The operating system 122 or the communication application 124 can perform a function or an operation in response to the detected input. For example, the operating system 122 can determine that the input corresponds to the user selecting one of the selectable controls 134, and in response, can send an indication of the corresponding selected option to the originating system 106.

[0032] During operation, the operating system 122 or the communication application 124 can automatically generate selectable controls 134 corresponding to the selectable options of the IVR system 110 provided by the sender system 106. The computing device 102 can obtain audio data from the audio mixer or sound engine of the operating system 122. The audio data generally includes the audible part of a voice call or a video call, including the IVR options provided by the IVR system 110.

[0033] Configuration example In this section, configuration examples of a system that provides selectable controls for an IVR system are described, and all or some of these may occur separately or together. In this section, various configuration examples are described, and for ease of reading, each configuration example is described in association with the drawings.

[0034] FIG. 2 shows an example of a device 200 of a computing device 202 that can provide selectable controls for an IVR system (e.g., IVR system 110). The computing device 202 is an example of the computing device 102, with some additional details.

[0035] As shown in FIG. 2, the computing device 202 may be a smartphone 202-1, a tablet device 202-2, a laptop computer 202-3, a desktop computer 202-4, a computerized wristwatch 202-5 or other wearable device, a voice assistant system 202-6, a smart display system, or a computing system installed in a vehicle.

[0036] In addition to communication unit 112 and user interface component 114, computing device 202 includes one or more processors 204 and a computer-readable storage media (CRM) 206.

[0037] Processor 204 may include any combination of one or more controllers, microcontrollers, processors, microprocessors, hardware processors, hardware processing units, digital signal processors, graphics processors, and graphics processing units. For example, processor 204 can be, by way of non-limiting example, an integrated processor and memory subsystem that includes a SoC, a CPU, a graphics processing unit, or a tensor processing unit. A SoC generally integrates many of the components of computing device 202, including a central processing unit, memory, and input / output ports, into a single device. A CPU generally executes the commands and processing required for computing device 202. A graphics processing unit performs operations for displaying the graphics of computing device 202 and can perform other specific computational tasks. A tensor processing unit generally performs symbolic match operations in neural network machine learning applications. Processor 204 may include a single core or multiple cores.

[0038] CRM206 can provide persistent and non-persistent storage for executable instructions (e.g., firmware, recovery firmware, software, applications, modules, programs, functions) and data (e.g., user data, operation data) to support the execution of the executable instructions to computing device 202. For example, when executed by processor 204, CRM206 includes instructions to execute operating system 122 and communication application 124. Examples of CRM206 include volatile and non-volatile memories, fixed media devices and removable media devices, and any suitable memory device or electronic data storage that holds executable instructions and support data. CRM206 excludes propagated signals. CRM206 can be a solid-state drive (SSD) or a hard disk drive (HDD).

[0039] Operating system 122 can include or control audio mixer 208 and caption module 210. Audio mixer 208 and caption module 210 can be dedicated hardware components, software components, or combinations thereof. In other examples, audio mixer 208 and caption module 210 are separate from operating system 122 (e.g., as a system plugin or additional add-on service installed locally on computing device 202).

[0040] The audio mixer 208 can obtain and integrate audio data generated by an application, including a communication application 124, executed on the computing device 202. The audio mixer 208 obtains an audio stream from an application such as the communication application 124, and when integrated and output from the integrated audio component 116, generates an audio output signal that reproduces the sound encoded in the audio stream. The audio mixer 208 can adjust the audio signal in other ways, such as controlling focus, intent, and volume. The audio mixer provides an interface between an application source that generates content and an audio component 116 that generates sound from the content. The audio mixer 208 can manage raw audio data, analyze it, and direct it to be output as an audio signal by the audio component 116 or transmitted as an audio signal to another computing device (e.g., the caller-side system 106) via the communication unit 112.

[0041] The caption module 210 is configured to analyze the raw audio data received by the audio mixer 208 (e.g., as a byte stream). For example, the caption module 210 can perform speech recognition on the audio data to determine whether the audio data includes selectable options of the IVR system, requests for user information, or transmission information related to the call context. Instead of processing each audio signal, the caption module 210 can identify individual pre-mixed audio data streams suitable for captioning. For example, the caption module 210 can automatically caption the audio data of spoken words, but not for notification or sonification audio data (e.g., system beep sounds, ringing tones). The caption module 210 can apply a filter to the byte stream received by the audio mixer 208 to identify audio data suitable for captioning. The caption module 210 can use a machine learning model to determine a description of the audio data from the audible portion of a voice call or video call.

[0042] Rather than captioning all audio data, the operating system 122 can use metadata to focus captioning on specific portions of the audio data. For example, the caption module 210 can focus on audio data related to providing selectable controls of the IVR system, user information in response to requests, or transmission information related to the call context. In other words, the operating system 122 can identify "captionable" audio data based on the metadata and refrain from captioning all audio data. Some examples of metadata include context indicators that specify the content of a voice call or video call. The audio mixer can use the context indicator to control routing, focusing, and captioning decisions regarding the audio data.

[0043] Some computing devices can transcribe voice or video calls. However, the transcription typically provides a direct transcription of the audible part of the call and cannot determine whether the conversation includes selectable options of an IVR system, requests for user information, or communicated information related to the context of the call. The user still has to read the transcript to determine the desired menu option, the requested user information, or the communicated information. Thus, even if a computing device provides a transcription, the user may still find it difficult to navigate an IVR system and select the desired option. In contrast, the systems and techniques described assist the user in navigating an IVR system, providing user information in response to requests, and managing communicated information from voice and video calls by displaying selectable controls and message elements along with relevant information.

[0044] Computing device 202 also includes one or more sensors 214. Sensors 214 obtain context information indicating the physical operating environment of computing device 202 or the characteristics of computing device 202 while it is operating in the physical operating environment. For example, caption module 210 can use this context information as metadata to focus on audio data processing. Examples of sensors 214 include motion sensors, temperature sensors, position sensors, proximity sensors, ambient light sensors, moisture sensors, and pressure sensors, among others.

[0045] During operation, the operating system 122 or the caption module 210 determines whether the audio data is for captioning. For example, the caption module 210 can determine whether the audio data includes selectable options of the IVR system, requests for user information, or transmission information related to the call context. In response to determining that the audio data is for captioning, the operating system 122 determines a description of the audio data. For example, the operating system 122 can execute a machine learning model (e.g., an end-to-end recurrent neural network transducer automatic speech recognition model) trained to generate a description of the audible portion of a voice call or a video call. The machine learning model can be any type of model suitable for learning a description of the voice, including a transcription of the spoken voice. Since the machine learning model used by the operating system 122 only needs to be trained to identify the audible portions of voice calls and video calls, it may be smaller and less complex than other machine learning models. The machine learning model can avoid processing all audio data transmitted to the audio mixer 208. Thus, the systems and techniques described can avoid using remote processing resources (e.g., machine learning models in remote computing devices) to avoid unnecessary privacy risks and potential processing latency.

[0046] Rather than depending on the audio signal generated by the audio component 116, by depending on the original audio data, the machine learning model can generate a description that more accurately represents the audible portions of voice calls and video calls. Before using the machine learning model, by determining whether the audio data is for captioning, the operating system 122 can avoid wasting resources by overly analyzing all the audio data output by the communication application 124. By making such a caption determination, the computing device 202 can execute a more efficient, smaller, and less complex machine learning model. In this way, the machine learning model can locally execute automatic speech recognition technology and automatic speech classification technology to maintain privacy.

[0047] The operating system 122 receives the description of the machine learning model and displays it using the display component 118. The display component 118 can also display other visual elements related to the description (for example, selectable controls that enable the user to perform actions on the computing device 202). For example, the operating system 122 can present visual elements (such as selectable control 134) as part of the user interface 126. The description can include a transcription or summary of the audible portions of voice calls and video calls (for example, a phone call). The description can also identify the context of the audible portion of the audio data. The details and operation of the machine learning model are described in more detail with respect to FIG. 3.

[0048] FIG. 3 is a diagram 300 showing an example of a machine learning model 302 of a computing device 202 that can provide text descriptions for selectable controls in response to an IVR system. In other implementations, the computing device 202 may be the computing device 102 of FIG. 1 or a similar computing device.

[0049] As shown in FIG. 3, the machine learning model 302 can be part of the caption module 210. The machine learning model 302 can convert the audio data 304 into a text description 306 of the audible portion of a voice call or video call (e.g., a text description of selectable options provided by the IVR system 110) without converting the audio data 304 into sound. The audio data 304 can include different types, forms, or variations of data from the communication application 124. For example, the audio data 304 can include raw, pre-mixed voice byte stream data or processed byte stream data. The machine learning model 302 can include multiple machine learning models combined into a single model that provides the text description 306 in response to the audio data 304.

[0050] An application including the communication application 124 can use the machine learning model 302 to process the audio data 304 into the text description 306. For example, the communication application 124 can communicate with the machine learning model 302 via the operating system 122 or the caption module 210 using an application programming interface (API) (e.g., a public API across all applications). In some implementations, the machine learning model 302 can process the audio data 304 within a secure section or enclave of the operating system 122 or the CRM 206 to ensure user privacy and security.

[0051] The machine learning model 302 can perform inferences. In particular, the machine learning model 302 can be trained to receive audio data 304 as input and provide a text description 306 of the audible portion of a call as output data. By using the machine learning model 302 to perform inferences, the caption module 210 can locally process the audio data 304. The machine learning model 302 can also perform classification, regression, clustering, anomaly detection, generation of recommendations, and other tasks.

[0052] Engineers can train the machine learning model 302 using supervised learning techniques. For example, engineers can use training data 308 (e.g., ground truth data) that includes examples of descriptions inferred from examples of audio data 304 from a series of voice calls and video calls to train the machine learning model 302. The inferences can be applied manually by an engineer or other expert, generated through crowdsourcing, or provided by other techniques (e.g., sophisticated speech recognition algorithms and content recognition algorithms). The training data 308 can include audio data from voice calls and video calls for the audio data 304. As an example, assume that the audio data 304 includes a voice call with an IVR system used in a clinic. The training data 308 for the machine learning model 302 can include a number of audio data files from a wide range of voice calls and video calls with the IVR system. As another example, assume that the audio data 304 includes a voice call with a corporate customer service representative. The training data 308 may include many audio data files from a wide range of similar voice calls and video calls. Engineers can also train the machine learning model 302 using unsupervised learning techniques.

[0053] The machine learning model 302 is trained in a training computing system and can then be provided for storage and implementation on one or more computing devices 202. For example, the training computing system can include a model trainer. It is also possible to include the training computing system in, or separately from, the computing device 202 that implements the machine learning model 302.

[0054] An engineer can train the machine learning model 302 either online or offline. In offline training (e.g., batch learning), the engineer trains the machine learning model 302 on the entire static set of training data 308. In online learning, the engineer continuously trains the machine learning model 302 as new training data 308 becomes available (e.g., while the machine learning model 302 is being used on the computing device 202 to perform inferences). For example, the engineer can first train the machine learning model 302 to replicate descriptions applied to audible portions of voice calls and video calls (e.g., a captioned IVR system, a captioned conference call). When the machine learning model 302 infers a text description 306 from the audio data 304, the computing device 202 can provide the text description 306 (and the corresponding portion of the audio data 304) as new training data 308 back to the machine learning model 302. In this way, the machine learning model 302 can continuously improve the accuracy of the text description 306. In some implementations, a user of the computing device 202 can provide an input to the machine learning model 302 to flag an error in a particular description. The computing device 202 can use this flag to train the machine learning model 302 and improve future predictions.

[0055] An engineer or trainer can perform centralized training of multiple machine learning models 302 (e.g., based on a centrally stored dataset). In other implementations, a trainer or engineer can use distributed training techniques, including distributed training or federated learning, to train, update, or customize a machine-learned model 302. An engineer may use user information to customize a machine learning model 302 only after receiving explicit permission from the user. For example, in situations where a computing device 202 may collect user information, an opportunity may be provided to individual users to provide an input for controlling whether the program or function of the machine learning model 302 can collect and use user information. Additionally, individual users may be provided with an opportunity to control what can or cannot be done with user information by the program or function.

[0056] The machine learning model 302 can be or include one or more artificial neural networks. In such an implementation, the machine learning model 302 can include a group of connected or not fully connected nodes (e.g., neurons). An engineer can also organize the machine learning model 302 into one or more layers (e.g., a deep network). In an example of a deep network, the machine learning model 302 can include an input layer, an output layer, and one or more hidden layers disposed between the input layer and the output layer.

[0057] The machine learning model 302 may also include one or more recurrent neural networks. For example, the machine learning model 302 may be an end-to-end recurrent neural network transducer automatic speech recognition model. Examples of recurrent neural networks include long short-term memory (LSTM) recurrent neural networks, gated recurrent units, bidirectional recurrent neural networks, continuous time recurrent neural networks, neural history compression programs, echo state networks, Elman networks, Jordan networks, recursive neural networks, Hopfield networks, fully recurrent networks, and sequence-to-sequence architectures.

[0058] At least some of the nodes of the recurrent neural network can form a cycle. When configured as a recurrent neural network, the machine learning model 302 can be particularly useful for processing sequential input data (e.g., audio data 304). For example, a recurrent neural network can use recurrent or directed cyclic node connections to pass or store information from a previous portion of the audio data 304 to a subsequent portion of the audio data 304.

[0059] The audio data 304 can also include time series data (e.g., audio data over time). As a recurrent neural network, the machine learning model 302 can detect or predict the speech of spoken words and related non-spoken words in order to analyze the audio data 304 over time and generate a text description 306 of at least a portion of the audio data 304. For example, continuous sounds from the audio data 304 can indicate spoken words in a sentence (e.g., natural language processing, speech detection, or processing).

[0060] The machine learning model 302 may also include one or more convolutional neural networks. A convolutional neural network may include multiple convolutional layers that perform convolutions on input data using learned filters or kernels. Engineers generally use convolutional neural networks to diagnose visual problems in still images or videos. Engineers can also apply convolutional neural networks to natural language processing of audio data 304 to generate text descriptions 306.

[0061] In this specification, the operations of the caption module 210 and the machine learning model 302 will be described in more detail with respect to FIG. 4.

[0062] Example method FIG. 4 is a flowchart illustrating an example of an operation 400 of a computing device that can provide selectable controls and user data related to voice calls and video calls. The operation 400 will be described below in the context of the computing device 202 of FIG. 2. In other implementations, the computing device 202 may be the computing device 102 of FIG. 1 or a similar computing device. The operation 400 may be performed in an order different from that shown in FIG. 4, and may be performed with additional operations or fewer operations.

[0063] At 402, the computing device optionally obtains content that includes user information of the computing device user. The computing device can use the user information to assist in searching for information requested by the user or saving communication information related to voice calls and video calls. Before obtaining the user information or before performing the options described below, the computing device 202 may obtain consent from the user to use the user information for voice calls and video calls. For example, the computing device 202 may use the user information only after receiving explicit consent. The computing device 202 can obtain user information from the user's input to an application on the computing device 202 (for example, input of contact information to a user profile, input of an account number via a third-party application), or it can learn it from information received by the application (for example, an account number included in a specification sent by email, a saved calendar item).

[0064] At 404, the computing device displays a graphical user interface of a communication application. For example, in response to the user making or receiving a voice call or a video call, the computing device 202 may instruct the display component 118 to display the user interface 126 of the communication application 124.

[0065] At 406, the computing device obtains audio data output from a communication application running on the computing device. The audio data includes the audible portion of a voice call or video call. For example, communication application 124 enables a user of computing device 202 to initiate and receive voice calls and video calls. Audio mixer 208 obtains audio data 304 output from communication application 124 during a voice call and during a video call. Audio data 304 includes the audible portion of a voice call or video call between a user of computing device 202 and a third party. To provide selectable controls and other information to the user during a voice call or video call, caption module 210 can extract audio data 304 from audio mixer 208.

[0066] At 408, the computing device determines whether the audio data contains relevant information using the audible portion of a voice call or video call. The relevant information can be two or more selectable options of an IVR system (e.g., phone tree options), a request for user information (e.g., a request for a credit card number, address, account number), or transmitted information (e.g., reservation details, contact information, account information). For example, the caption module 210 can determine whether the audio data 304 contains relevant information using the machine learning model 302. The relevant information can include two or more selectable options of an IVR system, a request for user information, or transmitted information. The user or a third party audibly provides the relevant information during a voice call or video call. The caption module 210 or the machine learning model 302 can filter out audio data 304 that does not require processing, such as notification sounds and background noise. Examples of the machine learning model 302 determining whether the audio data 304 contains two or more selectable options are shown in FIGS. 6A and 8A. Examples of the machine learning model 302 determining whether the audio data 304 contains a request for user information are shown in FIGS. 6B, 6C, 7A, and 8B. Examples of the machine learning model 302 determining whether the audio data 304 contains transmitted information are shown in FIGS. 6D, 7B, 7C, and 8C.

[0067] If the audio data does not contain relevant information, at 416, the computing device displays the user interface of the communication application. For example, in response to determining that the audio data 304 does not contain relevant information, the computing device 202 displays the user interface 126 of the communication application 124.

[0068] If it is determined that the audio data contains relevant information, at 410, the computing device determines the text description of the relevant information. The text description transcribes the relevant information. For example, the caption module 210 can perform speech recognition on the audio data 304 using the machine learning model 302 to determine the text description 306 of the relevant information. The text description 306 provides a transcription of at least a part of two or more selectable options, user information requests, or communicated information. Examples of the machine learning 302 determining the text description 306 of two or more selectable options are shown in FIGS. 6A and 8A. Examples of the machine learning model 302 determining the text description 306 of user information requests are shown in FIGS. 6B, 6C, 7A, and 8B. Examples of the machine learning model 302 determining the text description of communicated information are shown in FIGS. 6D, 7B, 7C, and 8C.

[0069] The caption module 210 can improve the accuracy of the text description 306 in various ways, including biasing the machine learning model 302 based on the context of the computing device 202. For example, the caption module 210 can bias the machine learning model 302 based on the identity of a third party in a voice call or video call. Suppose the user of the computing device 202 makes a voice call to a clinic. The caption module 210 can bias the machine learning model 302 using common words from the clinic conversation. In this way, the computing device 202 can improve the text description 306 of this voice call. The caption module 210 can use other context information types, including location information obtained from the sensor 214 and information from other applications, to bias the machine learning model 302.

[0070] In some implementation examples, the computing device 202 can translate the text description 306 into another language before displaying it. For example, the caption module 210 can determine the user's desired language from the operating system 122 and translate the text description 306 into the desired language. In this way, a Japanese user can view the text description 306 in Japanese even if the audio data 304 is in a different language (e.g., Chinese or English).

[0071] In 412, the computing device, optionally, identifies user data in response to a request for user information. The computing device does not perform this operation if the audio data does not include a request for user information. For example, in response to determining that a third party has requested user information, the computing device 202 can identify user data in response to the user information request. The computing device 202 can retrieve user data from the CRM 206, the communication application 124, another application on the computing device 202, or a remote computing device associated with the user or the computing device 202. Consider the above clinic call scenario. The clinic receptionist can request that the user provide insurance information. In response, the computing device 202 can retrieve the medical insurance company and the user account number from an email previously received by the user and stored on the computing device 202. Examples of the computing device 202 identifying a user data response to a request for user information are shown in FIGS. 6B, 6C, 7A, and 8B.

[0072] A computing device may use information in response to a request for user information only after receiving explicit permission from the user of the computing device. For example, in the above-described situation where the computing device may collect user data, individual users may be provided with the opportunity to provide an input for controlling whether the program or function of the computing device can collect and use user data. Further, individual users may be provided with the opportunity to control what can or cannot be done with user data by the program or function.

[0073] At 414, the computing device displays user data or selectable controls. The selectable controls are user-selectable and include text descriptions. Suppose a request for user information is included in the audio data. In this scenario, the computing device can display the specified user data. Suppose two or more selectable options of an IVR system are included in the audio data. In this scenario, the user can use the selectable control to indicate to a third party an option selected from the two or more selectable options. Suppose transmission information is included in the audio data. In this scenario, the user can use the selectable control to save the transmission information to the computing device, a communication application, or another application. For example, the computing device 202 can cause the display component 118 to display user data or a selectable control 134. The display component 118 can provide the user data as a text notification on the user interface 126. Consider the above clinic call scenario. The display component 118 can display medical insurance company and user account information as a text box on the user interface 126 during a voice call. The display component 118 can also provide a selectable control 134. The display component 118 can provide a text description 306 or requested information as part of a button on the user interface 126 of the communication application 124. Examples of the display component 118 that displays the selectable control 134 are shown in FIGS. 6A and 8A. Examples of the display component 118 that displays user data are shown in FIGS. 6B, 6C, 7A, and 8B. Examples of the display component 118 that displays the selectable control 134 and user data in response to transmission information are shown in FIGS. 6D, 7B, 7C, and 8C.

[0074] Suppose a clinic uses the IVR system 110 to direct a voice call to a receptionist. The display component 118 can display selectable controls 134. The selectable controls 134 provide text descriptions 318 for each of two or more selectable options provided by the IVR system 110. A user can use the selectable controls 134 to indicate to the clinic an option selected from the two or more selectable options.

[0075] Also, suppose a user wants to make a reservation at the clinic. The display component 118 can display selectable controls 134. The selectable controls 134 include text descriptions of the reservation. A user can use the selectable controls 134 to save reservation details to a calendar application.

[0076] At 416, the computing device displays a user interface of a communication application. For example, the display component 118 can display a user interface 126 associated with the communication application 124. The user interface 126 can include user data and selectable controls 134.

[0077] FIG. 5 shows an example of an operation 500 for providing selectable controls for an IVR system. The operation 500 is described in the context of the computing device 202 of FIG. 2. The operation 500 may be performed in a different order or with additional or fewer operations.

[0078] At 502, the computing device obtains audio data output from a communication application running on the computing device. The audio data includes the audible portion of a voice call or video call between the user of the computing device and a third party. For example, the audio mixer 208 of the computing device 202 can obtain the audio data 304 output from the communication application 124 running on the computing device 202. The caption module 210 can receive the audio data 304 from the audio mixer 208. The audio data 304 includes the audible portion of a voice call or video call between the user of the computing device 202 and a third party (e.g., a person, a computerized IVR system).

[0079] At 504, the computing device determines, using the audible portion, whether the audio data includes two or more selectable options. The third party audibly provides two or more selectable options during a voice call or video call. For example, the machine learning model 302 of the caption module 210 can determine, using the audible portion of the audio data 304, whether the audio data 304 includes two or more selectable options (e.g., IVR menu or numbered options of a phone tree). The third party audibly provides two or more selectable options during a voice call or video call.

[0080] In 506, in response to determining that the audio data includes two or more selectable options, the computing device determines a text description of the two or more selectable options. The text description provides a transcription of at least a portion of the two or more selectable options. For example, in response to determining that the audio data 304 includes two or more selectable options, the machine learning model 302 determines a text description 306 of the two or more selectable options. The text description 306 provides a transcription of at least a portion of the two or more selectable options. In some implementations, the text description 306 includes a word-by-word transcription of the two or more selectable options. In other implementations, the text description 306 provides a paraphrase of the two or more selectable options.

[0081] In 508, the computing device displays two or more selectable controls. The two or more selectable controls are selectable by a user to indicate a selected option among the two or more selectable options to a third party. Each of the two or more selectable controls provides a text description of a respective selectable option. For example, the display component 118 displays two or more selectable controls 134 on the display of the computing device 202. The display includes the user interface 126. The two or more selectable controls 134 are selectable by a user to provide an indication of a selected option among the two or more selectable options to a third party. Each of the two or more selectable controls provides the text description 306 of a respective selectable option.

[0082] Implementation example In this section, illustrative examples of systems and techniques are described that can assist a user in voice calls and video calls, all or some of which may occur separately or together. In this section, various illustrative examples are described and, for ease of reading, each is outlined in association with a particular drawing.

[0083] Figures 6A - 6D are diagrams showing examples of user interfaces of a computing device that assist a user in voice calls and video calls. Figures 6A - 6D are described in the context of the computing device 202 of FIG. 2 in sequence. The computing device 202 may provide a different user interface with fewer functions or additional functions than those shown in Figures 6A - 6D.

[0084] In FIG. 6A, the computing device 202 causes the display component 118 to display the user interface 126. The user interface 126 is associated with the communication application 124. The user interface 126 includes a caller - side box 128, a keypad icon 130, a speakerphone icon 132, selectable controls 134, and a call - end icon 136.

[0085] Suppose a user calls a hospital, which is a new medical service provider. In this implementation example, the user makes a voice call using the communication application 124. In other implementation examples, the user can make a video call using the communication application 124 or other applications on the computing device 202. The caller ID box 128 displays the name of a third-party business (e.g., the hospital) and the phone number (e.g., (111) 555-1234). The hospital uses the IVR system 110 to audibly provide a menu of selectable options. The IVR system 110 can direct the caller to the appropriate personnel and staff at the hospital. When the IVR system 110 answers the voice call, it provides a dialog such as "Thank you for calling the hospital. Please listen to the following options and select the option that best suits the purpose of your call today. If you need a prescription refill, press 1. If you need to make an appointment, press 2. If you have a billing question, press 3. If you would like to speak with a nurse, press 4."

[0086] When the IVR system 110 audibly provides selectable options, the caption module 210 obtains the audio data 304 output from the communication application 124. As described above, the audio mixer 208 can send the audio data 304 to the caption module 210. The caption module 210 then determines that the audio data 304 includes multiple selectable options. In response to this determination, the caption module 210 determines the text description 306 of the selectable options. For example, the machine learning model 302 can transcribe at least some of the selectable options. The transcription can be a word-for-word transcription of each of the selectable options or a paraphrase.

[0087] The caption module 210 then causes the display component 118 to display selectable controls 134 on the user interface 126. The selectable controls 134 include selectable controls associated with each of the selectable options provided by the IVR system 110, namely, a first selectable control 134-1, a second selectable control 134-2, a third selectable control 134-3, and a fourth selectable control 134-4. The selectable controls 134 include text descriptions 306 associated with the respective selectable options. For example, the first selectable control 134-1 includes the text "1 - Prescription refill". The number "1" indicates that the first selectable control 134-1 is associated with the first selectable option provided by the IVR system 110. The second selectable control 134-2 provides the text "2 - Appointment". The third selectable control 134-3 displays the text "3 - Billing". And the fourth selectable control 134-4 includes the text "4 - Talk to a nurse". In some implementations, the selectable controls 134 can omit the numbers associated with each selectable option.

[0088] As described above, the selectable controls 134 can be presented in various forms on the user interface 126. For example, the selectable controls 134 can be buttons, toggles, selectable text, sliders, checkboxes, or icons. The user can select a selectable control 134 to cause the computing device 202 to instruct the IVR system 110 of the selected option among the plurality of selectable options.

[0089] In response to the IVR system 110 providing selectable options, the user can select the numeric keypad icon 130 to display the numeric keypad and select the number associated with the desired selectable option. For example, the user can select the number "2" on the numeric keypad to make a reservation. In response, the computing device 202 can send DTMF tones to the IVR system 110. In other implementations, the IVR system 110 may enable the selected option to be provided when the user audibly says the number "2". Also, with the systems and techniques described, the user can select a selectable control 134 associated with the desired option. In this example, the user selects the second selectable control 134-2 to make a new reservation. In response to the user selecting the second selectable control 134-2, the input component 120 causes the computing device 202 to send a DTMF tone associated with the number "2" or an audible communication of the number "2" to the IVR system 110. Thus, the systems and techniques described assist the user in navigating selectable IVR menu options and selecting the desired option.

[0090] In some implementations, the computing device 202 can provide a series of selectable controls 134 depending on the different levels of the IVR menu. The computing device 202 can update the selectable controls 134 to correspond to the current selectable options. In other implementations, the computing device 202 can provide the option to display a previous menu of selectable options from a previous voice call or video call.

[0091] Figure 6B is an example of the user interface 126 that responds to a request for user information. In response to the user selecting the second selectable control 134-2 in the previous scenario, the IVR system 110 directs the user to the hospital receptionist. Since the user is a new patient, the receptionist may ask a series of questions to set up an account or profile related to the user. For example, the receptionist may ask for the user's medical insurance information. In such a situation, the audio data 304 may include the question "Are you covered by medical insurance?" The machine learning model 302 can use the audible portion of the voice call with the hospital to determine whether the audio data 304 contains a request for user information. In this example, the machine learning model 302 can use the word "medical insurance" along with other parts of the conversation and the context that the third party is a clinic to determine that the audio data 304 contains a request for user information.

[0092] In response, the machine learning model 302 can determine a text description 306 of the request for user information. In this example, the machine learning model 302 or the caption module 210 determines that the text description 306 includes "medical insurance". The caption module 210 or the computing device 202 can then identify user data in response to the request for medical insurance information in the CRM 206 and display it on the user interface 126 on the display component 118. In this example, the user data may include an insurance company, an insurance contract number, or an account identifier. The computing device 202 can also retrieve medical insurance information from an email in an email application or profile information stored in a contacts application. In some implementations, the computing device 202 can store and retrieve sensitive user data from a secure enclave of the CRM 206 or other memory within the computing device 202.

[0093] The display component 118 can display user data (e.g., insurance company and insurance contract number) on the message element 600 on the user interface 126. The message element 600 can be an icon, a notification, a message box, or a similar user interface element for displaying text information. The message element 600 can also include a text description 306 of the request for user information to provide context. In this example, the message element 600 provides the text "Your insurance company: Apex Medical Insurance Company" and "Your insurance number: 123456789-0". In the illustrated implementation, the message element 600 provides both user data sets with a single message element 600. In other implementations, the display component 118 can include user data in a plurality of message elements 604.

[0094] The display component 118 displays the message element 600 on the user interface 126 immediately after the receptionist asks the question. In some implementations, the computing device 202 can determine from the audio data 304 that the user is a new patient at the hospital. In response to this context, the machine learning model 302 or the caption module 210 can predict that the receptionist will ask for medical insurance information and retrieve this user data. In other implementations, the machine learning model 302 or the caption module 210 can predict that medical insurance information may be requested when the user calls the clinic. In such a situation, the medical insurance information can be displayed in response to the request for this information.

[0095] Computing device 202 can use sensor 214 to determine the context of computing device 202. In response to determining that the user is not looking at the display, computing device 202 can cause audio component 116 to provide an audio signal or haptic feedback. The audio signal can warn the user that user data related to a user information request is being displayed. For example, if computing device 202 determines that the user has the computing device 202 against their ear (such as by using a proximity sensor, gyroscope, or accelerometer), computing device 202 can cause audio component 116 to provide an audio signal (such as a soft tone) that only the user can hear. In other implementations, computing device 202 can provide haptic feedback to the user as a warning.

[0096] In response to reading the message element 600 containing medical insurance information, the user can audibly provide this information to the receptionist. Depending on the situation, the user may be in a public place and may not want to audibly provide user data. As a result, the user can select one of the plurality of selectable controls 134. The display component 118 displays a fifth selectable control 134-5 and a sixth selectable control 134-6. The fifth selectable control 134-5 includes the text "Read my insurance company". The sixth selectable control 134-6 reads the text "Read my insurance number". In response to the user selecting one of the selectable controls 134, the computing device 202 causes the audio mixer 208 to audibly read each user data to the receptionist without asking the user to audibly provide this information. In other embodiments, the computing device 202 can provide the user with additional selectable controls 134 for sending user data (e.g., medical insurance information) to the receptionist by email, text, or other means. Thus, the described techniques and systems provide a secure and private way to share sensitive user data with another person or entity during voice and video calls.

[0097] In FIG. 6C, computing device 202 provides user data in response to the proposed reservation time. Consider the previous voice call to the hospital. After the user provides medical insurance information, the receptionist proposes a reservation for Tuesday at 11:00 am. For example, audio data 304 includes a question from the receptionist, "Is Tuesday at 11:00 am next week okay?" In response to the proposed time, computing device 202 can check the user's calendar information in the calendar application and identify the possibility of schedule conflicts. In this example, the user has a dental appointment at 11:15 am on Tuesday. Computing device 202 causes display component 118 to display this information in message element 600. For example, display component 118 can display the text "Dental appointment at 11:15 am". In some implementations, computing device 202 can also automatically propose alternative times based on the user's calendar information. Display component 118 can display the text "There is a schedule conflict, so how about these alternative times: Tuesday at 9:30 am [or] Wednesday at 1:00 pm". In this way, computing device 202 helps the user make a new reservation at the hospital. While speaking with the receptionist, the user should not call the previously made dental appointment or open the calendar application on computing device 202. Also, after remembering the schedule conflict, the user can avoid calling the hospital again to make a reservation.

[0098] In FIG. 6D, computing device 202 displays communication information related to a voice call. Consider the previous voice call to the hospital. The receptionist noted that there was an available reservation slot at 1:00 PM on Wednesday and confirmed the reservation saying, "You have made a reservation for 1:00 PM on Wednesday, November 4th." In response, computing device 202 can cause display component 118 to display the details of the reservation in message element 600. For example, message element 600 can provide communication information such as "Medical appointment at the hospital on Wednesday, November 4, 2020, at 1:00 PM."

[0099] Computing device 202 can also provide the user with several selectable controls related to the communication information, including a seventh selectable control 134-7 and an eighth selectable control 134-8. In this example, the seventh selectable control 134-7 displays the text "Save to Calendar". When selected, the seventh selectable control 134-7 causes computing device 202 to save the reservation information to the calendar application. The eighth selectable control 134-8 displays the text "Send to Spouse". When selected, the eighth selectable control 134-8 causes computing device 202 to send the reservation information to the user's spouse. The user can also cause computing device 202 to save the reservation information to the calendar application via an audible command.

[0100] Computing device 202 can leave reservation-related message element 600 and selectable control 134 on user interface 126 on display component 118 until the voice call ends and for several minutes thereafter. In other implementations, the user can retrieve this information including message element 600 and selectable control by selecting the conversation with the hospital in the history menu of communication application 124. Thus, the user can save the transmitted information from a voice call or video call without writing down the schedule, recalling the schedule later, or separately entering the schedule into the calendar application. With the features and functions described with respect to FIGS. 6A-6D, computing device 202 can provide a more user-friendly experience in voice calls and video calls.

[0101] FIGS. 7A-7C show other examples of the user interface of a computing device that assists a user in voice calls and video calls. FIGS. 7A-7C are described in the context of computing device 202 in sequence. Computing device 202 may provide a different user interface with fewer or additional functions than those shown in FIGS. 7A-7C.

[0102] In FIG. 7A, computing device 202 causes user interface 126 to be displayed on the display component. Assume that the user uses communication application 124 to make a voice call to friend Aimee. The caller box 128 provides Aimee's name and phone number (e.g., (111) 555-6789). During the voice call, Aimee asks the user for the user's new address. As shown in FIG. 7A, the audio data 304 includes the phrase "What's your new address?".

[0103] In response to determining that the audio data 304 includes a request for user information (e.g., a user address), the computing device 202 determines the description of the request. In this example, the caption module 210 determines that the text description 306 of the request includes the user's home address. The computing device 202 locates the home address within the CRM 206 and displays it on the user interface 126. For example, the display component 118 can cause the message element 700 to be provided with the text description 306 and the corresponding user data. The message element 700 provides information such as "Your address: 100 First Street, San Francisco, California 94016". In most cases, the user remembers this user data, but may need help remembering specific details (such as the postal code).

[0104] Computing device 202 can also cause display component 118 to display selectable controls 702. The user can audibly provide their home address to Amy. Depending on the situation, the user may be in a public place and may not want to audibly provide their address. As a result, the user can select one of the selectable controls 702. In this example, the selectable controls 702 include a first selectable control 702-1, a second selectable control 702-2, and a third selectable control 702-3. The first selectable control 702-1 includes the text "Read my address". When selected, the first selectable control 702-1 causes audio mixer 208 to audibly read the user's home address to Amy without requiring the user to provide this information audibly. The second selectable control 702-2 includes the text "Send address as text". When selected, the second selectable control 702-2 causes communication application 124 or another application to use communication unit 116 to send a text message with the user's home address to Amy. The third selectable control 702-3 includes the text "Send address by email". When selected, the third selectable control 702-3 causes the email application to send an email with the home address to Amy. Computing device 202 can obtain Amy's email address from the contact application. Thus, computing device 202 provides a secure way for the user to share sensitive user data in a voice or video call without audibly streaming it to someone nearby.

[0105] In FIG. 7B, computing device 202 displays communication information related to a voice call. Consider a previous voice call with Amy and Amy providing new contact information (e.g., her new work email address). In response, computing device 202 provides the communication information to the user. Caption module 210 determines that the audio data 304 includes Amy providing the new email address "My email address is amy@email.com". Next, display component 118 displays the new email address in message element 702. The message element provides the text "Amy's email address: amy@email.com".

[0106] In some implementations, computing device 202 can confirm that the new email address is not saved in computing device 202 (e.g., in a contacts application or an email application). If the new email address is saved, computing device 202 may prevent caption module 210 from displaying this communication information. If the new email address is not saved, computing device 202 may allow caption module 210 to display this communication information.

[0107] Computing device 202 can display a fourth selectable control 702-4. The fourth selectable control 702-4 includes the text "Save to Contact". When the fourth selectable control 702-4 is selected, it causes computing device 202 to save the email address in the contacts application.

[0108] In FIG. 7C, computing device 202 provides additional selectable controls in response to communicated information during a voice call. Consider a previous voice call with Amy where the user and Amy agreed to meet for lunch. Audio data 304 includes the phrase "Let's meet at Merry's Restaurant in 20 minutes." spoken audibly by the user. In response to this communicated information, computing device 202 can display the address of Merry's Restaurant in message element 700. Message element 702 includes the text "Address of Merry's Restaurant: 500 20th St, San Francisco, CA 94016". Computing device 202 can also display a fifth selectable control 702-5. The fifth selectable control 702-5 displays the text "Directions to Merry's Restaurant". When selected, the fifth selectable control 702-5 causes computing device 202 to initiate navigation instructions from a navigation application.

[0109] In some implementations, the fifth selectable control 702-5 can be a slice window of a navigation application that provides a subset of the functionality of the navigation application related to the communicated information. For example, the slice window of the navigation application can enable the user to select walking directions, driving directions, or public transportation directions to Merry's Restaurant.

[0110] FIGS. 8A-8D show other examples of a user interface of a computing device that supports a user's voice calls and video calls. FIGS. 8A-8D are described in the context of computing device 202 of FIG. 2 in sequence. Computing device 202 can provide a different user interface with fewer or additional features than that shown in FIGS. 8A-8D.

[0111] In FIG. 8A, in response to selectable options of the IVR system 110, the computing device 202 causes the display component 118 to display a user interface 126 having a message element 800 and selectable controls 802. Suppose a user makes a voice call to a new utility company. In the caller box 128, the name of the called party's business (e.g., a utility company) and the phone number (e.g., (111) 555-2345) are displayed.

[0112] The IVR system 110 uses a voice response system that prompts the caller to provide voice responses to a series of questions and statements. Suppose the audio data 304 includes the statement "Thank you for contacting us regarding new customer registration. Please state the type of service you are interested in." The IVR system 110 can listen for phrases that match or closely match the list of services provided. For example, the utility company can pick up one of the selectable options such as home Internet service, home phone, or TV service. The computing device 202 can determine that the audio data 304 includes an implicit list of two or more selectable options. The display component 118 can display in the message element 800 the text "The following is a list of common responses provided by new customers." In this example, the selectable controls 802 can include a first selectable control 802-1 (e.g., "Home Internet service"), a second selectable control 802-2 (e.g., "Home phone"), and a third selectable control 802-3 (e.g., "TV service"). The selectable controls 802 can include additional or fewer suggestions. The user can select one of the selectable controls 802 and cause the audio mixer 208 to audibly provide the selected option to the IVR system 110.

[0113] Computing device 202 can determine potential proposals based on audio data 304 by interpreting services available from the audible portion of a voice call. Additionally, computing device 202 can determine selectable options based on data obtained from other computing devices that have been given similar requests by the same utility or similar companies. In this way, computing device 202 can assist the user in navigating open-ended IVR prompts, avoiding ineffective responses, or restarting the system.

[0114] FIG. 8B is an example of a user interface 126 that responds to a request for user information (e.g., payment information). In response to the user selecting a home internet service, the IVR system 110 directs the user to an account specialist to set up a new account to initiate the home internet service. Since the user is a new account owner, the account specialist collects payment information including a credit card number to set up the account. For example, the audio data 304 may include a request from the specialist such as "Please provide your preferred payment method for the new service". In response to determining that the audio data 304 includes a request for user information, the computing device 202 determines a text description 306 of the request. In this example, the caption module 210 determines that the text description 306 requests credit card information. The computing device 202 identifies the credit card information within the CRM 206 and displays the user data on the user interface 126. The response element 800 includes information such as "Your credit card information: #-#-#-1234, [expiration date] 01 / 21, [PIN] 789".

[0115] The computing device 202 can also determine whether user data includes confidential information. In response to determining that a portion of the user data is confidential information, the computing device 202 can obfuscate a portion of the confidential information (e.g., replace at least several digits of a credit card number with different symbols including "#" or "*", or omit them). In this way, the computing device 202 can maintain the confidentiality of the confidential information and make it less visible to others.

[0116] The display component 118 can display a selectable control 802 to maintain the confidentiality of the user data. In this example, the display component 118 displays a fourth selectable control 802-4 that includes the text "Please read the credit card information". When selected, the fourth selectable control 802-4 causes the computing device 202 to audibly read the entire credit card number, expiration date, and PIN to the account specialist. In this way, the computing device 202 provides a secure way for the user to share highly confidential credit card information with the account specialist.

[0117] In FIG. 8C, computing device 202 displays transmission information related to a voice call. Consider a previous voice call to a utility company. The account specialist provides the user with account information (e.g., an account number and a personal identification number (PIN)). In this situation, audio data 304 includes the sentence "Your new account number is UTIL12345 and the PIN associated with your account is 6789." In response, computing device 202 displays the account number and the PIN in message element 800. Specifically, message element 802 displays "Your account number: UTIL12345, your PIN: 6789". Computing device 202 can provide the user with a fifth selectable control 802-5 and a sixth selectable control 802-6. The fifth selectable control 802-5 includes the text "Save to contacts". When selected, the fifth selectable control 802-5 causes computing device 202 to save the account number and the PIN to the contacts application. The sixth selectable control 802-6 includes the text "Save to secure memory". When selected, the sixth selectable control 802-6 causes computing device 202 to save the account number and the PIN to a secure memory that requires an application or special privilege by the user to access.

[0118] In FIG. 8D, computing device 202 displays transmission information related to a previous voice call. Consider a previous voice call to a utility company. In this example, the user was unable to view the transmission information displayed on the user interface during or immediately after the voice call. Computing device 202 can store message element 802 related to the voice call, the fifth selectable control 802-5, the sixth selectable control 802-6, or a combination thereof. In this way, the user can later access the text description 306 of the transmission information.

[0119] The call history can provide a user interface 126 associated with each voice call or video call. For example, the user interface 126 associated with the history of a voice call with a public utility company may include a history element 804. The history element 804 may include history information regarding the voice call, including text such as "Outgoing on November 2".

[0120] Depending on the situation, the user may need to make another voice call or video call immediately after the end of a voice call with a public utility company, or may need to execute another function on the computing device 202. The computing device 202 can store a message element 800 and a selectable control 802 associated with each voice call or video call in a memory associated with the communication application 124. The communication application 124 may include a call history. In this way, the user can retrieve the message element 800 and the selectable control 802 associated with the voice call or video call at a convenient time later.

[0121] Example The following sections describe examples.

[0122] Example 1: A method, including a computing device obtaining audio data output from a communication application executed on the computing device, the audio data including an audible portion of a voice call or video call between a user of the computing device and a third party, the method further including the computing device determining, using the audible portion, whether the audio data includes two or more selectable options, the two or more selectable options being audibly provided by the third party during the voice call or video call, the method further including, in response to determining that the audio data includes two or more selectable options, the computing device determining a text description of the two or more selectable options, the text description providing a transcription of at least a portion of the two or more selectable options, the method further including the computing device displaying two or more selectable controls on a display of the computing device, the two or more selectable controls being configured to be selectable by the user to provide an indication of a selected option of the two or more selectable options to the third party, each of the two or more selectable controls providing a text description of a respective selectable option.

[0123] Example 2: The method of Example 1, further including the computing device receiving a selection of one of the two or more selectable controls associated with the selected option, the selection being made by the user during the voice call or video call, the method further including, in response to receiving the selection of the one selectable control, the computing device communicating the selected option to the third party.

[0124] Example 3: The method according to Example 2, wherein transmitting the selected option to a third party includes the computing device sending an audio response or a DTMF (Dual-Tone Multi-Frequency) tone to the third party without the user audibly transmitting the selected option.

[0125] Example 4: The method according to Example 2 or 3, further comprising the computing device obtaining additional audio data output from a communication application in response to transmitting the selected option to a third party, the additional audio data including two or more additional selectable options audibly provided by the third party during a voice call or a video call in response to the selected option.

[0126] Example 5: The method according to any one of the preceding examples, further comprising the computing device determining, using the audible portion, whether the audio data includes a request for user information, the request for user information being audibly provided by the third party during a voice call or a video call, the method further comprising the computing device identifying user data in response to the request for user information using the audible portion, and during the voice call or the video call, the computing device displaying the user data on a display or the computing device providing the user data to a third party.

[0127] Example 6: The method further includes a computing device determining, using an audible portion, whether audio data includes transmission information, where the transmission information is related to the context of a voice call or a video call and is audibly provided by a third party or a user during the voice call or the video call, the method further includes, in response to determining that the audio data includes transmission information, the computing device determining a text description of the transmission information, where the text description of the transmission information provides a transcription of at least a portion of the transmission information, the method further includes displaying on a display other selectable controls, where the other selectable controls provide the text description of the transmission information and are configured to be selectable by a user to save the transmission information to at least one of a computing device, an application, or another application on the computing device, the method according to any one of the preceding examples.

[0128] Example 7: Determining a text description of two or more selectable options includes the computing device executing a machine learning model to determine the text description of the two or more selectable options, where the machine learning model is trained to determine a text description from audio data, and the audio data is received from an audio mixer of the computing device, the method according to any one of the preceding examples.

[0129] Example 8: The method according to Example 7, where the machine learning model includes an end-to-end recurrent neural network transducer automatic speech recognition model.

[0130] Example 9: Two or more selectable options are a menu representing options of an Interactive Voice Response (IVR) system or a Voice Response Unit (VRU) system, and the IVR system or VRU system is configured to interact with a user and direct the user to at least one of another menu of the IVR system or VRU system, a person related to a third party, a department related to a third party, a service related to a third party, or information related to a third party, the method according to any one of the preceding examples.

[0131] Example 10: Two or more selectable controls include at least one of a button, a toggle, selectable text, a slider, a checkbox, or an icon, and are included in a user interface of a communication application, the method according to any one of the preceding examples.

[0132] Example 11: The text description includes numbers associated with each of two or more selectable options, and each of the selectable controls includes a visual representation of the numbers associated with each of the two or more selectable options, the method according to any one of the preceding examples.

[0133] Example 12: A display of a computing device includes a touch-sensitive screen, and the selectable controls are presented on the touch-sensitive screen, the method according to any one of the preceding examples.

[0134] Example 13: The computing device includes a smartphone, a computerized watch, a tablet device, a wearable device, or a laptop computer, the method according to any one of the preceding examples.

[0135] Example 14: A computing device comprising at least one processor configured to execute any one of the methods described in Examples 1 - 13.

[0136] Example 15: A computer-readable storage medium including instructions that, when executed, configure a processor of a computing device to execute any one of the methods described in Examples 1-13.

[0137] Conclusion Although various configurations and methods for providing selectable controls on a computing device for an IVR system have been described in terms of features and / or language specific to the methods, it should be understood that the subject matter of the appended claims is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as non-limiting examples for providing selectable controls on a computing device for an IVR system. Further, while various examples have been described above and each has specific features, it should be understood that a particular feature of one example need not be used exclusively with that example. Instead, any of the features described above and / or shown in the drawings may be combined with, added to, or alternatively used in place of any of the other features of those examples.

Claims

1. A method comprising: a computing device obtaining audio data output from a communication application running on the computing device, the audio data including an audible portion of a voice call or video call between a user of the computing device and a third party, the method further comprising: the computing device using the audible portion to determine whether the audio data includes two or more selectable options, the two or more selectable options being audibly provided by the third party during the voice call or the video call, the method further comprising: in response to determining that the audio data includes the two or more selectable options, the computing device determining a text description of the two or more selectable options; and displaying two or more selectable controls on a display of the computing device, the two or more selectable controls being configured to be selectable by the user to provide an indication of a selected option of the two or more selectable options to the third party, each of the two or more selectable controls providing the text description of a respective selectable option, the method further comprising: the computing device discriminating a request for user information audibly provided by the third party during the voice call or the video call; the computing device, in response to the request for user information, identifying user data from one or more data sets stored in the computing device, the one or more data sets including data held by an application different from the communication application, the method further comprising: during the voice call or the video call, the computing device displaying the user data on the display or the computing device providing the user data to the third party.

2. The method further comprises: receiving a selection of one of the two or more selectable controls associated with the selected option, the selection being made by the user during the voice call or the video call, the method further comprising responding to receiving the selection of the one selectable control, the computing device transmitting the selected option to the third party, the method of claim 1. **Claim 3** Transmitting the selected option to the third party includes the computing device transmitting an audio response or DTMF (Dual-Tone Multi-Frequency) tone to the third party without the user audibly transmitting the selected option, the method of claim 2. **Claim 4** The method further comprises responding to transmitting the selected option to the third party, the computing device obtaining additional audio data output from the communication application, the additional audio data including two or more additional selectable options audibly provided by the third party during the voice call or the video call in response to the selected option, the method of claim 2 or 3. **Claim 5** The method further comprises the computing device determining whether the audio data includes transmission information using the audible portion, the transmission information being related to the context of the voice call or the video call and being audibly provided by the third party or the user during the voice call or the video call, the method further comprising responding to determining that the audio data includes the transmission information, the computing device determining a text description of the transmission information, the text description of the transmission information providing a transcription of at least a portion of the transmission information, the method further comprising including displaying other selectable controls on the display, the other selectable controls providing the text description of the transmission information, and being configured to be selectable by the user to store the transmission information in at least one of the computing device, the application, or another application on the computing device, the method according to any one of claims 1 to 4.

6. Determining the text description of the two or more selectable options includes the computing device executing a machine learning model to determine the text description of the two or more selectable options, the machine learning model being trained to determine a text description from the audio data, the audio data being received from an audio mixer of the computing device, the method according to any one of claims 1 to 5.

7. The method according to claim 6, wherein the machine learning model includes an end-to-end recurrent neural network transducer automatic speech recognition model.

8. The two or more selectable options are a menu representing options of an interactive voice response (IVR) system or a voice response unit (VRU) system, the IVR system or the VRU system interacting with the user and guiding the user to at least one of another menu of the IVR system or the VRU system, a person related to the third party, a department related to the third party, a service related to the third party, or information related to the third party, the method according to any one of claims 1 to 7.

9. The two or more selectable controls include at least one of a button, a toggle, selectable text, a slider, a checkbox, or an icon, and are included in a user interface of the communication application, the method according to any one of claims 1 to 8.

10. The text description includes numbers associated with each of the two or more selectable options, and each of the selectable controls includes a visual representation of the numbers associated with each of the two or more selectable options, the method according to any one of claims 1 to 9.

11. The display of the computing device includes a touch-sensitive screen, and the selectable control is presented on the touch-sensitive screen, the method according to any one of claims 1 to 10.

12. The computing device includes a smartphone, a computerized watch, a tablet device, a wearable device, or a laptop computer, the method according to any one of claims 1 to 11.

13. A computing device comprising at least one processor configured to execute the method according to any one of claims 1 to 12.

14. A program comprising instructions that, when executed, configure a processor of a computing device to execute the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Interactive voice response system crawler

    JP2017188886A

  • Call processing method and device

    JP2017538327A

  • JPP6783492B

  • Electronic device and method for displaying phone call content

    US20160080558A1

  • Acoustic signal processing device, acoustic signal processing method, and hands-free calling device

    WO2018163328A1