Selecting and visually displaying voice menus for phone calls
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2022-07-19
- Publication Date
- 2026-08-03
Smart Images

Figure 0007899305000001 
Figure 0007899305000002 
Figure 0007899305000003
Abstract
Description
Background Art
[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 236,651, filed Aug. 24, 2021, entitled "Determination and Visual Display of Spoken Menus for Calls", and claims priority to U.S. Patent Application No. 17 / 540,895, filed Dec. 2, 2021, entitled "Determination and Visual Display of Spoken Menus for Calls", both of which are hereby incorporated by reference in their entirety.
[0002] Background Many businesses and other organizations provide automated telephone menus for callers who call the business, also known as an Interactive Voice Response (IVR) system. Generally, a caller who calls a business receives an automated voice that describes, in spoken words, a menu of multiple options that the caller can select. Often, a hierarchy of such a set of options is provided to allow the caller to move through the options to a desired result. For example, a caller may wish to receive certain information, request a company's product or service, speak to a human agent, etc. The caller can select an option in the call menu by speaking a number, word, or phrase associated with the option that is detected and recognized by the automated system, or by pressing a key.
[0003] The background description provided here is for the purpose of generally providing the background of the disclosure. The achievements of the inventors, now named, in the scope described in this background section, and aspects of the description that would not otherwise qualify as prior art at the time of filing, are not admitted as prior art to this disclosure, either expressly or implicitly.
Summary of the Invention
[0004] overview The present invention relates to the determination and visual representation of an audio menu for a call. In some embodiments, the method performed by a computer includes receiving audio data output in a call between a calling device and a device associated with a target entity. The audio data includes utterances indicating one or more selection options for the user of the calling device to navigate through a call menu provided by the target entity in the call. Text is determined by programmatically analyzing the audio data, and the text represents the utterances in the audio data. The selection options are determined based on programmatically analyzing at least one of the text or audio data. At least a portion of the text is displayed by the calling device during the call, and the text is displayed as one or more visual options corresponding to the selection options. Each visual option is selectable via user input to produce a corresponding navigation through the call menu.
[0005] Various implementations and examples of the method are described. For example, in some implementations, the method further includes causing a device associated with a target entity to send a selection instruction in response to receiving a selection of a particular visual option from one or more visual options, the instruction being an utterance provided by a calling device in a call, including a signal corresponding to pressing a key on a keypad associated with the particular visual option, or an indicator associated with the particular visual option. In some implementations, each of the one or more visual options is selectable via touch input on the calling device's touchscreen.
[0006] In some implementations, the audio data is first audio data, and in response to receiving a selection of a particular visual option, the method further includes receiving second audio data in a call, the second audio data includes a second utterance indicating one or more second selection options, the method further includes programmatically analyzing the second audio data to determine a second text representing the second utterance in the second audio data, determining one or more second selection options based on programmatic analysis of at least one of the second text or the second audio data, and causing the calling device to display at least a portion of the second text as one or more second visual options corresponding to the second selection options, each of the one or more second visual options being selectable via a second user input to result in a corresponding movement through a call menu. In some implementations, the one or more selection options are multiple selection options, and the method further includes programmatically analyzing at least one of the text or audio data to determine a hierarchical structure of the multiple selection options in a call menu. In some implementations, the method further includes storing one or more selection options in the storage of the calling device and / or the storage of a remote device communicating with the calling device on a communication network, and retrieving one or more selection options for the next call between the calling device and the target entity.
[0007] In some implementations, the method further includes obtaining selection option data containing one or more selection options before receiving audio data, and causing the calling device to display one or more visual options corresponding to one or more selection options before the calling device receives audio data containing utterances indicating one or more selection options. In some examples, the selection options in the selection option data are determined by programmatically analyzing audio data received during a previous call. For example, in some implementations, the obtained selection option data is cached by the calling device before the start of a call, and the obtained selection option data is related to entity identifiers previously called by the caller in the geographic area of the calling device, and related to entity identifiers that have been called at least a threshold number of times previously, or more times previously than other entity identifiers not related to the obtained selection option data.
[0008] In some implementations, a visual indicator is displayed during a call, the visual indicator highlighting a specific portion of the text of a visual option displayed during the call, the specific portion of the text being currently received during the call in the utterance in the audio data. In some implementations, the method further includes comparing selection option data with one or more selection options determined from the audio data, and determining whether a mismatch exists between the selection option data and one or more selection options determined from the audio data. In various implementations, in response to determining a mismatch, the method further includes causing a mismatch notification by the calling device and / or modifying the selection option data to match one or more selection options determined from the audio data. In some implementations, comparing selection option data with one or more selection options includes comparing the text of the selection option data with the text of one or more selection options, and / or comparing the audio data of the selection option data with audio data received during the call.
[0009] In some implementations, a calling device for displaying selection options for a call includes a memory storing instructions, a display device, and at least one processor coupled to the memory, the at least one processor being configured to access instructions from memory to perform an operation. The operation includes receiving audio data in a call between the calling device and a device associated with a target entity, the audio data including utterances indicating one or more selection options for a user of the calling device to navigate through a calling menu provided by the target entity in the call, further including programmatically analyzing the audio data to determine text representing the utterances in the audio data, determining one or more selection options based on programmatically analyzing at least one of the text or audio data, and displaying at least a portion of the text by the display device during the call, the portion of the text being displayed as one or more visual options corresponding to one or more selection options, each of which is selectable via user input to result in a corresponding navigation through the calling menu.
[0010] In various implementations of a calling device, the processor performs further operations, including, in response to receiving a selection of a particular visual option from one or more visual options, causing a selection instruction to be sent to a device associated with a target entity, the instruction being an utterance provided by the calling device in a call, including a signal corresponding to pressing a key on a keypad associated with a particular visual option, or an indicator associated with a particular visual option. In some implementations, the processor performs further operations, including, before receiving audio data, obtaining selection option data including one or more selection options and a hierarchical structure of one or more selection options in a calling menu, and causing a display device to display one or more visual options corresponding to one or more selection options before the calling device receives audio data including an utterance indicating one or more selection options.
[0011] In some implementations, the processor includes displaying a visual indicator during a call, the visual indicator highlighting a specific portion of the text of one or more visual options displayed during the call, and performing further operations on the specific portion of the text being currently spoken in the utterance in the audio data during the call. In some implementations, the processor performs further operations including comparing selection option data with one or more selection options determined from the audio data, and determining whether there is a mismatch between the selection option data and the one or more selection options determined from the audio data. In various implementations, the operations performed by the processor may include one or more features of the above method.
[0012] In some implementations, a non-temporary computer-readable medium, when executed by a processor, stores instructions that cause the processor to perform an operation. The operation includes receiving audio data in a call between a calling device and a device associated with a target entity, the audio data including utterances indicating one or more selection options for the user of the calling device to navigate through a call menu provided by the target entity, further including programmatically analyzing the audio data to determine text representing the utterances in the audio data, determining one or more selection options based on programmatically analyzing at least one of the text or audio data, and having at least a portion of the text displayed by the calling device during the call, the portion of the text being displayed as one or more visual options corresponding to one or more selection options, each of which is selectable via user input to result in a corresponding navigation through the call menu. In various implementations, the operation performed by the processor may include the above method or one or more features of the calling device. [Brief explanation of the drawing]
[0013] [Figure 1] This is a block diagram of an exemplary system that may be used for one or more implementations described herein. [Figure 2] This flowchart illustrates exemplary methods for determining and visually displaying voice menus for phone calls, based on several implementations. [Figure 3] This flowchart illustrates an exemplary method for obtaining entity selection options based on acquired data and / or calls, through several implementations. [Figure 4] This flowchart illustrates, through several implementations, exemplary methods for processing audio data from a call and displaying or updating visual options based on the audio data. [Figure 5] This is a schematic diagram of the user interface displayed by a calling device that can initiate a call, based on several implementations. [Figure 6] This is a schematic diagram of the user interface displayed by the calling device, showing the selection options for the call menu during a call, as seen in several implementations. [Figure 7] This is a schematic diagram of the user interface displayed by the calling device, showing the selection options for the call menu during a call, as seen in several implementations. [Figure 8] This is a schematic diagram of the user interface displayed by the calling device, showing the selection options for the call menu during a call, as seen in several implementations. [Figure 9] This is a schematic diagram of the user interface displayed by the calling device, showing the selection options for the call menu during a call, as seen in several implementations. [Figure 10] This is a schematic diagram of the user interface displayed by the calling device, showing the selection options for the call menu during a call, as seen in several implementations. [Figure 11] This is a schematic diagram of the user interface displayed by a calling device, where, in several implementations, the corresponding selection options are shown as visual options in the call menu before they are spoken during a call. [Figure 12] This is a schematic diagram of the user interface displayed by a calling device, where, in several implementations, the corresponding selection options are shown as visual options in the call menu before they are spoken during a call. [Figure 13] This is a schematic diagram of the user interface displayed by a calling device, where, in several implementations, the corresponding selection options are shown as visual options in the call menu before they are spoken during a call. [Figure 14]This is a schematic diagram of the user interface displayed by a calling device, where, in several implementations, the corresponding selection options are shown as visual options in the call menu before they are spoken during a call. [Figure 15] This is a block diagram of an exemplary device that may be used for one or more implementations described herein. [Modes for carrying out the invention]
[0014] Detailed explanation One or more implementations described herein relate to the determination and visual display of voice menus for phone calls. In various implementations, audio data, including utterances, is obtained from a call between a user's calling device and a target entity (e.g., a person or a company). The target entity may be an automated voice system (e.g., using an Interactive Voice Response (IVR) or answering machine) or a human agent. The utterances include selection options in a call menu, to which the user can navigate to obtain a desired outcome (e.g., receive information, speak to a human agent). Text is recognized from the call audio data, and the text represents the utterances describing the selection options. The selection options are discovered based on analysis of the text and / or audio data. At least a portion of the text is displayed by the calling device during the call as visual options corresponding to the selection options. Each visual option is selectable via user input to result in a corresponding navigation through the call menu.
[0015] Various additional features are described. For example, in some implementations, the selection by a user of a particular visual option can cause this selection to be sent to a target entity, where this selection can be a signal corresponding to pressing an appropriate key on the keypad of the calling device or can be an utterance provided by the calling device that selects the visual option. Audio data and / or text can be analyzed to determine a hierarchical structure of selection options in a call menu.
[0016] In some implementations, selection option data is obtained by the calling device prior to a call, e.g., received by the calling device from a server or other remote device that stores selection option data for various entities. In some examples, the selection option data may be determined from audio data of a previous call by the calling device to an entity. The calling device can download and cache selection option data for various entities and / or entity identifiers (e.g., the phone number, email address, instant message, or over-the-top (OTT) service identifier of an entity, etc.) prior to a call. In some examples, the cached selection option data can be for entity identifiers that are called more frequently by the user (e.g., the most-called in a set of entity identifiers) or for entity identifiers that are called at least a threshold number of times by the user in a geographic area (or threshold distance) of the calling device.
[0017] Using the cached selection options, corresponding visual options can be displayed before or during a call, before the selection option is spoken by the target entity in the call. Some implementations can compare the selection options spoken during the call with the cached selection option data, and if a mismatch is detected between these option versions, the user can be notified of the mismatch and / or the selection option data can be corrected to match the selection option determined from the current call's utterance data. In some implementations, a visual indicator is displayed during the call, and the visual indicator highlights a specific portion of the text of the visual option currently being spoken during the call.
[0018] The described techniques and features have several advantages. The described implementations can provide a visual representation of the audio call menu during a call. This can greatly assist the user when navigating through the call menu. Because audio call menus are often long and impose a significant cognitive load on the user to listen through long audio messages to find the options they need. Providing the corresponding visual version of the call menu on the call device can greatly assist the user in determining which options are available and which options are of interest to the user. Additionally, since the visual options presented are directly actionable and selectable by the user, the user can select the visual options using simple selection of the option, for example, via a touch on a touch screen. Thus, a complex audio experience is transformed into a simple visual experience by the described features.
[0019] In addition, some implementations can provide call menu options in a visual format before these options are uttered by the target entity during the call. This allows the user to see the call menu in advance, and in some call menus, the user can select a menu option before that option is uttered, instantly advancing the call menu to the target entity without the user having to wait for the user to hear the remaining options uttered during the call. The visual format of the menu allows the user to select an option much more quickly than if they were to look through the menu before the utterance, find the desired option, listen to the option in audio format, and then find and select the desired option.
[0020] The technical effect of one or more of the described implementations is that the device expends less computing resources to obtain results. For example, the technical effect of the described technology is a reduction in the consumption of system processing resources and power resources compared to a conventional system that does not offer one or more of the described technologies or features. For example, such conventional systems may require the user to spend a significant amount of time during a call listening to the output of available options before deciding which option best suits the user's needs. In some cases, in such conventional systems, the user may forget which menu options were previously offered due to the length of the spoken option messages, and may have to replay the menu or call to understand the available options, thereby spending more time. Such long call times waste system resources. The features described herein can mitigate such drawbacks by, for example, displaying selection options for a call menu, allowing the user to see the available call options and select the desired option more quickly, reducing call duration and initiating fewer calls, thereby lowering the overall processing and power requirements of the calling device, target entity devices, and other devices that communicate with the calling device to enable the call.
[0021] Furthermore, in some implementations, visual call menu options are displayed before these options are spoken during the call. Users can browse the visual options before the corresponding spoken options, find the desired option, and select it much more quickly than if the options were only available in audio format. Such features reduce call duration and conserve processing resources on calling and entity devices by allowing users to navigate through call menus at a faster pace, including quickly moving through call menus they have never heard or encountered before.
[0022] Furthermore, in some implementations, selection option data that provides selection options before a call can be downloaded to the calling device before the call begins and cached by the calling device, thereby reducing the consumption of processing and networking resources during the call. Furthermore, in some implementations, spoken selection options can be compared with cached selection option data to determine whether the options detected and displayed during the call may differ from the spoken options, thereby detecting errors or discrepancies that could otherwise waste processing and network resources on the calling device if the user sees and selects an inaccurate or undesirable option. Furthermore, some implementations of the described technology can provide displayed selection options in a call menu before a call and / or before these options are spoken during a call, based on data derived from previous calls by the user and user calling device (e.g., client device) to entities, without requiring that selection option data be received directly from entities or related entities that may not be available, for example.
[0023] In addition to the descriptions herein, users may be provided with controls that enable them to make choices regarding whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's call history specifying the entities called and entity identifiers, social networks, social behavior, or activities, occupation, user preferences including for call menus, user's current location, user's messages, outgoing calls made by the user, audio data of calls, or the user's device), as well as whether the user receives content or communications from the server. In addition, some data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information cannot be determined for the user, or a user's geographical location may be generalized where location information is obtained so that the user's specific location cannot be determined (e.g., at the city, zip code, or state level). Thus, users may have control over what information is collected about them, how that information is used, and what information is provided to them.
[0024] Figure 1 shows a block diagram of an exemplary network environment 100 that may be used in some of the embodiments described herein. In some embodiments, the network environment 100 includes one or more server devices, for example, a server system 102 in the example of Figure 1. The server system 102 may communicate, for example, over a network 130. The server system 102 may include a server device 104 and a database 106 or other storage devices. The network environment 100 also includes one or more client devices, for example, client devices 120, 122, 124, and 126, which may communicate with the server 102 via the network connection 130 to each other and / or to other devices. The network 130 can be any type of communication network, including one or more of the Internet, a local area network (LAN), a wireless network, a switch or hub connection, etc. In some implementations, network 130 may include peer-to-peer communication between devices 120-126, for example, by using a peer-to-peer wireless protocol (e.g., Bluetooth®, Wi-Fi Direct, etc.) or by having one client device that acts as a server to the other client device. One example of peer-to-peer communication between two client devices 120 and 122 is shown by arrow 132.
[0025] For ease of explanation, Figure 1 shows one block for server system 102, server device 104, and database 106, and four blocks for client devices 120, 122, 124, and 126. Server blocks 102, 104, and 106 may represent a number of systems, server devices, and network databases, and the blocks may be provided in configurations different from those shown. For example, server system 102 may represent a number of server systems that can communicate with other server systems, for example, via network 130. In some implementations, server system 102 may include, for example, a cloud hosting server or a server that provides telephone services (e.g., Voice over Internet Protocol (VOIP)). In some examples, database 106 and / or other storage devices may be provided in a server system block separate from server device 104, which can communicate with server device 104 and other server systems via network 130. There may also be any number of client devices. In some examples, the server system 102 communicates wirelessly with client devices over a network connection 130, and the client devices provide various features that can be activated or acquired by signals from the server mobile device.
[0026] The server system 102 and client devices 120-126 can be any type of device used in various applications, such as desktop computers, laptop computers, portable or mobile devices, mobile phones, smartphones, tablet computers, televisions, TV set-top boxes or entertainment devices, wearable devices (e.g., display glasses or goggles, head-mounted displays (HMDs), earpieces, earbuds, fitness bands, watches, headsets, armbands, jewelry, etc.), virtual reality (VR) and / or augmented reality (AR) enabling devices, personal digital assistants (PDAs), media players, game devices, etc. Some client devices may also have a local database similar to database 106 or other storage. In other implementations, the network environment 100 may not have all of the components shown and / or may have other elements, including other types of elements, instead of or in addition to those described herein.
[0027] In various implementations, client devices 120-126 may interact with the server system 102 through applications running on each client device and / or on the server system 102. For example, each client device 120, 122, 124, and 126 may communicate data to and from the server system 102. In some implementations, the server system 102 may send various data, such as content data (e.g., audio, images, video, messages, emails, etc.), notifications, and commands, to all or specific client devices. Each client device can send appropriate data to the server system 102, such as acknowledgments, data requests, notifications, user commands, and call requests. In some examples, the server and client devices may communicate data in various forms, including text data, audio data, video data, image data, or other types of data.
[0028] In various implementations, end users U1, U2, U3, and U4 may communicate with and / or with each other using their respective client devices 120, 122, 124, and 126. In some examples, users U1, U2, U3, and U4 may interact with each other through applications running on their respective client devices and / or on the server system 102, and / or through network services implemented on the server system 102, such as social networking services or other types of network services. In some implementations, the server system 102 may provide appropriate data to the client devices so that each client device can receive communicated content or shared content uploaded to the server system 102 and / or network services. In some implementations, “user” may include one or more programs or virtual entities, and people who interface with the system or network.
[0029] The user interface on client devices 120, 122, 124, and / or 126 can enable the display of user content and other content, including images, videos, data, and other content, as well as communications (e.g., for phone or internet calls, video conferencing, synchronous or asynchronous chat, etc.), privacy settings, notifications, and other data. Such a user interface can be displayed using software on the client device, software on the server device, and / or a combination of client software and server software running on server device 104, for example, application software or client software that communicates with server system 102. The user interface can be displayed by a display device on the client device or server device, for example, a touchscreen or other display screen, a projector, etc. In some implementations, an application program running on the server system can communicate with the client device to receive user input on the client device and to output data such as visual data and auditory data on the client device.
[0030] Various applications and / or operating systems running on the server and client devices can enable a variety of functions, including communication applications (e.g., connecting and providing audio calls or voice calls, video conferencing, chat, or other communications), email applications, content data display, privacy settings, notifications, browsers, etc. The user interface can be displayed on the client device using applications or other software running on the client device, software on the server device, and / or a combination of client software and server software running on server 102, for example, application software or client software communicating with server 102. The user interface can be displayed by a display device on the client device or server device, for example, a display screen, projector, etc. In some implementations, an application program running on the server can communicate with the client device to receive user input on the client device and to output data such as visual data and auditory data on the client device. In some implementations, one or more devices in the network environment 100, for example, one or more servers in the server system 102, may maintain electronic encyclopedias, knowledge graphs, one or more databases, corpora of words, phrases, symbols and other information, social network applications (e.g., social graph, social network for friends, social network for businesses, etc.), websites for places or locations (e.g., restaurants, car dealerships, etc.), mapping applications (e.g., websites for finding map locations), call characteristics and other call data, etc. In some implementations, the server system 102 may include classifiers for specific types of content items (e.g., text or images) and can determine whether any of a specific class is detected in the received content item.
[0031] Some implementations may provide one or more of the features described herein on a client or server device that is disconnected from or intermittently connected to a computer network. In some implementations, the client device may provide features and results for asynchronous communication, such as those described herein, for example, via chat or other messaging.
[0032] A machine learning model can be used by the server system 102 and / or one or more client devices 120-126 described herein. In some implementations, the machine learning model may be a neural network comprising, for example, one or more nodes arranged in one or more layers, various nodes connected via the network architecture, and associated weights, according to a network architecture. For example, in the training phase of the model, the model may be trained using training data, and then in the inference phase, the trained model may determine an output based on input data. In some implementations, the model may be trained offline, for example, on a test device in a laboratory or in other setting, and the trained model may be provided to a server running the model. In some implementations, the trained model may be retrained or updated locally on-device, or an untrained model may be trained on-device. In some implementations, federated learning may be used to update one or more trained models, for example, if individual server devices may each perform local model training, with user permission, and updates to the model may be aggregated to update one or more central versions of the model.
[0033] Figure 2 is a flowchart illustrating exemplary method 200 for determining and visually displaying a voice menu for a call, in several implementations. In some implementations, method 200 can be implemented on a server, for example, a server system 102 as shown in Figure 1. In some implementations, some or all of the blocks of method 200 can be implemented on one or more client devices (for example, client devices 120, 122, 124, and / or 126 as shown in Figure 1), one or more server devices, and / or both server devices and client devices. In the examples described, a system implementing a block of method 200 includes one or more processor hardware or processing circuits ("processors") and can access one or more storage devices, such as a database 106 or other accessible storage. In some implementations, different components of one or more server systems can execute different blocks or parts of blocks.
[0034] In some implementations, Method 200 or a portion thereof may be initiated based on user input. The user may, for example, select to initiate Method 200 or a specific block of Method 200 from a displayed user interface. In some implementations, Method 200 or a portion thereof may be executed with user guidance via user input. In some implementations, Method 200 or a portion thereof may be automatically initiated by a device. For example, a Method (or a portion thereof) may be initiated periodically or based on the occurrence of one or more specific events or conditions. For example, such events or conditions may include obtaining selection option data indicating one or more selection options provided in a call to an entity (for example, for block 208 to be executed), a predetermined period of time elapsed since the last execution of Method 200 or a portion thereof, and / or the occurrence of one or more other events or conditions that may be specified in the settings of the device implementing Method 200. In some examples, a device (server or client) may execute Method 200 with access to selection option data in a call (if user consent is received).
[0035] In block 202, it is checked whether user consent (e.g., user permission) has been obtained for the use of user data in the implementation of method 200. For example, user data may include user preferences, user-selected responses (e.g., dialer applications, communication applications, or other applications), user call characteristic data (e.g., call duration, number of calls and location, audio data received during a call), other content in the device's user interface, or other content data items in content collection (e.g., calls related to the user), messages sent or received by the user, information about the user's social networks and / or social contacts, content ratings, the user's geographical location, historical user data, etc. One or more blocks of the methods described herein may use user data in several implementations.
[0036] To that end, if user consent is obtained from the relevant user for which user data may be used in Method 200, then in block 204, it is determined that blocks of the Method herein may be carried out with possible uses of user data as described for those blocks, and the Method proceeds to block 208. If user consent is not obtained, in block 206, it is determined that the blocks will be carried out without the use of user data, and the Method proceeds to block 208. In some implementations, if user consent is not obtained, the remainder of Method 200 is not carried out, and / or any particular blocks that use user data are not carried out. In some implementations, if user consent is not obtained, blocks of Method 200 are carried out without the use of user data, with general or publicly accessible and publicly available data.
[0037] In block 208, selection options provided by various entities in previous calls are determined for an entity's entity identifier based on the data obtained and / or calls made to the entity identifier. Selection options are options provided to the user (e.g., the calling device) by the entity (e.g., a server device configured to automatically answer calls and provide audio data) via audio data during a call between the entity and the calling device. In some implementations, selection options can be options in a call menu provided by the target entity. In some implementations, a call menu can include multiple levels (e.g., hierarchical levels) of sets of options, for example, one menu level representing one set of options, and another set of different options in different menu levels based on the options selected in the previous menu level.
[0038] In some implementations, the selection options may be selectable elements or areas represented in the user interface that are not included in a call menu or other menus, such as selectable buttons, links, or other elements that produce an action by a target entity (e.g., the example described for block 232). Such selection options may be determined during a call, for example, from utterances provided before, during, or after call menu options are spoken in the call, or from utterances in a call that do not provide any call menu options.
[0039] In some implementations, structured (or annotated) information can be determined from utterances provided by various entities during a call, and structured information can be provided as selection options and / or visual information. For example, outlines, trees, formatted text (e.g., text with system-added paragraph breaks, sentence breaks and / or page breaks, punctuation, etc.), or other structured information can be determined from utterances. In further examples, structured information may include, for example, selectable selection options, which can be executable by user input causing the system to perform one or more operations such as searching for and displaying information or web pages, opening or running a program, etc., and may include uniform resource locators (URLs), hyperlinks, email addresses, dates, locations, confirmation numbers, account numbers, etc. Some structured information may be represented as visual information that is either executable or selectable by user input. Structured information can be determined and displayed by the calling device during a call as selection options, or in addition to the selection options described in the examples herein.
[0040] In a call, the active target entity can be an automated system such as the entity's Interactive Voice Response (IVR) system, an answering machine that provides a call menu and can receive selections from the caller, or, in some cases, the entity's human agent that speaks options in a call and receives voice selections of these options from the calling device. In a call that provides selection options, the user can select one or more options to achieve a desired outcome, such as navigating (e.g., forward or backward) through one or more hierarchical menu levels of the call menu and / or receiving specific information, requesting a specific product or service, or requesting to speak to a live human agent who can answer questions. The call can be a telephone call, a voice call, or any other call (e.g., made via instant messaging or over-the-top (OTT) services, etc.) connected to a calling device used by the user and initiated or answered by the calling device. The selection options are determined in block 208 for the various entity identifiers of the entity that provides such options in a call with the entity. As referred to herein, a target entity is an entity that is being called or is active in a call (e.g., after a calling device has called the target entity, or vice versa) and is associated with one or more entity identifiers (e.g., a target entity identifier), such as a telephone number or other address information (e.g., a user or entity name, a user identifier, etc.), which can be used to connect the target entity to a call that enables voice communication. An entity may include any of various individuals, organizations, companies, groups, etc.
[0041] In some implementations, the selection options are determined based on entity data received from entities (including from relevant entities, such as call centers or other entities handling calls for the entities). In some implementations, the selection options are determined based on previous calls made to the entities by the calling device.
[0042] Block 208 can be executed as a preprocessing block to determine and store selection option data before initiating a call to the target entity and before determining and visually displaying the selection options for the current call, as described below. Several examples of retrieving selection options provided by the entity are illustrated with reference to Figure 3. The method continues to Block 210.
[0043] In block 210, the entity identifier of the target entity is obtained by the calling device. The calling device is a device that can be used to make a call to the entity, for example, client devices 120-126 in Figure 1, or alternatively, a server or other device. The entity identifier can be, for example, a telephone number, other call name, address, or other entity identifier that enables the calling device to initiate a call to the target entity. The entity identifier can be obtained in any of several ways in various implementations and / or cases. For example, the entity identifier can be obtained through user input from the user of the calling device. Such user input may include the user selecting a key on a physical or virtual keypad or keyboard to enter the identifier. In some examples, the entity identifier can be obtained in response to the user selecting a contact entry in a contact list stored on the calling device, which automatically retrieves the entity identifier associated with the contact entry from storage and provides it for use. For example, the entity identifier can be entered into or provided to an application running on the calling device, such as a dialer or calling application that initiates a call, or another application that can initiate a call. In several other examples, entity identifiers can be obtained from another application running on the calling device or from a remote device on the network.
[0044] In some examples, the entity identifier is received independently of a call, for example, to view the displayed selection options provided when calling the entity, in preparation for one or more subsequent calls to the target entity, without initiating a call when the identifier is received and / or when the selection options are displayed. In other examples, the entity identifier is received to immediately initiate a call to the target entity at the current time, for example, the identifier is received in a dialer or calling application, and control to initiate the call is selected by the user, or the call is automatically initiated by the calling device. In some implementations, the entity identifier is received during a (current) call to the target entity or a different entity. For example, the entity identifier may be received by the calling device to initiate the current call, which may already be in progress, or to initiate a second call to a different entity. The method may continue to block 212.
[0045] In block 212, it is determined whether the selection option data is retrieved for the target entity's entity identifier. For example, it can be determined whether the selection option data is available for retrieval. In some examples, the selection option data for various entities may have been previously retrieved or determined in block 208 by a system that can access a communication device (e.g., a server or other remote device connected on a network) and / or by a communication device as illustrated in the example in Figure 3. In some implementations, a portion of the complete set of selection option data for the target entity identifier may have been retrieved or determined in block 208 and is available for retrieval. In some cases, the selection option data retrieved in block 208 does not include data for the target entity identifier, and the selection option data is not available for retrieval.
[0046] In some implementations, the selection option data may already be stored by the calling device, so that the selection option data does not need to be retrieved in block 212. For example, the selection option data for the target entity identifier may have been determined beforehand based on one or more previous calls to the target entity identifier by the calling device (e.g., before block 210). If only some of the selection options in the call menu for the target entity identifier have been determined and stored beforehand by the calling device, the remaining selection option data can be retrieved.
[0047] In another example, the selection options or a subset thereof for the target entity identifier may have been previously retrieved by the calling device as selection option data from one or more remote devices. In some implementations, the selection options or a subset thereof may be retrieved and stored in the calling device's local storage before receiving the target entity identifier (or any part thereof) in block 210. This may allow the calling device to access and view the selection options more quickly than retrieving the selection options from remote devices on the network (e.g., from a server, client server, or other device) when the entity identifier is obtained.
[0048] In some implementations, a subset of the select option data available on the remote device may be received and stored (e.g., cached) in the local storage of the calling device before the start of block 210. For example, select option data for popular entity identifiers called by a user may be retrieved and stored locally by the calling device. In some examples, these popular entity identifiers may be the most frequently called entities in a set of entity identifiers (e.g., the most called entities in a set of entity identifiers, or called more times than other entity identifiers not cached in local storage), the most frequently called entities within a specific period (as described above), and / or the most frequently called entities (as described above) by users located in the same geographical location or region as the calling device (or having other similar characteristics to the user / calling device). In another example, select option data for entity identifiers previously called by a caller at least a threshold number of times in the geographical area (or threshold distance) of the calling device may be retrieved and stored locally by the calling device. In yet another example, select option data may be downloaded for entity identifiers located in the user's country or the country where the calling device is currently located.
[0049] If no selection option data is found for the target entity identifier, the method may proceed to block 216, as described below, as determined in block 212. If selection option data is found, the method proceeds to block 214. In block 214, available selection option data is found and cached (or otherwise stored) in the call device's local storage. The cached selection option data can be used to display selection options during a call, as described below. For example, the call device may find selection option data associated with the target entity identifier over the network from a remote device (e.g., from a repository of call menu selection options for various entities), such as a server or other device, and cache the selection option data in the call device's local storage. In some implementations, the cached selection option data may include data indicating the structure of the call menu in which the selection options are organized. In some implementations, along with user authorization, portions of the original audio data (or its signature) from the call analyzed to determine the selection options may be cached in relation to the selection options.
[0050] In some implementations, the calling device may request that one or more selection options be prefetched from a remote device before the entity identifier is fully obtained by the calling device in block 210, for example, before the user has finished entering the entity identifier into the calling device. For example, the calling device may request and download selection options from a remote device for a number of candidate entities that match the portion of the entity identifier entered so far (e.g., the most frequently called entity or entity identifier in the geographical area of the calling device, as described above). The calling device can then select and use the set of selection options associated with the entity identifier after the identifier has been fully specified. Such prefetching may allow the selection options to be displayed to the calling device more quickly after the entity identifier has been specified, because the download of the selection option data begins before the identifier input is completed and the selection options are displayed from local storage.
[0051] In some implementations, prefetching of selective option data is performed when a threshold portion of the complete entity identifier is received. In some examples, if the complete entity identifier is 10 digits long, prefetching can be performed after, rather than before, the 8th (or alternatively, 9th) digit of the partial identifier is received. This allows the number of candidates to be narrowed down to the amount of data that can be received by the calling device in a relatively short time sufficient to determine matching selective option data after the complete identifier has been received. In some implementations, a larger subset of selective option data is determined by the calling device or remote device, for example, to be prefetched by the calling device. For example, the subset of data may be associated with an entity identifier that is the most likely identifier entered by the user, so as to be determined, along with user permissions, based on one or more factors such as historical data indicating which entities were mentioned in entities the user has previously called (e.g., the most frequently and / or recent entities previously called) and / or user data (e.g., recent messages accessed with user permission).
[0052] In some implementations, cached selection option data stored by a calling device may be periodically updated with newer or modified data based on specific conditions that occur, such as in response to data modifications based on calls made by the calling device, in response to data updates on a remote device (e.g., adding new entities or entity identifiers based on recent calls made by other users), or periodically after each specific period.
[0053] In some implementations, the selection option data (or a portion thereof) can be determined on the calling device during the call (as described below) and not downloaded from a different device. The method may continue to block 216.
[0054] In block 216, it is detected that a call has been initiated between the calling device and the target entity using the acquired entity identifier. In some implementations, the call connects the calling device to a device associated with the target entity. The call can be any connection to the target entity, including audio, e.g., telephone, call via OTT application, call via application program (e.g., browser, banking app, browser, etc.). In some implementations, the call may optionally be a video call, in which video data is transmitted to produce the display of video images of the caller and / or called party on the calling device and / or the target entity device connected to the call. In some examples, the user of the calling device may initiate the call, for example, by selecting call control in the user interface of an application such as a dialer application or calling application to cause the calling device to dial the entity identifier and initiate a call with the target entity. In some examples, the call may be automatically initiated by the application on the calling device after the entity identifier has been acquired, for example in block 210. In these cases, the user and the calling device are the callers. In some other examples, the call may be initiated by a target entity, in which case the target entity device is the caller and the user and calling device are the non-callers. Hereinafter, an automated system (e.g., an IVR system or answering machine) and / or human agent that may be active in a call and represent a target entity is referred to as the target entity. The method continues to block 218.
[0055] In block 218, it is determined whether the cached selection options are available for display in a call with the target entity (e.g., cached selection options suitable for display at the current menu level or other stages of the current call). As described above with respect to blocks 212 and 214, selection options for the entity identifier of the target entity may be cached in the call device's storage. In some implementations, for example, from one or more previous iterations of blocks 222-230 below, there may be cached selection option data in local memory that was previously determined and stored in the current call (or in a previous call with the same call device), and this cached data may be suitable for display at the current stage (e.g., if the user returns to a previous menu level in the call menu in the current call). If the cached selection options are not available, or if it is determined that the available cached selection options are not associated with or suitable for the current stage of the call (e.g., a particular hierarchical level of the call menu to which the user has moved), the method proceeds to block 222, as described below.
[0056] If relevant cached selection options are available for display, the method then proceeds to block 220, where one or more visual options (e.g., of a call menu) are displayed based on one or more corresponding cached selection options. Visual options are items displayed in the user interface of a calling device that can correspond to selectable options of a call menu that are typically spoken to the user during a call. For example, visual options can be displayed within the interface of a dialer application or other application, or in a message or notification displayed on the calling device. In some implementations, visual options can be displayed in a separate window or display area and / or in response to the user selecting a control to command the display of the visual options.
[0057] Visual options can include text, symbols, images, emojis, icons, and / or other information, providing user-selectable options. In some implementations, visual options are selectable by the user by providing touch input, for example, by touching or otherwise touching a touchscreen at a location corresponding to the display of the visual option. In some implementations, one or more of the visual options are associated with an indicator (e.g., a number, name, keyword, etc.), which is typically spoken or entered by the user during a call (e.g., by pressing a key) to select the option associated with the indicator. Several examples of displaying visual options are described with respect to Figure 4 (see block 224 in Figure 2) and with respect to Figures 5 to 13, which are described below.
[0058] Visual options may be displayed in block 220 before the corresponding selectable options are uttered in the call, for example, by the target entity, for example, by an automated system (e.g., an IVR system or answering machine) or by a human agent. Thus, the user can instantly see one, more, or all of the selectable options available to the user without having to wait to hear the options through a slower method of utterance. In some implementations, only the selectable options for the current level in a hierarchical call menu may be displayed, or in other implementations, for example, selectable options from multiple levels of the call menu may be displayed so that the user can see the selection path through the levels of the call menu. In some implementations, or when commanded by the user or user settings of the call device, visual options may be displayed before the call is initiated in a call device using an entity identifier, based on cached selectable options. In some implementations, visual options are displayed after the call has been initiated. For example, selectable options may be displayed in the interface of a dialer application or other application (or as an operating system notification) so that the user can see the visual options before the call is initiated.
[0059] In some implementations, visual options may also, or alternatively, include other selectable items. For example, visual options may provide information related to a target entity, and may or may not be related to the options in the call menu. Visual options may include selectable items or parts, such as buttons or checkboxes, that can be selected by the user to send a specific selection or information to the target entity. In some examples, visual options may include web links or other types of links to various sources of information. For example, when selected by the user, such links may cause a web page, window, or other display area on the call device, for example, in a browser application or other application, to open, and allow information to be downloaded for display there. In some implementations, for example, in addition to visual options, other visual information that is not selectable as described above (e.g., structured information) may be displayed. The method may continue to block 222.
[0060] In block 222, audio data is received from the call, including audio data that shows or represents utterances made in the call by the target entity (e.g., an automated system or a human agent) and the user. The method may then proceed to block 224.
[0061] In block 224, audio data is processed, and visual options are displayed and / or updated based on the audio data. For example, text is determined from the utterances represented in the audio data, where the text represents the utterances. The selection options (and / or other selection options) of the call menu are determined based on the text, and the selection options allow the user of the call device to navigate through the call menu. In some implementations or cases, the selection options are displayed as visual options by the call device. In some implementations or cases, the visual options are already displayed based on cached selection options (e.g., based on block 220), and these visual options and their corresponding selection options can be updated based on the processed audio data, where appropriate. In some implementations, the structure of the hierarchical call menu, including the selection options, can also be determined based on the audio data and the text derived therefrom. Several examples of displaying and / or updating visual options are described below with reference to Figure 4.
[0062] In some implementations, for example, if a cached selection option is displayed by the calling device in block 220, block 224 may be skipped or omitted. In some implementations, if the cached selection option is more likely to be recently determined and therefore current, block 224 may be skipped.
[0063] The audio data received during a call is also output by the calling device after the calling device's audio system has processed the audio data, for example, so that the utterances in the audio data are played back through the device speaker, headphones, or other audio devices within or connected to the calling device. The method may continue to block 226.
[0064] In block 226, it is determined whether one or more visual options have been selected by the user. Various implementations may allow one or more methods for selecting visual options. For example, visual options may be selectable by the user via a touchscreen interface, voice commands, a physical input device (such as a mouse, joystick, or trackpad), or other user input devices. If none of the visual options are selected in block 226, the method proceeds to block 218 to receive additional audio data from the call. If one or more of the visual options are selected, the method proceeds to block 228.
[0065] In block 228, the selection options corresponding to the selected visual option are sent to the target entity. In some implementations, a selection instruction is sent to the target entity, where the instruction corresponds to input provided as if the user had made a standard selection of options in a call. In some examples, if the selected option can normally be selected via user utterance (e.g., uttering an indicator such as a number or word related to the selection option), the transmitted instruction may be an appropriate utterance uttered by the call device in a call, for example, in a recording or in speech synthesized by the call device uttering an appropriate indicator. In some examples, the user can select a visual option via non-voice input (e.g., touching a button or area displayed on a touchscreen), and the call device may output an utterance that selects the corresponding selection option via utterance. In another example, if the selected option can normally be selected via pressing a key on a keypad or keyboard, the call device may send an instruction that is a signal corresponding to the user pressing that key on the device. For example, such signals may include touch tones (e.g., dual-tone multi-frequency or DTMF signals) or their encoding, or other in-band signals corresponding to a specific key being pressed. In some implementations, alternatives to touch tones or key press inputs may be used, such as out-of-band signals provided via a Session Initialization Protocol (SIP), Real-Time Transport Protocol (RTP), H323, etc. The method may continue to block 230.
[0066] In block 230, it is determined whether there are further selection options to display. For example, the user's selection in block 226 may result in a move to the next level of the call menu (e.g., moving forward or backward in the call menu). The target entity may then begin to provide that next level in the call by, for example, uttering a new set of selection options for the user based on the previously selected option, where the new set of options can be displayed by the call device. In some cases or implementations, the new set of selection options may be at a previous level of the call menu that the user moved to previously, and these selection options may have been cached in a previous iteration of method 200. Such cached selection options can be retrieved from a cache in the call device's local memory. If there are further selection options to display, the method proceeds to block 218 to check whether the cached selection options are available for the new set of selection options.
[0067] If no selection options exist to be displayed in block 230 in response to a user selection, the method proceeds to block 232, where a result is obtained based on one or more actions by the target entity. The target entity can perform any type and / or number of actions in response to receiving a selected option. For example, an action could be providing information received from the target entity, for example, if the user's selection is the last option in a particular path through a call menu. For example, the target entity may provide (e.g., make an utterance) information in the call requested by the user, which is received by the call device in block 230. In some implementations or cases, the call may be terminated after such information has been received. In some implementations or cases, the target entity may request the user to utter information, for example, the user's name or other information (address, account number, etc.). In some implementations, the target entity may connect a human agent of the target entity to the call to make an utterance to the user, which can be detected by the call device (e.g., using speech recognition technology) and a notification provided to the user.
[0068] In some cases, the target entity's actions may include putting the calling device on hold, for example, waiting for a human agent to become available. In some implementations, the calling device (and / or connected device) may automatically determine, without user input or intervention, whether the calling device is on hold in the current call by using speech recognition technology to determine whether the entity's automated system indicates that the call is on hold, for example, through specific words (e.g., "An agent will be able to take your call in 10 minutes. Thank you for waiting") or through the playback of music indicating that the call is on hold. In some implementations, if the call is put on hold by the target entity, the calling device may display an indication of that hold status, such as a message indicating music playback. In some implementations, the calling device may detect whether a human agent has connected to the call while it was on hold, for example, through specific words spoken by an agent or user, the end of hold music, or an automated voice message, indicating that the call is no longer on hold. In some implementations, the calling device may output a notification indicating that the call is no longer on hold and a human agent has connected to the call.
[0069] In some implementations or cases, the target entity provides an option to return to the call menu after an action has been taken by the target entity, in which case the process can proceed to block 218.
[0070] In some implementations, after a call ends, the selection options determined and displayed by Method 200 can be stored in the call device's cache (or other storage) and / or sent (over a network connection) to the storage of a remote device, such as a server, which stores selection option data for various entities and is accessible by multiple call devices. When the same target entity identifier is called again by a user using a call device (or other user device or client device), the selection option data stored in the call device's cache and / or stored on the remote device can be used for the new call, for example, to display the selection options before these options are spoken in the call. In some implementations, some of the selection options can be retrieved from the call device's local storage (e.g., these selection options were previously stored in local storage based on the call in which they were selected) and / or some selection options can be retrieved from a remote device, as described above. Similarly, any updates or modifications to cached selection options can be cached on the call device and / or sent to the storage of selection option data for various entities on a remote device, such as a server.
[0071] In some implementations, along with user authorization, data indicating the call event and / or outcome can be stored as metadata, along with the selection options that are remembered after the call. For example, along with user authorization, outcome data may include instructions on which selection options were selected during the call, whether the user was able to connect to a human agent after dialing a particular selection option, the duration before the user disconnected from the call, and the selected selection options. When such data is accumulated from a large number of calls and calling devices, it can be used, for example, to determine whether to modify the selection options provided by the entity in future calls to improve the effectiveness and efficiency of the caller menu provided.
[0072] In some implementations, if user consent is obtained, call records and / or selection options made by the user during a call may be stored and made available for the user to view, for example, from the call log on the calling device or other user interface.
[0073] Figure 3 is a flowchart illustrating exemplary methods 300 for obtaining entity selection options based on acquired data and / or calls, in several implementations. For example, method 300 may be implemented as part of block 208 or block 208 in Figure 2 to obtain entity selection options before a call to a target entity in which such selection options may be used. In some implementations, method 300 may be performed by a server or other device other than a calling device to obtain selection options that can be downloaded or accessed by the calling device (e.g., a client device or other device) before or during a call to a target entity, for example, as described with respect to Figure 2. In some implementations, method 300 may be performed by a calling device, e.g., a client device, or different parts of the method may be performed by a server device and / or a client device, respectively.
[0074] The method begins in block 302. In block 302, entity data is retrieved from a set of entities, and the entity data includes selection option data for entity identifiers associated with entities in the set of entities. In some implementations, entity data can be made available by entities to provide selection options in a call to those entities. For example, entity data may include instructions for selection options uttered in the call menu of the relevant entity during a call using the entity identifier of the entity identifier, including the option text and other details, as well as / or the hierarchical structure of the call menu in which the selection options appear. In some examples, entity data can be explored for a particular set of entities (or entity identifiers) that meet certain criteria. For example, the set may include the number of entities that have the most popular entity identifiers called in a region or area of the calling device and / or within a particular period of time. For example, popular entity identifiers may be those that are called most frequently and / or recently, as described above, and calls may be made by callers within a particular period of time and / or within a threshold distance or geographic area of the calling device. In some implementations, entity data may be retrieved periodically from entities so that the retrieved entity data includes more recent updates. In some implementations, entity data may be retrieved from entities associated with entities represented by entity identifiers, such as call centers or other entities associated with the entity. The method continues in block 304.
[0075] In block 304, it is determined whether entity data is unavailable for one or more entity identifiers in the set of entities from which entity data is searched. In various examples, entity data may not be provided by an entity for any of the following reasons: security restrictions, the tendency for selection options to change and quickly become obsolete, technical issues, etc. In some implementations, entity data for an entity identifier may be considered unavailable if the available entity data is known to be old and / or inaccurate. If entity data is available from the set of entities, the method proceeds to block 210 in Figure 2 as described above. If entity data is not available from one or more entities in the set of entities, the method proceeds to block 306.
[0076] In block 306, entity identifiers for which entity data is unavailable from the entity are selected from entity identifiers associated with the set of entities. In some implementations, this may include entity identifiers for which entity data is known to be incomplete, for example, entity data may specify some but not all of the selection options in a call menu. In some implementations, incomplete entity data may be determined from user feedback indicating that one or more selection options are missing (or inaccurate) in the cached selection options displayed during a call based on the retrieved entity data, and therefore the entity data is likely to be incomplete. The method continues to block 308.
[0077] In block 308, it is determined whether selection option data is available for the selected entity identifier from one or more previous calls that include the selected entity identifier. For example, one or more users may have made calls using the selected entity identifier on previous occasions, and the selection options received during these calls may be retained, e.g., discovered and / or stored, with user permission. In some implementations, other call characteristics of these calls (e.g., entity identifier, call time, call location, call duration, etc.) may also be retained, with user permission, in which case the call characteristics are decoupled from the user who made the call so that only the call characteristics are known. Some or all of such data may be available to method 300. For example, previous calls may have been made by a group of users on a communication network using a calling device from which call characteristics are retrieved, with user consent. If such data is not available, the process may proceed to block 312, as described below. If such selection options (and / or other data) are available, the process proceeds to block 310.
[0078] In block 310, the selection options are determined for the selected entity identifier based on the previous call. One or more selection options may be automatically determined by the system based on an analysis of speech data in audio data recorded from the previous call using techniques such as speech recognition via machine learning models or other techniques (as described below with reference to Figure 4), if user consent has been obtained. For example, the selection option data determined from the previous call for the selected entity identifier may represent the text of the selection options provided in that call, and / or structural data of the call menu containing the selection options provided in the call (e.g., the hierarchical structure of the selection options in the call menu and the dependency of specific options on previously selected options, indicating which selections of previous options are required to access these options). In some implementations, a particular call may navigate to some of the selection options in the call menu, rather than all of them, and these may be logged. For example, in a logged call, the user may have selected a single navigation path of consecutive selection options through the call menu without descending any other paths or branches of options. In some implementations, block 310 may include testing multiple previous calls to a selected entity identifier, traversing different branches of the selection options in the call menu, until, if possible, all selection options in each branch of the call menu have been determined. In some implementations, along with user authorization, portions of the audio data (or their signatures) analyzed to determine the selection options may be stored in relation to the determined selection options. The method may then proceed to block 318, as described below.
[0079] In block 312, after it is determined that the selection option data is not available from the previous call for the selected entity identifier, one or more calls are initiated using the selected entity identifier. In some implementations, an automated system may be used to make one or more calls using the selected entity identifier. In some implementations, calls may be made at a specific time during business hours, for example, if the entity is a company. In some implementations, multiple calls may be made at various times outside of business hours, for example, thereby determining different selection option data that may be available at such various times. The method continues to block 314.
[0080] In block 314, the uttered selection options in a call are determined, for example, detected and stored, so that the selection options for the selected entity identifier are determined. In some implementations, the selection options are detected using one or more speech recognition techniques, for example, machine learning models or other techniques. Several examples of detecting selection options and menu structures from audio utterance data are described below with reference to Figure 4, and similar techniques can be used in block 314. In some implementations, block 314 includes selecting the selection options provided in a call, thereby navigating to further hierarchical levels of the call menu, receiving audio data at these levels, and detecting further selection options. In some implementations, different navigation paths of selection options through the call menu can be selected in each call to the selected entity identifier to determine each available selection option in the provided call menu. In some implementations, the same path of selection options can be traversed across multiple calls, for example, to provide additional data for comparison and to check for errors in the detection of selection options. In some implementations, if several selection options were available before block 314 or before an iteration of block 314, the undecided portion of the menu (e.g., a branch) may be selected to determine the selection options offered, and any available options or portions may be skipped. In some implementations, along with user permission, the portion of audio data (or its signature) analyzed to determine the selection options may be stored in relation to the determined selection options. The method continues to block 316.
[0081] In block 316, the menu structure for the call menu of the selected entity identifier can be determined based on the detected selection options in block 314. For example, the detected selection options are stored, and a data structure (e.g., graph, table, etc.) is created that provides relationships and dependencies between the selection options at different hierarchical levels of the call menu. The selection option data from numerous calls to the selected entity identifier can be tested to form the most complete call menu structure possible from the available data. In some implementations, the structure of the selection options in the call menu may have been determined beforehand, for example, based on partially complete entity data from block 302, or on a previous iteration of method 300. The selection option data from calls made in blocks 312 and 314 can be added to such an existing data structure. The method may then proceed to block 318.
[0082] In block 318, it is determined whether entity data is unavailable and whether there are more entity identifiers to select for which selection option data can be determined. If so, the process proceeds to block 306, where another entity identifier is selected for which selection option data can be determined. If there are no more entity identifiers to select, the process may proceed to block 210 in Figure 2.
[0083] Figure 4 is a flowchart illustrating exemplary methods 400 for processing audio data from a call and displaying or updating visual options based on the audio data, in several implementations. For example, method 400 may be implemented in block 224 of Figure 2, after block 222, in which audio data is received in a call initiated with a target entity using an acquired entity identifier.
[0084] The method begins in block 402. In block 402, the text representing utterances in the audio data of the call is determined. In some implementations, the text is determined using one or more speech recognition techniques, for example, one or more machine learning models and / or other techniques. In some implementations, for example, if the user has given permission and / or set relevant user settings, the calling device may provide a call recording that displays the recognized text of all words spoken in the call, including introductory utterances, user responses, etc. The method continues to block 404.
[0085] In block 404, one or more current selection options and / or menu structures are determined based on the text determined in block 402 and / or the audio data received in block 222. In some implementations, each selection option generally includes a written option for selection, which may be accompanied by a selection indicator for the option that the user attempts to input into the call to select, for example, by saying an indicator or by pressing a corresponding key or button on the calling device (for example, to provide a tone as provided by a push-button phone, or other signals as illustrated with reference to Figure 2). In some examples, selection options may be detected based on specific voice words (or other voice indicators) that may indicate or describe the selection options being offered. For example, selection options may generally include the word “to” followed by a verb (e.g., “to speak to a representative” or “for your account balance”) or the word “for” followed by a noun (e.g., “for Spanish” or “for your account balance”). Selection options may typically begin or end with a phrase containing “press” or “say,” followed by an indicator such as a number or word, for example, “press or say 2.” In some implementations, speech recognition technology can be adapted or trained to recognize such words for detecting selection options.
[0086] In some implementations, the menu structure for a call menu of a selected entity identifier may also be determined or added based on the detected selection options, for example, if audio data and / or determined selection options indicate that different levels of the call menu have been accessed in the call. This may occur, for example, after a user has selected a provided selection option. In some examples, a data structure (e.g., graph, table, etc.) can be created that provides relationships and dependencies between selection options at different hierarchical levels of the call menu. In some implementations, determined and user-selected selection options may be tested to form a call menu structure that is added as the call progresses, with further options being selected. In some implementations, the call menu structure may be compared to a cached call menu structure (e.g., which may be similar to cached selection options as described herein), and / or the call menu structure may be stored as described herein and accessed in future calls to the target entity to provide a call menu structure for those calls.
[0087] In some implementations, one or more models can be used to detect selection options and / or menu structures from audio utterance data and / or text determined from audio utterance data. In various implementations, these models may differ from the models used to determine text representing utterances in audio data, such as those used in block 402, or the functionality of these models may be included in the same models used in block 402. In some implementations, a model for detecting selection options may be trained on call characteristics of a previous call, including audio data providing speech selection options, text selection options, call menu structures, etc. In some examples, a model may be trained with training data that provides examples of words corresponding to selection options, and / or non-text data (e.g., audio data fragments or signatures corresponding to selection options). In some implementations, the model is a machine learning model, for example, a neural network having one or more nodes arranged in one or more layers according to a network architecture, for example, having various nodes connected via a network architecture and having associated weights. For example, in the training phase of the model, the model may be trained using training data, and then in the inference phase, the trained model may provide an output based on input data. Additional examples of features that can be included in the model are described below with respect to Figure 15. Other types of models or techniques may also be used, or used as alternatives, to detect selection options.
[0088] In several exemplary implementations, a system comprising one or more machine learning models processes audio data from a call in streaming format, passes the audio data through a speech recognition model to provide text (as in block 402), and then passes it through a dedicated neural network pre-trained from BERT (Bidirectional Encoder Representations from Transformers) or other suitable encoding to detect selection options and / or call structure. In addition, the audio data can be processed directly by an audio-to-intent architecture, and the results may be based on a combination of outputs. The outputs provide a set of selection options and a call menu structure (e.g., hierarchical structure) of these options as detected from the audio data.
[0089] Some implementations may use any of several other features. For example, some systems may have streaming speech recognition, receiving a stream of audio data and processing the text into speech in real time. A machine learning model may modify the recognized text over time, changing its recognition of the streaming audio data as additional audio data is received. The model may determine confidence in speech recognition. The model may use audio and non-text cues or data portions such as timing (e.g., pauses between words) to help recognize speech. The method may continue to block 406.
[0090] In block 406, to show the user the available selection options before those options are uttered by the target entity, it is determined, for example, from block 220 in Figure 2, whether cached selection options have been displayed in the current call. If cached selection options have not been displayed, the method may proceed to block 414, which is described below. If cached selection options have been displayed, the method may proceed to block 408.
[0091] In block 408, it is determined whether there is a mismatch between the current selection option determined in block 404 and the displayed cached selection option, for example, whether there is a significant enough difference between these options to satisfy one or more thresholds. The current selection option is compared with the cached selection option, and in some implementations, the menu structures of the current and cached selection options are compared.
[0092] In various implementations, the current choice option can be compared to a cached choice option using one or more of various techniques. In some examples of the first technique, the text of the cached choice option can be compared to the corresponding determined text of the current choice option. The texts of these options may often not match exactly due to errors in speech recognition, for example, from poor acoustic characteristics in the call affecting the audio data or for other reasons. In some implementations, the magnitude or severity of the mismatch between the current and cached choice options can be determined, for example, using a text comparison technique. If the magnitude of the mismatch falls below a threshold, the cached and current choice options can be considered to match.
[0093] In some examples of other techniques for comparing selection options, the audio data of a call received in block 222 of Figure 2 can be compared (if available) with the corresponding audio data portion of a cached selection option to determine differences in the audio data. In some implementations, the cached selection option can be based on specific audio data from a call to a target entity identifier made prior to the current call (e.g., in a call initiated in block 312 of Figure 3 as described above), provided user permission is obtained. Such audio data may be available, for example, from use in training a machine learning model used to detect the selection option. For example, with user permission, the audio data (e.g., from a previous call) used to determine the cached selection option can be stored in relation to the cached selection option (or the audio signature derived from the audio data can be stored), and can be stored, for example, in the local memory of the calling device or retrieved from a remote device. The corresponding audio data (or corresponding audio signature) for the cached and current selection options can be compared to examine the differences. In current and cached audio data, discrepancies may exist if significant differences are observed in the audio data (e.g., differences exceeding a threshold). The comparison technique may be chosen to be robust to compensate for possible variations in audio quality between different calls.
[0094] In some examples of techniques for determining the accuracy of the text of the current choice option, both the audio data and text of the current choice option can be used to determine the accuracy of the text by aligning the audio and text. For example, a machine learning model can be trained on input of audio data and / or recognized text from a choice option in a previous call to output an indication of the likelihood that the text is accurately recognized from the audio data in the current call. This likelihood that the text is accurate is determined based on the audio data of the current call, for example, on audio data for the words corresponding to the text and on surrounding words that form context for the words. Such a model can be used to provide the accuracy of the text of the current choice option determined in block 404. In some implementations, the corresponding cached audio data and / or text of the choice option can also be provided as input to the model to provide further criteria or comparisons for the model to improve the accuracy of the model output (for example, the model may be trained on such supplemental input). In some implementations, this technique can be used in block 402 as a speech recognition technique for the current choice option.
[0095] If enough current selection options (and user selections) have been received, the menu structures of cached and current selection options can also be compared to determine at least part of the call menu structure from the current call. For example, selection options that can be moved from previous selection options can be compared between the cached and current call menus.
[0096] The comparison in block 408 can determine whether, in some implementations, the cached selection options (and / or menu structure) may be inaccurate. For example, the target entity may have changed its call menu, and the cached selection options may have been retrieved at a previous time before the selection options provided by the target entity were changed. The selection options determined in block 404 are generally more up-to-date because the call options have been discovered in the current call.
[0097] If the current selection option matches a displayed cached selection option (and the call menu structure matches) (for example, based on one or more thresholds), the method may proceed to block 418 as described below. In this case, the displayed visual option is not changed or updated because there is no significant mismatch with the current selection option. If there is a mismatch between the current and cached selection options, for example, if any of the current selection options differ from a displayed cached selection option (or the call menu structure differs) based on one or more thresholds, the method proceeds to block 410.
[0098] In block 410, cached selection options that differ from the current selection option are modified based on the current selection option. For example, a cache or other storage that stores different cached selection options can be updated by selectively replacing the different cached selection option (considered to be inaccurate) with the corresponding current selection option that is considered accurate. In some examples, the cached selection option "To receive information about your order, say 3 or press 3" may be detected as "To receive information about your order, say 4 or press 4" for the corresponding current selection option, provided that each word is matched except for the number. Thus, the instance before "3" is changed to "4" in the storage that stores this selection option in order to modify the option. In some implementations, the entire selection option or a larger portion of the cache may be discarded and replaced with the corresponding current selection option. In some implementations, if a mismatch similarly exists in the call menu structure, the structure element prior to the mismatch may be replaced with the element determined in method 400.
[0099] In some implementations, corrections to outdated selection options (and / or call menu structures) can also be sent to other devices that may store these selection options in order to update selection options stored by those other devices, or can be sent instead. For example, a server (or other remote device) may store the current selection options, such as those obtained in block 208, and the server can then send the correct updated selection options (and / or call menu structures) determined in block 404. In some implementations, the server may determine whether other call devices have sent such corrections to the server to determine the accuracy of the corrections. For example, if a threshold number of call devices have sent a correction to a particular selection option, the server can assume the correction is accurate and apply the correction to its corresponding stored selection option.
[0100] In some implementations, cached selection options (and / or call menu structures) may not be modified based on differences in the current selection options if, for example, one or more specific conditions are met. For example, in some implementations, if the current selection options (and / or their structure in the call menu) are recognized by speech recognition technology with a confidence level below a certain threshold, the current selection options may not be correctly recognized, and the cached selection options will not be adjusted. In some implementations, cached selection options may be adjusted if they have a creation time older than a threshold period prior to the current time, and therefore are likely to be obsolete or outdated. The method may continue to block 412.
[0101] In block 412, the visual options displayed by the calling device are updated based on the current selection options determined in block 404 from the current call. For example, a visual option corresponding to a selection option found to be inaccurate or outdated in block 410 may be replaced by a visual option corresponding to the corresponding (e.g., current) selection option that replaced the inaccurate option. In some examples, the text of the inaccurate visual option is changed to the text of the accurate visual option. In some implementations, a notification may also be displayed on the user interface of the calling device indicating that a correction has been made and / or specifically which correction has been made. In some implementations, the correction is not made under certain conditions, for example, if the confidence value of speech recognition for the text of the selection option falls below a threshold. In some implementations, a notification may be displayed indicating that there may be a mismatch between the displayed visual option and the utterances in the call (e.g., indicating that the information in the visual option may not have been spoken by the target entity in the call). In some implementations, no correction is made, and inaccurate (and / or all) visual options may be removed from the screen in response to a determination that one or more of the corresponding selection options are inaccurate. The method may then proceed to block 418, which is described below.
[0102] In block 414, after block 406 determines that a cached selection option is not available and has not been displayed for the current call, the current selection option determined in block 404 may be cached in the call device's local memory. In some implementations, such cached selection options may be retrieved later for display in the current call, and / or displayed in subsequent calls, for example, if the menu level is revisited by the user in the current call. In some cases, one or more of the current selection options may have already been cached in a previous iteration of method 400, for example, in the current call or a previous call. The method may then proceed to block 416.
[0103] In block 416, the visual options are determined and displayed for the current call based on the current selection options determined in block 404. In some implementations, the selection options are displayed after an utterance explaining that the selection options have been completed, and each additional selection option may be displayed after the corresponding utterance has finished explaining it (for example, in a later iteration of method 400 during the current call). In some implementations, if this is the first iteration of block 416 for the current call, the visual options may be the first visual options displayed in the user interface for the current call. In later iterations, the visual options displayed in block 416 may be added to existing visual options displayed in the previous iteration. The method may continue to block 418.
[0104] In block 418, an indicator of the currently spoken text is displayed and / or updated in the user interface to point to a portion of the visual options currently being spoken by the target entity in a call. In some embodiments, the display indicator visually indicates which word, phrase, or entire selection option is currently being spoken in a call. This feature can be used to show the user which of the previously displayed visual options is currently represented by the utterance in the call. In various examples, the indicator can take various forms, such as bold text for the currently spoken visual option (or portion thereof), changing the font, color, size, or other visual properties of such text relative to the displayed visual option and other text of other visual options, or adding a pointer to the interface that is visually associated with the word currently being spoken in a call. For example, the pointer can be an icon, arrow, or other object that appears over the word currently being spoken in a call.
[0105] In some implementations, if the confidence level of the text recognition determined from the audio data (as in block 402) falls below a threshold, the calling device may output a notification when displaying visual options in block 412 or 416. In some implementations, the text recognized from the call audio data (for example, in block 402) may be determined to be in a language different from the user's standard language used in the calling device, and this text may be automatically translated so that the call menu selection options are displayed in the user's language.
[0106] In various implementations, one or more blocks may be omitted from Method 400, for example, if certain features of those blocks are not provided in a particular implementation. For example, blocks 406-412 may be omitted in some implementations where cached selection options are not used. In another example, blocks in Figure 3 that are not used to retrieve entity data in a particular implementation may be omitted.
[0107] The methods, blocks, and operations described herein may be performed in a different order than those shown or described in Figures 2–4, and / or may be performed concurrently (partially or completely) with other blocks or operations, where appropriate. For example, block 220 in Figure 2 may be performed at least partially concurrently with blocks 222 and / or 224. In another example, blocks 414 and 416 in Figure 4 may be performed in a different order and / or at least partially concurrently. Some blocks or operations may be performed for one part of the data and then performed again later for another part of the data, for example. Not all of the described blocks and operations must be performed in various implementations. In some implementations, blocks and operations may be performed multiple times, in different orders, and / or at different times in the method.
[0108] One or more methods disclosed herein can operate in multiple environments and platforms, for example, as a standalone computer program that can run on any type of computing device, or as a mobile application ("App") that runs on a mobile computing device.
[0109] One or more methods described herein (e.g., 200, 300, and / or 400) can be run as standalone programs that can run on any type of computing device, as programs that run on a web browser, or as mobile applications ("Apps") that run on mobile computing devices (e.g., mobile phones, smartphones, tablet computers, wearable devices such as watches, armbands, jewelry, and headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, head-mounted displays, laptop computers, etc.). In one example, a client / server architecture can be used, for example, where the mobile computing device (as a client device) sends user input data to a server device and receives final output data for output (e.g., for display) from the server. In another example, all computations of the method can be performed within a mobile app (and / or other app) on a mobile computing device. In yet another example, the computations can be divided between the mobile computing device and one or more server devices.
[0110] In one example, a client / server architecture can be used, for example, where a mobile computing device (as a client device) sends user input data to a server device and receives final output data (for example, for display) from the server. In another example, all calculations can be performed within a mobile app (and / or other app) on the mobile computing device. In yet another example, calculations can be divided between the mobile computing device and one or more server devices.
[0111] The methods described herein can be implemented by computer program instructions or code that can be executed on a computer. For example, the code can be implemented by one or more digital processors (e.g., microprocessors or other processing circuits) and can be stored in computer program products, including non-temporary computer-readable media such as magnetic, optical, or electromagnetic media (e.g., storage media), or semiconductor storage media including semiconductor or solid-state memory, magnetic tape, removable computer diskettes, random-access memory (RAM), read-only memory (ROM), flash memory, rigid magnetic disks, optical disks, solid-state memory drives, etc. The program instructions can also be contained in and provided as electronic signals, for example, in the form of software as a service (SaaS) supplied from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more methods can be implemented in hardware (e.g., logic gates) or in a combination of hardware and software. Exemplary hardware may include programmable processors (e.g., field-programmable gate arrays (FPGAs), complex-programmable logic devices), general-purpose processors, graphics processors, and application-specific integrated circuits (ASICs). One or more methods may be executed as part of or as a component of an application running on a system, or as an application or software running in relation to other applications and operating systems.
[0112] Figure 5 is a schematic diagram of an exemplary user interface 500 displayed on the display screen of a calling device that can initiate a call, in several implementations. For example, the interface 500 can be displayed on a touchscreen by the display device of a different device, such as a client device, e.g., one of the client devices 120-126 shown in Figure 1, or a server device, e.g., a server system 102.
[0113] In some implementations, the user interface 500 can be associated with a calling application program that initiates calls to other devices, answers incoming calls from other devices, and communicates with other devices via a call connection. In this example, the name of the target entity 502 is displayed, where the target entity is selected by the user, for example, directly from a web page, contact list, or other information display, as a result of searching for or navigating to the entity, or as a result from another user application or application process running on the calling device or other device. A calling interface 504 is also displayed, which includes a numeric keypad 508, an identifier input field 510, and call control 512. The keys on the numeric keypad 508 can be selected by the user (for example, via a touchscreen or other input device) to enter the identifier 514 into the input field 510, for example, one character at a time, or in other implementations, multiple characters. The entity identifier 514 is associated with the entity indicated by the entity name 502. A call can be initiated to the entity by using the entity identifier 514. For example, entity identifier 514 is shown as a telephone number, but other types of identifiers may also be entered to enable a call to an entity associated with an identifier (e.g., an email address or other address). In some implementations, identifier 514 can be automatically entered in input field 510 by a calling device in response to a user selection to call a target entity from a different application (e.g., a map application, a web browser, etc.), for example, by displaying interface 500 on the calling device. Call control 512, when selected by the user, can cause the calling device to dial identifier 514 of the target entity indicated by name 502 and initiate a call to the target entity.
[0114] Figure 6 is a schematic diagram of a call interface 600 in which selection options for a call menu are displayed in a call, in several implementations. The call interface 600 can be displayed by a call device (e.g., a client device such as one of the client devices 120-126, or a display device of a different device such as a server system 102). For example, the call interface 600 can be displayed after a call to a target entity 502 has been initiated via the entity identifier 514 shown in Figure 5. In some examples, a call can be initiated in response to a user selection of the call control 512 of interface 500 in Figure 5. Alternatively, a call can be initiated in various other ways, for example, by an application in response to a user selection of a target entity or entity identifier, in response to another event, automatically based on a scheduled command from the user, etc.
[0115] The target entity name 602 may be displayed to indicate the called party in the current call. The duration may also be displayed to indicate the time elapsed since the call began.
[0116] In some implementations, if permission and / or commands are obtained from the user, rewrites of all utterances made by the user and target entity during a call may be rewritten by the calling device and displayed on the user interface 600. For example, rewritten utterances may be displayed in the display area 604 of the user interface 600. The interface 600 may also include various user controls that, when selected by the user, result in control of the call or functions associated with the interface 600. For example, a disconnection control 606 may cause the calling device to disconnect from the call; a keypad control 608 may cause a numeric keypad (or keyboard) to be displayed within or on the interface 600 (similar to a keypad 508); a speaker control 610 may cause the audio output of the calling device to be output as a speakerphone; and a mute control 612 may prevent the user's utterances and other sounds from the calling device from being transmitted to the called party in the call.
[0117] In the example in Figure 6, a call to the target entity is answered by an Interactive Voice Response (IVR) system (as the target entity) which speaks an automated voice and provides selection options in a call menu for the user (caller). Text 614 is rewritten from utterances detected and recognized during the call by the user's calling device. The first portion 616 of the rewritten text 614 corresponds to an utterance from the automated system that provides preface information that is detected as neither a selection option in the call menu nor part of such a selection option. In other implementations, text that is detected as not being included in the selection options, such as the first portion 614, is not displayed by the calling device unless the user has set the device's priority or setting to do so.
[0118] The second portion 618 of text 614 corresponds to an utterance detected in the call as being in a language different from the default language or the language of the first portion 616. The second portion 618 corresponds to a selection option in the call menu provided by the target entity ("Press 9 for Spanish"). In Figure 6, the second portion 618 has not yet been detected as a selection option.
[0119] Figure 7 is a schematic diagram of a call interface 600 in which selection options for a call menu are detected and displayed in several implementations. In the example of Figure 7, a second portion 618 of the rewritten text 614 (shown in Figure 6) is detected as a selection option in the call menu, for example, using speech recognition technology. For example, after detection of a word that designates it as a selection option, the second portion 618 of the text is converted into a selection option having a corresponding visual option 702 displayed on the interface 600 (for example, by the call device or a connected remote device). Furthermore, the text portion 614 is removed from the screen and replaced by the visual option 702. In some implementations, a portion of the text in the second portion 618 may be removed when converting the text to a visual option, as shown for the visual option 702, so that the visual option may be provided more clearly.
[0120] In some implementations, as shown, a border, outline, or other visual separator may be displayed around the text of the visual option 702 to indicate that the visual option 702 is a contoured option that can be selected like a button. In this example, the number specified in the spoken selection option ("nueve") is translated into a selection indicator 704 displayed on (or associated with) the visual option 702, which indicates a number that can be selected on the keypad (e.g., via the keypad control 608) (or, in some implementations, spoken in a call) to select the visual option 702. In this example, when the visual option 702 is selected by the user, the language spoken by the automated system in the current call is translated into the indicated language (Spanish in this example). The selection of the visual option 702 also causes the text displayed on the interface 600, such as the rewritten text and the selection option, to be translated into the selected language.
[0121] The visual option 702 can be selected by the user, for example, by touching the visual option 702 on a touchscreen, by operating an input device, or by user input such as voice commands. For example, if the visual option 702 is selected by touch input via a touchscreen, the selection of this option is transmitted to the target entity by the calling device. In some examples, the selection is transmitted by the calling device, which outputs a signal in a call (e.g., a tone) that provides an equivalent of the signal output when a specified numeric key is pressed by the user, such as the "9" key on the keypad in the example of Figure 7. This causes the target entity to receive a signal indicating that the user has selected the numeric key "9" and the corresponding selection option associated with the visual option 702.
[0122] In the example in Figure 7, visual option 702 is not selected, and the automated system of the target entity continues speaking, describing further selection options in its call menu. The utterance is detected and rewritten as text portion 706, which in this example is displayed as raw text below visual option 702 in display area 604 before being detected as a selection option.
[0123] Figure 8 is a schematic diagram of a call interface 600 in which additional selection options for a call menu are detected and displayed in several implementations. In the example in Figure 8, the text portion 706 (shown in Figure 7) is detected as a selection option. The visual option 802 corresponding to this selection option is displayed in a display area 604 after the visual option 702. The text portion 706 is removed from the screen and replaced by the visual option 802.
[0124] Visual option 802 is selectable by the user. In this case, selection option 802 instructs the user to say a word ("travel") as an indicator to select the option, rather than pressing a key on the keypad as in the case of visual option 702. When the user utters the word to the target entity, the called entity detects the word and the selection of the associated selection option. In some implementations, utterance-selection options can be visually designated in interface 600 to distinguish them from selection options selected by pressing a keypad key. In this example, icon 804 is displayed on visual option 802 (or visually associated with visual option 802) to indicate that the utterance-selection indicator can select this selection option. In some call menus, selection options can be selected either by utterance or by the user pressing a key. In some implementations, a visual option for such a selection option can be displayed with a selection indicator that shows a key identifier and indicates the possibility of selection by utterance.
[0125] In the example in Figure 8, additional selection options are detected after additional utterances from the target entity. These selection options are displayed as visual options sequentially in display area 604. Each visual option can be determined and displayed in the same way as visual options 702 and 802 described above. In some implementations, when the text of each selection option is detected and recognized from the utterances in the call, it is provided as a visual option, as shown in Figures 6 and 7. In some implementations, the utterances in the call may be displayed as raw text until all call menu selection options at the current menu level have been uttered in the call, at which point the text is converted into the visual options that are displayed instead of the text.
[0126] In some implementations, if the user selects one of the visual options before all of the selection options are detected and displayed at the current menu level, the remaining selection options at the current menu level will not be displayed as visual options (for example, upon receiving a selection, the target entity may interrupt uttering further selection options at the current menu level and begin uttering selection options at the next menu level).
[0127] In Figure 8, the user has not selected any of the selection options displayed on interface 600, and the additional selection options are detected and displayed as related visual options, similar to visual options 702 and 802. For example, visual options 806 and 808 are detected as selectable by pressing a key, and selection indicators 810 and 812 are displayed with numbers corresponding to the keypad keys for selecting these options, respectively. Visual option 814 is detected as selectable by pressing the "asterisk" key on the keypad, and selection indicator 816 is displayed with an asterisk symbol.
[0128] Figure 9 is a schematic diagram of a call interface 600 in which a visual option of the call menu is selected by the user, in several implementations. In the example in Figure 9, the user taps the touchscreen of the call device at the location of visual option 808 to select that visual option. In this example, in response to the selection, the call device displays a selected icon 902 instead of a selection indicator 812 (shown in Figure 8) to indicate that visual option 808 has been selected, and the other visual options of the call menu 702, 802, 806, and 814 are displayed more dimmed to highlight the selected visual option 808 (for example, their brightness and / or color are changed to be closer to the brightness / color of the background). Various implementations may provide other methods for highlighting the selected visual option against the other visual options of the call menu that are displayed.
[0129] In response to the selection of visual option 808, the calling device transmits a signal to the target entity in the call indicating the selected digit ("2"). The target entity receives the selected digit and responds accordingly, as described below.
[0130] Figure 10 is a schematic diagram of a call interface 600 with selected visual options in the call menu, in several implementations. In the example of Figure 10, after receiving a selection corresponding to visual option 808 as shown in Figure 9, the target entity changes the call menu to a different level (which may be the second, third, or subsequent levels of the call menu) based on the selected option. In this example, the next menu level in this navigation path of the call menu contains a number of selection options that are translated into visual options 1002, 1004, and 1006, which are spoken by the target entity in the call, detected and displayed by the call device. If further selection options exist at the current level of the call menu, additional visual options may also be displayed. One of these visual options can be selected by the user in the same way as described above with respect to the previous visual options. In some implementations, as shown, the display screen can be scrolled down to display further visual options detected in the call menu.
[0131] Figure 11 is a schematic diagram of a call interface 1100 displayed by a call device, in which, in several implementations, visual options of a call menu are displayed before corresponding selection options are spoken in the call. The call interface 1100 may be similar to the call interface 600 shown in Figure 6. In some implementations, the call interface 1100 may be displayed after the call device has initiated a call to a target entity, such as target entity 502 in Figure 5, via an entity identifier. In some implementations, the call interface 1100 (or a similar interface, such as user interface 500 in Figure 5) may be displayed before initiating a call to a target entity. For example, a selection option may be displayed for the target entity before the call is initiated, showing which options will be available to the user in the call after the call has been initiated via the target entity's entity identifier.
[0132] In this example, a call is initiated, for example, in response to the user selecting the call control 512 of the interface 500 in Figure 5, or by one of the other methods. The name of the target entity 1102 to which the current call's target entity is associated can be displayed, and the duration can be displayed to indicate the time elapsed since the call was initiated. In some implementations, once permission and / or a command is obtained from the user, rewrites of all utterances made by the caller and the called party during the call can be rewritten by the call device and displayed in the display area 1104 of the user interface 1100, as in the case of the call interface 600 in Figure 6. The disconnection control 1106, keypad control 1108, speaker control 1110, and mute control 1112 can be similar to the corresponding controls described above.
[0133] In the example in Figure 11, the call menu 1120 is displayed immediately after or during the start of a call (or may be displayed before the start of a call as described above). In this example, the call menu 1120 includes five visual options 1122, 1124, 1126, 1128, and 1130, similar to the selection options described above for Figures 6 to 10. The selection options for these visual options are accessible to the calling device before the target entity speaks, based on selection option data received before the call, as described herein. By displaying the visual options of the call menu 1120 before the target entity speaks the corresponding selection option, the user can see the call menu in advance, and in some call menu implementations, the user can select a selection option to advance the call menu to another level for the target entity without having to speak the remaining options of the menu.
[0134] In some implementations, as shown in Figure 11, other portions of the spoken content of a target entity in a call to that target entity can be retrieved before the call, as well as the selection options in the call menu, as described herein, and can be displayed before the target entity utters its text during the call. In the example in Figure 11, text 1132 is displayed before the target entity utters its text and is displayed above the call menu selection options in the display area 1104. For example, text 1132 may include introductory information, as in the example in Figure 6, which is detected or known in advance as not being part of the selection options in the call menu based on the selection option information retrieved before the call.
[0135] During a call, the target entity utters utterance information that is detected and recognized by the calling device and / or other connected devices. Generally, the utterance information should match the displayed text and visual options (the utterance information may not exactly match the displayed visual options by converting some of the utterance information into a visual option format such as selection indicator icons or numbers). As described above with respect to Figure 4, if the uttered information does not match the text of a visual option, the visual option may be corrected, and the corrected version will be displayed in place of the original version. In some implementations or cases where correction is not made, a notification may be displayed indicating that an error may exist in the visual option, and / or one or more of the visual options may be removed from the display screen.
[0136] In some implementations shown, an indicator can be displayed to show the portion of the displayed text (including the text in the selection options) that is currently being spoken by the target entity in a call. In this example, the indicator highlights the currently spoken text 1134 in bold font. The subsequent text 1134 and selection options 1122-1130 have not yet been spoken in the call and are displayed in a normal (e.g., non-bold) font and / or are displayed in a less visible font (e.g., with higher or lower brightness depending on the background brightness and / or color). In this example, previously spoken text in the call remains bold, with the spoken portion of the call menu continuing to highlight the new text, so that the new, preceding bold text indicates the text currently being spoken in the call. In some implementations, as shown in Figure 12, previously spoken text that is not part of a selection option may be shown with reduced emphasis relative to the text in the selection option. In some implementations, the currently spoken text may be highlighted in other ways, such as by displaying it in a different color from other displayed text, or by displaying another pointer, arrow, or other visual indicator above or near the currently spoken text.
[0137] The displayed instructions in the currently spoken text by the called party allow the user to see at a glance the progress of the spoken call menu, which may, for example, allow the user to see whether the target entity is currently waiting for the user to select one of the offered options. In some implementations of the call menu, the target entity may not respond to the selection of a selection option until a certain amount of progress has been made in speaking the call menu. For example, a selection option may have to be spoken completely or partially before it becomes selectable. In some of these implementations, providing an indicator of the currently spoken text in the call may allow the user to estimate when a visual option will become eligible to be selected, which may potentially reduce wasted attempts by the user to select an option when the target entity is unresponsive.
[0138] Figure 12 is a schematic diagram of the call interface 1100 of Figure 11, in several implementations, where the indicator of the currently spoken text has progressed to the visual options of the call menu. In this example, the target entity is speaking the remainder of the introductory text 1132 and the selection options provided by the visual options 1122. Thus, all of the text 1132 and the visual options 1122 are displayed in a highlighted format, for example, in bold text and / or more visible. In addition, the initial portion 1202 of the visual option 1124 is currently being spoken by the called party, and as a result, portion 1202 is displayed in a highlighted format compared to the other portions of the visual option 1124 (in some implementations, one or more portions of the visual option may also be highlighted when at least part of the visual option is being spoken, such as selection indicators and / or option borders as shown). Visual options 1126, 1128, and 1130 have not yet been spoken by the target entity and are displayed in a less visible format.
[0139] Figure 13 is a schematic diagram of the call interface 1100 of Figure 11, showing further advancements in the call menu of the currently spoken text indicator in several implementations. In this example, the target entity is speaking the selection options represented by the introductory text 1132 and visual options 1122, 1124, and 1126. Thus, text 1132 and these visual options are displayed in a more visible and emphasized format than before they were spoken. In addition, the initial part 1302 of visual option 1128 is currently spoken by the called party, and as a result, part 1302 is displayed in a more emphasized format compared to the other parts of visual option 1128. Visual option 1130 has not yet been spoken by the target entity and is displayed in a less visible format.
[0140] Figure 14 is a schematic diagram of the call interface 1100 of Figure 11, showing several implementations in which the indicator of the currently spoken text has progressed to the next level in the call menu. In this example, the target entity has spoken the introductory text 1132 and all of the initial level selection options in the call menu. Therefore, text 1132 and these visual options are displayed in a highlighted format (by scrolling the display screen, only visual options 1126, 1128, and 1130 are currently visible in Figure 14). In addition, the user has selected visual option 1128, as indicated by the selection indicator 1402.
[0141] Following the selection of visual option 1128, the next level of the call menu is displayed by the calling device. As with the previous level shown in Figure 11, the next level of visual options are known in advance by the calling device accessing the data indicating these options as described herein, and are displayed before they are uttered in the call. The next level of visual options are displayed as visual options 1404, 1406, 1408, and 1410. In the example in Figure 14, the initial portion 1412 of visual option 1404 is currently uttered by the called party, and as a result, portion 1412 is displayed in a highlighted form compared to the other portions of visual option 1404 (in some implementations, one or more portions of the visual option, such as the selection indicator of the associated key number, may also be highlighted, as shown). Visual options 1406–1410 have not yet been uttered by the target entity and are displayed less visible.
[0142] In the examples in Figures 6 to 14, the user and / or calling device that initiated the call is the caller, and the target entity is the called party in the exemplary call (the called entity). In other examples, the target entity may call the user and / or calling device, in which case the target entity is the caller and the user and / or calling device is the called party.
[0143] Figure 15 is a block diagram of an exemplary device 1500 that may be used to implement one or more features described herein. In one example, device 1500 may be used to implement a client device, for example, one of the client devices 120-126 shown in Figure 1. Alternatively, device 1500 may implement a server device, for example, server device 104, etc. In some implementations, device 1500 may be used to implement a client device, a server device, or a combination of the above. Device 1500 may be any suitable computer system, server, or other electronic or hardware device as described herein.
[0144] In some implementations, device 1500 includes a processor 1502, memory 1504, and an I / O interface 1506. The processor 1502 may be one or more processors and / or processing circuits for executing program code and controlling the basic operations of device 1500. "Processor" includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. A processor may include a system with a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), a multiplexing unit (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex-programmable logic device (CPLD), dedicated circuitry for achieving a function, a dedicated processor for performing neural network model-based processing, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems.
[0145] In some implementations, processor 1502 may include one or more coprocessors that perform neural network processing. In some implementations, processor 1502 may be a processor that processes data to produce a probabilistic output, for example, the output produced by processor 1502 may be inaccurate or accurate within a predetermined range from the expected output. For example, the processor may perform its functions in "real-time," "offline," "batch mode," etc. The processing parts may be performed by different (or the same) processing systems at different times and locations. The computer may be any processor that communicates with memory.
[0146] Memory 1504 is generally provided in device 1500 for access by processor 1502, and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), or flash memory, which is suitable for storing instructions for execution by the processor and is located separately from and / or integrated with processor 1502. Memory 1504 can store software that runs on server device 1500 by processor 1502, including an operating system 1508, a machine learning application 1530, other applications 1512, and application data 1514. Other applications 1512 may include applications such as a data display engine, communication applications (e.g., dialer or calling applications, over-the-top calling applications, other applications with calling capabilities such as applications related to specific entities such as banks, restaurants, or other organizations / providers that provide apps), a web hosting engine, an image display engine, a notification engine, and a social networking engine. In some implementations, the machine learning application 1530 and / or other applications 1512 may each include instructions that enable the processor 1502 to perform some or all of the functions described herein, for example, the methods in Figures 2, 3 and / or 4. The application data 1514 may include call menu data such as selection option data and other entity data, audio data from a call (by user permission), audio data from the call menu, a text record of the call menu, a timestamp of the call selection options and call menu structure indicating recency, call characteristics including the call time of the previous call (by user permission), call duration and other characteristics, and / or data structures (e.g., tables, lists, graphs) that can be used to determine the call selection options described herein.
[0147] The machine learning application 1530 may include one or more named entity recognition (NER) implementations, for which supervised learning and / or unsupervised learning may be used. The machine learning models may include multitask learning-based models, residual task bidirectional LSTMs (long-short-term memory) with conditional random fields, statistical NER, and the like. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application ("App") that runs on a mobile computing device, and so on.
[0148] In various implementations, the machine learning application 1530 may utilize Bayesian classifiers, support vector machines, neural networks, or other learning techniques. In some implementations, the machine learning application 1530 may include a trained model 1534, an inference engine 1536, and data 1532. In some implementations, data 1532 may include training data, for example, data used to generate the trained model 1534. For example, the training data may include any type of data suitable for training a model to determine selection options for a call, such as speech data showing utterances made during a previous call, call menu data showing selection options provided in a call by an entity, and call characteristics of a previous call by the user (if user consent has been obtained). The training data may be obtained from any source, for example, a data repository specifically marked for training, or data for which permission has been granted for use as training data for machine learning. In implementations where one or more users permit the use of their respective user data to train a machine learning model, for example, the trained model 1534, the training data may include such user data. In implementations where a user authorizes the use of each user data, data 1532 may include the authorized data.
[0149] In some implementations, the training data may include synthetic data generated for training purposes, such as data not based on user input or activity in the context being trained, e.g., data generated from simulations or models. In some implementations, the machine learning application 1530 excludes the data 1532. For example, in these implementations, the trained model 1534 may be generated, for example, on a different device and provided as part of the machine learning application 1530. In various implementations, the trained model 1534 may be provided as a data file containing the model structure or format and associated weights. The inference engine 1536 may read the data file for the trained model 1534 and run the neural network with node connectivity, layers, and weights based on the model structure or format specified in the trained model 1534.
[0150] The machine learning application 1530 also includes one or more trained models 1534. For example, such a model may include a trained model for recognizing utterances and determining selection options from utterances received as audio data in a call as described herein. In some implementations, the trained model 1534 may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, e.g., a linear network, a deep neural network that implements multiple layers (e.g., "hidden layers" between input and output layers, each of which is a linear network), a convolutional neural network (e.g., a network that divides or segments input data into many parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results from the processing of each tile), a sequence-to-sequence neural network (e.g., one that takes input sequential data such as words in a sentence or frames in a video, and generates a result sequence as output), etc.
[0151] The model format or structure may specify connectivity between various nodes and the organization of nodes into layers. For example, nodes in the first layer (e.g., the input layer) may receive data as input data 1532 or application data 1514. Such data may include, for example, utterance data from a call, entity data indicating selection options for a call, call characteristics of a previous call, and / or user feedback regarding the previous call and the provided selection options. Subsequent intermediate layers may receive inputs and outputs of nodes in the previous layer for each connectivity specified in the model format or structure. These layers may be called hidden layers. The final layer (e.g., the output layer) generates the output of the machine learning application. For example, the output may be a set of selection options provided in an interface. In some implementations, different layers or models may be used to recognize utterances, for example, by receiving audio data as input and providing an output that is text representing the utterances in the input audio data. In some implementations, the model format or structure also specifies the number and / or types of nodes in each layer.
[0152] In different implementations, one or more trained models 1534 may include multiple nodes arranged in layers according to the model structure or form. In some implementations, a node may be a memoryless computation node configured, for example, to process one unit of input to produce one unit of output. The computations performed by the node may include, for example, multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output.
[0153] In some implementations, the computations performed by the nodes may include applying a step function / activation function to a tuned weighted sum. In some implementations, the step function / activation function may be a nonlinear function. In various implementations, such computations may include operations such as matrix multiplication. In some implementations, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multicore processor, using individual processing units of a GPU, or using dedicated neural circuits. In some implementations, nodes may include memory, for example, to store and use one or more previous inputs when processing subsequent inputs. For example, a node with memory may include a Long-Short-Term Memory (LSTM) node. An LSTM node may use memory to maintain “states” that enable the node to function like a finite state machine (FSM). Models with such nodes may be useful when processing sequential data, such as words in a sentence or paragraph, frames in a video, speech or other audio, etc.
[0154] In some implementations, one or more trained models 1534 may include embeddings or weights for individual nodes. For example, a model may be initialized as a group of nodes organized into layers as specified by the model form or structure. During initialization, each connected node for each model form, for example, the connection between each pair of nodes in a successive layer of a neural network, may be assigned a respective weight. For example, each weight may be assigned randomly or initialized to a default value. The model may then be trained, for example, using data 1532 to produce results.
[0155] For example, training may involve applying supervised learning techniques. In supervised learning, training data may include multiple inputs (e.g., audio data and / or entity data) and corresponding expected outputs for each input (e.g., a set of selection options for a call menu and / or text representing utterances in the audio data). Based on a comparison of the model's output with the expected output, the weight values are automatically adjusted, for example, in a way that increases the probability that the model will produce the expected output when similar inputs are provided.
[0156] In some implementations, training may involve applying unsupervised learning techniques. In unsupervised learning, only input data may be provided, and the model may be trained to distinguish the data, for example, to group the input data into several groups, each group containing similar input data in some way. For example, the model may be trained to determine or collect call characteristics that are similar to one another.
[0157] In another example, a model trained using unsupervised learning may collect utterance or choice option features based on the use of utterances and choice options in the data source. In some implementations, unsupervised learning may be used, for example, to generate knowledge representations that may be used by machine learning application 1530. In various implementations, the trained model includes a set of weights or embeddings corresponding to the model structure. In implementations where data 1532 is omitted, machine learning application 1530 may include a trained model 1534 based on previous training, for example, by the developer of machine learning application 1530, a third party, etc. In some implementations, one or more of the trained models 1534 may each include a fixed set of weights, for example, downloaded from a server that provides weights.
[0158] The machine learning application 1530 also includes an inference engine 1536. The inference engine 1536 is configured to apply a trained model 1534 to data such as application data 1514 in order to provide inferences such as a set of selection options in a call menu and the structure of the call menu. In some implementations, the inference engine 1536 may include software code to be executed by a processor 1502. In some implementations, the inference engine 1536 may specify a circuit configuration (e.g., for a programmable processor, for a field-programmable gate array (FPGA), etc.) that enables the processor 1502 to apply the trained model. In some implementations, the inference engine 1536 may include software instructions, hardware instructions, or a combination thereof. In some implementations, the inference engine 1536 may provide an application programming interface (API) that can be used by the operating system 1508 and / or other applications 1512 to invoke the inference engine 1536, for example, to apply the trained model 1534 to application data 1514 to generate inferences.
[0159] The machine learning application 1530 may offer several technical advantages. For example, if the trained model 1534 is generated based on unsupervised learning, the trained model 1534 can be applied by the inference engine 1536 to generate knowledge representations (e.g., numerical representations) from input data, e.g., application data 1514. For example, a model trained to determine selection options and / or menu structures may generate such representations. In some implementations, such representations may help reduce processing costs (e.g., computational costs, memory usage, etc.) to generate outputs (e.g., labels, classifications, estimated characteristics, etc.). In some implementations, such representations may be provided as input to different machine learning applications that generate outputs from the output of the inference engine 1536.
[0160] In some implementations, the knowledge representations generated by the machine learning application 1530 may be provided to different devices that perform further processing on a network, for example. In such implementations, providing knowledge representations rather than data offers technical advantages, such as enabling faster data transmission while reducing costs.
[0161] In some implementations, the machine learning application 1530 may be implemented offline. In these implementations, the trained model 1534 may be generated in the first stage and provided as part of the machine learning application 1530. In some implementations, the machine learning application 1530 may be implemented online. For example, in such implementations, an application launching the machine learning application 1530 (e.g., one or more of the operating system 1508 and other applications 1512) may utilize the inferences generated by the machine learning application 1530, for example, by providing the inferences to the user and generating system logs (e.g., actions taken by the user based on the inferences, if permitted by the user, or the results of further processing, if used as input for further processing). The system logs may be generated periodically, for example, every hour, every month, every three months, etc., and may be used, with user permission, to update the trained model 1534, for example, by updating the embeddings for the trained model 1534.
[0162] In some implementations, the machine learning application 1530 may be implemented in a form that can adapt to a specific configuration of the device 1500 on which the machine learning application 1530 is executed. For example, the machine learning application 1530 may determine the computation graph that utilizes available computing resources, e.g., the processor 1502. For example, if the machine learning application 1530 is implemented as a distributed application across multiple devices, the machine learning application 1530 may determine which computations are performed on each device in a form that optimizes the computations. In another example, the machine learning application 1530 may determine that the processor 1502 includes a GPU with a specific number (e.g., 1000) GPU cores, and the inference engine may be implemented accordingly (e.g., as 1000 individual processes or threads).
[0163] In some implementations, the machine learning application 1530 may perform an ensemble of trained models. For example, the trained models 1534 may include multiple trained models, each applicable to the same input data. In these implementations, the machine learning application 1530 may select a particular trained model based on, for example, available computational resources, success rate with previous inferences, etc. In some implementations, the machine learning application 1530 may run the inference engine 1536 so that multiple trained models are applied. In these implementations, the machine learning application 1530 may combine the outputs from providing individual models by using a voting technique that records the individual outputs from providing each trained model, or by selecting one or more specific outputs. Furthermore, in these implementations, the machine learning application may apply a time threshold (e.g., 0.5 ms) for providing individual trained models and may utilize only these individual outputs available within the time threshold. Outputs not received within the time threshold may not be utilized, for example, may be discarded. For example, this approach may be suitable when there is a time limit specified while launching a machine learning application, for instance, by the operating system 1508 or one or more other applications 1512.
[0164] In different implementations, the machine learning application 1530 may produce different types of output. In some implementations, the machine learning application 1530 may produce output based on a format specified by the launching application, e.g., the operating system 1508 or one or more other applications 1512. In some implementations, the launching application may be another machine learning application. For example, such a configuration may be used in a generative adversarial network, where the launching machine learning application is trained using the output from the machine learning application 1530, and vice versa.
[0165] Any software in memory 1504 may, alternatively, be stored in any other suitable storage location or computer-readable medium. In addition, memory 1504 (and / or other connected storage devices) may store one or more messages, one or more classifications, electronic encyclopedias, dictionaries, glossaries, knowledge bases, message data, grammars, user selections, and / or other instructions and data used in the features described herein. Memory 1504 and any other type of storage (magnetic disks, optical disks, magnetic tapes or other tangible media) may be considered as “storage” or “storage devices.”
[0166] The I / O interface 1506 can provide functionality to enable the server device 1500 to interface with other systems and devices. Interfaced devices can be included as part of device 1500 or are separate and capable of communicating with device 1500. For example, network communication devices, storage devices (e.g., memory 1504 and / or database 106), and input / output devices can communicate via the I / O interface 1506. In some implementations, the I / O interface can be connected to interface devices such as input devices (keyboards, pointing devices, touchscreens, microphones, cameras, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, motors, etc.).
[0167] Some examples of interfaced devices that can be connected to the I / O interface 1506 include one or more display devices 1520 and one or more data stores 1538 (as described above). Display device 1520 can be used to display content, for example, a user interface for an output application as described herein. Display device 1520 can be any suitable display device, and can be connected to device 1500 via a local connection (e.g., a display bus) and / or a network connection. Display device 1520 can include any suitable display device, such as an LCD, LED, or plasma display screen, a CRT, a television, a monitor, a touchscreen, a 3D display screen, or any other visual display device. For example, display device 1520 could be a flat display screen provided in a mobile device, a multi-display screen provided in a goggles or headset device, a projector, or a monitor screen for a computer device.
[0168] The I / O interface 1506 can interface with other input and output devices. Some examples include display devices, printer devices, scanner devices, etc. Some implementations can provide a microphone for capturing sound, voice commands, etc., an audio speaker device for outputting sound, or other input and output devices.
[0169] For ease of explanation, Figure 15 shows one block for each of the following: processor 1502, memory 1504, I / O interface 1506, and software blocks 1508, 1512, and 1530. These blocks may represent one or more processors or processing circuits, operating systems, memory, I / O interfaces, applications, and / or software modules. In other embodiments, device 1500 does not have to have all of the components shown and / or may have other components, including other types of elements, instead of or in addition to those shown herein. While some components are described to perform blocks and operations as described in some embodiments herein, any suitable component or combination of components of environment 100, device 1500, a similar system, or any suitable processor of one or more suitable processors associated with such a system may perform the described blocks and operations.
[0170] The explanations describe specific implementations, but these specific implementations are merely illustrative and not limiting. Concepts shown in the examples may apply to other examples and implementations.
[0171] In addition to the above description, users may be provided with controls that enable them to choose whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., the user's social network, social behavior or activities, occupation, user choices, or the current location of the user or user device), and whether the user receives content or communications from the server. In addition, some data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information cannot be determined for the user, or a user's geographical location may be generalized (to the city, zip code, or state level, etc.) where the location information is obtained so that the user's specific location cannot be determined. Thus, users may have control over what information is collected about them, how that information is used, and what information is provided to them.
[0172] It should be noted that the functional blocks, operations, features, methods, devices, and systems described herein may be integrated or divided into different combinations of systems, devices, and functional blocks as known to those skilled in the art. Any suitable programming language and programming technique may be used to perform the routines of a particular implementation. Different programming techniques, such as procedural or object-oriented, may be employed. The routines may be executed on a single processing device or on multiple processors. Steps, operations, or calculations may be represented in a particular order, but the order may be changed in different particular implementations. In some implementations, multiple steps or operations shown herein as sequential may be executed simultaneously.
Claims
1. A method by which a computer performs an action. In a call between a calling device and a device associated with a target entity, the acquisition of selection option data including one or more selection options for the user of the calling device to navigate through a call menu provided by the target entity, The calling device displays one or more visual options corresponding to the one or more selection options before receiving audio data including utterances indicating the one or more selection options, each of the one or more visual options being selectable via user input to cause a corresponding movement through the calling menu, and each of the one or more visual options including its respective text, The aforementioned method, Receiving audio data in the aforementioned call, In order to determine the text representing the utterance in the aforementioned audio data, the audio data is analyzed by a program, A method performed by a computer, further comprising displaying a visual indicator during the call, wherein the visual indicator highlights a particular portion of the text of one or more visual options displayed during the call, and the particular portion of the text is currently being received during the call in the utterance in the audio data.
2. The further includes, in response to receiving a selection of a particular visual option from the one or more visual options, causing the device associated with the target entity to transmit an instruction for the selection, the instruction being: A signal corresponding to pressing a key on a keypad associated with the aforementioned specific visual option, or Utterances provided by the calling device in the call, including indicators related to the specific visual options. One of these is the method performed by the computer described in claim 1.
3. The computer method according to claim 1, wherein each of the one or more visual options can be selected via touch input on the touchscreen of the communication device.
4. The audio data is first audio data, and the method responds to receiving a selection of a specific visual option from the one or more visual options. In the aforementioned call, the call device receives second audio data which includes a second utterance indicating one or more second selection options for the user, In order to determine the second text representing the second utterance in the second audio data, the second audio data is analyzed by a program, Determining one or more second selection options based on programmatic analysis of at least one of the second text or the second audio data, A method performed by a computer according to claim 1, further comprising displaying at least a portion of the second text by the calling device, wherein the at least portion of the second text is displayed as one or more second visual options corresponding to one or more second selection options, and each of the one or more second visual options is selectable via second user input to cause a corresponding movement through the calling menu.
5. The method performed by a computer according to claim 1, wherein the one or more selection options are a plurality of selection options, and further comprises programmatically analyzing at least one of the text or the audio data to determine the hierarchical structure of the plurality of selection options in the call menu.
6. The method performed by a computer according to claim 1, wherein the one or more selection options in the selection option data are determined by programmatically analyzing audio data received during a previous call.
7. The method performed by a computer according to claim 6, wherein the selection option data is cached on the calling device before the start of the call, the selection option data relates to entity identifiers previously called by a caller in the geographic area of the calling device, and the entity identifiers have been called at least a threshold number of times, or more times than other entity identifiers not related to the selection option data.
8. The selection option data is compared with one or more selection options determined from the audio data. To determine whether there is a mismatch between the aforementioned selection option data and one or more selection options determined from the audio data, A method performed by a computer according to claim 1, further comprising causing the calling device to generate a notification of the mismatch in response to determining a mismatch between the selection option data and one or more selection options determined from the audio data.
9. Comparing the selection option data with one or more selection options determined from the audio data, To determine whether there is a mismatch between the aforementioned selection option data and one or more selection options determined from the audio data, A method performed by a computer according to claim 1, further comprising: modifying the selection option data to match one or more selection options determined from the audio data, in response to determining a mismatch between the selection option data and one or more selection options determined from the audio data.
10. Comparing the aforementioned selection option data with one or more selection options determined from the audio data is: Comparing the text of the selection option data with the text of one or more selection options determined from the audio data, or The audio data of the selected option data is compared with the audio data received during the call. A method performed by a computer according to claim 8, which includes one of the following:
11. Comparing the aforementioned selection option data with one or more selection options determined from the audio data is: Comparing the text of the selection option data with the text of one or more selection options determined from the audio data, or The audio data of the selected option data is compared with the audio data received during the call. A method performed by a computer according to claim 9, which includes one of the following.
12. The one or more selection options are stored in at least one of the storage of the calling device or the storage of a remote device that communicates with the calling device on the communication network. To search for the one or more selection options for the next call between the communication device and the target entity. A method performed by a computer according to claim 1, further comprising:
13. A method performed by a computer, In a call between a calling device and a device associated with a target entity, the acquisition of selection option data including one or more selection options for the user of the calling device to navigate through a call menu provided by the target entity, The calling device displays one or more visual options corresponding to the one or more selection options before it receives audio data including utterances indicating the one or more selection options, each of the one or more visual options being selectable via user input to cause a corresponding movement through the calling menu. The one or more selection options in the selection option data are determined by programmatically analyzing audio data received during the previous call. A method performed by a computer, wherein the selection option data is cached on the calling device before the start of the call, the selection option data is related to entity identifiers previously called by the calling party in the geographical area of the calling device, and the entity identifiers have been called at least a threshold number of times, or more times than other entity identifiers not related to the selection option data.
14. A method performed by a computer, In a call between a calling device and a device associated with a target entity, the acquisition of selection option data including one or more selection options for the user of the calling device to navigate through a call menu provided by the target entity, The calling device displays one or more visual options corresponding to the one or more selection options before receiving audio data including utterances indicating the one or more selection options, each of the one or more visual options being selectable via user input to cause a corresponding movement through the calling menu, and each of the one or more visual options including its respective text, The aforementioned method, Receiving audio data in the aforementioned call, The selection option data is compared with one or more selection options determined from the audio data. To determine whether there is a mismatch between the aforementioned selection option data and one or more selection options determined from the audio data, A method performed by a computer, further comprising causing the calling device to generate a notification of the mismatch in response to determining a mismatch between the selection option data and one or more selection options determined from the audio data.
15. A method performed by a computer, In a call between a calling device and a device associated with a target entity, the acquisition of selection option data including one or more selection options for the user of the calling device to navigate through a call menu provided by the target entity, The calling device displays one or more visual options corresponding to the one or more selection options before receiving audio data including utterances indicating the one or more selection options, each of the one or more visual options being selectable via user input to cause a corresponding movement through the calling menu, and each of the one or more visual options including its respective text, The aforementioned method, Receiving audio data in the aforementioned call, The selection option data is compared with one or more selection options determined from the audio data. To determine whether there is a mismatch between the aforementioned selection option data and one or more selection options determined from the audio data, A method performed by a computer, further comprising: modifying the selection option data to match one or more selection options determined from the audio data, in response to determining a mismatch between the selection option data and one or more selection options determined from the audio data.
16. A calling device for displaying selection options for making a call, wherein the calling device is The memory where the instructions are stored, Display devices and, A calling device for displaying selection options for a call, comprising: at least one processor coupled to the memory, wherein the at least one processor is configured to access the instructions from the memory in order to perform the method according to any one of claims 1 to 15.
17. A program that, when executed by a processor, includes instructions causing the processor to perform the method according to any one of claims 1 to 15.