Call screening with smart reply
By using a caller application on a computing device to filter incoming calls through natural language dialogue and generate candidate responses, the problem of mobile device users receiving calls at inconvenient times is solved, achieving response without user interaction and privacy protection.
Patent Information
- Application Number
- CN202480032698.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-15
- Filing Date
- 2024-05-10
- Publication Date
- 2025-12-12
AI Technical Summary
Mobile device users often receive calls at inconvenient times, especially spam calls, and existing technologies struggle to effectively screen and answer these calls, particularly bot calls and calls from spoofed numbers.
The computing device conducts natural language dialogue through a caller application, filters incoming calls, determines caller information, and generates candidate responses. Users can choose to reply and send, reducing user interaction.
It enables call filtering and answering without user interaction, reducing the need for user responses, protecting user privacy, and is suitable for various scenarios.
Smart Images

Figure CN121128159A_ABST
Abstract
Description
CLAIM OF PRIORITY
[0001] This application is a PCT of U.S. Provisional Patent Application No. 63 / 502,298, filed May 15, 2023, which claims provisional priority to the benefit of the filing date. The entire contents of this application are incorporated herein by reference. BACKGROUND
[0002] The widespread use of mobile devices, such as smartphones, has enabled users to place and receive calls at any location and at any time. As such, users of mobile devices can often receive calls at inopportune times, such as when the user is busy. Further, users of mobile devices can also receive spam calls, such as robot calls, and such spammers can spoof or fake their numbers in order to bypass anti-spam tools. As a result, users of mobile devices can be reluctant to answer incoming calls and can instead allow incoming calls to go to voicemail. SUMMARY
[0003] Generally, the technology of the present disclosure relates to technology for enabling a computing device to screen an incoming call received by the computing device. The computing device can screen the incoming call to gather information related to the call, such as to determine the identity of the calling party and determine the purpose of the call. Such information gathered by the computing device can enable a user of the computing device to determine whether to take the call, whether the call is a spam call, and the like.
[0004] According to aspects of the present disclosure, the computing device can screen the incoming call without the involvement of a user of the computing device. Instead, the computing device can be able to answer the call by engaging in a natural language conversation with the calling party without user input to determine information related to the call. While the computing device screens the incoming call, the computing device can output a real-time transcription of the conversation with the calling party so that a user of the computing device can be able to follow the conversation.
[0005] While the computing device screens the call, the computing device can determine one or more candidate replies for replying to what the calling party is saying in the call based on the context of the call. The computing device can determine one or more candidate replies related to the conversation, such as to answer a question or inquiry received in the conversation. A user of the computing device can select a candidate reply from the one or more candidate replies, and the computing device can generate and send a reply corresponding to the selected candidate reply in the conversation.
[0006] In one example, the disclosure describes a method that includes establishing, by one or more processors of a computing device, a call with a remote computing device, and having a conversation, by the one or more processors, with the remote computing device in the call. The method can further include determining, by the one or more processors and based at least in part on contextual information associated with the call, one or more candidate replies, receiving, by the one or more processors, an indication of user input that selects a candidate reply from the one or more candidate replies, and in response to receiving the indication of the user input that selects the candidate reply, sending, by the one or more processors, a reply corresponding to the candidate reply in the conversation.
[0007] In another example, the disclosure describes a computing device that includes a memory and one or more processors implemented in circuitry in communication with the memory. The one or more processors can be configured to establish a call with a remote computing device, have a conversation with the remote computing device in the call, determine one or more candidate replies based at least in part on contextual information associated with the call, receive an indication of user input that selects a candidate reply from the one or more candidate replies, and in response to receiving the indication of the user input that selects the candidate reply, send a reply corresponding to the candidate reply in the conversation.
[0008] In another example, the disclosure describes a non-transitory computer-readable storage medium encoded with instructions that, when executed by one or more processors, cause the one or more processors to establish a call with a remote computing device, and have a conversation, by the one or more processors, with the remote computing device in the call. The instructions can further cause the one or more processors to determine one or more candidate replies based at least in part on contextual information associated with the call, receive an indication of user input that selects a candidate reply from the one or more candidate replies, and in response to receiving the indication of the user input that selects the candidate reply, send a reply corresponding to the candidate reply in the conversation.
[0009] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1A and Figure 1B is a conceptual diagram illustrating an example environment for performing call screening with intelligent replies, in accordance with one or more aspects of the present disclosure.
[0011] Figure 2 is a block diagram illustrating further details of an example computing device, in accordance with one or more aspects of the present disclosure.
[0012] Figure 3A and Figure 3B Techniques for call screening for incoming calls according to various aspects of this disclosure are shown.
[0013] Figure 4 Additional techniques for call screening according to various aspects of this disclosure are shown.
[0014] Figure 5 This is a flowchart illustrating example techniques for determining one or more candidate responses according to various aspects of this disclosure.
[0015] Figure 6 Example candidate responses related to use cases of a call are shown in accordance with various aspects of this disclosure.
[0016] Figure 7 This is a flowchart illustrating example operations performed by an example computing device configured to perform call screening, according to one or more aspects of this disclosure. Detailed Implementation
[0017] Figure 1A and Figure 1B This is a conceptual diagram illustrating an example environment for performing call screening with intelligent response, according to one or more aspects of this disclosure. Figure 1A and Figure 1B In one example, environment 100 may include computing device 102 connected to network 130 to make and receive calls to and from other remote communication devices, such as remote computing device 136.
[0018] like Figure 1A As shown, computing device 102 can represent an individual mobile or non-mobile computing device. Examples of computing device 102 include mobile phones, tablet computers, laptop computers, desktop computers, servers, mainframes, set-top boxes, televisions, wearable devices (e.g., computerized watches, computerized eyewear, computerized headsets, computerized gloves, etc.), home automation devices or systems (e.g., smart thermostats or home assistant devices), personal digital assistants (PDAs), gaming systems, media players, e-book readers, mobile TV platforms, car navigation or infotainment systems, or any other type of mobile, non-mobile, wearable, and non-wearable computing device.
[0019] Computing device 102 may be connected to network 130 to make and receive calls to remote computing devices such as remote computing device 136. Network 130 refers to a public or private communication network used for transmitting data between computing systems, servers, and computing devices, such as the Internet, Wi-Fi, one or more wireless wide area networks (e.g., wireless cellular networks, satellite networks, and / or free-space optical communication networks), one or more telephone networks (such as one or more public switched telephone networks (PTSNs)), one or more VoIP services, and / or other types of networks. Network 130 may include one or more network hubs, network switches, network routers, or any other network equipment operatively coupled to each other to provide information exchange between computing device 102 and remote computing devices such as remote computing device 136. Computing device 102 and remote computing devices such as remote computing device 136 may use any suitable communication technology to transmit and receive data across network 130. Computing device 102 may be operatively coupled to network 130 using a corresponding network link such as Ethernet, Wi-Fi, cellular connection, or any other type of wired and / or wireless network connection.
[0020] Network 130 can implement any suitable technology and may include any suitable network enabling computing device 102 to make calls to and receive calls from remote computing devices. In some examples, network 130 may implement an IP Multimedia Subsystem (IMS) that manages call sessions, including call routing, authentication, and accounting. The IMS may also act as a Session Initiation Protocol (SIP) server that uses SIP to perform call setup and teardown functions and to perform signaling and messaging protocols for calls between computing devices. In some examples, network 130 may implement the functionality of an Evolved Packet Group (EPG) core or a System Architecture Evolution (SAE) core to handle communication between devices connected to network 130 and networks outside of network 130.
[0021] Remote computing device 136 refers to a device connected to network 130 for making and receiving calls (such as voice calls and / or video calls). Examples of remote computing device 136 include landline telephones, mobile phones, tablet computers, laptop computers, desktop computers, satellite phones, or any other type of device capable of communicating with network 130.
[0022] Computing device 102 includes a user interface component (UIC) 104, a user interface module 106 (“UI module 106”), a caller application 108, and a dialogue model 152. The UIC 104 of computing device 102 can function as both an input device and an output device for computing device 102. UIC 104 can be implemented using various technologies. For example, UIC 104 can function as an input device using a presence-sensitive input screen, such as a resistive touchscreen, surface acoustic wave touchscreen, capacitive touchscreen, projected capacitive touchscreen, pressure-sensitive screen, acoustic impulse recognition touchscreen, or another presence-sensitive display technology, such as radar-based presence-sensitive technology, millimeter-wave-based presence-sensitive technology, ultra-wideband presence-sensitive technology, etc. In some examples, UIC 104 can function as an input device using one or more audio input devices, such as one or more microphones. UIC 104 can function as an output (e.g., display) device using any one or more display devices, such as a liquid crystal display (LCD), a dot matrix display, a light-emitting diode (LED) display, a micro LED, an organic light-emitting diode (OLED) display, electronic ink, or a similar monochrome or color display capable of outputting visible information to a user of computing device 102. In some examples, UIC 104 can function as an audio output device and can include one or more speakers, one or more headphones, or any other audio output device capable of outputting audible information to a user of computing device 102.
[0023] In some examples, the UIC 104 of computing device 102 may include a presence-sensitive display that can receive tactile input from a user of computing device 102. UIC 104 can receive tactile input indication by detecting one or more gestures from the user of computing device 102 (e.g., the user touching or pointing at one or more locations of UIC 104 with a finger or stylus). UIC 104 may present output to the user, for example, at the presence-sensitive display. UIC 104 may present the output as a graphical user interface (e.g., graphical user interfaces 114A-114D), which may be associated with functionality provided by computing device 102. For example, UIC 104 may present various user interfaces of components of a computing platform, operating system, application (e.g., caller application 108), or service (e.g., email messaging application, internet browser application, mobile operating system, etc.) that are executed at or accessible by computing device 102. The user can interact with the appropriate user interface to cause computing device 102 to perform function-related operations.
[0024] UI module 106, caller application 108, and dialogue model 152 can perform the operations described herein using software, hardware, firmware, or a mixture of hardware, software, and firmware residing in or executing on computing device 102 or on one or more other remote computing devices. In some examples, UI module 106, caller application 108, and dialogue model 152 can be implemented as hardware, software, and / or a combination of hardware and software. Computing device 102 can utilize one or more processors to execute UI module 106, caller application 108, and dialogue model 152. Computing device 102 can execute any of UI module 106, caller application 108, and dialogue model 152 as a virtual machine executing on the underlying hardware or within a virtual machine executing on the underlying hardware. UI module 106 and caller application 108 can be implemented in various ways. For example, any of UI module 106, caller application 108, and dialogue model 152 can be implemented as a downloadable or pre-installed application or "app". In another example, any of UI module 106, caller application 108, and dialogue model 152 can be implemented as part of the operating system of computing device 102. Other examples of implementing the techniques of this disclosure on computing device 102 may include Figure 1A Additional components not shown.
[0025] exist Figure 1A In the example, dialogue model 152 may include a hardware device with various hardware, firmware, and software components. However, Figure 1AOnly one specific example of dialogue model 152 has been shown, and many other examples of dialogue model 152 can be used according to the techniques of this disclosure. In some examples, components of dialogue model 152 may be located in a single location. In other examples, one or more components of dialogue model 152 may be located in different locations (e.g., connected via network 130). That is, in some examples, dialogue model 152 may be part of a conventional computing device, while in other examples, dialogue model 152 may be part of a distributed or “cloud” computing system. Further, dialogue model 152 may be located in a computing system remote from computing device 102 and communicatively coupled to and executed by that computing system. In such examples, computing device 102 may generate and receive messages, data, or otherwise exchange information with a remote computing system, such that the remote computing system provides the functionality of dialogue model 152 to computing device 102. In some examples, dialogue model 152 may be functionally partitioned between computing device 102 and one or more remote computing systems, such that a portion of the functionality provided by dialogue model 152 is executed locally at computing device 102, and other portions of the functionality are executed by one or more remote computing systems.
[0026] UI module 106 can interpret input detected at UIC 104. UI module 106 can relay information about the input detected at UIC 104 to one or more associated platforms, operating systems, applications, and / or services executing at computing device 102, causing computing device 102 to perform functions. UI module 106 can also receive information and instructions from one or more associated platforms, operating systems, applications, and / or services executing at computing device 102 (e.g., caller application 108) for generating a GUI. Additionally, UI module 106 can act as an intermediary between one or more associated platforms, operating systems, applications, and / or services executing at computing device 102 and various output devices of computing device 102 (e.g., speakers, LED indicators, vibrators, etc.) to utilize computing device 102 to generate output (e.g., graphical output, audible output, haptic output, etc.).
[0027] Caller application 108 may include functionality for making calls to and receiving calls from remote computing devices, such as remote computing device 136, via network 130. Examples of caller application 108 may include telephone dialer applications, IP voice (VoIP) applications, messaging applications with voice calling functionality, video conferencing applications, video calling applications, or any other application that includes functionality for making and receiving telephone calls.
[0028] exist Figure 1AIn the example, application 112 may send data to UI module 106 causing UIC 104 to generate a user interface (GUI) and its elements, such as GUIs 114A-114D (collectively, "GUI 114"). In response, UI module 106 may output instructions and information to UIC 104 based on information received from the application, causing UIC 104 to display the user interface (GUI 114A). When processing input detected by UIC 104, UI module 106 may receive information from UIC 104 in response to input detected at the location of elements displaying the user interface on UIC 104's screen. UI module 106 propagates information about the input detected by UIC 104 to other components of computing device 102 for interpreting the input and for causing computing device 102 to perform one or more functions in response to the input.
[0029] According to various aspects of this disclosure, computing device 102 can establish a call with remote computing device 136, such as by making a call to or receiving a call from remote computing device 136 via network 130. Examples of calls made and received by computing device 102 may include voice calls such as telephone calls or Voice over Internet Protocol (VoIP) calls, video calls such as video conferencing calls, real-time mixed reality sessions, real-time augmented reality sessions, media calls, or any other call between two or more devices. Once computing device 102 has established a call with remote computing device 136, computing device 102 and remote computing device 136 may be able to exchange audio data. Computing device 102 may send audio data, such as spoken words and phrases, to remote computing device 136 during a call. Similarly, computing device 102 may receive audio data, such as spoken words and phrases, from remote computing device 136 during a call.
[0030] The computing device 102 may, in response to receiving an incoming call, determine whether to alert the user of the computing device 102 to the incoming call, such as by determining whether to ring (e.g., audibly output a ringtone and / or output a haptic pattern) to alert the user of the computing device 102 to the incoming call. The computing device 102 may determine whether to alert the user of the incoming call based at least in part on determining whether the incoming call is a spam call. The computing device 102 may use any suitable spam detection technology to determine whether the incoming call is a spam call, such as by determining whether the phone number associated with the incoming call is on a list of known spammers, whether the user of the computing device 102 has previously marked the phone number associated with the incoming call as a spammer, etc. If the computing device 102 determines that the incoming call is a spam call, the computing device 102 may reject the call and / or may send the incoming call to voicemail without alerting the user to the incoming call.
[0031] In some examples, computing device 102 may determine whether to alert the user of the incoming call based at least in part on whether the user is available to answer the call. Computing device 102 may determine whether the user is available to answer the call based on contextual information such as whether the Do Not Disturb feature of computing device 102 is enabled, the user's schedule stored in the calendar of computing device 102 (e.g., whether the user is currently in a scheduled meeting), the current time and / or date, the current location of computing device 102, and the user's current state (e.g., whether the user is currently driving or sleeping).
[0032] If computing device 102 determines that the user of computing device 102 is unavailable to answer an incoming call, computing device 102 may send the incoming call to a voicemail without notifying the user. In some examples, computing device 102 may execute caller application 108 to perform automatic call screening of incoming calls to collect certain information from the party calling from remote computing device 136, such as the identity of that party (e.g., the caller's name and / or the identity of the entity being called), the purpose of the call, and / or any other relevant information. As described below, computing device 102 may perform call screening to conduct natural language conversations with callers associated with remote computing device 136, and may store transcripts and records of the conversations for later viewing by the user of computing device 102.
[0033] exist Figure 1AIn the example, if computing device 102 determines that an incoming call needs to be alerted to the user of computing device 102, computing device 102 may, in response to receiving an incoming call from remote computing device 136, execute caller application 108 to send data to UI module 106 causing UIC 104 to display GUI 114A to alert the user of computing device 102 to the incoming call. For example, UIC 104 may display GUI 114A to alert the user of computing device 102 to the incoming call while the phone is ringing (e.g., audibly outputting a ringtone and / or outputting a haptic mode).
[0034] GUI 114A may include call information 124 associated with an incoming call, such as the telephone number from which the incoming call originated, the name of the person or entity associated with that telephone number, or whether the name of the person or entity associated with that telephone number is unknown. GUI 114A may include a call answering UI element 120 that the user can (e.g., by providing user input at UIC 104) select to answer the call.
[0035] GUI 114A may also include a call filtering UI element 122 that can be selected by the user (e.g., by providing user input at UIC 104) to cause the computing device 102 to perform call filtering for incoming calls. The computing device 102 may execute a caller application 108 to perform call filtering for incoming calls to collect certain information from the party calling from the remote computing device 136, such as the identity of that party (e.g., the caller's name and / or the identity of the entity being called), the purpose of the call, and / or any other relevant information.
[0036] Computing device 102 can perform call screening of incoming calls without the user of computing device 102 having to answer incoming calls and without the user of computing device 102 having to converse with the party calling from remote computing device 136. Instead, caller application 108 can perform call screening of incoming calls by answering incoming calls to establish a call between computing device 102 and remote computing device 136, and by engaging in natural language conversation with the party calling from remote computing device 136 using a human-like voice with human-like voice characteristics. That is, caller application 108 can receive speech from remote computing device 136, such as words, phrases, and sentences spoken by the user of remote computing device 136, and generate natural language speech (such as by audibly outputting natural language speech during a call), such as spoken words, phrases, and sentences, that is sent to remote computing device 136.
[0037] Caller application 108 may be able to engage in natural language conversation with remote computing device 136 to collect information from the party calling from remote computing device 136, such as the identity of that party (e.g., the caller's name and / or the identity of the entity being called) and the purpose of the call. Caller application 108 may be able to record the audio of the call and save a transcript of the call at computing device 102, so that a user of computing device 102 may be able to listen to the recording of the call and / or read the transcript of the call later.
[0038] Caller application 108 can engage in natural language conversations with remote computing device 136 without user interaction. That is, caller application 108 can generate utterances during a call and send such utterances to remote computing device 136 without user interaction. For example, caller application 108 can determine an appropriate response to utterances received from remote computing device 136 without user input instructing caller application 108 how to respond to the received utterances. In this way, caller application 108 can be able to engage in multi-turn natural language conversations with remote computing device 136.
[0039] When caller application 108 performs call screening, it can suppress the output of audio of the natural language conversation taking place with remote computing device 136. Caller application 108 can also suppress the transmission to remote computing device 136 of any audio that can be captured by an audio input device (e.g., a microphone) of UIC 104, and / or disable such an audio input device of UIC 104. Conversely, to enable the user of computing device 102 to follow the natural language conversation taking place between computing device 102 and remote computing device 136, when caller application 108 performs call screening of calls received from remote computing device 136 by engaging in natural language conversation with remote computing device 136 during a call, caller application 108 can output a real-time text transcription of the natural language conversation that occurs during the call.
[0040] like Figure 1A As shown, as part of performing call screening for incoming calls from remote computing device 136, caller application 108 can send data to UI module 106 causing UIC 104 to display GUI 114B, which includes a real-time transcript 116A of the conversation occurring during the call between computing device 102 and remote computing device 136. As can be seen in the real-time transcript 116A of the conversation, caller application 108 can begin the conversation by greeting the user of remote computing device 136 and asking for the purpose of the call (e.g., “Go ahead and say why you’re calling”).
[0041] Caller application 108 can detect prolonged silences during a call while the conversation is ongoing, such as by determining that caller application 108 has not received any speech from remote computing device 136 for a certain period of time (e.g., 5 seconds, 10 seconds, etc.). In response to detecting a prolonged silence, caller application 108 can prompt the caller to speak (e.g., “I’m sorry I didn’t catch that. What did you say?”). The caller application can engage in conversation to gather information about the call, such as the caller’s name and the purpose of the call. Thus, if caller application 108 determines, based on the conversation already conducted, that the caller has identified themselves but has not stated their purpose for calling, caller application 108 can inquire about the purpose of the call (e.g., “Go ahead and say why you’re calling”).
[0042] Caller application 108 can use dialogue model 152 to determine one or more words, phrases, sentences, etc., to be spoken as part of the dialogue, and generate utterances of such words, phrases, and sentences as part of the dialogue. Dialogue model 152 may include one or more neural networks, such as generative adversarial networks (GANs), recurrent neural networks (RNNs), etc., trained via machine learning on a corpus of anonymous telephone dialogue data to determine responses to utterances received from remote computing device 136.
[0043] For example, caller application 108 may, in response to receiving utterances from remote computing device 136 during a call, use automatic speech recognition to convert the utterances into text, and may input the converted text along with any other relevant contextual information into dialogue model 152. This other relevant contextual information may include prior utterances in the conversation during the call (e.g., words, phrases, and / or sentences previously spoken by the parties on the call), the vocal characteristics (e.g., tone of voice) of the utterances received from remote computing device 136, whether the identity of remote computing device 136 is listed in the contacts of computing device 102, the location of computing device 102 and / or remote computing device 136, the current time and / or date, events listed in the calendar application of computing device 102, previous conversations with the party using remote computing device 136, or any other relevant contextual information. Dialogue model 152 may, based on the input data, determine one or more words, phrases, and / or words that caller application 108 can convert (e.g., via text-to-speech) into utterances that caller application 108 can send as part of a conversation to remote computing device 136.
[0044] According to various aspects of this disclosure, when computing device 102 is in dialogue with remote computing device 136, caller application 108 can determine one or more candidate responses related to the dialogue. Caller application 108 can output an indication of one or more candidate responses for display at UIC 104, enabling a user of computing device 102 to select a candidate response. Caller application 108 can generate and send a response corresponding to the selected candidate response during the dialogue.
[0045] Caller application 108 can determine one or more candidate responses related to the dialogue. In the case of a dialogue, a relevant response to the dialogue can be a response related to responding to a utterance received most recently from the remote computing device 136 as part of the dialogue, and / or a response that is highly likely or may be selected by the user to respond to a utterance received most recently from the remote computing device 136 as part of the dialogue.
[0046] Caller application 108 can determine one or more candidate responses based on contextual information such as contextual information associated with the call. Such contextual information may include previous utterances in the conversation during the call (e.g., words, phrases, and / or sentences previously spoken by the parties on the call), vocal characteristics (e.g., tone of voice) of utterances received from remote computing device 136, whether the identity of remote computing device 136 is listed in the contacts of computing device 102, whether the caller is from a business or other entity, the location of computing device 102 and / or remote computing device 136, the current time and / or date, events listed in the calendar application of computing device 102, previous conversations with the party using remote computing device 136, the use case of the call, or any other relevant contextual information.
[0047] Caller application 108 can use dialogue model 152 to determine one or more candidate responses. For example, caller application 108 can, in response to receiving a utterance from remote computing device 136 during a call, input context information into dialogue model 152 to determine one or more candidate responses relevant to the dialogue based on the input data.
[0048] The caller application 108 can output an indication of each of one or more candidate responses, such as to be displayed at UIC 104, allowing the user to interact with UIC 104 to select a candidate response from one or more candidate responses. Figure 1A and Figure 1B In the example, as computing device 102 continues to communicate with remote computing device 136, caller application 108 can send data to UI module 106 causing UIC 104 to display GUI 114C including UI elements 132A and 132B, each corresponding to a candidate response determined by computing device 102.
[0049] GUI 114C also includes an updated real-time transcription 116B of the dialogue occurring during a call between computing device 102 and remote computing device 136. As shown in the updated real-time transcription 116B, one party at remote computing device 136 is calling to confirm a doctor's appointment for tomorrow at 2 PM. Caller application 108 can determine one or more candidate responses related to the dialogue, such as by determining that one party at remote computing device 136 is calling to confirm a doctor's appointment for tomorrow at 2 PM. Caller application 108 can therefore determine one or more candidate responses for responding to the request to confirm the appointment, such as a candidate response to confirm the appointment and a candidate response not to confirm the appointment, and caller application 108 can cause GUI 114C to include UI element 132A corresponding to the candidate response to confirm the appointment and UI element 132B corresponding to the candidate response not to confirm the appointment.
[0050] The computing device 102 can receive instructions from UIC 104 regarding user input to select a candidate response from one or more candidate responses, and in response, can send a response corresponding to the selected candidate response in the conversation. For example, if a user interacts with UIC 104 to select UI element 132A, UI module 106 can receive instructions from UIC 104 regarding user input to select UI element 132, and caller application 108 can receive instructions from UI module 106 that the user has selected UI element 132 corresponding to the candidate response confirming the appointment.
[0051] The caller application 108 may not necessarily send a reply that is word-for-word identical to the candidate reply. Instead, the caller application 108 may generate a reply (e.g., one or more words, phrases, and / or sentences) that has the same or similar meaning as the selected candidate reply, and may send the spoken version of the reply as part of the conversation to the remote computing device 136.
[0052] like Figure 1BAs shown, in response to a user interacting with UIC 104 to select a UI element 132A corresponding to a candidate response confirming an appointment, caller application 108 can generate a response confirming the appointment and can send this response as part of the conversation to remote computing device 136. Caller application 108 can send data to UI module 106 causing UIC 104 to display GUI 114D, which includes an updated real-time transcript 116C of the conversation. As can be seen in the updated real-time transcript 116C of the conversation, caller application 108 can generate a response “Yes we confirm the appointment for 2PM tomorrow” and can send this response as part of the conversation to remote computing device 136.
[0053] As caller application 108 continues to perform call screening and engage in dialogue with telecomputing device 136, caller application 108 can continue to update the candidate responses output for display by UIC 104. That is, the candidate responses determined by caller application 108 do not remain static during the call screening process, but rather adapt to the context of the call and / or dialogue, allowing the user to select candidate responses relevant to the current context of the call.
[0054] Although Figure 1A and Figure 1B The call screening techniques described herein relate to incoming call screening, but the call screening techniques described herein can also be applied to outgoing calls. For example, a user of computing device 102 can guide caller application 108 to make an outgoing call to perform a task, such as making a restaurant reservation. Caller application 108 may be able to make such a call and engage in natural language conversation during the call, as described above, to complete the task.
[0055] The technology disclosed herein enables computing devices to reduce the amount of user interaction required to respond to calls such as telephone calls or other voice calls. By performing call screening of incoming calls and engaging in natural language dialogue with the caller, the technology disclosed herein eliminates the need for the user of the computing device to speak or listen to the call. Furthermore, by determining candidate responses to the dialogue in relation to the dialogue, and by enabling the user to select a candidate response to reply to the ongoing dialogue, the technology disclosed herein allows the user to respond to questions or inquiries received during a call without having to speak or listen to the call.
[0056] The ability to engage in natural language conversation with the caller and respond to questions or inquiries received during a call without requiring the user to speak or listen can enable people with disabilities (such as those with social anxiety, hearing and / or speech impairments, or those who dislike speaking on the phone) to use computing devices to answer incoming calls without the need for additional specialized devices and / or services, such as voice-to-voice trunking services. This ability also prevents malicious actors from knowing what a user's voice sounds like and from potentially capturing and using the user's voice for malicious purposes.
[0057] Figure 2 This is a block diagram illustrating further details of an example computing device according to one or more aspects of this disclosure. The following will... Figure 2 The computing device 202 is described as follows: Figure 1A and Figure 1B An example of the computing device 102 shown.
[0058] Figure 2 The computing device 202 may be a mobile phone, tablet computer, laptop computer, desktop computer, server, mainframe, set-top box, television, wearable device, home automation device or system, gaming system, media player, e-book reader, mobile TV platform, car navigation or infotainment system, or configured to connect to a network (such as... Figure 1A and Figure 1B Examples of any other type of mobile, non-mobile, wearable, and non-wearable computing devices communicating via network 130 shown. Figure 2 This is just one specific example of computing device 202, and many other examples of computing device 202 can be used in other situations, and may include a subset of the components included in the example computing device 202, or may include... Figure 2 Additional components not shown.
[0059] like Figure 2 As shown in the example, computing device 202 includes a user interface component (UIC) 204, one or more processors 240, one or more input components 242, one or more communication units 244, one or more output components 246, and one or more storage components 248. The storage component 248 of computing device 202 also includes a user interface (UI) module 206, a caller application 208, and a dialogue model 252. The UI module 206 is... Figure 1A and Figure 1B The example of UI module 106, and the caller application 208 is... Figure 1A andFigure 1B The caller application 108 is an example. Dialogue model 252 is... Figure 1A and Figure 1B An example of dialogue model 152.
[0060] Communication channel 250 can interconnect each of components 240, 204, 244, 246, 242, and 248 for inter-component communication (physically, communicatively, and / or operationally). In some examples, communication channel 250 may include a system bus, a network connection, an inter-process communication data structure, or any other method for transmitting data.
[0061] One or more input components 242 of computing device 202 can receive input. Examples of input are tactile input, audio input, and video input. In one example, one or more input components 242 of computing device 202 include a presence-sensitive display, a touch-sensitive screen, a mouse, a keyboard, a voice response system, a video camera, a microphone, or any other type of device for detecting input from a person or machine.
[0062] One or more output components 246 of computing device 202 can generate output. Examples of output are haptic output, audio output, and video output. In one example, one or more output components 246 of computing device 202 include a presence-sensitive display, a sound card, a video graphics adapter card, a speaker, a liquid crystal display (LCD), a light-emitting diode (LED) display, a miniLED, a microLED, an organic light-emitting diode (OLED) display, a light field display, a haptic motor, a linear actuator, or any other type of device for generating output to a human or machine.
[0063] One or more communication units 244 of computing device 202 can communicate with external devices via one or more wired and / or wireless networks by transmitting and / or receiving network signals on one or more networks. Examples of one or more communication units 244 include network interface cards (e.g., Ethernet cards), optical transceivers, radio frequency transceivers, GPS receivers, or any other type of device capable of transmitting and / or receiving information. Other examples of one or more communication units 244 may include shortwave radios, cellular data radios, wireless network radios, and Universal Serial Bus (USB) controllers.
[0064] The UIC 204 of the computing device 202 can be hardware used as an input and / or output device for the computing device 202. For example, the UIC 204 may include a display component, which may be a screen on which the UIC 204 displays information, and a presence-sensitive input component that can detect objects at and / or near the display component.
[0065] One or more processors 240 may implement functionality and / or execute instructions within computing device 202. For example, one or more processors 240 on computing device 202 may receive and execute instructions stored in storage component 248, which perform the functionality of UI module 206, caller application 208, and dialogue model 252. Instructions executed by one or more processors 240 may cause computing device 202 to store information in storage component 248 during program execution. Examples of one or more processors 240 include application processors, display controllers, sensor hubs, and any other hardware configured to operate as processing units. One or more processors 240 may execute instructions from UI module 206, caller application 208, and dialogue model 252 to perform actions or functions. That is, UI module 206, caller application 208, and dialogue model 252 may be operated by one or more processors 240 to perform various actions or functions of computing device 202.
[0066] One or more storage components 248 within computing device 202 may store information for processing during operation of computing device 202. That is, computing device 202 may store data accessed by UI module 206, caller application 208, and dialogue model 252 during execution at computing device 202. In some examples, storage component 248 is temporary memory, meaning that its primary purpose is not long-term storage. Storage component 248 on computing device 202 may be configured as volatile memory for short-term storage of information and therefore does not retain the stored contents in the event of a power outage. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), and other forms of volatile memory known in the art.
[0067] In some examples, storage component 248 also includes one or more computer-readable storage media. Storage component 248 can be configured to store a larger amount of information than volatile memory. Storage component 248 can be further configured as a non-volatile memory space for long-term storage of information and to retain information after power-on / power-off cycles. Examples of non-volatile memory include magnetic hard disks, optical disks, floppy disks, flash memory, or in the form of electrically programmable memory (EPROM) or electrically erasable programmable memory (EEPROM). Storage component 248 can store program instructions and / or information (e.g., data) associated with UI module 206, caller application 208, and dialogue model 252.
[0068] One or more processors 240 are configured to execute UI module 206, caller application 208, and dialogue model 252 to perform any combination of the technologies described in this disclosure. For example, one or more processors 240 are configured to execute caller application 208 to receive data from a remote computing device (e.g., Figure 1A One or more processors 240 are configured to execute caller application 208 to determine whether to alert the user of computing device 202 to the incoming call, such as by determining whether to ring (e.g., audibly output a ringtone and / or output a haptic mode) to alert the user of computing device 202 to the incoming call. One or more processors 240 are configured to execute caller application 208 to determine whether to alert the user of computing device 202 to the incoming call based at least in part on determining whether the incoming call is a spam call. One or more processors 240 are configured to execute caller application 208 to determine whether the incoming call is a spam call using any suitable spam detection technology, such as by determining whether the phone number associated with the incoming call is on a list of known spam callers, whether the user of computing device 202 has previously marked the phone number associated with the incoming call as a spam caller, etc. If computing device 202 determines that the incoming call is a spam call, one or more processors 240 are configured to execute caller application 208 to reject the call and / or may send the incoming call to voicemail without alerting the user to the incoming call.
[0069] In some examples, one or more processors 240 are configured to execute caller application 208 to determine whether to alert the user of computing device 202 to an incoming call, at least in part, based on whether the user is available to answer the incoming call. One or more processors 240 are configured to execute caller application 208 to determine whether the user is available to answer the call based on contextual information such as whether the do-not-disturb feature of computing device 202 is enabled, the user's schedule stored in the calendar of computing device 202 (e.g., whether the user is currently in a scheduled meeting), the current time and / or date, the current location of computing device 202, and the user's current state (e.g., whether the user is currently driving or sleeping).
[0070] If caller application 208 determines that the user of computing device 202 is unavailable to answer an incoming call, one or more processors 240 are configured to execute caller application 208 to send the incoming call to voicemail without alerting the user. In some examples, one or more processors 240 are configured to execute caller application 208 to perform automatic call screening of incoming calls to gather information from the party calling from the remote computing device, such as the identity of that party (e.g., the caller's name and / or the identity of the entity being called), the purpose of the call, and / or any other relevant information. As described throughout this disclosure, one or more processors 240 are configured to execute caller application 208 to perform call screening to engage in natural language conversations with callers associated with the remote computing device, and may store transcripts and records of the conversations for later viewing by the user of computing device 202.
[0071] If caller application 208 determines to alert a user of computing device 202 to an incoming call, one or more processors 240 will cause computing device 202 to ring (e.g., audibly output a ringing sound and / or output a haptic mode) to alert the user of computing device 202 to the incoming call and enable the user to guide computing device 202 to perform call screening for the incoming call. If the user of computing device 202 guides computing device 202 to perform call screening for the incoming call, such as by interacting with UIC 204 to provide user input corresponding to the selection of call screening user interface elements, one or more processors 240 are configured to execute caller application 208 to perform call screening for the incoming call.
[0072] One or more processors 240 are configured to execute caller application 208 to perform call screening of incoming calls, collecting information from the party calling from the remote computing device, such as the party's identity (e.g., the caller's name and / or the identity of the entity being called), the purpose of the call, and / or any other relevant information. Caller application 208 can perform call screening of incoming calls without requiring the user of computing device 202 to answer the incoming call and without requiring the user of computing device 202 to converse with the party calling from the remote computing device. Instead, one or more processors 240 are configured to execute caller application 208 to establish a call between computing device 202 and the remote computing device by answering the incoming call and to perform call screening of incoming calls by engaging in natural language conversation with the party calling from the remote computing device. That is, caller application 208 is capable of receiving utterances from the remote computing device, such as words, phrases, and sentences spoken by the user of the remote computing device, and generating natural language utterances, such as spoken words, phrases, and sentences, that are sent to the remote computing device.
[0073] Caller application 208 may be able to engage in natural language conversation with a remote computing device to collect information from the party calling from the remote computing device, such as the identity of that party (e.g., the caller's name and / or the identity of the entity being called) and the purpose of the call. Caller application 208 may be able to record the audio of the call and save a transcript of the call at computing device 202, so that a user of computing device 202 may be able to listen to the recording of the call and / or read the transcript of the call later.
[0074] Caller application 208 can engage in natural language conversations with a caller at a remote computing device without user interaction. That is, caller application 208 can generate utterances during a call and send such utterances to a remote computing device without user interaction. For example, one or more processors 240 are configured to execute caller application 208 to determine an appropriate response to utterances received from a remote computing device without user input instructing caller application 208 how to respond to the received utterances. In this way, caller application 208 can engage in multi-round natural language conversations with a caller at a remote computing device.
[0075] When call screening is performed by caller application 208, one or more processors 240 are configured to execute caller application 208 to suppress the output of audio of the natural language conversation taking place with the remote computing device. One or more processors 240 are also configured to execute caller application 208 to suppress any audio transmitted to the remote computing device that can be captured by the audio input device (e.g., microphone) of UIC 204, and / or to disable such audio input device of UIC 204. Conversely, to enable the user of computing device 202 to follow the natural language conversation taking place between computing device 202 and remote computing device 136, one or more processors 240 are configured to execute caller application 208 to output a real-time text transcription of the natural language conversation occurring during the call for display at UIC 204. For example, one or more processors 240 are configured to execute caller application 208 to perform speech-to-text transcription of the call to generate a real-time transcription of the conversation occurring during the call.
[0076] One or more processors 240 are configured to execute a caller application 208 to conduct a dialogue to gather information about the call, such as the caller's name and the purpose of the call. The caller application 208 can use a dialogue model 252 to determine one or more words, phrases, sentences, etc., to be spoken as part of the dialogue, and generate utterances of such words, phrases, and sentences as part of the dialogue. The dialogue model 252 may include one or more neural networks, such as generative adversarial networks (GANs), recurrent neural networks (RNNs), etc., trained via machine learning on a corpus of anonymous telephone dialogue data to determine responses to utterances received by the caller application 208 during the call.
[0077] For example, one or more processors 240 are configured to execute caller application 208 to perform automatic speech recognition to convert the speech into text in response to receiving utterances during a call, and to input the converted text along with any other relevant contextual information into dialogue model 252. This other relevant contextual information includes prior utterances in the conversation during the call (e.g., words, phrases, and / or sentences previously spoken by the parties on the call), the vocal characteristics of the utterances received during the call (e.g., tone of voice), whether the identity of the remote computing device is listed in the contacts of computing device 202, the location of computing device 202 and / or the remote computing device, the current time and / or date, events listed in the calendar application of computing device 202, previous conversations with the party using the remote computing device, or any other relevant contextual information. One or more processors 240 are configured to execute dialogue model 252 to determine, based on the input data, one or more words, phrases, and / or words that caller application 208 can convert (e.g., via text-to-speech) into utterances that caller application 208 can send as part of a conversation to the remote computing device.
[0078] In some examples, one or more processors 240 are configured to execute dialogue model 252 to generate words, phrases, and sentences in different dialogue styles based on user preferences and / or contextual information associated with the call. For example, one or more processors 240 are configured to execute dialogue model 252 to determine the dialogue style of a conversation in a call, such as by determining whether the dialogue style is formal or casual based on conversations that have already occurred during the call, and to generate words, phrases, and sentences according to the determined dialogue style. In another example, one or more processors 240 are configured to execute dialogue model 252 such that if dialogue model 252 determines that the call is with a private friend of the user of computing device 202, words, phrases, and sentences are generated in a casual dialogue style, and if dialogue model determines that the call is a business call, words, phrases, and sentences are generated in a formal dialogue style.
[0079] When computing device 102 is in dialogue with a remote computing device, one or more processors 240 are configured to execute caller application 208 to determine one or more candidate responses related to the dialogue, and may output indications of one or more candidate responses for display at UIC 204, so that the user of computing device 202 can select a candidate response. One or more processors 240 are configured to execute caller application 208 to generate and send a response corresponding to the selected candidate response in the dialogue in response to the user of computing device 202 selecting a candidate response.
[0080] One or more processors 240 are configured to execute caller application 208 to determine one or more candidate responses related to the dialogue. In the case of a dialogue, a relevant response to the dialogue may be a response related to responding to a utterance received most recently from a remote computing device as part of the dialogue, and / or a response that is highly likely or may be selected by the user to respond to a utterance received most recently from a remote computing device as part of the dialogue.
[0081] Caller application 208 may determine one or more candidate responses based on contextual information such as contextual information associated with the call. Such contextual information may include previous utterances in the conversation during the call (e.g., words, phrases, and / or sentences previously spoken by the parties on the call), vocal characteristics of utterances received from the remote computing device (e.g., tone of voice), whether the identity of the remote computing device is listed in the contacts of computing device 202, the location of computing device 202 and / or the remote computing device, the current time and / or date, events listed in the calendar application of computing device 202, previous conversations with the party using the remote computing device, or any other relevant contextual information.
[0082] Caller application 208 may use dialogue model 252 to determine one or more candidate responses. For example, one or more processors 240 are configured to execute caller application 208 to input context information into dialogue model 252 in response to receiving a utterance from a remote computing device during a call, and one or more processors 240 are configured to execute dialogue model 252 to determine one or more candidate responses relevant to the dialogue based on the input data.
[0083] One or more processors 240 are configured to execute caller application 208 to output an indication of each of one or more candidate responses, such as to be displayed at UIC 204, allowing a user to interact with UIC 204 to select a candidate response from one or more candidate responses. A user of computing device 202 can provide user input at UIC 204 to select a candidate response from one or more candidate responses displayed by UIC 204, and one or more processors 240 are configured to execute UI module 206 to send an indication of the selected candidate response. One or more processors 240 are configured to execute caller application 208 to, in response to receiving an indication of the selected candidate response, send a response in the conversation corresponding to that candidate response, at least in part based on the selected candidate response. That is, caller application 208 can generate a response (e.g., one or more words, phrases, and / or sentences) with the same or similar meaning as the selected candidate response, and can send a spoken version of the response as part of the conversation to a remote computing device.
[0084] Figure 3A and Figure 3B Techniques for call screening for incoming calls according to various aspects of this disclosure are shown. Figure 3A and Figure 3B It is about Figure 2 The computing device 202 is described herein.
[0085] As described above, when computing device 202 receives an incoming call, it can perform call screening, which may include engaging in a natural language conversation with the caller to gather certain information. During call screening, computing device 202 can determine and output candidate responses related to the conversation, and the user of computing device 202 can select a candidate response, which can cause computing device 202 to formulate and send a response corresponding to the selected candidate response during the call.
[0086] According to various aspects of this disclosure, computing device 202 can enable a user of computing device 202 to determine a greeting that computing device 202 may send during a natural language conversation with a caller during call screening. Computing device 202 may, in response to receiving an incoming call, determine one or more candidate greetings associated with the incoming call, and may use these one or more candidate greetings to greet the caller of the incoming call. Computing device 202 may output an indication of each of the one or more candidate greetings, such as for display at UIC 204, and a user of computing device 202 may interact with UIC 204, such as by providing user input at UIC 204, to select a candidate greeting from the one or more candidate greetings. Computing device 202 may, in response to the selection of a candidate greeting, perform call screening of the incoming call, which includes sending a greeting corresponding to the selected candidate greeting as part of a natural language conversation.
[0087] The computing device 202 can determine one or more candidate greetings based on contextual information associated with an incoming call, such as the caller's identity (e.g., the caller's phone number), whether the caller is stored in the contacts of the computing device 202, the time of day, the location of the computing device 202, etc.
[0088] exist Figure 3A In the example, computing device 202 may execute caller application 208 in response to receiving an incoming call to send data to UI module 206 causing UIC 204 to display GUI 314A to alert the user of computing device 202 to the incoming call. For example, UIC 204 may display GUI 314A to alert the user of computing device 202 to the incoming call while the phone is ringing (e.g., audibly outputting a ringtone and / or outputting a haptic mode).
[0089] GUI 314A may include call information 324 associated with an incoming call, such as the phone number from which the incoming call originates, the name of the person or entity associated with that phone number, or whether the name of the person or entity associated with that phone number is unknown. As shown in GUI 314A, because the identity of the incoming call (e.g., phone number) is stored in the contacts of computing device 202, computing device 202 may be able to output the caller's name as part of the call information 324, and may also output the caller's photo in GUI 314A. GUI 314A may include a call answering UI element 320 that the user can (e.g., by providing user input at UIC 204) select to answer the call, and a call filtering UI element 322 that the user can (e.g., by providing user input at UIC 204) select to cause computing device 202 to perform call filtering for incoming calls, as described above regarding...Figure 1A and Figure 1B As described.
[0090] According to various aspects of this disclosure, computing device 202 may, in response to receiving an incoming call, also determine one or more candidate greetings associated with the incoming call. Computing device 202 may determine one or more candidate greetings associated with the incoming call based on the caller type of the caller (e.g., the person making the incoming call), such as whether the caller is listed in the contacts of computing device 202, whether the caller is a business or entity, etc. For example, computing device 202 may determine this based on information such as the caller's identity (e.g., the incoming call's telephone number), the caller type of the caller, and one or more candidate greetings associated with the incoming call.
[0091] exist Figure 3A In the example, because the caller's identity is stored in the contacts of computing device 202, computing device 202 can determine candidate greetings for "Is it urgent?" related to the incoming call, and computing device 202 can output instructions for the candidate greetings, such as by including the candidate greeting UI element 326 labeled "Is it urgent?" in the GUI 314A displayed by UIC 204. The user of computing device 202 can provide user input, such as touch input, to interact with UIC 204 to select the candidate greeting UI element 326. In response, UI module 206 can send, for example, instructions to caller application 208 corresponding to the selection of the candidate greeting corresponding to UI element 326, and in response, computing device 202 can begin performing call screening for the incoming call, including sending a greeting corresponding to the selected candidate greeting "Is this urgent?" as part of a natural language dialogue with the caller.
[0092] The computing device 202 can execute the caller application 208 to send data to the UI module 206, causing the UIC 204 to display the GUI 314B, which may be a call screening user interface that includes a real-time transcription 316A of the natural language conversation being conducted by the computing device 202 with the caller as part of the call screening.
[0093] The computing device 202 can initiate a conversation with the caller by determining a greeting corresponding to a selected candidate greeting for “Is this urgent?”. For example, the computing device 202 can determine a greeting that is a phrase or sentence asking the caller whether the call is urgent. The computing device 202 can then convert the greeting into speech and send the speech to the caller as part of a natural language conversation.
[0094] While performing call screening and engaging in dialogue with the caller, the computing device 202 can also determine one or more candidate responses related to the call. Figure 3A In the example, computing device 202 can receive the phrase "Hey this is your dad. I broke my leg" from the caller as part of the call, and in response, computing device 202 can determine one or more candidate responses to the phrase received from the caller. For example, computing device 202 can determine candidate responses "Hold on, I will answer" and "Call you later" related to the phrase from the caller, and can output candidate response UI element 332A associated with the candidate response "Hold on, I will answer" and candidate response UI element 332B associated with the candidate response "Call you later" in GUI 314B.
[0095] The user of computing device 202 can provide user input (such as touch input) to interact with UIC 204 to select a candidate response UI element 332A associated with the candidate response "Hold on, I will answer". In response, UI module 206 can send an instruction to, for example, caller application 208, of the user input corresponding to the selection of the candidate response UI element 332A, and in response, computing device 202 can formulate and send a response corresponding to the selected candidate response during the call. That is, computing device 202 can determine a response as an instruction for the user of computing device 202 to join the call. Computing device 202 can therefore convert the response corresponding to the selected candidate response into speech and can send the speech to the caller as part of a natural language conversation. Figure 3B As shown, computing device 202 can execute caller application 208 to send data to UI module 206 causing UIC 204 to display GUI 314C, which includes an updated real-time transcription 316B of the natural language dialogue being conducted by computing device 202, including a response to “Please hold on one second. I am connecting your call” corresponding to the selected candidate response.
[0096] Because the selected candidate response to "Hold on, I will answer" indicates the user of computing device 202's intention to join the call, computing device 202 can also end the call screening of the incoming call without interrupting the incoming call, in response to the user selecting the candidate response UI element 332A associated with the candidate response to "Hold on, I will answer". Instead, computing device 202 can maintain the call and enable the user of computing device 202 to answer the call and begin voice communication with the caller via the call (e.g., by speaking to the caller and listening to what the caller says).
[0097] Figure 4 Additional techniques for call screening according to various aspects of this disclosure are shown below. Figure 2 Description in the context of computing device 202 Figure 4 .
[0098] When computing device 202 performs call screening of incoming calls and engages in natural language dialogue with the caller, computing device 202 may be able to determine whether an incoming call is a spam call based on the ongoing dialogue with the caller. For example, computing device 202 may determine whether an incoming call is a spam call and / or the probability (e.g., the likelihood) that an incoming call is a spam call can be determined based on the pattern of the utterance received from the caller, keywords contained in the utterance received from the caller, or any other relevant contextual information.
[0099] If the computing device 202 determines that an incoming call may be a spam call, for example by determining the probability that the incoming call is above a probability threshold (e.g., greater than 80%, greater than 90%, etc.), the computing device 202 can output an indication that the incoming call may be a spam call for display at UIC 204. The computing device 202 can provide the user with the ability to report spam calls to an external computing system, which can add spam calls to a spam example repository that can be used to train and test the spam detection system.
[0100] like Figure 4As shown, when computing device 202 performs call screening for incoming calls, computing device 202 can execute caller application 208 to send data to UI module 206 causing UIC 204 to display GUI 414. GUI 414 can be a call screening user interface that includes real-time transcription 416 of the natural language conversation being conducted by computing device 202 with the caller as part of the call screening. Computing device 202 can output a user interface element 418 indicating that an incoming call may be spam in GUI 414 in response to detecting that an incoming call may be spam. Computing device 202 may also include a user interface element 432A that a user can select to cause computing device 202 to report spam calls to an external computing system that can add spam calls to a spam example repository that can be used to train and test the spam detection system. In some examples, computing device 202 may also output a user interface element 432B that a user can select to reply to the caller that the user will call the caller back later.
[0101] Figure 5 This is a flowchart illustrating example techniques for determining one or more candidate responses according to various aspects of this disclosure. Below is... Figure 2 Description in the context of computing device 202 Figure 5 .
[0102] As described throughout this disclosure, computing device 202 can perform call screening for incoming calls and can engage in natural language dialogue with the caller of an incoming call. While engaging in natural language dialogue, computing device 202 can determine one or more candidate responses related to the natural language dialogue. Computing device 202 can output a set of candidate responses for display at UIC 204. A user can select a candidate response from the candidate responses displayed at UIC 204, and in response, computing device 202 can formulate and send a response corresponding to the selected candidate response in the call.
[0103] exist Figure 5 In the example, computing device 202 may be able to determine one or more candidate responses related to the natural language dialogue, at least in part, based on the caller type (i.e., the party making the incoming call) and / or the use case of the incoming call. To this end, computing device 202 may determine the caller type (502) in response to receiving an incoming call. The caller type may be one of the following: contact / favorites / work, unknown number, spam, or business. A caller who is a contact / favorites / work caller type, a business type, or a spam type may be referred to as a known caller type, while a caller with an unknown number type may be referred to as an unknown caller type.
[0104] A caller classified as a contact / favorite / work caller can be a caller with contact details already stored in computing device 202, such as a caller with a phone number stored in the contacts of computing device 202. A caller classified as a business caller can be a caller that computing device 202 determines is from a business (e.g., a store, doctor's office, etc.) or another entity (e.g., a school, charity, etc.). Computing device 202 can determine that a caller is of the business type by determining that the caller's phone number is associated with a business or another entity. A caller classified as spam caller can be a caller that computing device 202 determines is making spam calls. Computing device 202 can determine that a caller is of the spam type by determining that the caller's phone number is associated with spam calls. If computing device 202 cannot determine that a caller is of the contact / favorite / work caller type, business type, or spam type, then computing device 202 can determine that the caller belongs to the unknown number type.
[0105] During call screening, computing device 202 can determine a set of candidate responses associated with the caller's caller type and can output the determined set of candidate responses, such as for display at UIC 204 (504). The user can select a candidate response, and computing device 202 can formulate and send a response corresponding to the selected candidate response during the call.
[0106] If computing device 202 determines that the caller is a contact / favorite / work caller type, then computing device 202 can determine the candidate response set as: "Is it urgent?", "Can you repeat?", "Tell me more", "I can't understand", "Can you text me?", "I'll text you", "Be right there", and "I'll call you back". If computing device 202 determines that the caller is a business caller type, then computing device 202 can determine the candidate response set as: "Report spam", "Can you repeat?", "Tell me more", "I can't understand", "Can you text me?", "Who is this?", "Call back later", and "Wrong number". If the computing device 202 determines that the caller is a spam caller, the computing device 202 can determine the candidate response set as: "Wrong number" and "Take me off your list".
[0107] If the computing device 202 determines that the caller is an unknown number, the computing device 202 can determine the candidate response set as: "Report spam", "Can you repeat?", "Tell me more", "I can't understand", "Who is this?", "Call back later", "I'll get back to you", and "Wrong number".
[0108] When computing device 202 engages in natural language dialogue during a call, it may be able to update the caller's caller type, at least in part, based on contextual information associated with the call, such as the context of the natural language dialogue. For example, computing device 202 may initially determine that the caller is an unknown number. When computing device 202 performs call screening and engages in natural language dialogue during the call, it may determine, based on contextual information such as the content of the dialogue, that the caller is not an unknown number but an enterprise caller.
[0109] The computing device 202 can therefore update the caller's caller type and, in response to updating the caller's caller type, update the candidate responses for the call to those associated with the enterprise caller type. For example, if the computing device 202 determines that the caller is not an unknown number but an enterprise caller type, the computing device 202 can update the candidate responses output in UIC 204 to those associated with the enterprise caller type: "Reportspam", "Can you repeat?", "Tell me more", "I can't understand", "Can you text me?", "Who is this?", "Call back later", and "Wrong number".
[0110] In another example, computing device 202 may initially determine that the caller is an unknown number. While computing device 202 performs call screening and engages in natural language dialogue during the call, computing device 202 may determine, based on contextual information such as the content of the dialogue, that the caller is not an unknown number but a spam caller. Computing device 202 may update the caller's caller type to spam caller type and, in response to updating the caller's caller type, may update the candidate responses for the call to those associated with the spam caller type. For example, if computing device 202 determines that the caller is not an unknown number but a spam caller, then computing device 202 may update the candidate responses output in UIC 204 to those associated with the spam caller type: "Wrong number" and "Take me off your list".
[0111] In addition to determining the caller type, computing device 202 can also determine the use case of the call (506). The use case of the call can be the purpose of the caller in calling computing device 202. Computing device 202 can determine the use case of the call based on any contextual information related to the call, such as the context of an ongoing natural language dialogue, keywords detected in the ongoing natural language dialogue, the caller's voice characteristics, or any other relevant contextual information. In some examples, computing device 202 can input such contextual information related to the call into dialogue model 252 to determine the use case of the call.
[0112] The computing device 202 can determine whether the use case of the call is known, wherein if the computing device 202 cannot determine the use case of the call, then the use case of the call is unknown (508). If the computing device 202 cannot determine the use case of the call, then the computing device 202 can continue to output a set of candidate responses associated with the caller's caller type (510). If the computing device 202 can determine the use case of the call, then the computing device 202 can determine a set of candidate responses associated with the use case of the call, such as... Figure 6 The process is described in more detail below, and a set of candidate responses associated with the use case of the call can be output for display at UIC 204 (512). The user can select candidate responses, and the computing device 202 can formulate and send a response corresponding to the selected candidate response during the call.
[0113] In some examples, computing device 202 may determine that the call includes dial pad options, which may include detecting a request to press certain numbers on the dial pad in the caller's utterance, such as "press 1 to for English". If computing device 202 determines that the call includes dial pad options, then computing device 202 may determine that the candidate response set includes "report spam" and "wrong number".
[0114] Figure 6 Example candidate responses related to use cases of a call are shown below, according to various aspects of this disclosure. Figure 2 Description in the context of computing device 202 Figure 6 .
[0115] like Figure 6As shown in Table 600, different use cases for calls that can be determined by computing device 202 include spam, appointment confirmation, callback, product replenishment, delivery, where are you, order ready, and notification / message. A call with a spam use case can be a spam call. A call with an appointment confirmation use case can be a call from a business or other entity to confirm an appointment. A call with a callback use case can be a call from a caller of the user of computing device 202. A call with a product replenishment use case can be a call from a business or other entity to notify the user of computing device 202 that a product of interest to the user has been replenished. A call with a delivery use case can be a call to notify the user of computing device 202 that an item ordered by the user (e.g., a food order) has arrived. A call with a where are you use case is a call to the user of computing device 202 to inquire about the user's location. A call with an order ready use case can be a call from a business or other entity to notify the user of computing device 202 that an item ordered by the user (e.g., a food order) is ready for pickup. A call with a notification / message use case can be a call for notifying the computing device 202 of a user-specific situation and / or a call for delivering a message to the user.
[0116] As can be seen in Table 600, each use case can be associated with a set of candidate responses. Therefore, computing device 202 can determine the set of candidate responses associated with the call in response to the use case determining the call, and can output an indication of the set of candidate responses associated with the call at UIC 204. The user can interact with UIC 204 to select a candidate response from the candidate responses associated with the call, and computing device 202 can formulate and send a response corresponding to the selected candidate response in the call in response to the user's selection.
[0117] Figure 7 This is a flowchart illustrating example operations performed by an example computing device configured to perform call screening, according to one or more aspects of this disclosure. Below is... Figure 1A Environment 100 and Figure 2 Description in the context of computing device 202 Figure 7 .
[0118] like Figure 7As shown, one or more processors 240 of computing device 202 can establish a call with remote computing device 136 (702). The one or more processors 240 can engage in a dialogue with remote computing device 136 during the call (704). The one or more processors 240 can determine one or more candidate responses based at least in part on context information associated with the call (706). The one or more processors 240 can receive an instruction from user input to select a candidate response from the one or more candidate responses (708). In response to receiving the instruction from user input to select a candidate response, the one or more processors 240 can send a response corresponding to the candidate response in the dialogue (710).
[0119] This disclosure includes the following examples.
[0120] Example 1. A method comprising: establishing a call with a remote computing device by one or more processors of a computing device; engaging in a dialogue with the remote computing device in the call by the one or more processors; determining one or more candidate responses by the one or more processors based at least in part on context information associated with the call; receiving an instruction from the one or more processors of user input to select a candidate response from the one or more candidate responses; and, in response to receiving the instruction from the user input to select the candidate response, sending a response corresponding to the candidate response in the dialogue by the one or more processors.
[0121] Example 2. The method described in Example 1, wherein the dialogue includes a multi-turn natural language dialogue.
[0122] Example 3. The method of any one of Examples 1 and 2, wherein conducting the dialogue with the remote computing device in the call further comprises: receiving a first utterance from the remote computing device by the one or more processors; determining a second utterance for responding to the first utterance by the one or more processors based at least in part on context information associated with the call; and sending the second utterance to the remote computing device by the one or more processors.
[0123] Example 4. The method as described in Example 3, wherein determining the second utterance for responding to the first utterance further comprises: the one or more processors using one or more neural networks trained on a corpus of telephone conversation data to determine the second utterance for responding to the first utterance.
[0124] Example 5. The method of any one of Examples 1-4, wherein conducting the dialogue with the remote computing device in the call further comprises: the dialogue being conducted by the one or more processors to determine the identity of the party making the call and the purpose of the call.
[0125] Example 6. The method of any one of Examples 1-5, wherein conducting the dialogue with the remote computing device in the call further comprises: detecting silence in the dialogue by the one or more processors; and prompting the one or more processors to speak in response to detecting silence in the dialogue.
[0126] Example 7. The method of any one of Examples 1-6, wherein determining the one or more candidate responses related to the conversation further comprises: the one or more processors determining the one or more candidate responses related to the latest utterance received in the call from the remote computing device.
[0127] Example 8. The method of any one of Examples 1-7, wherein determining the one or more candidate responses further comprises: determining, by the one or more processors, a caller type associated with the call; and determining, by the one or more processors, the one or more candidate responses based at least in part on the caller type associated with the call.
[0128] Example 9. The method as described in Example 8, wherein determining the caller type associated with the call further comprises: the one or more processors determining the caller type associated with the call based at least in part on a telephone number associated with the remote computing device.
[0129] Example 10. The method of any one of Examples 1-9, wherein determining the one or more candidate responses further comprises: determining a use case associated with the call by the one or more processors; and determining the one or more candidate responses by the one or more processors based at least in part on the use case associated with the call.
[0130] Example 11. The method of any one of Examples 1-10 further includes: outputting a real-time transcription of the conversation by the one or more processors for display on a display device.
[0131] Example 12. The method of any one of Examples 1-11, wherein the call includes an incoming call, and wherein establishing the call with the remote computing device and having the conversation with the remote computing device in the call further includes: receiving the incoming call by the one or more processors; and responding to the incoming call by the one or more processors to perform call screening of the incoming call.
[0132] Example 13. The method of Example 12, wherein responding to the incoming call further comprises: determining one or more candidate greetings by the one or more processors; receiving by the one or more processors an instruction from a second user input to select a candidate greeting as the selected candidate greeting from the one or more candidate greetings; and in response to receiving the instruction from the second user input, performing the call screening of the incoming call by the one or more processors, including sending a greeting corresponding to the selected candidate greeting as part of the conversation to the remote computing device.
[0133] Example 14. The method as described in Example 13, wherein the selected candidate greeting corresponds to asking whether the call is urgent, and wherein sending the greeting corresponding to the selected candidate greeting as part of the conversation further comprises: the one or more processors sending the greeting asking whether the call is urgent to the remote computing device as part of the conversation.
[0134] Example 15. The method of any one of Examples 1-14, wherein conducting the conversation with the remote computing device in the call further comprises: the one or more processors conducting the conversation with the remote computing device in the call without outputting audio of the ongoing conversation.
[0135] Example 16. The method of any one of Examples 1-15 further includes: in response to determining the one or more candidate responses, the one or more processors outputting an indication of the one or more candidate responses for display at a display device.
[0136] Example 17. A computing device includes: a memory storing instructions; and one or more processors that execute the instructions to: establish a call with a remote computing device, engage in a dialogue with the remote computing device in the call, determine one or more candidate responses based at least in part on context information associated with the call; receive an instruction from user input selecting a candidate response from the one or more candidate responses; and in response to receiving the instruction from user input selecting the candidate response, send a response corresponding to the candidate response in the dialogue.
[0137] Example 18. A computing device as described in Example 17, wherein the dialogue includes a multi-turn natural language dialogue.
[0138] Example 19. A computing device as described in any one of Examples 17 and 18, wherein, in order to conduct the dialogue with the remote computing device in the call, the one or more processors further execute the instructions to: receive a first utterance from the remote computing device by the one or more processors; determine a second utterance for responding to the first utterance by the one or more processors based at least in part on context information associated with the call; and send the second utterance to the remote computing device by the one or more processors.
[0139] Example 20. The computing device as described in Example 19, wherein, in order to determine the second utterance for responding to the first utterance, the one or more processors further execute the instruction to: use one or more neural networks trained on a corpus of telephone conversation data to determine the second utterance for responding to the first utterance.
[0140] Example 21. A computing device as described in any one of Examples 17-20, wherein, in order to conduct the dialogue with the remote computing device in the call, the one or more processors further execute the instructions to: conduct the dialogue to determine the identity of the party making the call and the purpose of the call.
[0141] Example 22. A computing device as described in any one of Examples 17-21, wherein, in order to conduct the dialogue with the remote computing device in the call, the one or more processors further execute the instruction to: detect silence in the dialogue; and, in response to detecting silence in the dialogue, prompt the party associated with the remote computing device to speak.
[0142] Example 23. A computing device as described in any one of Examples 17-22, wherein, in order to determine the one or more candidate responses related to the conversation, the one or more processors further execute the instruction to: determine the one or more candidate responses related to the latest utterance received in the call from the remote computing device.
[0143] Example 24. A computing device as described in any one of Examples 17-23, wherein, in order to determine the one or more candidate responses, the one or more processors further execute the instructions to: determine the caller type associated with the call; and determine the one or more candidate responses based at least in part on the caller type associated with the call.
[0144] Example 25. A computing device as described in Example 24, wherein, in order to determine the caller type associated with the call, the one or more processors further execute the instruction to: determine the caller type associated with the call based at least in part on a telephone number associated with the remote computing device.
[0145] Example 26. A computing device as described in any one of Examples 17-25, wherein, in order to determine the one or more candidate responses, the one or more processors further execute the instructions to: determine a use case associated with the call; and determine the one or more candidate responses based at least in part on the use case associated with the call.
[0146] Example 27. A computing device as described in any one of Examples 17-26, wherein the one or more processors further execute the instruction to: output a real-time transcription of the conversation by the one or more processors for display at a display device.
[0147] Example 28. A computing device as described in any one of Examples 17-27, wherein the call includes an incoming call, and in order to establish the call with the remote computing device and to conduct the conversation with the remote computing device in the call, the one or more processors further execute the instructions to: receive the incoming call; and respond to the incoming call to perform call screening of the incoming call.
[0148] Example 29. A computing device as described in Example 28, wherein, in response to the incoming call, the one or more processors further execute the instructions to: determine one or more candidate greetings; receive an instruction from a second user input to select a candidate greeting as the selected candidate greeting from the one or more candidate greetings; and, in response to receiving the instruction from the second user input, perform call screening of the incoming call, including sending a greeting corresponding to the selected candidate greeting as part of the conversation to the remote computing device.
[0149] Example 30. A computing device as described in Example 29, wherein the selected candidate greeting corresponds to asking whether the call is urgent, and wherein, in order to send the greeting corresponding to the selected candidate greeting as part of the conversation, the one or more processors further execute the instruction to: send the greeting asking whether the call is urgent to the remote computing device as part of the conversation.
[0150] Example 31. A computing device as described in any one of Examples 17-30, wherein, in order to conduct the conversation with the remote computing device in the call, the one or more processors further execute the instruction to: conduct the conversation with the remote computing device in the call without outputting audio of the ongoing conversation.
[0151] Example 32. A computing device as described in any one of Examples 17-31, wherein the one or more processors further execute the instruction to: in response to determining the one or more candidate responses, output an indication of the one or more candidate responses for display at a display device.
[0152] Example 30. A non-transitory computer-readable storage medium including instructions that, when executed, cause one or more processors of a computing device to: establish a call with a remote computing device and engage in a dialogue with the remote computing device in the call; determine one or more candidate responses based at least in part on context information associated with the call; receive an instruction from user input to select a candidate response from the one or more candidate responses; and, in response to receiving the instruction from user input to select the candidate response, send a response corresponding to the candidate response in the dialogue.
[0153] Example 31. A non-transitory computer-readable storage medium as described in Example 30, wherein the dialogue includes a multi-turn natural language dialogue.
[0154] Example 32. A non-transitory computer-readable storage medium as described in any one of Examples 30 and 31, wherein, in order to conduct the dialogue with the remote computing device in the call, the instruction further causes the one or more processors to: receive a first utterance from the remote computing device; determine a second utterance for responding to the first utterance based at least in part on context information associated with the call; and send the second utterance to the remote computing device.
[0155] Example 33. The computing device as described in Example 32, wherein, in order to determine the second utterance for responding to the first utterance, the instruction further causes the one or more processors to: use one or more neural networks trained on a corpus of telephone conversation data to determine the second utterance for responding to the first utterance.
[0156] Example 34. A computing device as described in any one of Examples 30-33, wherein, in order to conduct the dialogue with the remote computing device in the call, the instruction further causes the one or more processors to: conduct the dialogue to determine the identity of the party making the call and the purpose of the call.
[0157] Example 35. A computing device as described in any one of Examples 30-34, wherein, in order to conduct the dialogue with the remote computing device in the call, the instruction further causes the one or more processors to: detect silence in the dialogue; and, in response to detecting silence in the dialogue, prompt the party associated with the remote computing device to speak.
[0158] Example 36. A computing device as described in any one of Examples 30-35, wherein, in order to determine the one or more candidate responses related to the conversation, the instruction further causes the one or more processors to: determine the one or more candidate responses related to the latest utterance received in the call from the remote computing device.
[0159] Example 37. A computing device as described in any one of Examples 30-36, wherein, in order to determine the one or more candidate responses, the instructions further cause the one or more processors to: determine a caller type associated with the call; and determine the one or more candidate responses based at least in part on the caller type associated with the call.
[0160] Example 38. A computing device as described in Example 37, wherein, in order to determine the caller type associated with the call, the instruction further causes the one or more processors to: determine the caller type associated with the call based at least in part on a telephone number associated with the remote computing device.
[0161] Example 39. A computing device as described in any one of Examples 30-38, wherein, in order to determine the one or more candidate responses, the instructions further cause the one or more processors to: determine a use case associated with the call; and determine the one or more candidate responses based at least in part on the use case associated with the call.
[0162] Example 40. A computing device as described in any one of Examples 30-39, wherein the instruction further causes the one or more processors to output a real-time transcription of the conversation for display at a display device.
[0163] Example 41. A computing device as described in any one of Examples 30-40, wherein the call includes an incoming call, and in order to establish the call with the remote computing device and to conduct the dialogue with the remote computing device in the call, the instruction further causes the one or more processors to: receive the incoming call; and respond to the incoming call to perform call screening of the incoming call.
[0164] Example 42. A computing device as described in Example 41, wherein, in response to the incoming call, the instruction further causes the one or more processors to: determine one or more candidate greetings; receive an instruction from the one or more candidate greetings to select a candidate greeting as the selected candidate greeting; and, in response to receiving the instruction from the second user input, perform call screening of the incoming call, including sending a greeting corresponding to the selected candidate greeting as part of the conversation to the remote computing device.
[0165] Example 43. A computing device as described in Example 42, wherein the selected candidate greeting corresponds to asking whether the call is urgent, and wherein, in order to send the greeting corresponding to the selected candidate greeting as part of the conversation, the instruction further causes the one or more processors to: send the greeting asking whether the call is urgent to the remote computing device as part of the conversation.
[0166] Example 44. A computing device as described in any one of Examples 30-43, wherein, in order to conduct the conversation with the remote computing device in the call, the instruction further causes the one or more processors to conduct the conversation with the remote computing device in the call without outputting audio of the ongoing conversation.
[0167] Example 45. A computing device as described in any one of Examples 30-44, wherein the instruction further causes the one or more processors to: in response to determining the one or more candidate responses, output an indication of the one or more candidate responses for display at a display device.
[0168] By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that is accessible by a computer. Furthermore, any connection is appropriately referred to as computer-readable media. For example, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used herein, disks and optical discs include compact discs (CDs), laser discs, optical discs, digital universal discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0169] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Therefore, as used herein, the term "processor" can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules. Furthermore, the techniques can be fully implemented in one or more circuit or logic elements.
[0170] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless telephones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but these various components, modules, or units are not necessarily required to be implemented by different hardware units. Rather, as described above, various units may be combined within hardware units or provided as a collection of interoperable hardware units including one or more processors as described above, combined with suitable software and / or firmware.
[0171] Various embodiments have been described. These and other embodiments are within the scope of the appended claims.
Claims
1. A method comprising: A call is established between a computing device and a remote computing device by one or more processors of the computing device; The one or more processors communicate with the remote computing device during the call; The one or more processors determine one or more candidate responses based at least in part on context information associated with the call; The one or more processors receive an instruction from user input to select a candidate response from the one or more candidate responses; as well as In response to receiving the instruction from the user input that selects the candidate response, the one or more processors send a response corresponding to the candidate response in the conversation.
2. The method as described in claim 1: The dialogues mentioned therein include multi-turn natural language dialogues; as well as The dialogue with the remote computing device during the call further includes: The first utterance is received from the remote computing device by the one or more processors as part of the multi-turn natural language dialogue; The one or more processors determine a second utterance for responding to the first utterance, based at least in part on context information associated with the call; and The second message is sent from one or more processors to the remote computing device.
3. The method of claim 2, wherein determining the second utterance for responding to the first utterance further comprises: The one or more processors use one or more neural networks trained on a corpus of telephone conversation data to determine the second utterance for responding to the first utterance.
4. The method of any one of claims 1-3, wherein conducting the dialogue with the remote computing device during the call further comprises: The dialogue is conducted by one or more processors to determine the identity of the party making the call and the purpose of the call.
5. The method of any one of claims 1-4, wherein conducting the dialogue with the remote computing device during the call further comprises: Silence is detected in the dialogue by one or more processors; as well as In response to the detection of silence in the dialogue, the one or more processors prompt the party associated with the remote computing device to speak.
6. The method of any one of claims 1-5, wherein determining the one or more candidate responses related to the dialogue further comprises: The one or more processors determine the one or more candidate responses related to the latest utterance received in the call from the remote computing device.
7. The method of any one of claims 1-6, wherein determining the one or more candidate responses further comprises: The caller type associated with the call is determined by the one or more processors based at least in part on a telephone number associated with the remote computing device; as well as The one or more processors determine the one or more candidate responses based at least in part on the caller type associated with the call.
8. The method of any one of claims 1-7, wherein determining the one or more candidate responses further comprises: The one or more processors determine the use case associated with the call; as well as The one or more processors determine the one or more candidate responses based at least in part on the use case associated with the call.
9. The method of any one of claims 1-8, further comprising: The one or more processors output a real-time transcription of the dialogue for display on a display device.
10. The method of any one of claims 1-9, wherein the call includes an incoming call, and wherein establishing the call with the remote computing device and conducting the dialogue with the remote computing device in the call further comprises: The incoming call is received by one or more processors; The incoming call is answered by one or more processors to perform call filtering for the incoming call; One or more candidate greetings are determined by the one or more processors; The processor receives a second user input instruction from the one or more candidate greetings to select a candidate greeting as the selected candidate greeting; as well as In response to receiving the instruction from the second user input, the one or more processors perform the call screening of the incoming call, including sending a greeting corresponding to the selected candidate greeting to the remote computing device as part of the conversation.
11. The method of claim 10, wherein the selected candidate greeting corresponds to asking whether the call is urgent, and wherein sending the greeting corresponding to the selected candidate greeting as part of the conversation further comprises: The greeting, inquiring whether the call is urgent, is sent by one or more processors to the remote computing device as part of the conversation.
12. The method of any one of claims 1-11, wherein conducting the dialogue with the remote computing device during the call further comprises: The conversation is conducted by one or more processors with the remote computing device during the call without outputting audio of the ongoing conversation.
13. The method of any one of claims 1-12, further comprising: In response to determining the one or more candidate responses, the one or more processors output an indication of the one or more candidate responses for display on a display device.
14. A computing device comprising a component for performing any one of the methods described in claims 1-13.
15. A computer program product comprising at least one non-transitory computer-readable medium, said at least one non-transitory computer-readable medium comprising one or more instructions, which, when executed by at least one processor, cause said at least one processor to perform any of the methods described in claims 1-13.
Citation Information
Cited By
Context-based Anti-spam application enabled with human-computer interaction analysis
US20250006188A1