Method and apparatus for managing reservation

An automatic on-hold client monitors voice communication sessions using audio stream analysis to detect the end of a hold state, addressing the challenges of power consumption and user input, and ensuring efficient monitoring of voice communication sessions.

JP7682950B2Active Publication Date: 2025-05-26GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023097188
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-05-26
Estimated Expiration
2038-06-28

AI Technical Summary

Technical Problem

Users face challenges in monitoring voice communication sessions on hold, as they need to continuously monitor the call to determine when a service representative becomes available, leading to increased power consumption and user input requirements.

Method used

An automatic on-hold client monitors the audio stream of a voice communication session to detect when the hold state ends, using techniques such as voice activity detection, speaker diarization, and machine learning models to determine the presence of a human voice and respond with a request signal.

Benefits of technology

This solution allows for efficient monitoring of voice communication sessions on hold without user intervention, reducing power consumption and user input requirements, and accurately determining when a service representative becomes available.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682950000001
    Figure 0007682950000001
  • Figure 0007682950000002
    Figure 0007682950000002
  • Figure 0007682950000003
    Figure 0007682950000003
Patent Text Reader

Abstract

To provide a method and apparatus for managing holds.SOLUTION: Provided is an automated monitoring of a voice communication session, when the session is in an on hold status, to determine when the session is no longer in the on hold status. When it is determined that the session is no longer in the on hold status, user interface output is rendered that is perceptible to a calling user that initiated the session, and that indicates that the on hold status of the session has ceased. In some implementations, an audio stream of the session can be monitored to determine, based on processing of the audio stream, a candidate end of the on hold status. In response, a response solicitation signal is injected into an outgoing portion of the audio. The audio stream can be further monitored for a response (if any) to the response solicitation signal. The response (if any) can be processed to determine whether the end of the on hold status is an actual end of the on hold status.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Humans can participate in voice communication sessions (such as calls) using various client devices. When an individual (referred to herein as the "caller" or "user") calls a specific number and no one is currently answering the call, many organizations can put the caller on hold. The hold state indicates that the caller is waiting to interact with a living person (also referred to herein as a "user"). While the user is on hold, music is frequently played for the user. Additionally, the music can be interrupted by various human-recorded voices that can provide additional information such as information about the organization the user called (e.g., the organization's website, the organization's normal business hours, etc.). Additionally, an automated voice can provide the user with an estimated remaining wait time indicating how much longer the user will remain on hold.

[0002] When a call is on hold, the calling party must closely monitor the call to determine when a second user, such as a service representative, becomes active in the call. For example, if the music during hold is switched to human voice, the calling party must determine whether the voice the calling party is hearing is pre-recorded voice or an actual service representative. To enable close monitoring of a call on hold initiated via a client device, the calling party can increase the call volume, can leave the audio output of the call to the speakerphone modality, and / or (to ensure that the call is still active and on hold) can repeatedly activate the screen of the client device while the call is on hold. Those and / or other monitoring activities of the calling party on hold may increase the power consumption of the client device. For example, such activities may increase the power consumption of the mobile phone being used for the call, which may cause rapid depletion of the mobile phone's battery. Additionally, those and / or other monitoring activities on hold may require the calling party to perform a large number of inputs on the client device, such as inputs for increasing the volume, activating the speakerphone modality, and / or activating the screen. Summary of the Invention Means for Solving the Problems

[0003] The implementations described herein relate to automatic monitoring of a voice communication session while the session is on hold, for determining when the session is no longer on hold. When it is determined that the session is no longer on hold, a user interface output is rendered that is recognizable to the calling user who initiated the session and indicates that the hold state of the session has ended. In various implementations, the on-hold client (e.g., that operates at least in part on the client device that initiated the voice communication session) can be utilized to monitor at least the incoming portion of the audio stream of the session to determine when the session is no longer on hold. In some of those various implementations, the on-hold client determines a candidate for the end of the hold state based on processing of the audio stream. The candidate for the end of the hold state can be based on detecting the occurrence of one or more events within the audio stream. As some non-limiting examples, the candidate for the end of the hold state can be based on detecting a transition in the audio stream (e.g., any transition, or a transition from "on-hold music" to human speech), detecting any human speech (e.g., using voice activity detection), detecting a new human voice (e.g., using speaker diarization), detecting the occurrence of specific terms and / or phrases (e.g., "hello", "hi", and / or the name of the calling user), and / or other events.

[0004] In some variations of those various implementations, in response to detecting a candidate for ending the hold state, the holding client causes a response request signal to be inserted into the outgoing portion of the audio stream (such that it can be "heard" by the callee). The response request signal can be a recorded human voice speaking one or more words, or a synthetically generated voice speaking one or more words. The one or more words can be, for example, "Hello", "Are you there", "Hi, are you on the line", etc. The holding client further monitors for a response (if any) to the response request signal, and the response can be used to determine whether the candidate for ending the hold state indicates an actual end of the hold state. If so, the holding client can render a user interface output that is recognizable to the calling user who initiated the session, indicating that the hold state of the session has ended (i.e., that the voice communication session is no longer on hold). If not, the holding client can continue to monitor for another occurrence of a candidate for ending one hold state. In some implementations, the holding client converts the response to text (e.g., using a speech-to-text processor) based on determining the likelihood that the response is human voice, and determines whether the response indicates that the candidate for ending the hold state is an actual end of the hold state based on whether the text responds to the response request signal, based on determining that the response is prerecorded voice (including voice characteristics different from those of prerecorded voice for a voice communication session), and / or based on other criteria. The holding client can optionally utilize a trained machine learning model when determining the likelihood that the response is human voice.

[0005] In these and other methods, the on-hold client can monitor the incoming portion of the audio stream of the on-hold session and dynamically determine when to provide a response invitation signal. Further, the on-hold client can utilize a response (if any) to the response invitation signal when determining whether the hold state of the session has ended. These actions by the on-hold client can be performed without any intervention from the calling user and without requiring the client device to audibly render the audio stream of the voice communication session. Further, as described herein, in various implementations, the on-hold client can be initiated automatically (without any required user input) or with minimal user input (e.g., a single tap of a graphical element, or a single verbal command).

[0006] A voice communication session can utilize various protocols and / or infrastructures, such as Voice over Internet Protocol (VOIP), public switched telephone network (PSTN), private branch exchange (PBX), or any of various video and / or audio conferencing services. In various implementations, a voice communication session is between a calling user's client device (initiating the voice communication session) and one or more devices of the called party. The voice communication session enables two-way audio communication between the calling user and the called party. The voice communication session can be a direct peer-to-peer session between the calling user's client device and the called party's device, and / or can be routed through various servers, networks, and / or other resources. The voice communication session can occur between various devices. For example, the voice communication session can be between a calling user's client device (such as a mobile phone, stand-alone interactive speaker, tablet, laptop) and a called party's landline phone, between a calling user's client device and a called party's client device, between a calling user's client device and a called party's PBX, etc.

[0007] In some implementations described herein, a pending client that operates at least partially on a client device can be initiated in response to the client device detecting that a voice communication session (initiated by the client device) has been placed on hold. A client device, such as a mobile phone, can examine the audio stream of the voice communication session and can determine in various ways that the session is on hold. As an example, the client device can determine that the session is on hold based on detecting music within the incoming portion of the audio stream, such as typical "on-hold music". For example, the incoming portion of the audio stream can be processed and compared to a list of known on-hold music (e.g., the audio characteristics of the audio stream can be compared to the audio characteristics of known on-hold music) to determine whether the incoming portion of the audio stream is typical on-hold music. Such a list can be stored locally on the client device and / or stored on a remote server to which the client device can connect via a network (e.g., a cellular network). Additionally or alternatively, the incoming portion of the audio stream can be processed and compared to a list of known on-hold voices. As another example, the client device can determine that the session is on hold based on detecting any music within the incoming portion of the audio stream. As yet another example, the client device can additionally or alternatively determine that the session is on hold based on comparing the dialed number of the session to a list of phone numbers known to put callers on hold. For example, if a user calls a "Hypothetical Utility Company", the client device can store the phone number associated with the "Hypothetical Utility Company" as a number that typically puts callers on hold before the user can speak to an actual representative.Furthermore, a list of telephone numbers known to put callers on hold can have a corresponding list of known on-hold music and / or known on-hold voices used by that number. Additionally or alternatively, a user can provide to the client device telephone numbers that normally put the client device on hold. With the user's permission, these telephone numbers provided by the users can be shared among client devices and added to the list of numbers that normally put people on hold in other client devices.

[0008] In some implementations, the user can indicate to the client device that the user has been placed on hold. In some variations of those implementations, the client device can detect that the user may be on hold and provide a user interface output (e.g., selectable graphical elements and / or audible prompts) that prompts the user as to whether the user wishes to start a hold client. If the user responds with an affirmative user interface input (e.g., selection of a selectable graphical element and / or verbal affirmative input), the hold client can be started. In some other variations of those implementations, the user can start a hold client without the client device detecting that the user may be on hold and / or without the client device prompting the user. For example, the user can provide a verbal command (e.g., "Assistant, initiate on hold monitoring") to start the hold client and / or can select a selectable graphical element whose presence is not conditional on the user's determination that the user may be on hold. In many implementations, the client device can monitor the audio stream for the entire voice communication session and detect whether the user has been placed on hold at a point in time other than the start of the session. For example, the user can interact with the person who placed the user on hold while the person who placed the user on hold is transferring the session to a second person. In various implementations, the hold client can operate in the background to detect when the voice communication session has been placed on hold and can be "started" (e.g., transition to an "active" state), where it is noted that it then performs other aspects of the present disclosure (e.g., to detect when the voice communication session is no longer on hold).

[0009] When the on-hold client is started, the on-hold client can monitor at least the incoming part of the audio stream of the voice communication session to determine when the voice communication session is no longer on hold. When the session is no longer on hold, the calling user can interact with a live person such as a company representative or a receptionist at a clinic. Monitoring the audio stream of a voice communication session using an on-hold client can be performed without direct interaction from the user (e.g., the user does not need to listen to the session while it is on hold).

[0010] In some implementations, the on-hold client can determine when the on-hold music has changed to human voice. Since this human voice may sometimes be a human recording, the on-hold client determines whether a recording is being played or a live human has joined the session. In various implementations, the on-hold client can ask a question (referred to herein as a "response request signal") to the detected voice and check whether the voice responds to the question. For example, when human voice is detected in the audio signal of the session, the on-hold client can ask "Are you there?" and check whether the voice responds to the question. An appropriate response to the question when the on-hold client starts indicates that the hold has ended and a second person has joined the session. In other implementations, the question is ignored and the on-hold client can determine that a second person has not joined the session. For example, if the on-hold client transmits "Is anyone there?" as an input to the audio signal and does not receive a response (e.g., instead the on-hold music continues to play), it may indicate that the voice is a recording and the session is still on hold.

[0011] In some implementations, a "candidate for end of hold event" may be used to determine when a hold may have ended. In many implementations, this candidate for end of hold event can initiate a holding client that sends a request-for-response signal over the session's audio channel to confirm whether the voice is human. This candidate for end of hold event can be detected in various ways. For example, a client device can detect when music stops playing and / or when a person starts speaking. The change from music to a person speaking can be determined using various audio fingerprinting processes, including a discrete Fourier transform (DFT). The DFT monitors blocks of the hold session and can determine when a sufficient change from one block is detected compared to the previous block (e.g., detecting the block when music stops playing and the change from music to human voice in additional blocks). In various implementations, one or more machine learning models can be trained and used to determine when a hold session has changed from audio to human voice.

[0012] In many implementations, the threshold for determining when to ask a question (sometimes called a "response request signal") via an audio signal is low, and asking the question takes very few computational resources (and doesn't bother the person if they are not currently on the other side of the session), so the pending client asks questions frequently. In some of those implementations, a first machine learning model can detect candidates for the end of a pending event and be used to determine when to ask as an input to the audio signal. Determining whether a response has been detected may require additional computational resources, and in various implementations, a second machine learning model (in addition to the first machine learning model) can determine whether a person has responded to the question. The second machine learning model used to detect whether a person has responded to a response request signal can be stored locally on the client device and / or stored outside the client device, i.e., on one or more remote computing systems, often called the "cloud". In some implementations, the pending client can use a single machine learning model to combine all parts of handling a pending session. In some of those implementations, the machine learning model can be used to process an audio stream and provide an output indicating the likelihood that the session is pending. In some variations of those implementations, meeting a first, easier threshold regarding likelihood can be used to determine candidates for the end of the pending state, and meeting a second, harder threshold regarding likelihood can be used to determine the actual end of the pending state.

[0013] In various implementations, one or more machine learning models can utilize an audio stream as input, and one or more models can make determinations such as that a voice communication session is on hold, that the hold is potentially ending, that a response request signal should be sent as input to the audio stream of the voice communication session, and / or that the hold of the voice communication session has ended and there is no need to send a response request signal, and can generate various outputs. In some implementations, a single machine learning model can perform all audio stream analysis for a client on hold. In other implementations, the outputs of different machine learning models can be provided to the client on hold. Additionally or alternatively, in some implementations, a portion of the client on hold can provide inputs and / or receive outputs with one or more machine learning models, and a portion of the client on hold does not interact with any machine learning models.

[0014] Additionally or alternatively, some of the voice communication sessions on hold can verbally indicate the estimated remaining hold time. In many implementations, the on-hold client can determine the remaining hold time estimated by analyzing natural language within the audio stream of the voice communication session and can indicate the remaining hold time estimated to the user. In some such implementations, the estimated remaining hold time can be rendered to the user as a dialog box on a client device having a display screen, such as by pushing a pop-up message that says "Your on hold call with “Hypothetical Water Company" has provided an updated remaining estimated hold of 10 minutes". This message can be displayed on the client device in various ways, including as a new pop-up on the screen, as a text message, etc., as part of the on-hold client. Further, the client device can additionally or alternatively render this information to the user as a verbal indication using one or more speakers associated with the client device. In some implementations, the on-hold client can learn the average amount of time the user spends on holds with known numbers and can provide the average hold time (e.g., as a countdown) to the user when a more specific estimate is unknown. A machine learning model associated with the on-hold client can learn that the length of the estimated remaining hold is indicated within the audio stream when using the audio stream as an input and / or can learn the estimated hold times for known numbers.

[0015] Machine learning models can include feedforward neural networks, recurrent neural networks (RNNs), convolutional neural networks (CNNs), and the like. The machine learning model can be trained using a set of labeled training data having a labeled output corresponding to a given input. In some implementations, labeled audio streams of a set of previously recorded pending voice communication sessions can be used as a training set for the machine learning model.

[0016] In various implementations, the on-hold client can utilize speaker diarization, which can split the audio stream of the session to detect individual voices. Speaker diarization is the process of splitting an input audio stream into homogeneous segments according to speaker identification information. It answers the question "who spoke when" in a multi-speaker environment. For example, speaker diarization can be used to identify that the first segment of the input audio stream can be attributed to a first human speaker (without specifically identifying who the first human speaker is), the second segment of the input audio stream can be attributed to a different second human speaker (without specifically identifying who the first human speaker is), the third segment of the input audio stream can be attributed to the first human speaker, and so on. When a particular voice is detected, the on-hold client can query the voice to confirm whether a response has been received. If the voice does not respond to the on-hold client's question (e.g., "Hello, are you there?"), the on-hold client can determine that the identified voice is a recording and not an indication that the hold has ended. A particular voice can be learned as a recording by the on-hold client and, if heard again during the hold of a voice communication session, that voice is ignored. For example, the voice characteristics of a particular voice and / or the words spoken by a particular voice can be identified, and future occurrences of those voice characteristics and / or words in the voice communication session can be ignored. In other words, often when a person is on hold, the recording played for the user is a loop that includes music interrupted by the same recording (or one of several recordings). Voices identified as recordings within this on-hold recording loop are ignored (i.e., not prompted again with the same voice) if the hold of the voice communication session loops back to the same identified voice. In some such implementations, recorded voices can be shared among many client devices as known voice recordings.

[0017] In some implementations, the content detected in the audio signal is a strong indicator that the human user is on the phone without the need for a response request signal. For example, if the on-hold client detects one of a list of keywords and / or phrases such as the name of the caller, the surname of the caller, the full name of the caller, etc., the on-hold client can determine that a live human user is on the phone without asking any questions via the audio stream of the voice communication session. Additionally or alternatively, service representatives often follow a script when interacting with the user. The on-hold client can monitor for a greeting along a typical script from a particular company's service representative to identify that the hold has ended without sending questions via the audio stream. For example, assume the user called a "Hypothetical Utility Company" at a particular number. The on-hold client can learn the response along the script used by the service representative at the "Hypothetical Utility Company" when answering the voice communication session. In other words, the on-hold client can learn that the service representative at the "Hypothetical Utility Company" starts the voice communication session with the user with a message along the script such as "Hello, my name is [service representative's name] and I work with Hypothetical Utility Company. How may I help you today?" Detecting a message along the script can trigger ending the on-hold client without further needing to query the voice to confirm whether it is a live second user.

[0018] When the holding client detects the end of the hold, in various implementations, the holding client can send scripted messages to a second user who is currently also active in the session. For example, the holding client can send a message such as "Hello, I represent Jane Doe. I am notifying her and she will be here momentarily". This message helps keep the second user on the phone while the user who initiated the session is being notified of the end of the hold. Additionally or alternatively, the voice communication session can be passed to an additional client instead of returning to the user to interact with the session. In some such implementations, the additional client can use information known about the user and / or information the user has provided to the additional client regarding a particular voice communication session to interact with the voice communication session. For example, the user can provide the additional client with information about when they want to make a dinner reservation at the "Hypothetical Fancy Restaurant", and the additional client can interact with additional live human users to make a dinner reservation for the user.

[0019] In many implementations, the user who initiated the session is notified when the held client determines that the hold state has ended (i.e., the hold has ended and the person has picked up the phone). In some implementations, the user can select how they would like to be notified when the held client is started or around the same time. In other implementations, the user can select how they would like to be notified as a setting within the held client. The user can be notified using the client device itself. For example, the client device can notify the user by ringing a ringtone on the client device, vibrating the client device, providing speech output on the client device (e.g., "you are no longer on hold"), etc. For example, the client device can vibrate when the hold ends, and the user can press a button on the client device to initiate a conversation with the session.

[0020] Additionally or alternatively, the queued client can notify the user via, for example, one or more other client devices and / or peripheral devices (e.g., Internet of Things (IoT) devices) that are part of the same tuned ecosystem of client devices that are shared on the same network and / or under the user's control. The queued client can have knowledge of other devices on the network through the device topology. For example, if the queued client knows that the user is in a room with a smart light, the user can choose to be notified by changing the state of the smart light (e.g., turning the light on / off and blinking, dimming the light, increasing the light intensity, changing the color of the light, etc.). As another example, a user interacting with a display screen such as a smart TV can choose to be notified by a message that appears on the display screen of the smart TV. In other words, the user can watch the TV while the session is queued, and through the user's TV, be notified by the queued client that the queue has ended and the user can rejoin the session. As yet another example, a voice communication session can be conducted via a mobile phone, and the notification can be rendered via one or more smart speakers and / or other client devices. In various implementations, the client device used for the voice communication session can be a mobile phone. Alternative client devices can be used for the voice communication session. For example, the client device used for the voice communication session can include a dedicated automated assistant device (e.g., a smart speaker and / or other dedicated assistant device) having the ability to conduct a voice communication session for the user.

[0021] The implementations disclosed in this specification can improve the usability of a client device by shortening the time the client device spends interacting with a pending voice communication session. Instead of a client device fully interacting with a pending voice communication session, computing resources can be conserved by executing a pending client process in the background of a computing device. For example, many users output a pending voice communication session through a speaker associated with the client device. Background monitoring of the voice communication session requires less computational processing by the client device as compared to outputting the session at the speaker. Additionally or alternatively, executing a pending process in the background of the client device can conserve the battery life of the client device as compared to outputting a pending voice communication session through one or more speakers associated with the client device (which may further include outputting an audio stream that a user can hear when the client device is next to the user's ear and outputting an audio stream of the pending voice communication session by an external speaker).

[0022] The foregoing is provided as an overview of the various implementations disclosed in this specification. Additional details are provided herein with respect to those various implementations, as well as additional implementations.

[0023] In some implementations, a method implemented by one or more processors is provided, including detecting that a voice communication session is on hold. The voice communication session is initiated by a calling user's client device, and the step of detecting that the voice communication session is on hold is at least partially based on the audio stream of the voice communication session. The method further includes starting a holding client on the client device. The step of starting the holding client is during the voice communication session and is based on detecting that the voice communication session is on hold. The method further includes using the holding client to monitor the audio stream of the voice communication session for candidates for the end of the hold state. The step of monitoring the audio stream of the voice communication session occurs without direct interaction from the calling user. The method further includes detecting candidates for the end of the hold state based on the monitoring. The method further includes, in response to detecting a candidate for the end of the hold state, transmitting a response request signal from the client device as an input to the audio stream of the voice communication session, monitoring the audio stream of the voice communication session for a response to the response request signal, and determining that the response to the response request signal indicates that the candidate for the end of the hold state is the actual end of the hold state. The actual end of the hold state indicates that a human user is available to interact with the calling user in the voice communication session. The method further includes rendering a user interface output in response to determining the actual end of the hold state. The user interface output is recognizable by the calling user and indicates the actual end of the hold state.

[0024] These and other implementations of the technology disclosed herein can include one or more of the following features.

[0025] In some implementations, the step of detecting candidates for ending the hold state includes the step of detecting the voice of a human speaking in the audio stream of the voice communication session.

[0026] In some implementations, the client device is a mobile phone or a stand-alone interactive speaker.

[0027] In some implementations, the step of starting a held client responds to a user interface input provided at the client device by the calling user. In some variations of those implementations, the method further includes the step of rendering, at the client device, a proposal to start a held client in response to detecting that the voice communication session is in a held state. In those variations, the user interface input provided by the calling user is a positive user interface input provided in response to rendering the proposal at the client device.

[0028] In some implementations, the held client is automatically started by the client device in response to detecting that the voice communication session is in a held state.

[0029] In some implementations, the step of detecting that the voice communication session is in a held state includes the step of detecting music in the audio stream of the voice communication session and, optionally, the step of determining that the music is included in a list of known held music.

[0030] In some implementations, the step of detecting that the voice communication session is in a held state is further based on the step of determining that the phone number associated with the voice communication session is on a list of phone numbers known to put callers on hold.

[0031] In some implementations, the step of detecting candidates for ending the hold state includes using audio fingerprinting to determine at least a threshold change in the audio stream.

[0032] In some implementations, the step of determining that a response to a response request signal indicates that a candidate for ending the hold state is the actual end of the hold state includes processing the response using at least one machine learning model to generate at least one prediction output, and determining, based on the at least one prediction output, that a candidate for ending the hold state is the actual end of the hold state. In some variations of those implementations, the at least one prediction output includes predicted text regarding the response, and the step of determining, based on the prediction output, that a candidate for ending the hold state is the actual end of the hold state includes determining that the text responds to the response request signal. In some additional or alternative variations of those implementations, the at least one prediction output includes a prediction as to whether the response is human speech, and the step of determining, based on the prediction output, that a candidate for ending the hold state is the actual end of the hold state includes determining that the prediction as to whether the response is human speech indicates that the response is human speech.

[0033] In some implementations, after determining that a response to a response request signal indicates that a candidate for ending the hold state is the actual end of the hold state, the method further includes transmitting, from the client device, an end-hold message as an input to the audio stream of the voice communication session. The end-hold message is audible to a human user and indicates that the calling user has returned to the voice communication session. In some of those implementations, after determining that a response to a response request signal indicates that a candidate for ending the hold state is the actual end of the hold state, the method further includes ending the held client at the client device.

[0034] In some implementations, the user interface output indicating the actual end of the hold state is rendered via the client device, an additional client device linked to the client device, and / or a peripheral device (e.g., a networked light).

[0035] In some implementations, the method further includes identifying one or more pre-recorded voice characteristics of a pre-recorded human voice associated with a telephone number (or other unique identifier) associated with the voice communication session. In some variations of those implementations, the step of determining that a response to the response request signal indicates that a candidate for the end of the hold state is the actual end of the hold state includes determining one or more response voice characteristics for the response and determining that the one or more response voice characteristics are different from the one or more pre-recorded voice characteristics.

[0036] In some implementations, a method is provided that is performed by one or more processors and includes receiving user interface input provided via a client device. The user interface input is provided by the calling user while a voice communication session is on hold. The voice communication session is initiated by the client device, and the called party controls the hold state of the voice communication session. The method further includes, in response to receiving the user interface input, monitoring audio generated by the called party during the voice communication session for candidates for ending the hold state. The method further includes detecting a candidate for ending the hold state based on the monitoring. The method further includes, in response to detecting a candidate for ending the hold state, transmitting an audible output for inclusion within the voice communication session by the client device. The audible output includes recorded human voice speaking one or more words or synthetically generated voice speaking one or more words. The method further includes monitoring audio generated by the called party following the audible output and determining that the audio generated by the called party following the audible output meets one or more criteria indicating that the candidate for ending the hold state is the actual end of the hold state. The actual end of the hold state indicates that the human user is available to interact with the calling user in the voice communication session. The method further includes, in response to determining the actual end of the hold state, rendering a user interface output. The user interface output is recognizable by the calling user and indicates the actual end of the hold state.

[0037] These and other implementations of the technology can optionally include one or more of the following features.

[0038] In some implementations, the step of determining that the audio generated by the callee following the audible output meets one or more criteria includes generating text by performing a speech-to-text conversion of the audio generated by the callee following the audible output, and determining that the text responds to one or more words of the audible output.

[0039] In some implementations, the user interface input is a positive response to a graphical and / or audible proposal rendered by the client device, and the proposal is a proposal to initiate a pending client to monitor for the end of the hold state. In some of those implementations, the proposal is rendered by the client device in response to detecting that the call is on hold, based on audio generated by the callee during the voice communication session.

[0040] In some implementations, a method is provided that is performed by a client device that initiated a voice communication session, and while the voice communication session is on hold, includes monitoring the audio stream of the voice communication session for the occurrence of human speech in the audio stream, transmitting a response request signal as an input to the audio stream in response to detecting the occurrence of human speech during monitoring, monitoring the audio stream for a response to the response request signal, determining whether the response to the response request signal is a response from a human who responded to the response request signal, and if the response is determined to be a response from a human who responded to the response request signal, causing a user interface output that is recognizable by the calling user and indicates the end of the hold state to be rendered.

[0041] In addition, some implementations include one or more processors of one or more computing devices, the one or more processors being operable to execute instructions stored in associated memory, the instructions being configured to cause execution of any of the methods described above. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to perform any of the methods described above.

Brief Description of the Drawings

[0042]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0043] FIG. 1 shows an exemplary environment 100 in which various implementations may be implemented. The exemplary environment 100 includes one or more client devices 102. For brevity and simplicity, the term “pending client” as used herein to refer to what “serves” a particular user often refers to a combination of a pending client 104 operated by a user on a client device 102 and one or more cloud-based pending components (not shown).

[0044] The client device 102 may include, for example, a desktop computing device, a laptop computing device, a touch-sensitive computing device (e.g., a computing device capable of receiving input via a touch from a user), a cellular phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system), a stand-alone interactive speaker, a smart home appliance such as a smart TV, a projector, and / or a wearable device of the user including a computing device (e.g., a wristwatch of the user having a computing device, glasses of the user having a computing device, a virtual or augmented reality computing device, etc.). Additional and / or alternative computing devices may be provided.

[0045] In some implementations, the pending client 104 may participate in a dialog session in response to a user interface input even if the user interface input is not explicitly directed to the pending client 104. For example, the pending client 104 may examine the content of the audio stream of a voice communication session and / or the content of the user interface input and may participate in a dialog session. For example, in response to a particular term present in the audio stream of a voice communication session in the user interface input and / or based on other cues, the pending client may be able to participate in a dialog session. In many implementations, the pending client 104 utilizes speech recognition to convert the user's utterance to text and, in response thereto, may respond to the text, for example, by providing search results, general information, and / or taking one or more response actions (e.g., initiating hold detection, etc.).

[0046] Each client device 102 may execute respective instances of the pending client 104. In various implementations, one or more aspects of the pending client 104 may be implemented remotely from the client device 102. For example, one or more components of the pending client 104 may be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) communicatively coupled to the client device 102 via one or more local and / or wide area networks (e.g., the Internet). Each of the client computing devices 102 may include one or more memories for storing data and software applications, one or more processors for accessing the data and executing the applications, and other components for facilitating communication via a network. Operations performed by one or more of the computing devices 102 and / or the pending client 104 may be distributed across multiple computer systems. The pending client 104 may be implemented as a computer program executed on one or more computers operating at one or more locations coupled to each other via a network, for example.

[0047] In many implementations, the pending client 104 may include a corresponding audio capture / text-to-speech ("TTS") / speech-to-text ("STT") module 106, a natural language processor 108, an audio stream monitor 110, a pending detection module 112, and other components.

[0048] The pending client 104 may include the corresponding voice capture / TTS / STT module 106 described above. In other implementations, one or more aspects of the voice capture / TTS / STT module 106 may be implemented separately from the pending client 104. Each voice capture / TTS / STT module 106 may be configured to perform one or more functions, i.e., to capture the user's voice via, for example, a microphone (not shown) integrated into the client device 102, to convert the captured audio to text (and / or other representations or embeddings), and / or to convert text to voice. For example, in some implementations, since the client device 102 may be constrained with respect to computing resources (e.g., processor cycles, memory, battery, etc.), the voice capture / TTS / STT module 106 that is local to each client device 102 may be configured to convert a finite number of different utterance phrases, particularly phrases that call the pending client 104, to text (or other forms such as lower-dimensional embeddings). Other voice inputs may be sent to a cloud-based pending client component (not shown) that may include a cloud-based TTS module and / or a cloud-based STT module.

[0049] The natural language processor 108 of the pending client 104 may process natural language input generated by the user via the client device 102 and generate an annotated output for use by one or more components of the pending client 104. For example, the natural language processor 108 may process natural language free-form input generated by the user via one or more user interface input devices of the client device 102. The generated annotated output may include one or more annotations of the natural language input and, optionally, one or more (e.g., all) of the terms of the natural language input.

[0050] In some implementations, the natural language processor 108 is configured to identify and annotate various types of grammatical information in natural language input. For example, the natural language processor 108 may include a part-of-speech tagger configured to annotate terms with their grammatical roles. Also, for example, in some implementations, the natural language processor 108 may additionally and / or alternatively include a dependency parser (not shown) configured to determine syntactic relationships between terms within the natural language input.

[0051] In some implementations, the natural language processor 108 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references within one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), etc. The entity tagger of the natural language processor 108 may annotate references to entities at a high level of granularity (e.g., to enable identification of all references to an entity class such as a person) and / or at a low level of granularity (e.g., to enable identification of all references to a specific entity such as a particular person). The entity tagger may depend on the content of the natural language input to resolve a particular entity and / or may optionally communicate with a knowledge graph or other entity database to resolve a particular entity.

[0052] In some implementations, the natural language processor 108 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, the coreference resolver may be utilized to resolve the term "there" in the natural language input "I liked Hypothetical Cafe last time we ate there" to "Hypothetical Cafe".

[0053] In many implementations, one or more components of the natural language processor 108 may depend on annotations from one or more other components of the natural language processor 108. For example, in some implementations, the named entity tagger may depend on annotations from the coreference resolver and / or the syntactic parser when annotating all mentions of a particular entity. Also, for example, in some implementations, the coreference resolver may depend on annotations from the syntactic parser when clustering references to the same entity. In many implementations, when processing a particular natural language input, one or more components of the natural language processor 108 may use related prior inputs external to the particular natural language input and / or other related data to determine one or more annotations.

[0054] In many implementations, the queued client 104 can interact with the queued voice communication session without any required interaction from the user who placed the session. In some additional or alternative implementations, the queued client 104 can initiate the queued process, end the queued process, notify the user that the voice communication session is no longer queued, and / or pass the no-longer-queued voice communication session to an additional client on the client device 102.

[0055] In many implementations, the audio stream monitor 110 can be used by the client device 102 and / or the queued client 104 to monitor the incoming and / or outgoing portions of the audio stream of a voice communication session. For example, the incoming portion of the audio stream may include the audio portion (e.g., another person's voice, music, etc.) that the calling party hears after a voice communication session has been established. Similarly, the outgoing portion of the audio stream of a voice communication session may include what the calling party said to another calling party via the audio stream and / or other signals provided by the queued client (such as a request response query asking whether another person is on the phone). In some such implementations, the client device 102 can use the audio stream monitor 110 to detect when a voice communication session has been queued and pass the queued voice communication session to the queued client 104. Additionally or alternatively, the queued client 104 can monitor the audio stream of the voice communication session, and the queued client 104 itself can determine when the voice communication session has been queued. Signals within the audio stream detected by the audio stream monitor 110 indicating that a voice communication session has been queued may include detection of known queued music, detection of any music (since it is unlikely that users will play songs to each other via a voice communication session), transitions from human voice to music, transitions from music to human voice, etc.

[0056] The hold detection module 112 can use the determinations regarding the audio stream of the voice communication session made by the audio stream monitor 110 to determine when a voice communication session has been placed on hold, when the voice communication session is no longer on hold, and to determine an estimated remaining wait time, etc. The hold detection module 112 can provide an indication to the user of the client device 102 when the session is no longer on hold, and can pass the voice communication session to an additional client on the client device 102 to interact with the voice communication session (which may or may not require further interaction from the user).

[0057] Additionally or alternatively, the user can indicate to the client device 102 via the user interface that the voice communication session is on hold but the user wishes to start a hold process using the hold client 104. The hold detection module 112 can hold the session when it receives a positive indication from the user via the user interface on the client device 102 that the session is on hold, by recommending that the session be held and the user respond in an affirmative manner to start a hold process, and / or by the user directly indicating that the session is on hold via the user interface on the client device 102 using the hold detection module 112 to start a hold process. In other implementations, the hold detection module 112 can automatically start a hold process when it detects a session in a held state.

[0058] In many implementations, the hold detection module 112 can additionally or alternatively determine when a session is no longer on hold. In many implementations, the user can indicate a way in which the user desires to be notified when the user's held process ends. For example, the user may desire to receive a voice communication session on the mobile computing system that indicates that it is from the hold number. Additionally or alternatively, the user can require that a connected smart device within the same ecosystem as the client device 102, such as a smart light, respond in a particular way when the end of the hold is detected. For example, a smart light on the same network as the client device 102 can be instructed to blink on and off, dim in intensity, increase in intensity, change color, etc. to indicate the end of the hold of the voice communication session. Additionally or alternatively, a user watching a smart TV can require that a notification appear on the TV when the end of the hold is detected.

[0059] Figures 2, 3, and 4 each show an interaction between a held client (such as the held client 104 shown in FIG. 1) and a voice communication session. FIG. 2 shows an image 200 that includes a held client 202 that is interacting with a voice communication session 206 that is still on hold. In response to detecting a potential (also called "candidate") end of the hold of the voice communication session, the held client 202 can transmit a response request signal via the audio stream of the voice communication session 206 to determine whether an additional live user has become active in the session. In many implementations, the held client can determine a text phrase (e.g., "Are you there") to transmit as the response request signal. In some such implementations, a text-to-speech module (similar to the voice capture / TTS / STT module 106 shown in FIG. 1) can convert the text phrase to audio to provide as an input to the audio stream.

[0060] In various implementation forms, the potential end of the hold of a voice communication session can be detected by the holding client 202 by detecting any of various signals within the audio stream of the voice communication session, including changes in music, changes from music to human voice (not only raw voice but also potentially recorded voice), signals detected by various signal processing techniques such as discrete Fourier transform, the output of a neural network model, etc. Human voice can be analyzed as a signal, and additionally or alternatively, a voice text conversion module (similar to the voice capture / TTS / STT module 106 shown in FIG. 1) can convert human voice into text. The spoken language of the text within the audio stream can be further analyzed by a natural language processor (such as the natural language processor 108 shown in FIG. 1) to determine the meaning of what was spoken by the human voice detected within the audio stream. The output of the natural language processor can be further used to determine the potential end of the hold in the voice communication session. Additionally or alternatively, the output of the natural language processor can be used when determining that a live human user has entered the session. For example, the output of the natural language processor can provide input to one or more neural network models.

[0061] In some implementations, the neural network model can learn to identify one or more "voices" to be ignored within a voice communication session. The voices can include one or more individual speakers, background music, background noise, and the like. For example, one or more neural network models can include a recurrent neural network (RNN). The RNN can include at least one memory layer, such as a long short-term memory (LSTM) layer. The memory layer includes one or more memory units to which inputs can be sequentially applied, and in each iteration of the applied input, the memory units can be utilized to compute a new hidden state based on the input of that iteration and based on the current hidden state (which can be based on the input of the previous iteration). In some implementations, the model can be used to generate speaker diarization results for any of various lengths of audio segments. As an example, the audio stream of a voice communication session can be split into one or more data frames. Each data frame can be a portion of the audio signal, such as 25 milliseconds or other duration portion. Frame features (or the frames themselves) can be sequentially applied as input to a trained speaker diarization model to generate a series of outputs, each including the corresponding probability for each of N invariant speaker labels. For example, the frame features of audio frame 1 can be first applied as input to generate N probabilities, each of the N probabilities corresponding to one of the N speaker labels, and the frame features of audio data frame 2 can then be applied as input to generate N probabilities, each of the N probabilities corresponding to one of the N speaker labels, and so on. It should be noted that the N probabilities generated for audio data frame 2 are specific to audio data frame 2, but since the model can be an RNN model, they depend on the processing of audio data frame 1.

[0062] Additionally or alternatively, the N probabilities can indicate whether the session has been put on hold, whether the session is still on hold, and / or whether an indication of the end of a potential hold has been detected. In many implementations, an estimated remaining hold time can be determined for the voice communication session (through the knowledge that the on-hold client has of the typical hold length for a particular called number and / or the estimated remaining hold time indicated within the audio stream of the voice communication session). The estimated remaining hold time can be additionally input into a machine learning model by many implementations, and the shorter the remaining estimated hold time, the more likely the machine learning model can output that the hold has ended.

[0063] In other implementations, the on-hold client can use knowledge of the potential remaining hold time to increase and / or decrease the threshold used to send a response request signal (regardless of the use of one or more machine learning models). For example, if a voice communication session is expected to have 20 minutes of remaining hold, the on-hold client can have a higher threshold for sending a response request signal. Similarly, a voice communication session that is expected to have only a few minutes (e.g., 3 minutes) can have a lower threshold for sending a response request signal.

[0064] Detection of a potential end of a hold of a voice communication session can cause the on-hold client 202 to send a response request signal via the audio stream of the voice communication session to determine whether an additional user has joined the voice communication session and whether the hold has ended. For example, the on-hold client 202 can send a response request signal 204 such as "Are you there". Additionally or alternatively, the response request signal can be any of a variety of questions that prompt a response, such as "Is anyone there", "Hello, are you there", "Am I still on hold".

[0065] In many implementations, the response request signal can prompt a predictable response from an additional live human user who has resumed the voice communication session. For example, a response to the response request signal 204 of "Are you there" can include similar words or phrases that indicate "yes" and / or a positive response (e.g., "Yeah", "Yup", and phrases that can include a positive response). Transmitting the response request signal as an input to the audio stream of the voice communication session may be computationally inexpensive. Additionally or alternatively, since the likelihood of confusing the recording by repeating the same question (which can be played during the hold of the voice communication session) is low, the threshold for transmitting a response request query may be low. In other words, in many implementations, the holding client has few (if any) disadvantages in frequently transmitting the response request signal, so the response request signal is frequently transmitted. Further, if the holding client fails to transmit the response request signal when it should be transmitted, the voice communication session may potentially end, and the user may need to start the holding process again using the phone number.

[0066] In many implementations, the response request signal 204 can be transmitted via the audio stream of the voice communication session when the hold has not ended. If the response request signal is transmitted and the hold of the voice communication session has not ended, the response 208 is not detected by the holding client 202 in the audio stream of the voice communication session 206.

[0067] In many implementations, recorded voice may be played back while a voice communication session is on hold. In some such implementations, the recorded voice does not respond to a response request signal, and the on-hold client can learn not to send a response request signal to that voice in the future. For example, while on hold, a phone number may play a recording that includes information about the called number (e.g., website, business hours, etc.). This recording that includes information about the number may loop several times while the voice communication session is on hold. When the on-hold client determines that this voice does not respond to a response request signal, the on-hold client can learn not to send an additional response request signal to that particular voice. In many implementations, the on-hold client can learn to ignore a voice using one or more of various signals (e.g., voice fingerprinting) generated by a particular voice that includes the pitch of the voice, identification information of the voice itself, and / or a particular series of words the voice is speaking.

[0068] Figure 3 shows an image 300 that includes a pending client 302 that is interacting with a voice communication session 306. In many implementations, the pending client 302 can send a response request signal 304, such as "Is anyone there?", as an input to the audio stream of the voice communication session. The text response request signal provided by the pending client can be converted to audio using an SST module (such as the voice capture / TTS / STT module 106 shown in FIG. 1). For example, the pending client can provide the text phrase "Is anyone there?" as a response request signal. The STT module can convert this phrase to the spoken language that can be sent as an input to the audio signal of the voice communication session. Determining when to send the response request signal 304 was described above with respect to FIG. 2. The image 300 further shows a pending client that can receive a response 308 of "Yes, I am here" to the response request signal and determine that the voice communication session is no longer pending. When determining that the voice communication session is no longer pending, the pending client can convert the detected input to an audio stream and use an STT module (the voice capture / TTS / STT module 106 shown in FIG. 1) to convert the input to text. Further, a natural language processor (such as the natural language processor 108) can analyze the text response to the response request signal to provide the meaning of the text response.

[0069] As described above with respect to FIG. 2, in many implementations, the question 304 "Is anyone there?" generally elicits an affirmative response from a second user, such as "Yes, I am here." In other implementations, the response request signal can generally be phrased to elicit a negative response. For example, the question "Am I still on hold?" can elicit a negative response from a second user, such as "No, you are not on hold." In some implementations, a held client can utilize a typical response to a particular response request signal that is used in part when the session is determined to no longer be on hold. In many implementations, a user who placed an audio communication session can be notified when the held client 302 determines that the session is no longer on hold.

[0070] In some implementations, the user may be notified that the session is no longer on hold. For example, a mobile phone can ring and / or vibrate to simulate a new incoming session when the hold on a voice communication session is completed. Additionally or alternatively, a networked device near the user can be used as a notification that the hold on a voice communication session has ended. For example, a user who places a voice communication session may be near a smart light. The smart light can blink, dim in intensity, increase in intensity, change color, etc. to notify the user. Additionally or alternatively, a message can be pushed to a screen that the user is interacting with, including a mobile phone, a computing device, a television, etc. For example, a user watching a smart TV in the same device topography as the client device used to initiate the voice communication session can receive a notification on the TV when the hold on the session ends. In various implementations, the user can select how they are to be notified as a hold setting. Additionally or alternatively, the user can select how they are to be notified when the hold process is initiated.

[0071] FIG. 4 shows an image 400 that includes a held client 402 and a voice communication session 406. In many implementations, the held client can receive a very strong indication that the hold on the voice communication session has ended. In some such implementations, the held client can continue to notify the user that the session is no longer on hold without transmitting a request-for-response signal. Human speech detected in the audio stream can be converted to text output using an STT module (the voice capture / TTS / STT module 106 shown in FIG. 1), and the text output can be provided to a natural language processor (such as the natural language processor 108 shown in FIG. 1) to provide the meaning of the text to the held client. For example, a message 404 such as "Hello Ms. Jane Doe. My name is John Smith and I represent ‘Hypothetical Utility Company'. How may I help you today?" can include a strong indication that the voice communication session is no longer on hold. For example, detection of the user's name (such as Jane Doe and / or Ms. Doe), detection of a phrase indicating an additional user name ("My name is John Smith"), and detection of other phrases (such as "How may I help you today?") can all, individually and / or in combination, cause the held client to determine that the hold on the voice communication session has ended without transmitting a request-for-response signal. In many implementations, when the held client determines that the voice communication session is no longer on hold, the user can be notified as previously described.

[0072] Figure 5 is a flowchart showing an exemplary process 500 according to many implementations disclosed herein. For convenience, the operations of the flowchart of FIG. 5 will be described with reference to the system that performs the operations. This system can include various components of various systems, such as one or more components of the client device 102. Further, although the operations of process 500 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.

[0073] In block 502, the client device can optionally determine that a voice communication session is on hold. As described above with respect to the hold detection module 112 shown in FIG. 1, the client device can determine that a voice communication session is on hold in a variety of ways, including detecting known hold music, detecting any music, detecting a change from human voice to music, detecting direct input from the user that the session has been placed on hold, determining that the called number is a known number that generally places the user on hold, and using any of a variety of signal processing techniques including discrete Fourier transform, and in a determination by one or more machine learning models associated with the on-hold client in the client device.

[0074] In block 504, the client device starts an on-hold client similar to the on-hold client 104 described above with respect to FIG. 1.

[0075] In block 506, the on-hold client can monitor the incoming and / or outgoing portions of the audio stream of the voice communication session on hold. In many implementations, the on-hold client can monitor the audio stream in a manner similar to the audio stream monitor 110 described above with respect to FIG. 1.

[0076] In block 508, the on-hold client can determine when to send a response request signal via the audio stream of the voice communication session. Various ways for the on-hold client to determine to send a response request signal were described above with respect to FIG. 2. In many implementations, the on-hold client can send one or more response request signals until the voice communication session is no longer on hold and / or until the on-hold client receives an instruction from the user to end the on-hold process (e.g., the user gets tired of waiting on hold, ends the on-hold process, and wants to call the phone number again later). In other implementations, the on-hold client cannot send a response request signal. For example, a strong indicator that the session is no longer on hold (as described above with reference to FIG. 4) can be detected, and the on-hold client can determine that the voice communication session is no longer on hold without sending a response request signal.

[0077] In block 510, the on-hold client can determine that the voice communication session is no longer on hold. In various implementations, this determination can be made based on the received response to the response request signal. In other implementations, this determination can be made using the strength of the information monitored via the audio stream that is strong enough to indicate that the voice communication session is no longer on hold without sending a response request signal. Additionally or alternatively, the on-hold client can send one or more (unanswered) response request signals and then receive a strong indicator that the voice communication session is no longer on hold such that no additional response request signals are sent.

[0078] In block 512, the holding client notifies the user that the voice communication session is no longer on hold. The various ways in which the holding client can notify the user of the end of the hold of the voice communication session were described above with respect to FIG. 1. Additionally or alternatively, the holding client can pass the voice communication session to another client associated with the client device to process the voice communication session on behalf of the user. For example, when the holding client determines that the voice communication session is no longer on hold, the holding client can pass the voice communication session to a second client that can interact with additional people on the voice communication session on behalf of the user.

[0079] FIG. 6 is a block diagram of an exemplary computer system 610. Computer system 610 typically includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices can include, for example, a storage subsystem 624 that includes a memory 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with computer system 610. The network interface subsystem 616 provides an interface to an external network and is coupled to a corresponding interface device within another computer system.

[0080] The user interface input device 622 can include a keyboard, a mouse, a trackball, a touchpad, or a pointing device such as a graphics tablet, a scanner, a touch screen incorporated in a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 610 or a communication network.

[0081] The user interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 610 to a user or another machine or computer system.

[0082] The memory subsystem 624 stores programming structures and data structures that provide some or all of the functionality of some of the modules described herein. For example, the memory subsystem 624 can include logic for executing selected aspects of the client device shown in FIG. 1, the process 500 shown in FIG. 5, any operation discussed herein, and / or any other device or application discussed herein.

[0083] These software modules are generally executed either alone or in combination with other processors by processor 614. Memory 625 used in memory subsystem 624 can include several memories, including main random access memory (RAM) 630 for storing instructions and data during program execution and read-only memory (ROM) 632 in which fixed instructions are stored. File storage subsystem 626 can provide permanent storage for program files and data files and can include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular implementation can be stored by file storage subsystem 626 within memory subsystem 624 or within other machines accessible by processor 614.

[0084] Bus subsystem 612 provides a mechanism for enabling the various components and subsystems of computer system 610 to communicate with each other as intended. Although bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem can use multiple buses.

[0085] Computer system 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computer system 610 shown in FIG. 6 is intended only as a specific example for the purpose of illustrating some implementations. Many other configurations of computer system 610 with more or fewer components than the computer system shown in FIG. 6 are possible.

[0086] In situations where the systems described in this specification may collect or use personal information about a user (or, as often referred to herein, a "participant"), the user may be provided with the opportunity to control whether a program or function collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location), or the opportunity to control whether and / or how the user receives content from a content server that may be relevant to the user. Also, certain data may be processed in one or more ways before it is stored or used so that information that could identify an individual is removed. For example, a user's identifying information may be processed so that information that could identify an individual cannot be determined about the user, or a user's geographic location may be generalized at the location where the geographic location information is obtained (such as at the city, zip code, or state level) so that the user's specific geographic location cannot be identified. Thus, the user may control how information is collected and / or used about the user.

[0087] Although several implementations are described and illustrated herein, various other means and / or structures may be utilized to perform the functions and / or to obtain the results and / or one or more of the advantages described herein, and each such variation and / or modification is to be regarded as being within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be illustrative, and the actual parameters, dimensions, materials, and / or configurations will depend upon the particular application for which the teachings are used. One of ordinary skill in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the foregoing implementations are presented by way of example only, and it is to be understood that within the scope of the appended claims and their equivalents, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each of the individual features, systems, articles, materials, kits, and / or methods described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

Explanation of Signs

[0088] 100 Environment 102 Client Device 104 Pending Client 106 Audio Capture / Text-to-Speech (「TTS」) / Speech-to-Text (「STT」) Module, Audio Capture / TTS / STT Module 108 Natural Language Processor 110 Audio Stream Monitor 112 Pending Detection Module 200 Image 202 Pending Client 204 Response Request Signal 206 Voice Communication Session 208 Response 300 Image 302 Pending Client 304 Response Request Signal, Question 306 Voice Communication Session 308 Response 400 Image 402 Pending Client 404 Message 406 Voice Communication Session 610 Computer System 612 Bus Subsystem 614 Processor 616 Network Interface Subsystem 620 User Interface Output Device 622 User Interface Input Device 624 Memory Subsystem 625 Memory 626 File Storage Subsystem 630 Main Random Access Memory (RAM) 632 Read Only Memory (ROM)

Claims

**Claim 1** A method performed by one or more processors, comprising: identifying a telephone number associated with a voice communication session initiated by a client device of a calling user; determining whether the voice communication session is on hold based on comparing the identified telephone number with an entity known to put the calling user on hold; in response to determining that the voice communication session is on hold by comparing the identified telephone number with the entity known to put the calling user on hold, examining the voice in the audio stream of the voice communication session; comparing the audio characteristics of the voice in the audio stream with a list of known on-hold voices based on one or more previous voice communication sessions with the identified telephone number; determining that the voice communication session is on hold in response to determining that the voice in the audio stream is the known on-hold voice based on the comparison; The method includes. **Claim 2** The step of determining whether the identified telephone number is associated with the entity known to put the calling user on hold includes comparing the identified telephone number with a list shared among client devices of telephone numbers known to put the calling user on hold; determining that the identified telephone number is included in the list of telephone numbers known to put the calling user on hold; The method according to claim 1, including. **Claim 3** The step of determining whether the identified telephone number is associated with the entity known to put the calling user on hold includes comparing the identified telephone number with a list locally stored on the client device of telephone numbers known to have previously put the calling user on hold; determining that the identified telephone number is included in the list locally stored on the client device of telephone numbers known to have previously put the calling user on hold; The method according to claim 1, including. **Claim 4** In response to determining that the voice communication session is in the held state, starting a holding client in the client device, wherein starting the holding client during the voice communication session and further comprising starting the holding client based on determining that the voice communication session is in the held state, the method of claim 1. **Claim 5** Monitoring, using the holding client, an audio stream of the voice communication session for a candidate for ending the held state, wherein monitoring the audio stream of the voice communication session occurs without direct interaction from the calling-side user, the step of In response to detecting the candidate for ending the held state Transmitting, from the client device, a response request signal as an input to the audio stream of the voice communication session, the step of further comprising the method of claim 4. **Claim 6** Monitoring the audio stream of the voice communication session for a response to the response request signal, the step of Determining that the response to the response request signal indicates that the candidate for ending the held state is the actual end of the held state, wherein the actual end of the held state indicates that a human user is available to interact with the calling-side user in the voice communication session, the step of In response to the determination of the actual end of the held state, causing a user interface output to be rendered, wherein the user interface output is recognizable by the calling-side user and indicates the actual end of the held state, the step of further comprising the method of claim 5. **Claim 7** Monitoring, using the holding client, an audio stream of the voice communication session for a candidate for ending the held state, wherein monitoring the audio stream of the voice communication session occurs without direct interaction from the calling-side user, the step of In response to detecting the actual end of the hold state, causing a user interface output to be rendered, the user interface output being recognizable by the calling user and indicating the actual end of the hold state; The method according to claim 4, further comprising. **Claim 8** When executed by one or more processors, the one or more processors are caused to Identify a phone number associated with a voice communication session initiated by a client device of a calling user; Determine whether the voice communication session is on hold based on comparing the identified phone number with an entity known to put the calling user on hold; In response to determining that the voice communication session is on hold by comparing the identified phone number with the entity known to put the calling user on hold, Examine the voice within the audio stream of the voice communication session; Using the identified phone number, compare the audio characteristics of the voice within the audio stream with a list of known on-hold voices based on one or more previous voice communication sessions; In response to determining that the voice within the audio stream is the known on-hold voice based on the comparison, determine that the voice communication session is in a hold state; A computer-readable storage medium configured to store instructions for performing operations including. **Claim 9** Determining whether the identified phone number is associated with the entity known to put the calling user on hold, Comparing the identified phone number with a list shared among client devices of phone numbers known to put the calling user on hold; Including determining that the identified phone number is included in the list of phone numbers known to put the calling user on hold. The computer-readable storage medium according to claim 8. **Claim 10** Determining whether the identified phone number is associated with the entity known to put the calling user on hold, comparing the identified telephone number with a list locally stored on the client device of telephone numbers known to have previously placed the calling user on hold; determining that the identified telephone number is included in the list locally stored on the client device of telephone numbers known to have previously placed the calling user on hold; The computer-readable storage medium according to claim 8, comprising: [

11. ] The operation is initiating a held client at the client device in response to determining that the voice communication session is in the held state, the held client being initiated during the voice communication session and based on determining that the voice communication session is in the held state; The computer-readable storage medium according to claim 8. [

12. ] The operation is monitoring, using the held client, an audio stream of the voice communication session for candidates for ending the held state, the monitoring of the audio stream of the voice communication session occurring without direct interaction from the calling user; in response to detecting the candidate for ending the held state transmitting, from the client device, a response request signal as an input to the audio stream of the voice communication session; further comprising: The computer-readable storage medium according to claim 11. [

13. ] The operation is monitoring the audio stream of the voice communication session for a response to the response request signal; determining that the response to the response request signal indicates that the candidate for ending the held state is the actual end of the held state, the actual end of the held state indicating that a human user is available to interact with the calling user in the voice communication session; in response to the determination of the actual end of the held state, causing a user interface output to be rendered, the user interface output being recognizable by the calling user and indicating the actual end of the held state; further comprising: The computer-readable storage medium according to claim 12. **Claim 14** The operation is monitoring the audio stream of the voice communication session for candidates for ending the held state using the held client, wherein the monitoring of the audio stream of the voice communication session occurs without direct interaction from the calling user; in response to detecting an actual end of the held state, causing a user interface output to be rendered, the user interface output being recognizable by the calling user and indicating the actual end of the held state; further comprising The computer-readable storage medium according to claim 11.

Citation Information

Patent Citations

  • Double module music detection method

    CN101202992A

  • Communication apparatus and facsimile machine

    JP2004304299A

  • Method and apparatus for replacing telephone on-hold music at caller's side

    JP2015220755A

  • Alerting the user to changes in the audio stream

    JP2019525527A

  • Method for Determining the On-Hold Status in a Call

    US20090136014A1