Escalation to a human operator

An automated system addresses the challenge of collecting data from multiple locations by initiating calls and engaging in conversations, reducing human input and improving efficiency in data collection and task completion.

JP7839218B2Active Publication Date: 2026-04-01GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Users face challenges in efficiently collecting and verifying data from multiple locations without human intervention, as existing systems require manual selection and analysis of call recipients and often encounter automated phone trees that limit user responses.

Method used

An automated or semi-automated system, referred to as a 'bot', initiates calls and engages in conversations to perform tasks on behalf of users, detecting trigger events and following predefined workflows to communicate with humans and automated systems, reducing the need for human input and improving efficiency.

Benefits of technology

The system reduces the amount of human input required for data collection and task completion by automatically initiating calls and handling multiple tasks simultaneously, enhancing efficiency and scalability compared to human operators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839218000001
    Figure 0007839218000001
  • Figure 0007839218000002
    Figure 0007839218000002
  • Figure 0007839218000003
    Figure 0007839218000003
Patent Text Reader

Abstract

To enable a bot to have a proper conversation with a human.SOLUTION: In some implementations, a method includes analyzing, by a call initiating system, a real-time conversation between a first human and a bot during a phone call between the first human on a first end of the phone call and the bot on a second end of the phone call. The call initiating system can determine, based on the analysis of the real-time conversation, whether the phone call should be transitioned from the bot to a second human on the second end of the phone call. In response to determining that the phone call should be transitioned to the second human on the second end of the phone call, the call initiating system transitions the phone call from the bot to the second human.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to natural language processing.

Background Art

[0002] Users may need to collect types of information that are not easily accessible without human intervention. For example, in order to verify or collect data from multiple locations of a business or organization, a user may need to place calls to each of the businesses or organizations to collect information. A web search engine can assist a user in such tasks by providing contact information for services or businesses, but the user still needs to place calls to the services or businesses themselves to complete the task.

[0003] To maintain a database of information collected from multiple locations of a business or organization, a human operator can initiate automated calls to multiple businesses to collect data, but when performed manually, it can take time to select the called parties (e.g., all restaurants in a particular town that offer the same cuisine) and place the calls. Additionally, determining when to place a call and whether to place a call generally requires a human analysis of existing data to identify the need for verification, updating, or supplementary information.

[0004] Users may also want to perform tasks such as making a reservation or requesting a service. However, generally, there are people with whom the user must interact to complete the desired task. For example, to make a reservation at a small restaurant without a website, a user may need to place a call and speak to the hostess. In some cases, even when the user places the call themselves, they may encounter an automated phone tree that only accepts a limited set of user responses.

Summary of the Invention

Means for Solving the Problems

[0005] The system can assist users in a variety of tasks, including communicating with humans and automated systems operating over the telephone (e.g., IVR) via telephone calls, by determining from data received by the system whether to initiate a call to a specific number. Once a call is made, the system can retrieve information, provide it to third parties, and, for example, perform actions on behalf of the user. In certain examples, the system participates in a dialogue with a human on behalf of the user. This dialogue can occur via a telephone connection between the system and the human. In some examples, the system may include, operate, or form part of a search engine, following a workflow associated with a user intent that submits a query containing a task to be completed. The system may perform the user's tasks through the operation of at least one autonomous or semi-autonomous software agent ("bot").

[0006] In a typical embodiment, the method includes the steps of: initiating a call and receiving data indicating a first event by a call triggering module of the call initiation system for initiating a call and facilitating a conversation between a bot of the call initiation system and a human caller during the call; determining by the call triggering module and by using the data indicating the first event that the first event is a specific trigger event of a plurality of possible trigger events that trigger a workflow of the call initiation system initiated by initiating a telephone call; selecting a specific workflow from a plurality of possible workflows based on the determined trigger event, wherein the specific workflow corresponds to the determined trigger event; and in response to the selection, i) initiating a telephone call to a caller designated by the specific workflow; and ii) executing the workflow as a two-way conversation between the bot and the caller.

[0007] The implementation may include one or more of the following characteristics. For example, the determined trigger event is a mismatch between a value associated with a first data source and a corresponding value associated with a second data source. The data indicating the first event may be provided by the user. The determined trigger event may be a user request. The determined trigger event may be a specific type of event, such as a weather event, a recreational event, or a seasonal event. The determined trigger event may be a trend detected in a search request submitted to a search engine. The determined trigger event may be the passage of a predetermined period of time.

[0008] In another general embodiment, the method includes the steps of: determining by the task manager module that a triggering event has occurred in order to provide the current status of a user call request; determining by the task manager module the current status of the user call request; generating a representation of the current status of the user call request; and providing the generated representation of the current status of the user call request to the user.

[0009] The implementation may include one or more of the following characteristics. For example, the determined trigger event may be a user request for status. The determined trigger event may be an operator interaction to provide the status to the user after the operator has reviewed the session information associated with the user call request. The determined trigger event may be a status update event. The representation of the current status may be a visual representation. The representation of the current status may be an oral representation. The step of providing the user with a generated representation of the current status of the user call request may include a step of determining a convenient time and method for delivering the current status to the user.

[0010] In another common embodiment, a method for transferring a telephone call from a bot includes the steps of: having a call initiation system analyze a real-time conversation between a first human at a first end of the telephone call and a bot at a second end of the telephone call; having the call initiation system determine, based on the analysis of the real-time conversation, whether the telephone call should be transferred from the bot to a second human at a second end of the telephone call; and having the call initiation system transfer the telephone call from the bot to the second human in response to the determination that the telephone call should be transferred to a second human at a second end of the telephone call.

[0011] The implementation may include one or more of the following features. For example, the step of analyzing the real-time conversation between a first human and a bot during a phone call may include a step of determining the stress level during the phone call based on the first human's behavior, attitude, tone of voice, level of discomfort, language, or word choice. The method may include a step of determining an increase in stress level during the phone call when the bot repeats the same thing, apologizes, or asks for clarification. The method may include a step of determining an increase in stress level during the phone call when the human corrects the bot or complains about the quality of the call. The method may include a step of determining a decrease in stress level during the phone call when the bot responds appropriately to the first human's dialogue. The step of analyzing the real-time conversation between a first human and a bot during a phone call may include a step of determining the confidence level of the call initiation system that the phone call task will be completed by the bot. The step of analyzing the real-time conversation between a first human and a bot during a phone call may include a step of determining that the first human has requested that the phone call be transferred to another human. Steps to analyze the real-time conversation between a first human and a bot during a phone call may include determining whether the first human mocked the bot or asked whether the bot was a robot. Steps to determine whether a phone call should be transferred from the bot to a second human may include determining whether the tension exceeds a predefined threshold, and determining, in response to the determination that the tension exceeds the predefined threshold, that the phone call should be transferred from the bot to a second human. Steps to analyze the real-time conversation between a first human and a bot during a phone call may include tracking one or more events in the conversation. Steps to determine whether a phone call should be transferred from the bot to a second human may include using a feature-based rule set to determine whether one or more events in the conversation meet the criteria of a rule, and determining, in response to the determination that one or more events in the conversation meet the criteria of a rule, that the phone call should be transferred from the bot to a second human.

[0012] The step of analyzing the real-time conversation between a first human and a bot during a phone call may include the steps of identifying intents from the conversation and identifying past intents and past outcomes from previous conversations. The step of determining whether a phone call should be transferred from the bot to a second human may include the steps of sending intents, past intents, or past outcomes from the conversation to one or more machine learning models and determining whether the phone call should be transferred based on the intents, past intents, or past outcomes. The second human may be a human operator. The bot may use the same voice as a human operator so that the transfer from the bot to the second human is transparent to the first human. The second human may be the user with whom the bot is making a phone call. The method may include the step of terminating the phone call if the transfer of the phone call from the bot to the second human would take longer than a predetermined amount of time. The method may include the step of terminating the phone call instead of transferring the phone call to a human.

[0013] Other implementations of this embodiment and other embodiments include corresponding methods, apparatus, and computer programs configured to perform actions of an encoded method on a computer storage device. One or more computer programs may consist of having instructions that cause the apparatus to perform actions when executed by a data processing device.

[0014] Certain embodiments of the subject matter described herein may be implemented to achieve one or more of the following advantages: The amount of data storage required for various data sources is reduced because only one set of verified data is stored, rather than multiple sets of unverified data. For example, instead of storing three different unverified sets of opening hours for a particular grocery store (e.g., one set collected from the storefront, one set collected from the store's website, and one set collected from the store's answering machine), the data source can store one set of verified store hours obtained from calls to a human representative of the grocery store.

[0015] By automatically detecting trigger events that indicate a call is to be initiated in the call initiation system, the amount of human input required to perform actions such as collecting data from the callee, scheduling appointments, or providing information to third parties is reduced. Furthermore, because calls are initiated only when a trigger event occurs, the amount of computer resources required to maintain the information database is reduced by reducing the number of outgoing calls. The system automatically initiates calls to specific callees or sets of callees, reducing the amount of analysis that must be performed by humans and the amount of data that must be monitored by humans.

[0016] Furthermore, the system conducts conversations on behalf of human users, further reducing the amount of human input required to perform specific tasks. The call initiation system can coordinate multiple calls simultaneously. For example, a user may want to make a reservation for 30 minutes in the future. The system can make calls to each restaurant specified by the user and conduct conversations with representatives on other lines. An employee at the first restaurant to which the call was made may indicate that a reservation is possible, but the customer needs to be seated at the bar. An employee at the second restaurant to which the call was made may indicate that there is a 20-minute wait, and an employee at the third restaurant to which the call was made may inform the system that the customer needs to finish their meal within an hour at the third restaurant, and therefore the table will be ready within an hour. The system can make parallel calls to each of the three restaurants, consult with the user by presenting options and receiving responses, and based on his responses, can reserve the restaurant best suited to the user while declining all other reservations. The automated call initiation system is more efficient than a human one because the automated system can make these calls at once. A human assistant cannot easily perform all of these calls to restaurants in parallel.

[0017] Details of one or more embodiments of the subject matter described herein are given in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]

[0018] [Figure 1A] This is an exemplary block diagram of a system for a call initiation system, which initiates a call and facilitates a conversation between a call initiation system bot and a human during the call. [Figure 1B] This is an exemplary block diagram of a system for a call initiation system, which initiates a call and facilitates a conversation between a call initiation system bot and a human during the call. [Figure 1C]A diagram showing an exemplary user interface where a user can enter more details regarding a request. [Figure 1D] A diagram showing an example of a user speaking to a bot to make a request. [Figure 2A] A flowchart showing an example of a process for completing a task assigned by a user. [Figure 2B] A flowchart showing another example of a process for completing a task assigned by a user. [Figure 3] A diagram showing an exemplary workflow of a process executed by a system. [Figure 4] A block diagram of a triggering module. [Figure 5] A flowchart showing an example of a process for initiating a phone call. [Figure 6] A block diagram of a task manager module of a system. [Figure 7A] A diagram showing an operator dashboard indicating information about the progress of existing tasks. [Figure 7B] A diagram showing an operator review screen for reviewing one of the tasks requested by a user. [Figure 8] A flowchart showing an example of a process for providing the status of a task. [Figure 9A] A diagram showing the visual status of the haircut reservation request of Figure 1B while reservation scheduling is in progress. <​​​​​​​​​​This figure shows an exemplary process 1100 for transferring a telephone call from a bot to a human. [Figure 12] This is a schematic diagram showing examples of computing devices and mobile computing devices. [Modes for carrying out the invention]

[0019] Similar reference numbers and names in various drawings refer to the same elements.

[0020] This disclosure describes a technology that enables an automated or semi-automated system (automated system), referred to herein as a “bot,” to communicate with people by initiating calls and engaging in conversations with or independently of humans during those calls. The bot receives and monitors data to detect trigger events that indicate a call should be initiated. The bot operates through predefined workflows, or sequences of repeatable behavioral patterns linked by abstract descriptions of actions to be performed, or through intents. Essentially, the bot can use these workflows to determine how to respond to humans and what to tell them in order to perform tasks that are helpful to the user.

[0021] The system handles a variety of tasks that come in as queries, such as "Reserve a table for two at Yves Saint Thomas restaurant on Thursday" or "My sink is leaking and I need a plumber! It's after 10 p.m.!"

[0022] Users who wish to schedule reservations, purchase goods, or request services may need to perform multiple searches and make many calls before completing the tasks they have set out to accomplish. In the first use case, reserving a table at a restaurant, the user might search for the restaurant on a search engine. In some examples, if the restaurant is on a website or application, the query may be performed on that website or application (or through an integration with a website or application), or if it is not on a website or application, the user may make a call to the restaurant to negotiate a reservation.

[0023] For example, the system may be used to make calls on behalf of users. The system communicates with businesses and other services to complete tasks requested by users. In some examples, bots perform much of the communication. In some examples, human operators may review and verify the success of actions performed by bots. In some examples, human operators perform actions, and bots learn from the human operators' communication to improve their automated communication skills.

[0024] In a second use case, users may want to find a plumber outside of normal business hours. Such queries can be more difficult to process. For example, if a user searches for a plumber manually, they might use a search engine to find plumbers and call some of them. The user may need to explain their time constraints, location, and the nature of the problem to each plumber and obtain a price quote. This can be very time-consuming.

[0025] Similarly, in a third use case, you might need to search local stores to check if they have the item in stock, and then call each store to determine if they have the specific item or product you are looking for.

[0026] In addition to assisting users with specific tasks, the system can update indexes of information such as business hours and services offered. The system can be automatically triggered to update data in response to the detection of missing data, data that has changed over time, or inconsistent data. Generally, to obtain such information, users may need to check each business or data source individually.

[0027] The system offers numerous advantages, including reducing the amount of human input required to complete specific tasks, such as initiating phone calls. For example, the system can automatically initiate a phone call based on a determination that certain trigger criteria have been met, such as a mismatch between the services offered by a salon and those listed on a third-party booking website. The system can reduce friction in transaction queries by, for example, detecting human frustration or discomfort at one end of a phone call and terminating the call or altering how the conversation is conducted. The system can connect users in developing countries with services such as transportation or education. The system can also connect users with low-tech industries that lack a website or digital presence. Furthermore, the system is scalable to different applications, even compared to the largest aggregators.

[0028] Figure 1A shows an exemplary block diagram of a system for a call initiation system that initiates a call and facilitates a conversation between the call initiation system bot and a human 104 during the call. Each component shown in Figure 1A is described in detail below.

[0029] System 100 includes various components and subsystems that work together to enable a bot to communicate effectively with a human 104. System 100 may include a communication framework or platform 102, a dialer 106, a sound system 108, a call triggering module or trigger module 110, an audio package 112, a session recorder 114, session storage 116, a text-to-speech module 118, a speech endpoint detector 120, stored text-to-speech results or recordings 122, an intent-to-text module 124, a speech-to-text module 126, a speech application programming interface (API) 128, a text-to-intent module 130, a flow manager 132, an operator controller 134, and a relief module 136. In some implementations, the system includes all modules. In other implementations, the system includes combinations of these modules. For example, in one implementation, the text-to-intent layer is not required, and intents are provided directly to the speech synthesis module.

[0030] Figure 1B shows an exemplary block diagram of an alternative system for a call initiation system that initiates a call and facilitates a conversation between a call initiation system bot and a human during the call. In this example, the communication platform 102 is replaced by a client entry point for user requests and a telephone signaling server for other requests, namely inbound calls from businesses. The system sends both types of requests to a telephone server (196) that makes a call with a bot service (195) making a call from the other end. In some implementations, the bot service (195) includes a dialogue model (198) and a language model (199) to enable the bot service to engage in human-like telephone conversations. The telephone server (196) may include a TTS model (197). A speech recognition device (191) and / or an audio mixer (194) may provide information for the telephone server (196) to understand and respond to the human on the other end of the telephone call (190). The operator (134) monitors calls using the task user interface (160) and the curation user interface (170). The operator (134) can review recorded calls from the recording studio and the evaluation TTS (114). The call player (162) plays back and returns the calls to the operator (134). The operator can schedule calls using the local agent (175) to initiate telephone calls through the telephone server (196).

[0031] In the implementation shown in Figure 1A, the communication platform 102 enables the bot to contact external actors by performing tasks such as making calls, receiving inbound calls from businesses or users (104, 144), or contacting target businesses. The communication platform 102 also enables the bot to receive requests from users to make calls on their behalf.

[0032] In some implementations, users request calls to other users or businesses through interaction with the user interface or through speech requests. These user requests may be for assistant-type tasks such as making reservations, booking restaurants, finding someone to walk the dog, or finding stores that carry the goods the user wants to buy.

[0033] Figure 1C shows an exemplary user interface in which the user may enter more details about their request. The user may initiate a request by clicking a “Book Now” button or by interacting with the user interface in other ways. For example, if the user wants to book a haircut, the user may interact with a website associated with the salon where the user wants to get a haircut. Alternatively, the user may interact with a search results list that includes salons as results in the search results, or with a user interface that shows salons on a map. Either of these interfaces may allow the user to request a call. The user may enter details of their request, such as the professional stylist they want to meet, the category of service they want, and the date and time of the service. As shown in Figure 1B, the user may click a “Continue Booking” button or take some other action to indicate that a request for a hair salon appointment has been made.

[0034] Figure 1D shows an example of a user speaking to a bot to make a request. After the user speaks to the bot, the bot may confirm the request. The bot may also request additional information about the user's request. For example, if a user speaks to the bot to request a haircut, the bot may ask for information about where the user wants to get the haircut, the day the haircut should be scheduled, and what kind of haircut service the user wants to schedule.

[0035] A user may make a task request to the system when the task cannot be performed. For example, a user may request a call to schedule a haircut appointment at 11 p.m., when all hair salons have closed. Therefore, the system may store the request in task information storage so that it starts and completes at a later time, such as during the salon's business hours, which would normally be determined or obtained by system 100.

[0036] In some implementations, the system provides the user with initial feedback that there will be a delay in processing the request. For example, if a user makes a request to schedule a haircut appointment at 11 p.m. when the salon is closed, the system provides the user with a visual, audio, or some other indication that there will be a delay in completing the task from when the salon opens until the system reaches the salon, since the salon is closed.

[0037] In some implementations, the task information storage 150 stores information about each task, such as the name of the user requesting the task, one or more people or places to call, the type of task requested, how the task request was made, details about the specific type of task, details of the activities performed to complete the task, the task start date, the task completion date, the time of the last status update to the requesting user, the operator who double-checked the call task, the user who requested the task completion date, and the task's current status.

[0038] In some implementations, the task manager module 160 determines when to schedule calls to people or businesses. The task manager module 160 monitors tasks from the task information storage 150 and determines the appropriate time to schedule incoming tasks. Some tasks are scheduled immediately, while others are scheduled after a specific triggering event occurs.

[0039] In many situations, the other end of a call initiated by system 100 involves a human, such as human 104. Human 104 may be a representative of the organization the bot is trying to contact. In some examples, a communication platform is used to make calls to businesses. System 100 can be integrated with a communication platform. For example, system 100 can programmatically operate a web browser and use a framework for testing web applications to use web-based video conferencing services. System 100 can create and use several communication platform accounts. In some examples, system 100 can automatically switch between different communication platform accounts to avoid call speed throttling.

[0040] The dialer 106 facilitates the initiation or initiation of calls by the bot. The dialer 106 is communicatively connected to the communication platform 102. The dialer 106 provides commands to the communication platform 102 to initiate a telephone call to a specific recipient selected by the dialer 106. For example, the dialer 106 can play an audio tone corresponding to the digits of a telephone number. Once the call is made, the system 100 can converse with the human recipient at the other end of the line.

[0041] The dialer 106 can receive commands to initiate a call to a specific callee. For example, the dialer 106 can receive data containing commands from other modules in the system 100, such as the trigger module 110 or the flow manager 132.

[0042] The trigger module 110 detects a trigger event, or a specific event that indicates that system 100 should initiate a call to a particular callee. A trigger event can be of a predetermined type. For example, a user of system 100 can specify a particular type of trigger event. A trigger event can include an explicit action performed by a user of system 100, a detected pattern of data provided to the trigger module 110, a predetermined time period that has elapsed since a particular event occurred, and various other types of events. In response to the detection of a trigger event, the trigger module 110 provides instructions to the dialer 106 to initiate a call to a particular callee, or to the flow manager 132 to select a node in a particular workflow, or to provide instructions to the dialer 106.

[0043] The sound system 108 is used to record and play back audio. In some examples, three virtual streams are set up: (a) incoming audio from a telephone or video conferencing service to system 100, (b) output audio returning from system 100 to the communication platform, and (c) a mixed stream combining a and b, which are used to record the entire call. The sound system 108 uses the audio package 112 to perform communication through the communication platform 102.

[0044] The audio package 112 is used to communicate with the sound system 108. In some examples, the system 100 includes an audio module that encloses the audio package 112 and processes a continuous stream of incoming audio packets. The module also records all incoming packets and enables playback of pre-recorded audio files. The system 100 uses various bit depths, sampling frequencies, packet sizes, etc.

[0045] System 100 can record incoming and outgoing conversations made by the bot. Audio package 112 can enable System 100 to record a specific session or call using session recorder 114. In some examples, session recorder 114 can record a portion of a conversation made by the bot by recording the bot's speech as it is generated. In other examples, session recorder 114 can record a portion of a conversation made by the bot by externally recording the bot's speech as it is output to human 104 by communication platform 102. Session recorder 114 can also record human 104's responses.

[0046] The session recorder 114 stores the recorded session data in the session storage 116. The recorded session data may be stored as audio data or as feature data representing the audio data. For example, the recorded session data may be stored as a vector that stores the values ​​of specific features of the session's audio data. The session storage 116 may be a local database, a remote server, physical memory within the system 100, or various other types of memory.

[0047] The speech endpoint detector 120 simplifies the conversation between the bot and the human on the other end of the line. To simplify the conversation, the conversation is divided into individual sentences, which are switched separately between the human and the bot. The speech endpoint detector 120 receives a continuous input audio stream from the audio package 112 and converts it into separate sentences.

[0048] The speech endpoint detector 120 detects the end of a speech. In one implementation, the speech endpoint detector 120 operates in two states: speech standby and silence standby. The speech endpoint detector 120 alternates between these states as follows: Each audio packet is examined by comparing its mean squared deviation (RMSD) to a predefined threshold. If the RMSD is below this threshold, a single packet is considered "silent". When a non-silent packet is received, the module switches from the "speech standby" state to the "silent standby" state. Depending on the overall system state, the module switches back only after a period of consecutive silent packets lasting a predefined duration has been received.

[0049] In some implementations, during the "sound wait" period, the speech endpoint detector 120 creates pure silence packets (one packet for every 10 actual packets) and sends them to the speech-text module 126. The created packets can avoid disconnection from the speech API 128. During the "silence wait" period, the speech endpoint detector 120 sends silence packets from the stream for a predefined time period (useful for baseline noise estimation), and then sends all audio packets.

[0050] In other implementations, the speech endpoint detector 120 uses machine learning, a neural network, or some form of deep learning trained to observe intonation and language context to find endpoints.

[0051] In some examples, the speech endpoint detector 120 considers what is said, the speaker's intonation, etc., when deciding how to parse a particular stream of audio input. For example, the speech endpoint detector 120 may determine that a particular human being addressed 104 tends to end sentences with low intonation, and the speech endpoint detector 120 may predict the end of a sentence spoken by the addressed person 104 when a decrease in intonation is detected. The speech endpoint detector 120 can dynamically adjust its threshold between calls based on the signal-to-noise ratio within a time frame.

[0052] The speech-text module 126 converts the audio data parsed by the speech endpoint detector 120 into text that can be parsed for an intent used to select the bot's next response. The output of the speech-text module 126 is an ordered list of speech options, and in some cases, a confidence score for the best option is provided. The speech recognition process includes two main components: an acoustic module and a language module. For the acoustic module, the system can use a model trained from recordings of people speaking directly to their phones. A neural network may be used by the model, and in some examples, the first layer of the neural network may be retrained to account for a vocoder present in a phone call. A vocoder is a speech codec that produces sound from the parsing of speech input. The neural network may also be retrained to account for different background noises between business calls and personal phone calls. The language module may be built using a system that biases the language module based on the system's past experience. In some examples, the bias may be configured automatically. In some examples, this bias is set manually. In some examples, the language bias configuration is changed between verticals.

[0053] The speech-text module 126 uses the context of a call to bias the language module based on what the other person on the other end of the conversation is likely to say. For example, the system's bot might ask, "Are you open on Tuesdays?" Based on this question, the other person on the other end of the conversation is likely to respond with an answer such as, "No, we're closed," or "Yes, we're open." The bot learns possible responses based on past calls and uses predictions to understand the incoming audio. While the bot can predict full sentence responses, it can also predict phrases. For example, after the bot says, "There are seven of us in total," it might predict the phrase, "Did you say seven?" The bot might also predict the phrase, "Did you say eleven?" because seven and eleven sound similar. The bot might also predict that responses such as, "Did you say two?" are less likely. The bot can then assign probability weights to each phrase based on its predictions.

[0054] In some implementations, the speech-to-text module 126 uses the speech API 128 to convert audio data to text. In some examples, the speech API 128 uses machine learning to convert audio data to text. For example, the speech API 128 can use a model that accepts audio data as input. The speech API 128 can use any of a variety of models, such as decision trees, linear regression models, logistic regression models, neural networks, classifiers, support vector machines, inductive logic programming, ensembles of models (e.g., using techniques such as bagging, boosting, and random forests), genetic algorithms, and Bayesian networks, and can be trained using a variety of techniques such as deep learning, perceptrons, association rules, inductive logic, clustering, maximum entropy classification, and learning classification. In some examples, the speech API 128 can use supervised learning. In some examples, the speech API 128 uses unsupervised learning. In some examples, the speech API 128 can be accessed by the speech-to-text module 126 over a network. For example, the speech API 128 can be provided by a remote third party on a cloud server.

[0055] To address dialogue synchronization, such as determining the context in which a person was speaking to determine a natural opportunity for a bot response, system 100 can identify intents. An intent is a formal linguistic representation of a single semantic meaning in a sentence, uttered by either a human or a bot. In some implementations, system 100 ignores any intents received from a human between the last intent received and the bot's response, in order for the bot to generate a response related to the most recent sentence spoken by the human. However, system 100 can use previous intents to signal future responses. For example, the system can mark intents received before the most recent intent received as ANCIENT, parse the ANCIENT intent, and store it for offline evaluation. In some examples, various other forms of processing logic may be used.

[0056] While much of System 100 is use-independent, some parts of System 100 are manually configured or fully programmed for specific use cases, i.e., system industries. An industry essentially consists of an intent schema and business logic code. An intent within a schema is an internal formal linguistic representation of a single semantic meaning in a sentence, uttered by either a human or a bot. For example, in the business hours extraction industry, there is a bot intent, "Are you open {date:tomorrow}", and a corresponding human intent, "We are closed {date range:September}". The process by which incoming audio from a human is converted into an intent is referred to herein as intent resolution. The reverse process (converting a bot intent into speech) is called intent-speech. While schemas are configured per industry, most of the code that learns and classifies intents is general and used across industries. In some examples, only the language-specific parts of the system exist in the intent resolution and intent-speech configuration.

[0057] Logic code can be programmed industry-specific (sharing some common code) and determines the bot's behavior for all possible situations defined by the call context (input parameters and what has happened up to that point), as well as the incoming human intent. In one implementation, speech is converted to text and then interpreted as human intent. The human intent is used to determine the robot's intent. In some industries, the bot leads the conversation, while in others, the bot mostly responds to the human. For example, in data-gathering industries, the bot aims to extract certain information from a business. Typically, the bot will try to ask a series of questions until all the desired information is obtained. For example, in transaction-type industries where a bot targets, such as making a reservation, it answers questions primarily posed by the human ("What is your name?"... "And what is your phone number?"...). In such cases, the system takes the lead only if the human suddenly becomes silent, for example. The programmer can design the flow between human intent and robot intent so that the conversion makes logical sense. In some implementations, protocols exist for flows that can be controlled by non-engineers to modify or update the conversion from human intent to robot intent. These flows can also be learned automatically using machine learning.

[0058] In another implementation, the resolution of the input human intent may be a hidden layer, and machine learning may be used to learn the output robot intent directly from the input text. Human speech input may be converted to text, and then the robot intent may be determined directly from this text. In yet another implementation, the system can output intent directly from human speech. Both of these designs use machine learning to learn the context and the robot intent corresponding to each input.

[0059] The text-intent module 130 is constructed using a schema of possible incoming intents, example sentences for each such intent, and a language bias configuration. Essentially, the text-intent module 130 plays the role of "snapping" incoming sentences into a predefined list of (or "unknown") intents, while taking into account unfamiliar phrases and errors in the speech recognition process. For example, in some implementation forms, the text-intent module 130 can identify that the sentence (received from the speech recognition module) "We are open from 11am to 9am tomorrow, sorry, until 9:30am" is similar to the known example "We are open from +(hour, from) and close at +(hour, until)am, sorry," which is an example of the intent "We are open from 11am, until 9:30am." Fields like "from" and "until" are intent parameters.

[0060] The text-intent module 130 may consist of two main parts: (1) an annotator and (2) an annotated text-intent classifier. In some implementations, the system has a post-classification stage that classifies the argument. For example, in the phrase "Monday to Tuesday, sorry, Wednesday, we are closed," the annotator rewrites the text to "<Date:Monday> to <Date:Tuesday>, sorry, <Date:Wednesday> is closed." This example shows that the phrase has been rewritten by the annotator, which specifies annotations in the response text. The annotated text-intent classifier changes the annotated phrase to We are open {Sun1:Monday, Sun2:Tuesday, Sun3:Wednesday}. The post-classification stage rewrites the phrase to: We are open {Sun:Monday to, Sun:Wednesday, Wrong day;Tuesday}.

[0061] As soon as system 100 receives speech options, it uses the text-intent module 130 to annotate each of them for date, time, common name, etc. This is done for two purposes: (1) to extract intent parameters from the logical module (e.g., "Time: 10am"), and (2) to generalize the text to simply find matches with previously encountered sentences. The text-intent module 130 receives output from speech-text module 126 and annotates the list of speech options. The text-intent module 130 then uses the annotations to map the most likely options to intents used by flow manager 132 to select the next action within a particular workflow.

[0062] To reduce computation time between calls, system 100 can pre-build a library of known texts that should be annotated (for the current date and time). For example, on Tuesday, September 1, 2015, possible annotations for "Date: (2015, 9, 2)" might include "tomorrow," "this Wednesday," and "September 2nd." In real time, system 100 iterates through the words in the input sentence and searches for the longest match among the annotation candidates (after some normalization). System 100 then replaces the text with the annotations and returns an edited string where all candidates from left to right are replaced with the annotations. For example, "We open at 7am" would be replaced with "We open at @(time, 7am)."

[0063] The annotation or text-intent module 130 can also play a role in contraction. For example, system 100 might encounter a sentence like, "Um... yes, 4 o'clock, 4 pm." The text-intent module 130 can replace "4 o'clock, 4 pm" with a single annotation, "@(time, 4 pm)." Furthermore, the text-intent module 130 can contract smaller time corrections, such as from "We close at 10 o'clock, um, 10:30 pm" to "We close at @(time, 10:30 pm)."

[0064] In other implementations, the system may use other methods for annotating text, such as machine learning algorithms that can learn how to annotate text based on curated data, prefix trees that can be used to annotate text, or rule-based patterns that can be specifically derived for annotation.

[0065] The text-intent module 130 parses and annotates most new sentences of a call, and speech recognition often distorts many of the spoken words. System 100 has thousands of intents stored for each use case and classifies each sentence as having an intent from the stored intents determined to be most relevant to the sentence based on the sentence's parameters. For example, System 100 may classify a particular sentence as having an intent to ask for a name based on detecting words in that sentence that suggest a question asking for the caller's name. In some implementations, System 100 may not recognize the intent of a sentence and classify that sentence as having an unknown intent.

[0066] The text-intent module 130 uses machine learning algorithms to handle classification. For example, the system may use a combination of the conditional random field module and the logistic regression module. In one implementation, classification is performed at the sentence level, i.e., a string of text is converted into a set or list of intents. In another implementation, all tokens in the original string are classified into intents, and intent boundaries are also classified. For example, the sentence "We open at 7am on Mondays, um... we open at 8am on Tuesdays" is classified in the first implementation as containing the intents GiveDailyHours+AskToWait. In the second implementation, the substring "We open at 7am on Mondays" is classified as the boundary of the GiveDailyHours intent, the substring "um..." is classified as another intent of type AskToWait, and the substring "We open at 8am on Tuesdays" is classified as another intent of type GiveDailyHours.

[0067] In some implementations, the text-intent module 130 may not use machine learning algorithms, but instead uses a set of examples for each intent, and then uses a one-nearest neighbor (pattern recognition algorithm) between each speech option and all examples, where the algorithm's distance metric is the change in normalized edit distance of words in a sentence (how two different strings, such as words, fit together). The distance between two individual words is more complex and aims to be an approximation of the phonological distance. In some examples, the semantic distance may be determined by the annotated text-intent module 130.

[0068] In practice, the text-intent module 130 can also use cross-speech option signals (for example, a number present in only one of the speech options is likely to be a bad interpretation). In some cases, the text-intent module 130 pre-biases the results of the annotated text-intent module based on the system context. Finally, the text-intent module 130 has some customized extracts for ambiguously defined intents such as "ComplexOpeningHours," where the system can identify that a complex phrase of opening hours has been given, but the system could not extract the parameters precisely (for example, "...dinner is served until 9 pm, dessert can be ordered for another hour longer, the bar is open until 2 pm, but we do not accept customers after 1 pm").

[0069] In some cases, the examples used for classification are automatically inferred based on curated past calls, and can also be edited manually. The generalization process can replace text with annotations and omit questionable curations.

[0070] In some examples, humans speak not a single intent, but rather a series of intents. For example, "Would you like a haircut? What time?" The exemplary system supports any number of intents within a given sentence. In some examples, the Annotated Text-Intent module specifically determines positive and negative intents as prefixes to other intents (e.g., "No, we are closed today" => negation + WeAreClosed). In some examples, the Annotated Text-Intent module supports any chain of intents.

[0071] The system includes multiple modules that perform different functions of organizational logic, including a flow manager 132 which includes a common sense module 133 and a relief module 136.

[0072] The flow manager 132 can include industry-specific custom code that tracks each call and determines how to respond to each intent (or long silence) received from a human. However, in other implementation forms, the flow manager is common across industries. The response is a list of synthetic intents to convey to the human 104 (the bot can also choose to remain silent) and may be a command to terminate the call. The flow manager 132 also plays a role in generating the call outcome, which includes any information collected during the call. In some examples, system 100 learns how to respond to each input based on live calls initially created by a human and later created by "child" bots. System 100 maintains as much flexibility in its logic as possible to resolve any misunderstandings during the call.

[0073] System 100 has multiple flows, each tailored to a specific type of task, such as determining business hours for a business or making reservations for salon appointments. System 100 maintains a common library shared between different flows and can extract subflows from the history of outgoing calls, allowing the system to jumpstart new business types on a task-by-task basis. In some examples, System 100 may automatically learn the flows for different tasks based on manually outgoing calls.

[0074] Humans may skip some important details when speaking without confusing their conversation partner. For example, a human might say, "We are open from 10 am to 4 pm." A bot needs to understand whether the business opens at 10 am or 10 pm, and similarly, whether it closes at 4 pm or 4 pm. For example, if the business is a nightclub, the bot might be expected to assume the hours are from 10 pm to 4 pm, and if the business is a restaurant, the bot might be expected to assume the hours are from 10 am to 4 pm, and so on.

[0075] The flow manager 132 includes a common sense module 133 that obscures intents in the received speech input. In some examples, the flow manager 132 includes several types of common sense modules, such as a module that learns from statistics on several datasets (e.g., a baseline local database) and a module that is manually programmed. The first type of module takes an optional dataset (e.g., opening hours) and calculates a p-value for each option and sub-option (e.g., "2 a.m. to 4 a.m." or "just 2 a.m."). The second type of module uses a set of predefined rules to ensure that the system does not make any "common sense" errors that may exist in the dataset. Whenever there are multiple ways to interpret some variables, the flow manager 132 can combine two scores to determine the most likely option. In some examples, the flow manager 132 concludes that no option is sufficiently likely, and the system 100 relies on a human to explicitly ask for clarification of their meaning.

[0076] The common sense module 133 can use data from similar callers to select the most likely option. For example, if most bars in Philadelphia are open from 8 p.m. to 2 a.m., the common sense module 133 may determine that the most likely option for the ambiguous phrase "We're open from 10 p.m. to 2 a.m." is that the speaker means from 10 p.m. to 2 a.m. In some examples, the common sense module 133 may indicate to the flow manager 132 that further clarification is needed. For example, if most post offices in Jackson, Michigan have business hours from 10 a.m. to 5 p.m., and the system 100 believes that the caller responded that business hours are "from 2 p.m. to 6 p.m.," which is a different threshold amount than typical post offices, the common sense module 133 may instruct the flow manager 132 to ask for clarification.

[0077] In some cases, tension builds up between calls, typically due to high background noise, exceptional scenarios, strong accents, or bugs in the code. Tension can also arise from unexpected intents. For example, when calling a restaurant, the system might encounter unexpected phrases such as, "So, would you like to give a presentation?" or "Just so you know, we don't have a TV showing the Super Bowl." The system needs to process intents it hasn't encountered before. To identify a problematic situation for either party, the bot attempts to quantify the amount of stress exhibited during the call. Relief module 136 can mimic an operator supervising the call and choose when to implement manual intervention.

[0078] The operator controller 134 is communicatively connected to the flow manager 132, allowing a human operator to provide commands directly to the flow manager 132. In some examples, once a call is forwarded to a human operator for processing, the operator controller 134 places the flow manager 132 into a holding pattern or pauses or shuts down the flow manager 132.

[0079] When the flow manager 132 selects the next node in a particular workflow based on the intent determined from the text-intent module 130, the flow manager 132 provides instructions to the intent-text module 124. The instructions provided by the flow manager 132 include the next intent to be transmitted to the caller through the communication platform 102. The intent-text module 124 also generates a markup queue for speech synthesis, defining, for example, several different emphasis or prosody of a word. The intent-text module 124 can use manually defined rules or reinforcement learning to generate new text from the intent.

[0080] The output of the intent-text module 124 is text that is converted into audio data for output on the communication platform 102. The text is converted into audio by the text-speech module 118, which uses previously stored text-speech outputs and read values ​​122. The text-speech module 118 can select a previously stored output from the stored outputs / read values ​​122. In some implementations, the system uses a text-speech synthesizer between calls. For example, if a common response selected by the flow manager 132 for the bot to provide is "Great, thanks for the help!", the text-speech module 118 can select a previously generated text-speech output without generating an output at runtime. In some examples, the text-speech module 118 uses a third-party API accessed over the network connection, as well as the speech API 128.

[0081] As described above, in one example, a user may initiate a task in system 100 by interacting with search results provided to the user (e.g., a web search). For example, the user might search for "reserve a table for two at a Michelin-starred restaurant tonight." The task manager module 140 may receive the task and store the task information in the task information storage 150. The task manager module 140 may then decide when to schedule the task and set a triggering event. For example, if the user requests to reserve a table at a Michelin-starred restaurant before it opens, the task manager module 140 may determine when the restaurant opens and set a triggering event for that time. If the task manager module 140 knows there will be a delay in processing for the triggering event, it may warn the user of the delay by providing a visual, audio, or some other indication. In some implementations, the task manager module 140 may also provide the time it will take to complete the task, the time the task is scheduled to start, and further information about why the task is delayed.

[0082] The trigger module 110 can detect when a specific trigger event occurs (in this example, the opening time of a restaurant) and instruct the dialer 106 to place a call. In some examples, the system 100 can present the user with options for selecting a restaurant to call. In other examples, the system 100 can automatically place a call to a specific restaurant selected based on a set of characteristics. The user can define default settings for placing calls for a particular task. For example, the user can specify that the system 100 should select the restaurant closest to the user's current location to place a call, or that the system 100 should select the highest-rated restaurant to place a call.

[0083] In some examples, system 100 includes, forms part of, or is configured to communicate with, a communication application such as a messaging or chat application that includes a user interface in which a user provides the system with a request for assistance with a task. For example, a user might use the messaging application to text a number in their request, such as "Does Wire City have 20 AWG red wire in stock?". The system may receive the text message, parse the request to determine that a trigger event has occurred, and initiate a call to take appropriate action. For example, the system might make a call to the nearest Wire City to find out if they currently have 20 gauge red wire in stock.

[0084] Similarly, in some examples, system 100 includes, forms part of, or is configured to communicate with a virtual assistant system, which itself is a collection of software agents for assisting the user with various services or tasks. For example, the user might type (by voice or text input) "Is my dry cleaning ready?" into the virtual assistant. The virtual assistant might process this input and determine that communication with the business is necessary to satisfy the query, and accordingly, identify the intent, make a call, and communicate with the system to perform the appropriate workflow.

[0085] In a specific example, system 100 autonomously performs tasks through multiple dialogues with multiple people, collecting and analyzing the individual or cumulative results of the dialogues, taking action on them, and / or presenting them. For example, if system 100 is assigned the task of collecting data on when the busiest times are for several restaurants in a designated area, system 100 can autonomously make calls to each restaurant and ask how many customers are seated during a certain period of time in order to analyze the data and provide results.

[0086] Figure 2A shows an exemplary process 200 for completing a task assigned by a user. Briefly, process 200 may include the steps of: mapping a conversation to an initial node in a predefined set of workflows, each linked by an intent (202); selecting an outgoing message based on the current node in the workflow (204); receiving a response from a human user (206); mapping the response to an intent in the predefined workflow (208); selecting the next node as the current node in the workflow based on the intent (210); and repeating steps 204-210 until an end node of the set of linked nodes in the predefined workflow is reached. Process 200 may be performed by a call-blocking system such as system 100.

[0087] Process 200 may include a step (202) of mapping the conversation to an initial node in a predefined set of workflows, each linked by an intent. For example, the flow manager 132 described above with respect to Figure 1 can map the conversation to an initial node in a predefined set of workflows, each linked by an intent. In some examples, the conversation between system 100 and a human caller may be initiated by a user. In some examples, the conversation includes intents that map to nodes in a predefined set of workflows. For example, system 100 may store a set of predefined workflows that have actions to be performed. In some examples, the system may select a predefined workflow based on an identified intent. Each of the workflows may be linked by an intent. In some examples, system 100 may make a telephone call to a business specified by the user during the conversation. In some examples, the business may be a restaurant, a salon, a clinic, etc. In some examples, the system may consider the call to have been successfully made only if a human answers, and if no one answers, or if the system is directed to the telephone tree and does not successfully navigate the telephone tree, the system may determine that the call was not successfully made.

[0088] Process 200 may include a step (204) of selecting an outgoing message based on the current node in the workflow. For example, the flow manager 132 may select a message saying "Hi, I would like to schedule a haircut appointment" if the current node in the workflow indicates that the user wants to schedule such an appointment.

[0089] Process 200 may include a step (206) of receiving a response from a human user. For example, system 100 may receive a response from the human caller at the other end of the telephone call, such as, "Understood, what date and time would you like to schedule this appointment?" In some examples, system 100 may record the response (for example, using a session recorder 114). In some examples, system 100 may play back the response to a human operator. In some examples, a human operator may be monitoring the call (for example, using an operator controller 134).

[0090] Process 200 may include a step (208) of mapping responses to intents in a predefined workflow. The flow manager 132 can map responses to intents in a predefined workflow. In some examples, the system compares an identified intent with an intent to which a set of predefined workflows are each linked.

[0091] Process 200 may include a step (210) of selecting the next node as the current node in the workflow based on the intent. For example, flow manager 132 may use the intent to determine the next node in the workflow. Flow manager 132 may then specify the next node as the current node. Process 200 may include repeating steps 204-210 until an end node is reached. Thus, the specified current node is used in each repeated cycle of steps 204-210 to determine the next outgoing message until an end node is reached.

[0092] Figure 2B shows an exemplary process 250 for completing a task assigned by a user. Briefly, process 250 may include the steps of receiving a task associated with an intent from the user (252), identifying the intent (254), selecting a predefined workflow based on the intent from a set of predefined workflows linked by the intent (256), following the predefined workflow (258), and completing the task (260). Process 250 may be executed by a call initiation system such as system 100.

[0093] Process 250 may include a step (252) of receiving a task associated with an intent from a user. For example, a user might submit a search query to system 100 through the user interface, such as "make a haircut appointment." In some examples, the search query may be received by a trigger module 110, which detects that the query is a trigger event indicating that a call should be made to a specific callee. The task might be to make an appointment, and the intent might be to get a haircut. In some examples, the task or intent may not be explicitly entered. In some examples, a user may submit a task and intent without entering a search query. The task associated with the intent may be received by the system to assist with the task.

[0094] Process 250 may include a step (254) of identifying an intent. For example, system 100 may process an incoming task associated with an intent and identify the intent. In some examples, the intent may be explicitly entered and separated from the task. In some examples, the intent may be a characteristic of the task. In some examples, the input is provided as a speech input, the speech endpoint detector 120 provides the parsed output to the speech-text module 126, the speech-text module 126 sends the text to the text-intent module 130, and the text-intent module 130 identifies the intent.

[0095] Process 250 may include the step (256) of selecting a predefined workflow based on an intent from a set of predefined workflows linked by the intent. For example, system 100 may store a set of predefined workflows that have actions to be performed. In some examples, the system may select a predefined workflow based on an identified intent (254). For example, flow manager 132 may select a predefined workflow based on an intent (254) identified by text-intent module 130. In some examples, the system compares the identified intent with the intents to which each of the predefined workflow sets is linked.

[0096] Process 250 may include steps (258) that follow a predefined workflow. For example, system 100 may include modules that follow instructions included in a predefined workflow. In some examples, a bot in system 100 can follow instructions included in a predefined workflow. For example, an instruction may include instructing trigger module 110 to provide control data to dialer 106 in order to make a call and have a conversation with a human representative of the business.

[0097] Process 250 may include a step (260) to complete a task. For example, system 100 may complete an entire assigned task, such as paying a bill or changing a dinner reservation. In some examples, system 100 may complete parts of a task, such as making a call or navigating a phone tree, until it reaches a human. In some examples, system 100 may complete parts of a task specified by a user. For example, a user may specify that the system complete all tasks and transfer the call to the user for verification.

[0098] Many use cases may involve users who want to purchase something from a business but have difficulty making the purchase due to the complexity of the transaction, menu navigation, language issues, reference knowledge, etc. Transaction queries can accumulate support from vendor-side personnel who hope the system will help them successfully complete the transaction. In some examples, the system provides crucial support to developing countries and low-tech and service industries such as plumbing and roofing. Workflows can be used not only to help human users successfully navigate such transactions but also to encourage vendor-side systems to assist the user. The system is scalable to accommodate a variety of use cases. For example, a restaurant reservation application may partner with thousands of businesses worldwide, and the system disclosed herein can be configured to issue restaurant reservations at the required scale.

[0099] Figure 3 shows an exemplary workflow 300 of the process performed by the system. In this particular example, a simple Boolean question is asked by a bot in system 100. It is understood that the system can respond to more complex and higher-level problems, and workflow 300 is presented for the sake of simplicity in explanation.

[0100] Flow 300 shows an exemplary question posed by the bot: "Are you open tomorrow?" Possible responses provided by a human are laid out, and the bot's response to each of the human responses is provided. Depending on the human response, there are several stages in Flow 300 that System 100 may be led through. The stages indicated in double frames are the final stages in which System 100 exits Flow 300. For example, in response to the binary question posed by the bot, the human caller can confirm that the business is open tomorrow and exit Flow 300. The human caller confirms that the business is not open tomorrow and exits Flow 300. The human caller asks the bot to hold, sending the bot to a separate hold flow, and exits Flow 300.

[0101] To facilitate user access and promote the dissemination of system 100, system 100 is integrated with existing applications, programs, and services. For example, system 100 may be integrated with an existing search engine or application on a user's mobile device. Integration with other services or industries allows users to easily submit requests to have tasks completed. For example, system 100 may be integrated with a search engine knowledge graph.

[0102] In some use cases, real-time human judgment can be automated. For example, system 100 can automatically detect that a user is running 10 minutes late for a barber shop appointment and alert the barber shop before the user arrives.

[0103] System 100 can select specific parameters for the bot based on the context of the conversation taking place or data about a particular person being called stored in a knowledge database. For example, based on the person being called's accent, position, and other contextual data, System 100 can determine that the person being called is more comfortable in a different language than the one currently being spoken. System 100 can then switch to the language that the bot deems more comfortable for the person being called and ask the person if they prefer to speak in the new language. By reflecting the specific speech characteristics of a human person being called, System 100 increases the likelihood of a successful call. To reduce tension accumulated between calls, System 100 reduces potential sources of friction in the conversation due to speech characteristics. These characteristics may include the average length of words used, the complexity of sentence structure, the length of pauses between phrases, the language the person being called is most comfortable speaking, and various other speech characteristics.

[0104] Figure 4 is a block diagram 400 of the call triggering module of system 100. The trigger module 110 is communicatively connected to the dialer 106 and, based on detecting trigger events, provides instructions to the dialer 106 to initiate a call to a specific callee or a set of callees. In some examples, the trigger module 110 can communicate with the flow manager 132 to provide trigger event data that the flow manager 132 uses to select a node in a particular workflow or to provide instructions to the dialer 106.

[0105] The trigger module 110 receives input from various modules, including a mismatch detector 402, a third-party API 404, a trend detector 406, and an event identifier 408. The trigger module 110 can also receive input from the flow manager 132. In some examples, each of modules 402-408 is integrated with system 100. In other examples, one or more of modules 402-408 are separate from system 100 and connected to the trigger module 110 via a network such as a local area network (LAN), wide area network (WAN), the internet, or a combination thereof. The network can connect one or more of modules 402-408 to the trigger module and facilitate communication between components of system 100 (for example, between the speech API 128 and the speech-text module 126).

[0106] The mismatch detector 402 receives data from multiple different sources and detects inconsistencies between data values ​​from a first data source and corresponding data values ​​from a second source. For example, the mismatch detector 402 can receive data indicating the clinic's operating hours and detect if the clinic's operating hours listed on the clinic's website differ from those listed outside the clinic. The mismatch detector 402 can provide the trigger module 110 with data indicating the cause of the conflict, the type of data value inconsistent, the conflicting data values, and various other characteristics. In some examples, the mismatch detector 402 provides the trigger module 110 with a command to initiate a call to a specific callee. In other examples, the trigger module 110 determines, based on the data received from the mismatch detector 402, which callee to contact and which fields of data should be collected from that callee.

[0107] The trigger module 110 can detect trigger events based on data provided by the mismatch detector 402. A trigger event may include receiving user input indicating a discrepancy. For example, the trigger module 110 may receive user input through a user interface 410. The user interface 410 may be an interface for a separate application or program. For example, the user interface 410 may be a graphical user interface for a search engine application or a navigation application.

[0108] In some implementations, the user interface 410 can prompt the user to provide information. For example, if the system detects that the user is in the store after the store's advertised closing time, it can ask the user if the store is still open or ask the user to enter the time. The user can enter the requested data through the user interface 410, and the mismatch detector 402 can determine whether there is a discrepancy between the data entered through the user interface 410 and the corresponding data from a second source, such as a knowledge base 412. The knowledge base 412 can be a storage medium such as a remote storage device, a local server, or various other types of storage media. The mismatch detector 402 can determine whether the user is in the store for a predetermined amount of time outside of normal business hours (for example, more than 20 minutes, as stores may stay open for a few extra minutes, especially for late customers).

[0109] In another exemplary scenario, the mismatch detector 402 may determine that information on an organization's website is outdated. For example, based on data from the knowledge database 412, the mismatch detector 402 may detect that a bass fishing club's website indicates that its monthly meeting takes place on the first Wednesday of each month, while all of the club's more active social media profiles indicate that the monthly meeting takes place on the second Tuesday of each month. The mismatch detector 402 can then output data indicating this detected mismatch to the trigger module 110.

[0110] A trigger event may include determining that a particular set of data has not been updated for a predetermined amount of time. For example, a user of system 100 may specify an amount of time during which the data should be refreshed, regardless of any other trigger events that occur. A mismatch detector may compare the last updated timestamp for a particular data value and, based on the timestamp, determine whether a predetermined amount of time has elapsed. The characteristics of a particular data field, including the timestamp and the data value itself, may be stored in the knowledge database 412. A timer 414 may provide data to the knowledge database 412 to update the amount of time that has elapsed. Based on the timing data provided by timer 414, a mismatch detector 402 may determine that a predetermined time period has elapsed.

[0111] For example, the mismatch detector 402 can determine, based on data from the knowledge database 412, that the opening hours of a small coffee shop in Ithaca, New York, have not been updated for three months. The mismatch detector 402 can then provide the trigger module 110 with output data indicating the detected event.

[0112] A trigger event may include receiving a request to initiate a call from one or more users. For example, trigger module 110 may detect the arrival of a request from a user through a third-party API 404. The third-party API 404 is communicably connected to a user interface, such as user interface 416, through which the user can provide input indicating a request to initiate a call. For example, user interface 416 may be a graphical user interface for an application through which the user can request that a call campaign be scheduled and executed. The user can provide data indicating a specific callee or set of callees, and specific data requested for extraction. For example, the user may request that a call campaign be launched to each hardware store in Virginia that sells livestock supplies, and that the hardware store be asked whether it carries chick feed (for example, so that an index of the locations where supplies are delivered is available for later searching).

[0113] During a call, the called party can schedule different times for system 100 to call back to them. For example, if asked whether any changes have been made to a restaurant menu, the human called party can ask system 100 to call back an hour later or the next day for further action after they have had a chance to see the new menu. System 100 can then schedule a call at the requested time. In some examples, trigger module 110 can schedule future trigger events. In other examples, flow manager 132 can schedule intents or call events to be executed by dialer 106 to initiate a call.

[0114] A trigger event can include trends or patterns detected in stored data within a knowledge database or in data provided in real time. For example, a trend detected in search data received from the search engine 418 could be a trigger event. The search engine 418 can receive search requests from users and provide data representing those search requests to the trend detector 406. The trend detector 406 analyzes the received data and detects trends in the received data. For example, if searches for Cuban restaurants in Asheville, North Carolina, have increased by 500% in the past month, the trend detector 406 can detect the increase in searches and provide data representing the trend to the trigger module 110.

[0115] The trend detector 406 can output data to the trigger module 110 indicating a specific caller or set of callers based on the identified trend. In some implementations, the trend detector 406 provides data indicating the detected trend, and the trigger module 110 determines a specific caller or set of callers based on the identified trend. For example, the trend detector 406 may determine that searches for "Tornado Lincoln, Nebraska" have increased by 40% and provide the search keywords to the trigger module 110. The trigger module 110 may then determine that calls should be made to all stores that provide emergency supplies to check how much stock each store has of essential goods and their opening hours (for example, for indexing by search engine users and subsequent searches).

[0116] A trigger event can include a specific event of interest that has been identified as affecting the normal operation of a business, organization, or individual. The event identifier 408 receives data from various third-party sources, including third-party databases 420 and event database 422. The event identifier 408 can also receive data from other sources, such as a local memory device or a real-time data stream. The event identifier 408 identifies a specific event from databases 420 and 422 and outputs data indicating the identified event to the trigger module 110. In some examples, the trigger module 110 selects a specific caller or set of callers and the data requested between calls based on the data provided by the event identifier 408.

[0117] Certain events that could affect the operation of a business, organization, or individual include extreme weather conditions, federal holidays, religious holidays, sporting events, and various other events.

[0118] Third-party database 420 provides event identifier 408 with data from various third-party data sources, including weather services and government alerts. For example, third-party database 420 can provide event identifier 408 with a storm warning. Event identifier 408 can then determine that a winter storm is approaching the northeast corner of Minneapolis, Minnesota, and that a call should be made to a hardware store in the northeast corner of Minneapolis to determine the current inventory of available generators.

[0119] The event database 422 provides data from various data sources to the event identifier 408, specifically including data indicating known events. For example, the event database 422 may provide data indicating federal and state holidays, religious holidays, parades, sporting events, exhibition opening days, visits by high-ranking officials, and various other events.

[0120] For example, if a particular city is hosting a Super Bowl, the event database 422 can provide data to the event identifier 408, which in turn provides data indicating the event to the trigger module 110. Based on known information about the current Super Bowl and stored information about past Super Bowls, the trigger module 110 can determine that calls should be made to all hotels in the area to check availability and prices. The trigger module 110 can also determine that calls should be made to sporting goods stores to determine the availability of jerseys for each team participating in the Super Bowl. In such a situation, other information that may affect the operation of a business, organization, or individual that the trigger module 110 may request includes closures of office buildings or schools, changes to public transport schedules, special restaurant offerings, or various other information.

[0121] One or more of the various modules of system 100 can determine an estimated trigger event or requested information based on event information received from event identifier 408. For example, a South American restaurant, specifically a Mexican restaurant's Dia de Muertos, may have a special menu or opening hours for the celebration. In such an example, trigger module 110 can provide instructions to dialer 106 to call the South American restaurant to update its opening hours and menu for the day.

[0122] In some implementations, trigger events can be detected from calls initiated by the system 100 itself. Based on a portion of the conversation conducted by the system 100, the flow manager 132 can determine that an intent suggesting that a call should be initiated was expressed during the conversation. For example, if a human caller says, "Yes, we are still open until 8 p.m. every Thursday, but next week we switch to our summer schedule and are open until 9:30 p.m.," the flow manager 132 can identify an intent that provides further information about the data field.

[0123] In some implementations, a trigger event may include receiving an unsatisfactory result from a previously initiated call. For example, if a bot makes a call to a business to determine if the business has special holiday hours on the Independence Day holiday, and the truthfulness of the response provided by the business's human representative does not have at least a threshold amount of confidence, then system 100 may schedule a call on another specific day or time, such as July 1st, to determine if special holiday hours are scheduled. In such an example, the trigger module 110 may schedule a trigger event or provide information to the flow manager 132 in order to schedule an action. In some examples, the flow manager 132 schedules the start of a callback by scheduling the sending of an instruction to the dialer 106.

[0124] System 100 has a common sense module 133 that enables the flow manager 132 to intelligently schedule and select nodes for a particular workflow. For example, in the above situation, if there is a deadline for the usefulness of the requested information during the call, the common sense module 133 can also determine when to schedule the call and what information to request. In some examples, the common sense module 133 is a component of the flow manager 132, as illustrated in Figure 1. In other examples, the common sense module 133 is a component of the trigger module 110, facilitating the trigger module 110 to make an intelligent decision about whether a call should be initiated.

[0125] Figure 5 shows an example of a process 500 for initiating a telephone call. Briefly, the process 500 may include the steps of: (502) receiving data indicating a first event by a call triggering module of the call initiation system for initiating a call and for a conversation between a bot of the call initiation system and a human caller during the call; (504) determining, by the call triggering module and by using the data indicating the first event, that the first event is a trigger event that triggers a workflow of the call initiation system that is initiated by initiating a telephone call; (506) selecting a specific workflow based on the determined trigger event; and (508) initiating a telephone call to a caller specified by the specific workflow in response to the selection.

[0126] Process 500 may include step (502) receiving data indicating a first event by a call triggering module of the call initiation system for initiating a call and facilitating a conversation between the call initiation system's bot and a human callee during the call. For example, trigger module 110 may receive data from mismatch detector 402 indicating a discrepancy between the opening hours of Sally's Sloon of Sweets listed on the store's website and the opening hours stored in a search index related to the business.

[0127] Process 500 may include the step (504) of determining, by the call triggering module and by using data indicating the first event, that the first event is a trigger event that triggers a call initiation system workflow which is initiated by initiating a telephone call. In some examples, the determined trigger event is a mismatch between a value associated with a first data source and a corresponding value associated with a second data source. For example, the trigger module 110 may use a mismatch detected by the mismatch detector 402 to determine that the mismatch is a trigger event that triggers a workflow to determine what Sally's Saloon's actual business hours are.

[0128] In some cases, data indicating the first event is provided by the user. For example, a user may report a discrepancy between the opening hours listed on Sally's Saloon's website and the opening hours posted at Sally's Saloon's storefront.

[0129] In some examples, the determined trigger event is a user request. For instance, a user may provide input to a third-party API, such as a third-party API 404, through a user interface, such as a user interface 416, to request scheduling and execution of a call to a specific callee or a specific set of callees.

[0130] In some examples, the determined trigger event is a specific type of event, such as a weather event, a sporting event, a recreational event, or a seasonal event. For example, event identifier 408 can determine that the Head of Charles Regatta is taking place in Boston, Massachusetts, and can provide event data to trigger module 110. Trigger module 110 can then determine that the regatta is the trigger event.

[0131] In some examples, the determined trigger event is a trend detected in search requests submitted to the search engine. For example, the trend detector 406 can receive search engine data from the search engine 418 and determine that Spanish tapas restaurants are trending. The trend detector 406 can provide the trend-indicating data to the trigger module 110, which can then determine that the trend is a trigger event.

[0132] In some examples, the determined trigger event is the passage of a predetermined time period. For example, the mismatch detector 402 may determine, based on data in the knowledge database 412 from the timer 414, that the menu of a Cuban restaurant in Manhattan, New York, has not been updated for four months. The mismatch detector 402 can provide timing data to the trigger module 110, which can determine that the passage of four months without updating the menu data for the Cuban restaurant in Manhattan is the trigger event. The trigger module 110 can then provide data to the flow manager 132 indicating that the Cuban restaurant in Manhattan is making a call to retrieve the updated menu information.

[0133] Process 500 may include a step (506) of selecting a specific workflow based on a determined trigger event. The trigger module 110 can provide trigger event data to the dialer 106 or flow manager 132 for use in selecting a specific workflow or a node in a workflow. For example, the trigger module 110 can provide the flow manager 132 with trigger event data indicating a discrepancy in the listed business hours of Sally's Saloon for Sweets, and the flow manager 132 can use the data to select a specific workflow to call Sally's Saloon to resolve the discrepancy.

[0134] Process 500 may include a step (508) in which, in response to a selection, a telephone call is initiated to a callee designated by a particular workflow. The flow manager 132 can provide the dialer 106 with a command indicating a specific callee to be contacted. For example, the flow manager 132 can provide the dialer 106 with a command to initiate a call to Sally's Saloon.

[0135] The initiation of workflows by the systems and methods described herein, and more specifically, the initiation of calls, can be relatively automated by triggering events, but safeguards may be included in System 100 to prevent unwanted calls or calls that violate local government regulations. For example, if a called party indicates that they no longer wish to receive calls from the system, the system may recognize this and build a call inspection to the called party's number to prevent further calls.

[0136] Furthermore, to the extent that the systems and methods described herein collect data, the data may be processed in one or more ways before it is stored or used, such that personally identifiable information is removed or permanently obscured. For example, the identity of a defendant may be permanently removed or processed so that personally identifiable information cannot be determined, and, if necessary, the geographical location of the called party may be generalized so that the specific location of the user cannot be determined if location information is obtained. If personal, private, or confidential information is received during a call, whether requested as part of the workflow, voluntarily provided by the called party, or received in error, the workflow may include a step to permanently remove or obscure the information from the system.

[0137] In certain cases, system 100 may provide the user with the current status of its efforts to perform a task, either automatically or at the user's request. For example, system 100 may provide the user with the status of a task being performed through notifications about the device the user is using, such as a computer or mobile device. In some cases, system 100 may notify the user of the status of an ongoing task through other means, such as a messaging application or via telephone communication.

[0138] Figure 6 is a block diagram 600 of the task manager module of system 100. The task manager module 140 is connected to the communication platform 102, the trigger module 110, the task information storage 150, and the session storage 116. When a user communicates a task through the communication platform, the task information is stored in the task information storage 150, and the task manager module 140 determines when the task should be scheduled. The task manager can associate tasks with trigger events. A task may have a status initially set to "new," or some other indicator that processing has not yet been done for the request. When a trigger event occurs, the trigger module 110 starts the dialing process. In some implementations, the task manager module 140 monitors the session storage to update the status of each task as the task status changes from start to in progress to completed.

[0139] From the session information, the task manager module can determine the status and outcome of each call. For example, a bot might attempt to call a restaurant several times before connecting with someone to make a reservation. Session storage holds information about each call the bot makes. In some implementations, the task manager module may periodically poll the session storage to determine the status of the call task, i.e., whether the call is being initialized, in progress, or completed. In other implementations, the session storage may send the call result to the task manager module to update the status of the task in the task information storage.

[0140] In some implementations, calls are reviewed by operators through an operator dashboard that displays information about the call task and the progress of the task.

[0141] Figure 7A shows an operator dashboard displaying information about the progress of an existing call task. For example, Figure 7A shows a haircut appointment task. The operator dashboard may provide information about the appointment, including the appointment time, the requester's name, the requested service, the business name, the date, and the appointment time. The operator may review the request and associated session information from the call associated with the request to determine whether the requested appointment has been properly booked.

[0142] Figure 7B shows an operator review screen for reviewing one of the tasks requested by a user. The screen may show the operator the current status of the task. As shown in Figure 7B, the task is completed after a reservation is made. However, in some cases, the task may not be completed, or a reservation may not have been made. The operator may have the option to play back the recording associated with the task, or view other stored information from the call, such as transcriptions, extracted intents, etc., or to make a call to the business associated with the task, or to schedule an automated call for the future. Furthermore, the operator may have the option to provide the requesting user with the current status of the task.

[0143] The user can also request the task status via the communication platform 102. Additionally or alternatively, the task manager module 140 may decide when to send a status update to the user based on a task status change or other triggering event such as time.

[0144] Figure 8 is a flowchart showing an example of a process 800 for providing task status. Process 800 may include a step (802) in which the task manager module determines that a triggering event has occurred in order to provide the current status of a user call request. As described above, the triggering event may include a user request for a status, the passage of a certain amount of time, or a change in the status of a particular task. Process 800 then includes a step (804) in which the task manager module determines the current status of the user call request. The task manager module can determine the current status by checking the status in the task information storage. The status of a task is initialized when the task is added to the task information storage 150. When a call associated with the task is made and completed, the status of the task is updated. The task manager then generates a representation of the current status of the user call request (806). The representation may be a visual representation or an audio representation that conveys the current status of the task. Process 800 provides the user with the generated representation of the current status of the user call request (808).

[0145] Figure 9A shows the visual status of the haircut appointment request in Figure 1B while the appointment scheduling is in progress. The user may have access to a user interface to check the status of the task request, and the status may be sent to the user's device, such as a smartphone, smartwatch, laptop, personal home assistant device, or other electronic device. The status may be sent via email, SMS, or other mechanism.

[0146] Figure 9B shows the visual status of the haircut reservation request from Figure 1B after the reservation has been successfully scheduled. This status may be requested by the user and may be sent to the user without prompting once the reservation has been successfully completed.

[0147] Figure 10A illustrates the verbal status request and update for the restaurant reservation request in Figure 1C. As shown in Figure 10A, in response to the user asking whether the restaurant reservation has been made, the system may describe the steps it took to complete the task, such as making two calls to the restaurant. The system may also inform the user when it is scheduled to attempt the call again and may notify the user of the status after the call attempt.

[0148] Figure 10B illustrates a verbal status update provided by the system without user prompting for a restaurant reservation request in Figure 1C. Once the system recognizes that the user's task is complete, it can provide the user with a status update. In some implementations, the system provides the user with a status update immediately. In other implementations, the system determines a convenient time or method for notifying the user. For example, a user might request a dinner reservation in London, England. However, the user might currently be located in Mountain View, California, USA. The system could attempt to call the restaurant while the user is asleep. If the system confirms a reservation for 12:00 PM in London, the system might determine that sending a status update text message at 4:00 AM (PDT) would wake the user. The system could then choose an alternative status update method, namely email, or hold the status update for a time more convenient for the user. The system can use information from the user's schedule, time zones, habits, or other personal information of the user to determine an appropriate and convenient time and method for providing the user with a status update.

[0149] In some implementations, a system may use user information to determine the urgency of a task or whether to repeat efforts to complete the task. For example, a system might be trying to make a reservation for a user at a specific restaurant in Mountain View, California. The user's trip to Mountain View might end on May 15th. If the system has not yet succeeded by May 15th, it would not make sense for the system to continue requesting the reservation after May 16th, as the user's trip is over. However, it would make sense to call twice as often on May 14th to try and get in touch with someone at the restaurant to make a reservation. Tasks become more urgent as the deadline approaches, and may lose their urgency or become too late once the deadline has passed.

[0150] In some implementations, the relief module 136 in Figure 1B determines the type of intervention to be introduced into the call while the call is in progress. The relief module 136 may choose to manually relieve the bot's conversation in real time, allowing another party to take over the call. In other implementations, the module may allow a human operator to quietly take over the call. Additionally or alternatively, the relief module 136 may choose to politely terminate the telephone call between the bot and the human without manual intervention.

[0151] Figure 11 shows an exemplary process 1100 for transferring a telephone call from a bot to a human. Process 1100 may include the step of the call initiation system analyzing the real-time conversation between the first human and the bot during the telephone call between the first human at the first end of the telephone call and the bot at the second end of the telephone call (1102). The call initiation system may then determine, based on the analysis of the real-time conversation, whether the telephone call should be transferred from the bot to the second human at the second end of the telephone call (1104). In response to the decision that the telephone call should be transferred to the second human at the second end of the telephone call, the call initiation system transfers the telephone call from the bot to the second human (1106).

[0152] To determine the most appropriate type of intervention for a particular bot phone call, the relief module 136 may identify a tension event and look for other indications that the call should be terminated or handed over to a human operator.

[0153] In some implementations, the relief module 136 identifies tension events that indicate tension in either the human or the bot in order to respond appropriately to a human question. Each time the relief module 136 identifies a tension event, the memorized levels of both the local and global tension in the call increase. Whenever the conversation appears to be getting back on track, the relief module 136 resets the local tension level. For example, when a bot calls a restaurant to make a reservation for six people, the human might ask the bot, "How many high chairs do we need?" The bot might respond, "We all need chairs." The human might become slightly irritated based on the bot's response and respond, "Yes, I know we all need chairs, but how many high chairs do we need for the baby?" The system may detect intonation patterns, i.e., higher tones at the beginning, end, or throughout a human utterance. In some implementations, intonation patterns are pre-associated with stress or irritation. The system can match pre-associated patterns with patterns detected in real-time conversation. In some implementations, intonation patterns can detect repeated words, intentionally slower speech, or keywords or phrases ("Are you listening to me?", "Am I talking to a robot?").

[0154] The system increases the local tension level of the call when it detects a slightly irritated tone in the human voice. Local tension is a running score that reflects the amount of tension that may be associated with the current state. If any of the tension indicators appear in the human utterances in the real-time conversation, the tension score increases until the score reaches an intervention threshold. If no stress indicators appear, the system may indicate that the call is progressing according to the workflow, and the local tension score decreases or remains low (or is 0). If the bot responds appropriately to the question by providing a response expected by a human, such as "There are no children in our group," the system may reduce the local tension. If the system detects that a human has responded without any irritation in the voice, the relief module may determine that the call is back on track and may reset the local tension to its default value or to zero.

[0155] The global tension of a phone call is cumulative. Local tension, on the other hand, attempts to assess whether there is tension in the current interaction with the human, while global tension attempts to assess the total tension of the entire call. For example, a threshold may be set for three misunderstandings before a bot is rescued by a human operator. If the bot fails to understand the human three times in a row, the local tension will increase, leading to the bot being rescued. In a different call, if the bot fails to understand the other side twice in a row but understands the third sentence, the local tension may be reset in the third exchange, and the conversation may continue. Global tension still retains information indicating that there were two misunderstandings between the bot and the human. In a subsequent call, if the bot again fails to understand the human twice in a row, the global tension level will exceed the threshold, and the bot will be rescued, even if the local tension is still below the set threshold of three misunderstandings.

[0156] As described above, if either the local or global tension level reaches a certain threshold, the relief module 136 indicates to the system 100 that it is time for manual intervention or politely withdraws from the call. In some cases, the relief module 136 considers an event a tension event whenever it is necessary to repeat the same thing, apologize, or ask for clarification, or whenever a person corrects the system 100 or complains about the call (for example, "I can't hear you, can you hear me?").

[0157] In some cases, the relief module 136 will consider an event a tension event if a human asks whether the bot is a robot, mocks the bot by asking meaningless questions, or behaves in any other way that the system does not anticipate (for example, if the system is asked about a sporting event when trying to make a restaurant reservation).

[0158] In some implementations, the relief module 136 is a set of feature-based rules that determine when the system should intervene manually. One feature-based rule might be that the system should be relieved if two consecutive unknown input intents occur. Another rule might be that the system should be relieved by a manual operator if four unknown input intents occur at any point during a call. The system tracks events occurring within the conversation and determines whether an event has occurred that meets the criteria of the rule.

[0159] In other implementations, the relief module 136 uses machine learning to automatically predict when to bail out to a human operator. For example, the relief module 136 can receive intents from a human conversation as input to one or more machine learning models. The machine learning models can then decide whether to bail out to a human operator based on the received intents, as well as past intents and outcomes. The system can train the machine learning models on features from annotated recordings that indicate when relief occurred or should not have occurred. The machine learning module can then predict when relief will occur given a set of input features.

[0160] The relief module 136 uses many factors to determine the relief measures, including human behavior, human tone of voice, the determined level of human discomfort, the language used by the human, or the choice of human words.

[0161] System 100 can escalate conversations being conducted by the bot to a human operator for processing. For example, if there is a threshold amount of tension in a particular conversation, the relief module 136 can provide feedback data to the flow manager 132. The flow manager 132 can instruct the bot to hand over the call to a human operator providing input through the operator controller 134, with or without audibly alerting the human caller. For example, the bot might say, "Of course, thank you for your time today. This is my manager." The human operator can then complete the task that the bot was attempting to perform through the operator controller 134.

[0162] The recovery module 136 can also determine a confidence level that defines the system's confidence in the current task being accomplished. For example, a bot might be tasked with making a dinner reservation for a user. If the bot makes a call to the restaurant and the human asks several questions that the bot does not know the answers to, the system may have a low confidence level in the current task being accomplished. After the system receives questions for which it does not have an answer, the system's confidence level in accomplishing the task may be even lower. If the system recovers and determines that the conversation is moving in the direction of accomplishing the task, the system may increase its confidence level.

[0163] In some implementations, the system hands the phone conversation to a human operator who monitors the call. The system may alert the operator to the need to transfer the call using the operator's user interface or some other notification mechanism. Upon notification, the operator may have a finite amount of time to transfer the call before the system decides to terminate it. The system may use the same voice as the operator. In such cases, the transfer from bot to operator may be transparent to the other party because the voice remains the same.

[0164] In other implementations, the system hands over the phone conversation to the human user who requested the task. The system can alert the user to ongoing phone calls. The system can notify the user if there are problems completing the task or if the bot is asked a question it does not know the answer to. The bot can communicate details of the conversation that require user input via text, email, or some other method. In some implementations, the bot waits a threshold amount of time, i.e., 5 seconds, for the user to respond before continuing the conversation without user input. Because the conversation is happening in real time, the bot cannot wait for a long period of time for the user to respond. In some implementations, the system can attempt to hand over the phone call to the requesting user when the system determines that the phone call needs to be handed over from the bot. As mentioned above, the system may wait a threshold amount of time for the user to respond and take over the phone call. In some implementations, if the user does not take over the phone call within the threshold amount of time, the system hands over the phone call to an operator. In other examples, the system terminates the phone conversation. The system may also use the same voice as the human user so that the handover from the bot to the user is seamless from the other side of the conversation.

[0165] Figure 12 shows examples of computing device 1200 and mobile computing device 1250 that may be used to implement the techniques described above. Computing device 1200 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Mobile computing device 1250 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are intended to be illustrative only and are not intended to limit the forms of implementation of the invention described and / or claimed herein.

[0166] The computing device 1200 includes a processor 1202, memory 1204, storage device 1206, a high-speed interface 1208 connecting to memory 1204 and multiple high-speed expansion ports 1210, and a low-speed interface 1212 connecting to low-speed expansion port 1214 and storage device 1206. Each of the processor 1202, memory 1204, storage device 1206, high-speed interface 1208, high-speed expansion port 1210, and low-speed interface 1212 are interconnected using various buses and may be mounted on a common motherboard or in other appropriate ways.

[0167] To display graphical information for a GUI on an external input / output device such as a display 1216 coupled to the high-speed interface 1208, the processor 1202 can process instructions, including instructions stored in memory 1204 or storage device 1206, for execution within the computing device 1200. In other implementations, multiple processors and / or multiple buses may be used appropriately, along with multiple memories and several types of memory. Also, multiple computing devices may be connected, with each device providing some of the necessary operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0168] Memory 1204 stores information within the computing device 1200. In some implementations, memory 1204 is a volatile memory unit. In some implementations, memory 1204 is a non-volatile memory unit. Memory 1204 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0169] The storage device 1206 can provide high-capacity storage to the computing device 1200. In some implementations, the storage device 1206 may be an array of devices including computer-readable media such as floppy disk devices, hard disk devices, optical disk devices, or tape devices, flash memory or other similar solid memory devices, or devices in a storage area network or other configuration. The computer program product may be explicitly embedded in an information medium. The computer program product may also include instructions that, when executed, perform one or more of the methods described above. The computer program product may also be explicitly embedded in computer-readable or machine-readable media such as memory 1204, the storage device 1206, or memory on the processor 1202.

[0170] The high-speed interface 1208 manages the bandwidth-intensive operation of the computing device 1200, and the low-speed interface 1212 manages the low-bandwidth-intensive operation. Such function allocations are illustrative only. In some implementations, the high-speed interface 1208 is coupled to memory 1204, a display 1216 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 1210 that can accept various expansion cards (not shown). In this implementation, the low-speed interface 1212 is coupled to the storage device 1206 and the low-speed expansion port 1214. The low-speed expansion port 1214, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet®, Wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, or networking device such as a switch or router via a network adapter.

[0171] The computing device 1200 can be implemented in several different forms, as shown in the drawings. For example, it may be implemented as a standard server, and may be implemented multiple times within a group of such servers. Furthermore, the computing device 1200 may be implemented in a personal computer, such as a laptop computer 1222. The computing device 1200 may be implemented as part of a rack server system 1224. Alternatively, components from the computing device 1200 may be combined with other components (not shown) in a mobile device, such as a mobile computing device 1250. Each of such devices may include one or more of the computing device 1200 and the mobile computing device 1250, and the entire system may consist of multiple computing devices communicating with each other.

[0172] The mobile computing device 1250 includes, among other components, a processor 1252, memory 1264, input / output devices such as a display 1254, a communication interface 1266, and a transceiver 1268. The mobile computing device 1250 may also include a storage device such as a microdrive or other device to provide additional storage. Each of the processor 1252, memory 1264, display 1254, communication interface 1266, and transceiver 1268 is interconnected using various buses, and some of the components may be mounted on a common motherboard or in other appropriate ways.

[0173] The processor 1252 can execute instructions within the mobile computing device 1250, including instructions stored in memory 1264. The processor 1252 may be implemented as a chipset of a chip containing multiple separate analog and digital processors. The processor 1252 may be provided for coordinating other components of the mobile computing device 1250, such as user interface control, applications run by the mobile computing device 1250, and wireless communication by the mobile computing device 1250.

[0174] The processor 1252 can communicate with the user through a control interface 1258 and a display interface 1256 coupled to the display 1254. The display 1254 may be, for example, a TFT (thin-film transistor liquid crystal display) display, an OLED (organic light-emitting diode) display, or other suitable display technology. The display interface 1256 may include appropriate circuitry for driving the display 1254 to present graphical and other information to the user. The control interface 1258 may receive commands from the user and translate them for submission to the processor 1252. Furthermore, an external interface 1262 may provide communication with the processor 1252 to enable short-range communication between the mobile computing device 1250 and other devices. The external interface 1262 may be provided for wired communication in some implementations and for wireless communication in other implementations, and multiple interfaces may also be used.

[0175] Memory 1264 stores information within the mobile computing device 1250. Memory 1264 may be implemented as one or more computer-readable media, volatile memory units, or non-volatile memory units. Furthermore, extended memory 1274 may be provided to and connected to the mobile computing device 1250 through an expansion interface 1272, which may include, for example, a SIMM (Single In-Line Memory Module) card interface. Extended memory 1274 may provide extra storage space for the mobile computing device 1250 and may store applications or other information for the mobile computing device 1250. Specifically, extended memory 1274 may include instructions for executing or supplementing the processes described above, and may also include secure information. Therefore, for example, extended memory 1274 may be provided as a security module for the mobile computing device 1250 and may be programmed with instructions that enable secure use of the mobile computing device 1250. Furthermore, secure applications may be provided via a SIMM card along with additional information, such as placing identification information on the SIMM card in a hacker-proof manner.

[0176] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as described later. In some implementations, the computer program product is explicitly embedded in an information medium. When executed, the computer program product includes instructions that perform one or more of the methods described above. The computer program product may be in a computer-readable or machine-readable medium, such as memory 1264, extended memory 1274, or memory on processor 1252. In some implementations, the computer program product may be received in a propagating signal, for example, via transceiver 1268 or external interface 1262.

[0177] The mobile computing device 1250 can communicate wirelessly through a communication interface 1266, which may include digital signal processing circuitry as needed. The communication interface 1266 can provide communication under various modes or protocols, including, among others, GSM® voice call (Global System for Mobile Communications), SMS (Short Message Service), EMS (Extended Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA® (Wideband Code Division Multiple Access), CDM252000, or GPRS (General Purpose Packet Radio Service). Such communication may be performed, for example, through a transceiver 1268 using radio frequencies. Furthermore, short-range communication may be performed using Bluetooth®, Wi-Fi, or other such transceivers (not shown). Furthermore, the GPS (Global Positioning System) receiver module 2570 can provide the mobile computing device 1250 with additional navigation and location-related wireless data, which can be appropriately used by applications running on the mobile computing device 1250.

[0178] The mobile computing device 1250 may also communicate audibly using an audio codec 1260, which can receive voice information from a user and convert it into usable digital information. The audio codec 1260 may also generate audible sounds for the user, for example, through a speaker in the handset of the mobile computing device 1250. Such sounds may include sounds from voice phone calls, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications running on the mobile computing device 1250.

[0179] The mobile computing device 1250 can be implemented in several different forms, as shown in the drawings. For example, the mobile computing device 1250 can be implemented as a cellular phone 1280. The mobile computing device 1250 can also be implemented as part of a smartphone 1282, a personal digital assistant, a tablet computer, a wearable computer, or other similar mobile device.

[0180] Several implementation forms are described. Nevertheless, it will be understood that various modifications can be made without departing from the intent and scope of this disclosure. For example, steps can be rearranged, added, or removed to use various forms of the flow shown above.

[0181] All functional operations described herein may be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or one or more combinations thereof. The techniques disclosed may be implemented as one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of one or more computer program products, i.e., data processing devices. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a machine-readable propagation signal affecting a signal, or one or more combinations thereof. The computer-readable medium may be a non-transient computer-readable medium. The term “data processing device” encompasses all devices and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, a device may include code that generates the execution environment for the computer program, such as processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagating signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal that is produced to encode information for transmission to a suitable receiving device.

[0182] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, such as a standalone program or as modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs do not necessarily correspond to files in a file system. A program may be stored in part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files storing one or more modules, subprograms, or parts of code). A computer program can be deployed to run on one computer, or located in one site, or distributed across multiple sites and interconnected by a communication network.

[0183] The processes and logic flows described herein may be executed by one or more programmable processors that run one or more computer programs to perform their functions by manipulating input data and generating outputs. The processes and logic flows may also be executed by dedicated logic circuits such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the devices may be implemented as dedicated logic circuits.

[0184] Processors suitable for executing computer programs include, as an example, both general-purpose and dedicated microprocessors, and any one or more processors in any type of digital computer. Generally, a processor receives instructions and data from read-only memory or random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operablely coupled to receive data from there or transfer data thereto, or both. However, a computer does not necessarily have such devices. Furthermore, a computer may be embedded in another device, to name just a few examples, such as a tablet computer, mobile phone, personal digital assistant (PDA), mobile audio player, or Global Positioning System (GPS) receiver. Computer-readable media suitable for storing computer program instructions and data include, for example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROM disks; and all forms of non-volatile memory, media, and memory devices. The processor and memory may be supplemented by or incorporated into dedicated logic circuits.

[0185] To provide user interaction, the disclosed techniques may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, a keyboard on which the user can provide input to the computer, and a pointing device, such as a mouse or trackball. Other types of devices may also be used to provide user interaction; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input.

[0186] An implementation may include a computing system that includes, for example, a backend component such as a data server, or a middleware component such as an application server, or a frontend component such as a client computer having a graphical user interface or a web browser that allows the user to interact with the implementation of the disclosed technique, or any combination of one or more such backend components, middleware components, or frontend components. The components of the system may be interconnected by digital data communication in any form or medium, such as a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), such as the Internet.

[0187] A computing system may include a client and a server. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is established by computer programs running on each computer that have a client-server relationship with each other.

[0188] This specification contains many details, but these should not be interpreted as limitations, but rather as descriptions of features specific to particular implementations. Certain features described herein in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately or in any suitable subcombinations in multiple implementations. Furthermore, features are described above as working in a particular combination, and even if initially claimed as such, in some cases one or more features may be removed from the claimed combination, and the claimed combination may be led to subcombinations or variations of subcombinations.

[0189] Similarly, while the operations are shown in a specific order in the diagrams, this should not be understood as requiring that such operations be performed in a specific order or sequence shown, or that all illustrated operations be performed, in order to achieve the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the aforementioned implementation forms should not be understood as requiring such separation in all implementation forms, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.

[0190] Therefore, a specific implementation is described. Other implementations are within the scope of the following claims. For example, the actions described in the claims may be performed in a different order and still achieve the desired result. [Explanation of symbols]

[0191] 100 Systems 102 Communication framework or platform 104 humans 106 Dial 108 Sound System 110 Call triggering module or trigger module 112 Audio Packages 114 Session Recorder 114 Recording Studios and Evaluation TTS 116 Session Storage 118 Text-to-Speech Module 120 Speech endpoint detector 122 Stored text-speech results or recordings 122 Text-speech output and readings 122 Stored output / read values 124 Intent - Text Module 126 Speech-Text Module 128 Speech Application Programming Interface (API) 130 Text-Intent Modules 132 Flow Manager 133 Common Sense Module 134 Operator Controller 134 Operators 136 Relief Module 140 Task Manager Module 150 Task Information Storage 160 Task User Interface 160 Task Manager Modules 162 Call Player 170 Curation User Interface 175 Local Agent 190 phone call 191 Speech Recognition Device 194 Audio Mixer 195 Bot Services 196 Telephone Server 197 TTS model 198 Dialogue Model 199 Language Models 200 processes 250 processes 300 Workflows 400 Block Diagram 402 Mismatch Detector 404 Third Party API 406 Trend Detector 408 Event Identifier 410 User Interface 412 Knowledge Base 412 Knowledge Databases 414 Timer 416 User Interface 418 search engines 420 Third-Party Databases 422 Event Database 500 processes 600 Block Diagram 800 processes 1100 processes 1200 computing devices 1202 processors 1204 memory 1206 Storage Devices 1208 High-speed interface 1210 High-Speed ​​Expansion Ports 1212 Low-speed interface 1214 Low-speed expansion port 1216 displays 1222 Laptop Computer 1224 Rack Server System 1250 Mobile Computing Devices 1252 processors 1254 displays 1256 Display Interface 1258 Control Interface 1260 audio codecs 1262 External Interface 1264 memory 1266 Communication Interface 1268 Transceiver 1274 Expansion Memory 1280 Cellular Phone 1282 Smartphones 2570 GPS (Global Positioning System) Receiver Module

Claims

1. A computer implementation method, A step in which a call initiation system receives a request from a human user to perform a task, wherein the call initiation system is configured to initiate a telephone call between the call initiation system's bot and a human representative of the organization and to conduct a telephone conversation. In response to receiving the request to perform the task, the call initiation system determines to initiate a telephone call and conduct a telephone conversation between the bot of the call initiation system and the human representative of the organization in order to perform the task. The steps include: in response to a decision to initiate a telephone call between the bot of the call initiation system and the human representative of the organization and to have a telephone conversation, the steps include: initiating a telephone call between the bot of the call initiation system and the human representative of the organization and having a telephone conversation; The steps include: determining that the status of the task has changed based on a conversation that took place between the bot of the call initiation system and the human representative of the organization during the aforementioned telephone call; In response to a determination that the status of the task has changed, a step of generating a summary of the telephone conversation based on the conversation that took place between the bot of the call initiation system and the human representative of the organization during the telephone call, The above summary includes the steps of: during the conversation, the call initiation system identifies a question from the human representative for which it does not know the answer; The steps include providing the human user with an outline of the aforementioned telephone conversation for output, and Includes, The step of determining that the status of the task has changed includes determining that in order to complete the task, at least one additional telephone call should be made between the bot of the call initiation system and the human representative of the organization, and an additional telephone conversation should take place. A computer implementation method for generating the summary of the aforementioned telephone conversation, based on the determination that at least one additional telephone call should be made between the bot of the call initiation system and the human representative of the organization, and that an additional telephone conversation should take place, in order to complete the task.

2. The method according to claim 1, wherein the summary includes a one-line text summary of the status of the conversation.

3. The method according to claim 1, wherein the above summary summarizes one aspect of what was discussed between the bot and the human representative during the telephone call.

4. The method according to claim 1, wherein the summary outlines the current status of the conversation between the bot and the human representative.

5. The method according to claim 1, wherein the human user is not a party to the telephone call between the bot and the human representative.

6. The method according to claim 1, wherein the summary includes information collected during the telephone call.

7. It is a method, A step of receiving a request from an automated telephone call initiation system to perform a task that requests the bot of the automated telephone call initiation system to initiate a telephone call and conduct a telephone conversation between the bot and the person being called, wherein the automated telephone call initiation system (i) includes a task manager module, and (ii) is configured to initiate a telephone call and conduct a telephone conversation between the bot of the automated telephone call initiation system and the person being called, The task manager module of the automated telephone call initiation system determines that a triggering event has occurred to provide the current status of the task which requests the bot of the automated telephone call initiation system and the called party of the telephone call to initiate the telephone call and have a conversation over the telephone, wherein the determination that the triggering event has occurred is The automated telephone call initiation system includes the steps of initiating the telephone call between the bot of the automated telephone call initiation system and the person called in the telephone call, conducting a conversation over the telephone, and determining that the status of the task has changed. In response to the triggering event, the task manager module of the automated telephone call initiation system determines the current status of the task which requests the bot of the automated telephone call initiation system to initiate the telephone call and have a conversation over the telephone; The steps include generating a representation of the current status of the task which requests the bot of the automated telephone call initiation system to initiate the telephone call and have a conversation over the telephone between the bot of the automated telephone call initiation system and the person being called in the telephone call, A step of generating a summary of a telephone conversation based on the telephone conversation in response to the triggering event, wherein the summary identifies questions asked by the caller of the telephone call during the conversation, for which the automated telephone call initiation system does not know the answer. The steps include providing for output the current status of the task and an outline of the telephone conversation for the automated telephone call initiation system, which requests the bot of the automated telephone call initiation system and the called party of the telephone call to initiate the telephone call and have a telephone conversation; Includes, Determining that the status of the task has changed includes determining that, in order to complete the task, at least one additional telephone call should be initiated between the bot of the automated telephone call initiation system and the called party, and an additional telephone conversation should take place. A method for generating the summary of the aforementioned telephone conversation, based on the determination that at least one additional telephone call should be initiated between the bot of the automated telephone call initiation system and the called party, and that an additional telephone conversation should take place, in order to complete the task.

8. The method according to claim 7, wherein the triggering event is a user request regarding a status.

9. The method according to claim 8, wherein the triggering event is an operator interaction for providing a status to a user after the operator has reviewed session information associated with the request to perform the task.

10. The method according to claim 7, wherein the triggering event is a status update event.

11. The method according to claim 7, wherein the representation of the current status is a visual representation.

12. The method according to claim 7, wherein the expression of the current status is an audio expression.

13. Providing the user with the current status of the representation of the task requesting to initiate a telephone call, The method according to claim 7, comprising determining a convenient time and method for delivering the current status to the user.

14. The method according to claim 7, further comprising the step of notifying the user that a delay will occur in processing the request to perform the task.

15. A system comprising at least one processor and memory, wherein the memory records instructions that, when executed, enable the at least one processor to perform the method according to any one of claims 1 to 14.

16. A non-temporary computer-readable recording medium that records instructions, when executed, causing at least one processor to perform a procedure according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Vehicle use communication proxy device

    JP2003032329A

  • Method and apparatus for use in computer-to-human escalation

    JP2007532989A

  • System and method for making a reservation associated with a calendar appointment

    US20100094668A1