Outbound calling method and system
The outbound calling system uses CTI, STT, and generative AI to generate scripts for outbound calls, addressing the inefficiencies of existing systems by providing real-time support for communicators, enhancing their ability to recommend products or services effectively.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-31
AI Technical Summary
Existing outbound call systems in call centers lack real-time AI support, making it difficult for inexperienced communicators to effectively recommend products or services, as they struggle with explaining the purpose, evoking customer needs, analyzing responses, and providing appropriate recommendations, leading to inefficient operations.
An outbound calling system utilizing CTI, STT, vectorization, and generative AI to generate scripts based on customer attributes and utterances, supporting communicators with initial talks, short, and long responses, enabling real-time assistance even for inexperienced staff.
Enables inexperienced communicators to achieve average results by providing real-time support through AI-generated scripts tailored to customer interactions, improving communication efficiency and product/service recommendations.
Smart Images

Figure 2026055174000001_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an outbound call for recommending products or services to customers by phone from a call center.
Background Art
[0002] In call center operations, systems are provided where AI (artificial intelligence) supports communicators. However, most of them support inbound calls (calls from customers to the call center), and there is no known system that supports communicators in real time during conversations in outbound calls (calls from the call center to customers). In the case of inbound calls, since the customer calls the window selected according to the purpose, the content of the conversation can be predicted to some extent, and it is possible for AI to learn the content and results of past calls and support the communicator in real time.
[0003] On the other hand, in outbound calls, since the call is made from the communicator side, at the start of the conversation, there is no purpose for the customer to talk. Therefore, outbound calls need to clear several steps that inbound calls do not have. 1) Explain the purpose of the call and obtain permission to explain (initial talk). 2) Evoke needs from the customer (initial talk). 3) Analyze the customer's response to the evoked needs and provide an explanation suitable for the response. 4) Analyze the flow of the conversation and demonstrate services, products, etc. to be recommended to the customer. 5) Ask about the customer's questions and concerns and answer them accurately. 6) Close the conversation.
[0004] Regarding point 1), communicators take a long time to verify customer data. Regarding point 2), communicators with little success experience cannot select appropriate talk. Regarding points 3), 4), and 5), inexperienced communicators cannot grasp where the customer's interests lie, and cannot select the most suitable product or service to match the speed of the conversation and explain its features. Regarding point 6), it is difficult to explain all the important details necessary for contracting a product or service accurately. Furthermore, they cannot alleviate anxieties and questions that are likely to lead to cancellation. As described above, outbound calls are difficult, it is difficult to train communicators to become proficient, and the operation of the call center becomes inefficient. [Overview of the project] [Problems that the invention aims to solve]
[0005] The objective of this invention is to provide real-time support to communicators from the beginning to the end of outbound calls, enabling even inexperienced communicators to achieve average results. [Means for solving the problem]
[0006] In this invention's outbound calling method, a call center makes an outbound call to a customer, recommending that the customer purchase a product or service. In this invention, a CTI (Computer Telephony Integration) device 2 makes a call to the customer's communication terminal and displays the customer's attribute information on the communicator terminal 4, The STT (Text Conversion Device) 6 converts the customer's spoken audio into text, Administrator terminal 5, server 10, generating AI 12, vectorization device 20 that converts text-based voice data into vector data, An initial-word database 15 stores initial-word phrases from conversations with customers along with customer attribute information. A short sentence database 16 stores vectorized data of relatively short utterances from customers, along with scored response scripts. A long sentence database 17 stores vectorized data of relatively long utterances from customers, along with scored response scripts. An outbound call system is used, which includes a product database 18 of recommended products or services.
[0007] In this invention, the generated AI12 is used, Based on the customer's attribute information, the Initial Talk Database 14 is used to generate an Initial Talk script for the customer. For relatively short customer utterances, a short sentence database 15 is used to generate a response script for the customer. For relatively long customer utterances, a script for responding to the customer is generated using the long sentence database 16. The generated script is displayed on communicator terminal 4.
[0008] Furthermore, the outbound call system of this invention is a system for making outbound calls from a call center to customers and recommending that customers purchase products or services. A CTI (Computer Telephony Integration) device 2 makes a call to the customer's communication terminal and displays the customer's attribute information on the communicator terminal 4, The STT (Text Conversion Device) 6 converts the customer's spoken audio into text, Administrator terminal 5, server 10, generating AI 12, vectorization device 20 that converts text-based voice data into vector data, An initial-word database 15 stores initial-word phrases from conversations with customers along with customer attribute information. A short sentence database 16 stores vectorized data of relatively short utterances from customers, along with scored response scripts. A long sentence database 17 stores vectorized data of relatively long utterances from customers, along with scored response scripts. It includes a product database 18 of recommended products or services.
[0009] In this invention, the generated AI12 is Based on the customer's attribute information, the Initial Talk Database 14 is used to generate an Initial Talk script for the customer. For relatively short customer utterances, a short sentence database 15 is used to generate a response script for the customer. For relatively long customer utterances, a script for responding to the customer is generated using the long sentence database 16. Server 10 is configured to display the generated script on communicator terminal 4.
[0010] In this specification, descriptions of outbound calling methods also apply to the outbound calling system, and descriptions of the outbound calling system also apply to the outbound calling methods. "Script" refers to the content and outline of the response to the customer, and also serves as a suggestion to the communicator. The server is connected, for example, to a database containing communicator terminals, administrator terminals, generation and vectorization devices, initial talk, short sentences, and long sentences, as well as a product database describing the product or service, terms of trade, and contractual notes.
[0011] In this invention, an AI generates an initial-talk script based on customer attribute information and displays it on the communicator's terminal, allowing even inexperienced communicators to perform appropriate initial-talk.
[0012] For short customer utterances, the AI generates response scripts using a short sentence database. Since the short sentence database should contain vectors similar to the vectorized utterance, the system searches for similar vectors. Similar vectors correspond to similar customer interests, and by generating response scripts based on those with high scores, an appropriate response script can be produced. For example, one response script is generated for each short sentence of customer utterance. Alternatively, the communicator's terminal could display the top three response scripts based on their scores for each customer utterance, allowing the communicator to refer to multiple scripts to respond to the customer. This approach also applies to responses to long sentences.
[0013] For long customer utterances, the generating AI uses a long-sentence database to generate a script. In this case, since the utterance contains many keywords, for example, these keywords as a whole (context) are included in the vectorized data, and the generating AI considers response scripts with high scores among similar vectors to generate a response script. Also, for example, there is one response script for one long-sentence customer utterance. For these reasons, even inexperienced communicators can respond appropriately to customers, whether they are responding to short sentences or long sentences.
[0014] The distinction between short and long sentences is made by considering the number of sentences in a grammatical sense, as well as the number of keywords included in the customer's utterance. If an utterance is, for example, one sentence and contains few keywords, it is a short sentence. If an utterance consists of, for example, multiple sentences and contains many keywords, it is a long sentence. If a customer's utterance is longer than one sentence, but its length is below the upper limit, and it is a coherent utterance, it is a long sentence. The upper limit for the length of a long sentence is, for example, 5 to 10 sentences. It is unlikely that a customer will speak for more than 10 sentences without interruption, so most utterances can be divided into short and long sentences. If a customer speaks for more than 10 sentences without interruption, for example, every 10 sentences can be memorized as a separate long sentence.
[0015] By inputting customer utterance examples (e.g., model utterances) and exemplary response scripts with scores from the administrator terminal and vectorizing them, short sentence databases and long sentence databases can be generated initially. Alternatively, by initially making outbound calls to customers using the administrator terminal, vectorizing the customer's utterances, and assigning high scores to the communicator's (administrator's) responses, short sentence databases and long sentence databases can be generated. The response scripts can also be stored as vector data, for example.
[0016] When making an outbound call to a customer, adding the customer's speech and the script of the scored response to the database can grow the database. In this case, the score can be automatically assigned based on the result of the call or assigned by the communicator himself / herself. Preferably, for the response scripts regarding short sentences and long sentences displayed on the communicator terminal 4, the communicator changes the score from the terminal 4. That is, allowing the communicator to change the score for the response script displayed on the communicator terminal results in a more reliable score. For example, change the score in three levels of "○, no input (do not change the score), ×". The server 10 receives the score input from the terminal 4.
[0017] Support by the generative AI can be performed in real time, and in this invention, even an inexperienced communicator can achieve average results.
Brief Description of the Drawings
[0018] [Figure 1] Block diagram of the outbound call method and call system of the embodiment [Figure 2] Block diagram showing the generation of the initial talk in the embodiment [Figure 3] Block diagram showing the generation of short sentences in the embodiment [Figure 4] Block diagram showing the generation of long sentences in the embodiment [Figure 5] Block diagram showing the processing related to the database and AI in the embodiment [Figure 6] Diagram showing an example of the conversation suggested by the AI in the embodiment
Modes for Carrying Out the Invention
[0019] The following shows the optimal embodiments for carrying out the present invention.
Embodiment
[0020] Figures 1 to 6 show examples of the system. Figure 1 shows an overview of the outbound call system 1 of the example, where a call is made from system 1 to customer 02 (more accurately, the customer's communication terminal) via CTI2 (Computer Telephony Integration device). System 1 has two types of terminals: a communicator terminal 4 and an administrator terminal 5. In principle, the communicator (not shown) who speaks with customer 02 using headphones or similar means is the one looking at the monitor 8 on terminal 4. CTI2 sends customer 02's voice to terminal 4 and also displays customer 02's attribute information (customer information: age, gender, address, etc.) on monitor 8.
[0021] CTI2 inputs the voice of customer 02 and the voice of the communicator into the STT (text conversion device) 6 and converts it into text. System 1 inputs the text of the conversation with customer 02, at least the customer's utterances, preferably the utterances of both the customer and the communicator, into a virtual server 10, for example, a virtual server Lambda provided by Amazon (Lambda is a trademark of Amazon). A real server may be used instead of the virtual server 10, and the type of virtual server is arbitrary.
[0022] The virtual server 10 inputs the transcribed conversation into a generative AI (generative artificial intelligence) 12, a product database 18, and a vectorization device 20. The type of generative AI 12 is arbitrary; for example, any artificial intelligence capable of generating suggestions or scripts for the communicator using a random forest, neural network, etc. is used. Here, Amazon Bedlock Claude3 (a generative AI platform for using the generative AI model Claude provided by Antropic) provided by Amazon is used as the generative AI 12. The specific type of generative AI 12 is arbitrary. Also, "Bedlock" is a trademark of Amazon, and "Claude" is a trademark of Antropic.
[0023] The product database 18 stores data about products or services that System 1 recommends to customers. This data is entered, for example, from the administrator terminal 5, and includes product or service features, precautions, FAQs (answers to frequently asked questions), prohibited items, and other matters that should be explained in the contract with the customer.
[0024] The vectorization device 20 extracts keywords from the conversation and converts them into a sequence of keywords (a conversation vectorized, for example, sentence by sentence, with keywords as its components; hereinafter referred to as a vector). The conversation to be vectorized includes the utterances of customer 02, and preferably also includes the utterances of the communicator. The vectorization device 20 is, for example, Titan / Cohere provided by Amazon, but the specific type is arbitrary. Note that Titan / Cohere is a trademark of Amazon.
[0025] The vectorization device 20 vectorizes the conversation with customer 02 into two types: short sentences and long sentences. A short sentence is, for example, one sentence, but it is not limited to cases where it is grammatically one sentence; any sentence with a short interval between utterances, where there is a relationship between the spoken words, and where the number of words is within a predetermined range is considered a short sentence. If customer 02 utters multiple sentences in succession and the number of words in each sentence is large, the multiple sentences uttered in succession are considered a long sentence. Even if there are weak communicator utterances such as interjections between multiple sentences, in this embodiment it is considered a single long sentence. However, in such cases it may be processed as if there were multiple short sentences. The vectorized short sentences are mainly vectorized words with strong meaning (keywords) that were uttered. The vectorized long sentences contain many keywords, and the context that constitutes them as a whole is more meaningful than the individual keywords. The distinction between short sentences and long sentences is, for example, that an utterance of one sentence is a short sentence, while an utterance of multiple sentences with an upper limit of 5 to 10 sentences in length is a long sentence.
[0026] To support the communicator, a response script (story or suggestion) is displayed on monitor 8. Data for generating the script is stored in database 14, and based on this data, the generating AI 12 generates the response script. Database 14 is searchable by the virtual server 10 and stores short sentence vectors and long sentence vectors generated by the vectorization device 20. Furthermore, it stores short sentence vectors and long sentence vectors. Database 14 has three types of sub-databases: the initial talk database 15, the short sentence database 16, and the long sentence database 17.
[0027] Note that the short sentence database 16 and the long sentence database 17 may be treated as the same entity in memory. In this case, searching the database one sentence at a time will result in the short sentence database 16, while searching multiple sentences that are close together in the conversation will result in the long sentence database 17.
[0028] This section describes how to implement data (vector format) in databases 16 and 17. In one implementation method, a vector is generated in databases 16 and 17 by adding the customer's utterance to the response from the communicator terminal 4, and then adding a score from the administrator terminal 5. The score is an evaluation value given according to the outcome of the conversation; for example, a high score is given if the customer purchases a product or service. Alternatively, the communicator may assign a score to their own utterance, or the administrator may assign a score. A score is assigned to each pair of customer utterances and the communicator's response, and the generating AI 12 refers to the vector with the highest score to generate suggestions for the communicator.
[0029] In other implementation methods, the administrator terminal 5, etc., creates a set of customer utterances and exemplary responses to them, each assigned an initial score (in this case, a high score). These are then generated in both short and long sentence formats and stored in databases 16 and 17. In either case, the scores for the response scripts in databases 16 and 17 are changed along with the operation of system 1, so the result is essentially the same.
[0030] Figure 2 shows the processing related to initial talk. To construct the initial talk database 15, for example, the administrator terminal 5 vectorizes exemplary initial talk scripts (initial talk scripts) for each customer attribute (customer information, which is an abstraction of customer information by attributes) and inputs them into the database 15. Instead of inputting exemplary initial talk scripts, for example, the initial talk actually spoken from the communicator terminal 4 may be vectorized, and a score based on customer attributes and the conversation result may be added and stored in the database 15.
[0031] The generating AI 12 extracts initial talk scripts from the database 15, for example, those with the same customer attributes, or those with the same customer attributes and a high score, and displays them on the monitor 8. The degree to which the displayed script closely resembles the actual spoken text, or whether it is closer to just keywords, is optional. Based on the customer attributes and script displayed on the monitor 8, the communicator informs the customer of the purpose of the call and obtains permission to explain products and services. Once permission is obtained, the communicator stimulates needs by briefly introducing products and services.
[0032] The system displays high-scoring initial phrases selected based on customer attributes on Monitor 8, allowing communicators to use these as a reference when speaking to customers, making it easier to engage them in conversation. Furthermore, it enables even inexperienced communicators to produce above-average initial phrases, regardless of their experience or skills.
[0033] Once the initial conversation is successful (and permission is granted to explain the product or service in more detail), the process moves on to generating suggestions based on the following short sentences (suggestions from AI12, displayed on monitor 8, and supporting the communicator's conversation; equivalent to a script) and then to generating suggestions based on long sentences.
[0034] Figure 3 shows the processing related to short sentences. To construct database 16, text data representing the content of the conversation with the customer (customer's utterances and communicator's utterances) is input from STT6 to virtual server 10, and the results related to that conversation (evaluation of contract success or failure as a score) is input from administrator terminal 5 to virtual server 10. Note that inputting the score is optional. The input text data is combined with the results, classified into short sentences and long sentences, and vectorized by vectorization device 20. The vector data related to short sentences is stored in database 16. Similarly, data related to long sentences is stored in database 17 for long sentences (see Figure 4).
[0035] During conversations with customers, the customer's utterances (and preferably the communicator's utterances as well) are transcribed into text and input as vector data into the Generating AI 12. In addition to the input text data, the Generating AI 12 knows the customer's attributes and past conversations as context (data referenced as background knowledge). It extracts responses (communicator's responses to the customer) with high scores from the database 16 that are highly similar to the input text, generates a response script for the customer, and displays it on the monitor 8. The communicator answers the customer's questions and concerns by referring to the display on the monitor 8. The communicator can also input responses to the response script displayed on the monitor 8, such as ○ (increase score), no response (do not change score), or × (decrease score).
[0036] The data in database 16 consists of vectorized short sentences from past conversation examples, for example, with scores assigned based on the results, and stored in memory. The generating AI 12 then extracts past successful cases (cases with high scores) from conversations with similar vectors and contexts related to the customer's utterances, generates a response (things to say to the customer), and displays it on monitor 8. Alternatively, it can generate a response based on the generating AI 12's learning results for similar vectors and similar contexts, without using scores. In this case, even if the things to be explained to the customer are complex and require a lot of knowledge, the generating AI 12 can refer to the product database 18 and generate an accurate response in real time. This allows even inexperienced communicators to have conversations on par with veteran communicators.
[0037] Figure 4 shows the processing for long sentences, and is the same as the processing in Figure 3, except for the following points. Firstly, the conversations targeted are conversations with a large amount of information (many words) from the customer, and the criteria for generating the optimal response from the communicator to the customer are more complex. In particular, it is necessary to understand the customer's intentions and requests even when they are not explicitly stated and to provide accurate answers in real time. Furthermore, it is necessary to generate better answers by making suggestions that the customer will like or by conducting interviews.
[0038] Understanding customer needs and making appropriate suggestions requires insight into customer concerns and questions, as well as detailed knowledge of products or services. Furthermore, proceeding to a contract with a customer requires detailed explanations, including explanations of important matters. In such cases, Generative AI12 can generate appropriate responses by referencing high-scoring similar vector data. Thus, even communicators with insufficient skills can respond flexibly to customers with the support of Generative AI12, reducing problems caused by carelessness or lack of product knowledge.
[0039] The processes shown in Figures 2-4 are summarized in Figure 5. A simplified example of a conversation is also shown in Figure 6. The customer's attributes are, for example, a 67-year-old male, residing in Tokyo and working for a company. In Figure 6, the symbols i and i' represent initial talk, s and s' represent short sentences, and L and L' represent long sentences. The symbol ' indicates that it is the customer's utterance.
[0040] In the initial-word conversation scene in Figure 6, the initial-word conversation is tailored to the customer's attributes. In the subsequent conversation, in response to the customer's question, "Isn't diabetes a problem?", the communicator refers to the short-sentence database 16 and explains that enrollment is possible if treatment is continued, but in that case a doctor's certificate of continued treatment is required, and the premium rate will increase by 2%. For a communicator with limited product knowledge, it would be difficult to explain in real time and appropriately that enrollment is possible even with diabetes if treatment is continued, and what the additional conditions are in that case.
[0041] Regarding the long sentence, the generating AI 12 determines from the utterance, "I have several insurance policies, and the premiums must be high," that the customer is interested in insurance but concerned about the premiums, and suggests recommending term life insurance. If the conversation proceeds to the conclusion of an insurance contract, there is much for the communicator to explain, such as detailed contract contents and important matters. The generating AI 12 supports the communicator in these matters and provides appropriate suggestions to monitor 8. At the end of the conversation, it is important to understand the customer's potential questions and concerns, as this helps prevent cancellations later. These questions and concerns can often be inferred from the customer's overall utterance, so the generating AI 12 refers to the long sentence database 17 and generates a script that addresses the customer's questions and concerns, which is displayed on monitor 8 of the communicator's terminal 4. Also, if the customer makes a short utterance during the conversation, the short sentence is processed and responded to appropriately. [Explanation of Symbols]
[0042] 02 Customers 1. Outbound Call System 2. CTI (Computer Telephony Integration System) 4 Communicator terminals 5. Administrator terminal 6. STT (Text Conversion Device) 8,9 Monitor 10 Virtual Servers 12. Generative AI (Generative Artificial Intelligence) 14 Databases 15 Initial Talk Database 16 Short Sentence Database 17 Long Sentence Database 18 Product Database 20 Vectorization device
Claims
1. A method of making outbound calls from a call center to customers and recommending the purchase of products or services, A CTI (Computer Telephony Integration) device 2 makes a call to the customer's communication terminal and displays the customer's attribute information on the communicator terminal 4, STT (Text Conversion Device) 6 converts the customer's spoken audio into text, Administrator terminal 5, server 10, generating AI 12, vectorization device 20 that converts text-based voice data into vector data, An initial-word database 15 stores initial-word phrases from conversations with customers along with customer attribute information. A short sentence database 16 stores vectorized data of relatively short utterances from customers, along with a script of the response with a score. A long-sentence database 17 stores vectorized data of relatively long utterances from customers, along with a script of the response with a score. An outbound call system equipped with a product database 18 of recommended products or services, By generating AI12, Based on the customer's attribute information, the initial talk database 14 is used to generate an initial talk script for the customer. For relatively short customer utterances, a response script is generated using the short sentence database 15. For relatively long customer utterances, a response script is generated using the long sentence database 16. Display the generated script on communicator terminal 4. Outbound calling methods.
2. An outbound call method according to claim 1, characterized in that it accepts changes in the score from the communicator terminal 4 to the response scripts for short sentences and long sentences displayed on the communicator terminal 4.
3. A system for making outbound calls from a call center to customers and recommending the purchase of products or services, A CTI (Computer Telephony Integration) device 2 makes a call to the customer's communication terminal and displays the customer's attribute information on the communicator terminal 4, STT (Text Conversion Device) 6 converts the customer's spoken audio into text, Administrator terminal 5, server 10, generating AI 12, vectorization device 20 that converts text-based voice data into vector data, An initial-word database 15 stores initial-word phrases from conversations with customers along with customer attribute information. A short sentence database 16 stores vectorized data of relatively short utterances from customers, along with a script of the response with a score. A long sentence database 17 stores vectorized data of relatively long utterances from customers, along with a script of the response with a score. It includes a product database 18 of recommended products or services, The generated AI12 is, Based on the customer's attribute information, the initial talk database 14 is used to generate an initial talk script for the customer. For relatively short customer utterances, a response script is generated using the short sentence database 15. For relatively long customer utterances, a response script is generated using the long sentence database 16. Server 10 is configured to display the generated script on the communicator terminal 4. Outbound call system.