Call program, call system, and call method
The call system automates telemarketing by using AI to generate and switch speakers based on voice inputs, addressing efficiency limitations and delays in reaching intended call recipients.
Patent Information
- Application Number
- JP2025050049
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2044-09-11
AI Technical Summary
Telemarketing efficiency is limited by human capacity, and it takes time to reach intended call recipients, especially when receptionists transfer calls to senior executives.
A call system and method that includes a processor to make calls, acquire recipient voices, generate speech based on voice content, and switch speakers when certain conditions are met, using AI and rule-based systems to handle calls efficiently.
Enhances call efficiency by automating the process from initial contact to connecting with intended recipients, reducing time delays and improving human interaction efficiency.
Smart Images

Figure 0007758402000001_ABST
Abstract
Description
[Technical Field]
[0001] At least one embodiment of the present invention relates to a call program, a call system, and a call method. [Background technology]
[0002] Patent Document 1 describes an automatic call device that notifies a voice message. The automatic call device causes a computer device to execute the following functions: a function of acquiring a call recipient list, a function of selecting a call recipient from the call recipient list, a function of calling the call recipient, a function of providing a voice message when the call recipient is called, and a function of selecting another call recipient from the call recipient list when the call to the call recipient is unavailable. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2021-034757 Summary of the Invention [Problem to be solved by the invention]
[0004] Telemarketing is one sales method that utilizes the telephone. A human telephone appointment maker refers to a pre-prepared list of people to call, such as a customer list, and calls the people in order to set up appointments. With traditional methods, there is a limit to the amount of human activity that can be done, so the results that each person can achieve are limited.
[0005] Furthermore, if the intended call recipient is a company employee or a senior executive, that person may not always answer the phone first. For example, the company's receptionist may answer the call first and then transfer the call to the person in charge or senior executive. For this reason, it used to take a long time for the telephone appointment maker to start talking to the intended call recipient.
[0006] An object of at least one embodiment of the present invention is to provide a calling program, a calling system, and a calling method that solve the above-mentioned problems and improve the efficiency of calling a predetermined call recipient. [Means for solving the problem]
[0007] From a non-limiting perspective, a call program according to one embodiment of the present invention causes a processor to realize a call function for making a call to a recipient terminal that is the target of a call, a recipient voice acquisition function for acquiring a recipient voice input to the recipient terminal, a speech voice generation function for generating a speech voice to be transmitted to the recipient terminal based on the content of the recipient voice, a speech function for transmitting the speech voice to the recipient terminal, and a speaker switching function for switching the input channel of the speech voice to be transmitted to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies a predetermined condition.
[0008] From a non-limiting perspective, a call system according to one embodiment of the present invention is a call system comprising one or more processors, which implement a call function for making a call to a recipient terminal to be called, a recipient voice acquisition function for acquiring a recipient voice input to the recipient terminal, a speech voice generation function for generating a speech voice to be transmitted to the recipient terminal based on the content of the recipient voice, a speech function for transmitting the speech voice to the recipient terminal, and a speaker switching function for switching the input channel of the speech voice to be transmitted to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies a predetermined condition.
[0009] From a non-limiting perspective, a calling method according to one embodiment of the present invention is a calling method by a device equipped with a processor, and includes a calling step of making a call to a recipient terminal to be called, a recipient voice acquisition step of acquiring a recipient voice input to the recipient terminal, a voice generation step of generating a voice to be transmitted to the recipient terminal based on the content of the recipient voice, a speaking step of transmitting the voice to the recipient terminal, and a speaker switching step of switching the input channel of the voice to be transmitted to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies a predetermined condition. [Effects of the Invention]
[0010] Each embodiment of the present application addresses one or more of the deficiencies. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a block diagram showing an example of the configuration of a call system corresponding to at least one of the embodiments of the present invention. [Figure 2] FIG. 1 is a block diagram showing a configuration of a server according to at least one of the embodiments of the present invention. [Figure 3] 1 is a flowchart showing an example of processing of a call system corresponding to at least one of the embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, examples of embodiments of the present invention will be described with reference to the drawings. Note that the various components in the examples of the embodiments described below can be combined as appropriate to the extent that no inconsistencies or the like arise. Furthermore, content described as an example of one embodiment may be omitted in other embodiments. Furthermore, the content of operations and processes unrelated to the characteristic parts of each embodiment may be omitted. Furthermore, the order of various processes constituting the various flows and sequences described below may be in any order to the extent that no inconsistencies or the like arise in the processing content.
[0013] The following description will be given taking as an example a call program executed on a server, which is an example of a computer. However, the computer may be another device, such as a user terminal. Also, the call system as a whole may execute the call program.
[0014] Fig. 1 is a block diagram showing an example of the configuration of a call system corresponding to at least one embodiment of the present invention. The call system 1 comprises a server 10 and a user terminal 20 used by users of the call system 1. User terminals 20A, 20B, and 20C are each an example of the user terminal 20. The configuration of the call system 1 is not limited to this. For example, the call system 1 may be configured so that multiple users use a single user terminal. The call system 1 may also comprise multiple servers.
[0015] The server 10 and the user terminal 20 are examples of devices. The server 10 and the user terminal 20 are each communicatively connected to a communication network 30. The connection between the communication network 30 and the server 10, and the connection between the communication network 30 and the user terminal 20 may be a wired connection or a wireless connection. The communication network 30 may be an internet line, an analog telephone line, or a line that combines these.
[0016] The call system 1 includes a server 10 and a user terminal 20, and thereby realizes various functions for executing various processes in response to user operations.
[0017] The server 10 includes a processor 11, a memory 12, and a storage device 13. The processor 11 is, for example, a central processing unit such as a CPU (Central Processing Unit) that performs various calculations and controls. If the server 10 includes a GPU (Graphics Processing Unit), some of the calculations and controls may be performed by the GPU. The server 10 uses the data read into the memory 12 to execute various information processing operations using the processor 11, and stores the obtained processing results in the storage device 13 as necessary.
[0018] The storage device 13 functions as a storage medium for storing various types of information. The configuration of the storage device 13 is not particularly limited, but may be configured to store all of the various types of information necessary for the control performed by the call system 1, from the viewpoint of reducing the processing load on the user terminal 20. Examples of such a configuration include an HDD and an SSD. However, the storage device that stores the various types of information only needs to have a storage area accessible by the server 10, and may be configured to have a dedicated storage area outside the server 10, for example.
[0019] The user terminal 20 is managed by a user. The term "user" here refers to both the telephone appointment maker and the call recipient, who is the person the telephone appointment maker calls. Examples of the user terminal 20 include a fixed telephone terminal, a mobile phone terminal, a smartphone, a personal computer, and a tablet. The user terminal 20 may be a speaker terminal or a receiver terminal, which will be described later.
[0020] The user terminal 20 may be equipped with hardware and software for connecting to the communication network 30 and performing various processes by communicating with the server 10. The multiple user terminals 20 may also be configured to be able to communicate directly with each other without going through the server 10.
[0021] The user terminal 20 may include a processor 21, a memory 22, and a storage device 23. The processor 21 is, for example, a central processing unit such as a CPU (Central Processing Unit) that performs various calculations and controls. Furthermore, if the user terminal 20 includes a GPU (Graphics Processing Unit), some of the various calculations and controls may be performed by the GPU. The user terminal 20 uses the data read into the memory 22 to execute various information processing in the processor 21, and stores the obtained processing results in the storage device 23 as necessary. The storage device 23 functions as a storage medium that stores various types of information.
[0022] 2 is a block diagram showing the configuration of a server corresponding to at least one of the embodiments of the present invention. Server 10 includes calling unit 101, receiver voice acquisition unit 102, speech sound generation unit 103, speaking unit 104, and speaker switching unit 105. The processor included in server 10 refers to a call program stored in a storage device and executes the program to functionally realize calling unit 101, receiver voice acquisition unit 102, speech sound generation unit 103, speaking unit 104, and speaker switching unit 105.
[0023] The calling unit 101 has a function of making a call to the recipient terminal to be called. The recipient voice acquisition unit 102 has a function of acquiring the recipient voice input to the recipient terminal. The speech generation unit 103 has a function of generating the speech to be transmitted to the recipient terminal based on the content of the recipient voice. The speech unit 104 has a function of transmitting the speech to the recipient terminal. The speaker switching unit 105 has a function of switching the input channel of the speech to be transmitted to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies a predetermined condition.
[0024] The recipient terminal is a terminal that receives a call from a telephone appointment point. For example, a company recipient terminal may be a fixed telephone or a mobile phone terminal. The recipient terminal is an example of the above-mentioned user terminal 20. The telephone number of the recipient terminal to be called may be specified by a user of the call system 1. A list containing the telephone numbers of recipient terminals to be called may be stored in advance in memory 12, and the calling unit 101 may obtain the telephone number of the recipient terminal to be called from the list.
[0025] The receiver's voice input to the receiver's terminal may be converted into digital data before it reaches the server 10 from the receiver's terminal.
[0026] [Speech generation] When the call recipient is a person in charge or a manager at a company, the person in charge may not always answer the phone. For example, a company receptionist may temporarily receive the call and then transfer the call to the person in charge or manager who is the call recipient. For this reason, in the past, it took a long time for a telephone appointment maker to start a conversation with the intended call recipient. In an embodiment of the present disclosure, in order to improve work efficiency, the call system 1, rather than a human telephone appointment maker, handles the call from the time the call is placed to the intended call recipient until the call is connected to the intended call recipient (such as the person in charge or a manager). Therefore, the call system 1 generates a speech to respond to the recipient until the call is transferred to the telephone appointment maker.
[0027] The speech generation unit 103 generates speech to be transmitted to the listener terminal based on the content of the listener voice acquired from the listener terminal.
[0028] The speech generation unit 103 may generate speech using AI. For example, the speech generation unit 103 includes a large-scale language model (LLM). The speech generation unit 103 inputs the recipient's speech or converted information obtained by converting the recipient's speech into the LLM, thereby generating speech to be transmitted to the recipient terminal.
[0029] The LLM may generate text data to be transmitted to the receiver terminal. The speech generation unit 103 converts the text data into speech. The LLM may directly generate speech to be transmitted to the receiver terminal.
[0030] The speech generation unit 103 may generate speech using rule-based AI. In this case, the speech generation unit 103 analyzes the content of the receiver's speech acquired from the receiver's terminal. For example, the speech generation unit 103 calculates the similarity between text data obtained by converting the receiver's speech into text and a receiver's utterance scenario pre-registered in memory 12. For example, if the text data obtained by converting the receiver's speech into text is "This is XX Co., Ltd.", the text data has a high similarity to the receiver's utterance scenario "Yes, this is XX Co., Ltd." registered in memory 12. The speech generation unit 103 extracts from memory 12 a scenario that is highly similar to the text data obtained by converting the receiver's speech into text, and identifies a message stored in memory 12 in association with the extracted scenario as the speech text. Such a correspondence between the receiver's utterance scenario and the message as the speech text may be pre-stored in memory 12 in the form of a table or the like.
[0031] The speech generation unit 103 may determine the content of the speech based on a conditional branch implemented in the rule-based AI. For example, the speech generation unit 103 converts the speech of the listener into text and determines whether the text of the listener's speech contains a predetermined keyword. The content of the speech to be generated is changed depending on whether the predetermined keyword is included.
[0032] The speech voice does not have to be generated from scratch, but may be generated by combining parts of speech voices that have been recorded in advance according to the content of the received voice.
[0033] The speech unit 104 transmits the speech to the receiver's terminal. From the time the call is made to the call recipient until the call is connected to the call recipient (such as the person in charge or a manager mentioned above), the speech unit 104 transmits the speech generated by the speech generation unit 103 to the receiver's terminal. On the other hand, after the call is connected to the call recipient, the speech unit 104 transmits the speech uttered by the user of the call recipient's terminal to the receiver's terminal. The call recipient's terminal is, for example, a terminal used by a human telephone appointment maker. The speaker switching unit 105 switches the speaker on the call recipient's side.
[0034] The speaker switching unit 105 has a function of switching the input channel of the uttered voice to be transmitted to the receiver terminal to the user of the speaker terminal when the receiver voice satisfies a predetermined condition.
[0035] The predetermined condition for the speaker switching unit 105 to switch the input channel may be, for example, the following condition. (Condition 1) It must be possible to determine that the person inputting the receiver's voice is the person to whom the call is being made. (Condition 2) It can be estimated that the person inputting the receiver's voice will change from a person who is not the call target to a person who is the call target.
[0036] The above-mentioned (Condition 1) will be explained in more detail. For the sake of convenience, the call target will be denoted as A. Call target A is, for example, the manager of a company. A person who is not the call target will be denoted as B. Call target B is, for example, the receptionist of a company.
[0037] If the receiver's voice contains a keyword that can be used to determine that the person currently answering the phone is call recipient A, such as "Hello, this is A," the speaker switching unit 105 determines that predetermined condition 1 is met and switches the input channel of the spoken voice to the user of the speaker's terminal. Furthermore, voice characteristics of the call recipient may be registered in advance in the memory 12 of the server 10. The speaker switching unit 105 may compare the voice characteristics included in the receiver's voice with the voice characteristics of the call recipient registered in the memory 12, and determine that the person inputting the receiver's voice is the call recipient. The speaker switching unit 105 may combine the keyword-based determination and the voice characteristic comparison determination to determine that the person inputting the receiver's voice is the call recipient.
[0038] The above-mentioned (Condition 2) will be explained in more detail. When the receiver's voice contains a predetermined keyword, such as "Please wait a moment" or "I'll transfer you to A," which indicates that the person on the other end of the line will change from a person who is not the call target to a call target, the speaker switching unit 105 determines that the predetermined condition 2 is met and switches the input channel of the spoken voice to the user of the speaker's terminal. The predetermined keyword may be stored in memory 12 in advance.
[0039] When the person answering the phone changes from person B, who is not the call target, to call target A, a telephone hold tone may be played during that time. Therefore, when the receiver's voice contains the hold tone, the speaker switching unit 105 determines that predetermined condition 2 is satisfied and switches the input channel of the voice to the user of the speaker's terminal.
[0040] FIG. 3 is a flowchart showing an example of processing of a call system corresponding to at least one of the embodiments of the present invention.
[0041] The calling unit 101 makes a call to the receiver terminal to be called (St11). The receiver voice acquiring unit 102 acquires the receiver voice input to the receiver terminal (St12).
[0042] The speaker switching unit 105 determines whether the receiver's voice satisfies a predetermined condition (St13). If the receiver's voice satisfies the predetermined condition (St13: YES), the process proceeds to step St16. If the receiver's voice does not satisfy the predetermined condition (St13: NO), the process proceeds to step St14.
[0043] In step St14, the speech generation unit 103 generates a speech to be transmitted to the listener terminal based on the content of the listener's speech. The speech unit 104 transmits the speech to the listener terminal (St15). Thereafter, the process returns to step St12, and the next listener's speech is acquired from the listener terminal.
[0044] In step St16, the speaker switching unit 105 switches the input channel of the uttered voice to be transmitted to the receiver terminal to the user of the speaker terminal. After that, the telephone appointment maker, who is the user of the speaker terminal, starts a conversation with the receiver.
[0045] An example of a telephone conversation using the call system 1 according to an embodiment of the present disclosure will be shown below. [Conversation example 1] In Conversation Example 1, the intended call recipient A answers the phone at the receiver's terminal from the beginning. Therefore, the receiver's voice acquired in step St12 includes the utterance "Hello, this is A." Then, since the receiver's voice satisfies the above-mentioned (Condition 1), the process proceeds from step St13 to step St16. The speaker switching unit 105 switches the input channel of the uttered voice to be transmitted to the receiver's terminal to the user of the speaker's terminal. In other words, from the beginning, the speaker's side will be answered by a human user such as a telephone appointment maker.
[0046] [Example conversation 2] In conversation example 2, initially, person B, who is not the intended call recipient, answers the phone at the receiver's terminal. The receiver's voice acquired in step St12 includes the voice, "Hello. This is XX Co., Ltd." Because the receiver's voice does not satisfy either (Condition 1) or (Condition 2) above, the process transitions from step St13 to step St14. That is, the call system 1 answers the phone with person B, who is not the intended call recipient. For example, the speech generation unit 103 generates a speech such as, "Thank you for your help. Is Mr. A, the manager of the sales department, here?" The speech unit 104 transmits the generated speech to the receiver's terminal.
[0047] Person B, who is not the intended call recipient, says into the phone, "A, right? I'll take over, so please wait a moment," and switches the phone to hold mode. Then, the next receiver voice acquired in step St12 satisfies the above-mentioned (Condition 1) and (Condition 2), so the process transitions from step St13 to step St16. The speaker switching unit 105 switches the input channel of the uttered voice to be transmitted to the receiver's terminal to the user of the speaker's terminal. In other words, when person B, who is not the intended call recipient, switches the phone to hold mode, the entity handling the call on the speaker's side switches from the call calling system 1 to a human user such as a telephone appointment maker.
[0048] As described above, each embodiment of the present application solves one or more deficiencies. Note that the effects of each embodiment are non-limiting effects or examples of effects.
[0049] In each of the above-described embodiments, the user terminal 20 and the server 10 execute the above-described various processes in accordance with various control programs (e.g., a call program) stored in their own storage devices. Also, other computers, not limited to the user terminal 20 and the server 10, may execute the above-described various processes in accordance with various control programs (e.g., a call program) stored in their own storage devices.
[0050] Furthermore, the configuration of the call system 1 is not limited to the configuration described as an example of the embodiment above. For example, the server may execute some or all of the processes described as processes executed by the user terminal, or the user terminal may execute some or all of the processes described as processes executed by the server. Furthermore, the user terminal may be configured to have some or all of the storage unit (storage device) that the server has. In other words, the user terminal or the server in the call system 1 may be configured to have some or all of the functions that the other one has.
[0051] Furthermore, the program may be configured to cause a part or all of the functions described as examples of each of the above-mentioned embodiments to be realized by a single device that does not include a communication network.
[0052] [Note] The above-described embodiments have been described so that at least the following invention can be implemented by a person having ordinary skill in the art to which the invention pertains.
[0053] [1] The processor A call function for making a call to a call recipient terminal; a receiver voice acquisition function for acquiring a receiver voice input to the receiver-side terminal; a speech generation function that generates a speech to be transmitted to the receiver terminal based on the content of the receiver voice; a speech function for transmitting the speech to the receiver terminal; a speaker switching function that switches the input channel of the speech voice to be transmitted to the receiver terminal to the user of the speaker terminal when the receiver voice satisfies a predetermined condition; A phone call program that makes this possible.
[0054] According to the above-described call program, it is possible to increase the efficiency of calls to predetermined call recipients.
[0055] [2] The predetermined condition is It is possible to determine that the person who inputs the receiver's voice is the call target, [1] The call program described in [1]. This allows the caller to switch from the call system to a human answering the phone when the intended call recipient answers.
[0056] [3] The predetermined condition is It can be estimated that the input person of the receiver's voice changes from a person who is not the call target to a call target, [1] The call program described in [1].
[0057] [4] The predetermined condition is The listener voice includes a hold tone. [3] The call program described in [3].
[0058] According to these calling programs, if someone other than the intended call recipient answers the phone first, the call recipient on the caller's side can be switched from the call system to a human at the same time that the call recipient is replaced by the intended call recipient.
[0059] [5] A calling system comprising one or more processors, the processor, A call function for making a call to a call recipient terminal; a receiver voice acquisition function for acquiring a receiver voice input to the receiver-side terminal; a speech generation function that generates a speech to be transmitted to the receiver terminal based on the content of the receiver voice; a speech function for transmitting the speech to the receiver terminal; a speaker switching function that switches the input channel of the speech voice to be transmitted to the receiver terminal to the user of the speaker terminal when the receiver voice satisfies a predetermined condition; A call system that makes this possible.
[0060] According to the above-described call system, it is possible to increase the efficiency of calls to predetermined call recipients.
[0061] [6] A calling method by a device including a processor, a calling step of calling a call destination terminal of a call recipient; a receiver voice acquisition step of acquiring a receiver voice input to the receiver-side terminal; a speech generation step of generating a speech to be transmitted to the listener terminal based on the content of the listener voice; a speaking step of transmitting the uttered voice to the receiver terminal; a speaker switching step of switching an input channel of the uttered voice to be transmitted to the receiver terminal to a user of the speaker terminal when the receiver voice satisfies a predetermined condition; A calling method comprising:
[0062] According to the above-described calling method, it is possible to increase the efficiency of calling predetermined call recipients. [Industrial Applicability]
[0063] According to one embodiment of the present invention, it is useful as a calling program, a calling system, and a calling method that improve the efficiency of calling a predetermined call recipient. [Explanation of symbols]
[0064] 1. Call system 10 Servers 11 processors 12 Memory 13 Storage device 20, 20A, 20B User terminal 21 processors 22 Memory 23 Storage device 30 Communication Network 101 Calling Unit 102 Receiver voice acquisition unit 103 Speech generation unit 104 Speech Unit 105 Speaker switching unit
Claims
1. A program for telemarketing, The processor: When a predetermined condition is satisfied, the condition being that it can be determined that the person inputting the receiver voice input to the receiver terminal that is the call target is the call target, or the condition being that it can be estimated that the person inputting the receiver voice will change from a person who is not the call target to a call target, the input channel of the speech voice to be transmitted to the receiver terminal is switched to the user of the speaker terminal. program.
2. The processor: Transmitting a speech voice based on the receiver's voice input to the receiver's terminal to be called, to the receiver's terminal; The program according to claim 1.
3. The predetermined condition is: The listener voice includes a hold tone. The program according to claim 1 or 2.
4. A program as described in claim 1 or claim 2, wherein the specified condition is that the listener's voice contains a keyword.
5. A program as described in claim 4, wherein the keyword is a keyword that can determine that the person receiving the call is the person to whom the call is being made.
6. A program as described in claim 4, wherein the keyword is a keyword that can be used to infer that the person answering the phone will change from someone who is not the person being called to someone who is the person being called.
7. The predetermined condition is based on voice characteristics contained in the listener's voice. The program according to claim 1.
8. A system for telemarketing, comprising one or more processors, the processor: A system that switches the input channel of the spoken voice to be sent to the receiver's terminal to the user of the speaker's terminal when a predetermined condition is met, which is a condition that can determine that the inputter of the receiver's voice input to the receiver's terminal that is the call target is the call target, or a condition that can be estimated that the inputter of the receiver's voice will change from a person who is not the call target to a call target.
9. 1. A method for telemarketing by an apparatus including a processor, comprising: The method includes a step of switching the input channel of the speech voice to be transmitted to the receiver terminal to the user of the speaker terminal when a predetermined condition is met, the condition being a condition that can determine that the inputter of the receiver voice input to the receiver terminal to be the call target is the call target, or a condition that can be estimated that the inputter of the receiver voice will change from a person who is not the call target to a call target.
Citation Information
Patent Citations
Line connection changeover device
JP1990030269A
Operator's operation support system for call center
JP2007004000A
Automatic call device and automatic call method
JP2021034757A
Information processing system, information processing method, and information processing program
JP2022032967A
JPP7659936B