Calling program, calling system, and calling method
The calling system automates telemarketing interactions, efficiently connecting calls to intended recipients and switching to human operators when conditions are met, addressing limitations of conventional methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2026-03-24
AI Technical Summary
Conventional telemarketing methods are limited by human activity, and it takes a long time to establish a conversation with intended power supply targets, especially when receptionists or others initially answer calls intended for company officers.
A calling system with a processor that initiates calls, acquires recipient voice input, generates speech based on that input, and switches the voice channel to a human operator when predetermined conditions are met, using AI and rule-based systems to streamline interactions.
Enhances the efficiency of reaching intended targets by automating interactions from call initiation to connection, allowing seamless transitions to human operators when conditions are met.
Smart Images

Figure 2026052620000001_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment of the present invention relates to a power supply program, a power supply system, and a power supply method.
Background Art
[0002] Patent Document 1 describes an automatic power supply device that notifies a voice message. The automatic power supply device causes a computer device to execute functions of acquiring a power supply target list, selecting a power supply target from the power supply target list, making a call to the power supply target, providing a voice message when the power supply target is called, and selecting another power supply target from the power supply target list when a call to the power supply target is unavailable.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Telemarketing is one of the business methods that utilize telephones. A human telephone pointer refers to a prepared power supply target list such as a customer list, makes phone calls to the power supply targets in order, and makes appointments. In the conventional method, since there is a limit to the amount of human activity, the results that can be achieved per person are limited.
[0005] Also, when the intended power supply target is a person in charge or an officer in a company, etc., the person himself / herself of the power supply target does not always answer the phone first. For example, the company's reception may temporarily answer the phone and transfer it to the person in charge or officer who is the power supply target. Therefore, conventionally, it has taken a long time for the telephone pointer to start a conversation with the power supply target who is the intended party.
[0006] An object of at least one embodiment of the present invention is to provide a calling program, a calling system, and a calling method that solve the above problems and improve the efficiency of calling predetermined target persons. [Means for solving the problem]
[0007] In a non-limiting view, a calling program according to one embodiment of the present invention provides the processor with a calling function for making a call to a recipient terminal to be called; a recipient voice acquisition function for acquiring recipient voice input to the recipient terminal; a speech voice generation function for generating speech voice to be sent to the recipient terminal based on the content of the recipient voice; a speaking function for sending the speech voice to the recipient terminal; and a speaker switching function for switching the input channel of the speech voice to be sent to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies predetermined conditions.
[0008] In a non-limiting view, a calling system according to one embodiment of the present invention is a calling system comprising one or more processors, wherein the processors implement: a calling function for making a call to a recipient terminal to be called; a recipient voice acquisition function for acquiring recipient voice input to the recipient terminal; a speech voice generation function for generating speech voice to be transmitted to the recipient terminal based on the content of the recipient voice; a speaking function for transmitting the speech voice to the recipient terminal; and a speaker switching function for switching the input channel of the speech voice to be transmitted to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies predetermined conditions.
[0009] In a non-limiting view, a calling method according to one embodiment of the present invention is a calling method using a device equipped with a processor, comprising: a calling step of making a call to a receiver terminal to be called; a receiver voice acquisition step of acquiring receiver voice input to the receiver terminal; a speech voice generation step of generating speech voice to be transmitted to the receiver terminal based on the content of the receiver voice; a speaking step of transmitting the speech voice to the receiver terminal; and a speaker switching step of switching the input channel of the speech voice to be transmitted to the receiver terminal to the user of the speaker terminal when the receiver voice satisfies predetermined conditions. [Effects of the Invention]
[0010] Each embodiment of the present invention resolves one or more of the shortcomings. [Brief explanation of the drawing]
[0011] [Figure 1] A block diagram showing an example of the configuration of a power supply system corresponding to at least one embodiment of the present invention. [Figure 2] A block diagram showing the configuration of a server corresponding to at least one embodiment of the present invention. [Figure 3] A flowchart showing an example of a power supply system process corresponding to at least one embodiment of the present invention. [Modes for carrying out the invention]
[0012] Hereinafter, examples of embodiments of the present invention will be described with reference to the drawings. The various components in each embodiment described below can be combined as appropriate, provided that no inconsistencies arise. Furthermore, some aspects described in one embodiment may be omitted in other embodiments. Also, operations and processes unrelated to the characteristic features of each embodiment may be omitted. Moreover, the order of the various processes constituting the various flows and sequences described below is not necessarily in any particular order, provided that no inconsistencies arise in the processing content.
[0013] The following explanation uses a call-making program executed on a server, which is an example of a computer, as an example. However, the computer may be other devices such as a user terminal. Alternatively, the entire call-making system may execute the call-making program.
[0014] Figure 1 is a block diagram showing an example configuration of a call system corresponding to at least one embodiment of the present invention. The call system 1 comprises a server 10 and user terminals 20 used by users of the call system 1. User terminals 20A, 20B, and 20C are examples of user terminals 20. The configuration of the call system 1 is not limited thereto. For example, the call system 1 may be configured so that a single user terminal is used by multiple users. The call system 1 may also comprise multiple servers.
[0015] Server 10 and user terminal 20 are examples of devices. Server 10 and user terminal 20 are each connected to a communication network 30 in a communication-enabled manner. The connection between the communication network 30 and server 10, and the connection between the communication network 30 and user terminal 20, may be a wired or wireless connection. The communication network 30 may be an internet line, an analog telephone line, or a line combining these.
[0016] The telephone system 1, comprising a server 10 and a user terminal 20, realizes various functions for executing various processes in response to user operations.
[0017] Server 10 includes a processor 11, a memory 12, and a storage device 13. The processor 11 is a central processing unit such as a CPU (Central Processing Unit) that performs various operations and controls, for example. Also, when the server 10 includes a GPU (Graphics Processing Unit), a part of various operations and controls may be performed by the GPU. The server 10 executes various information processes using the data read into the memory 12 by the processor 11, and stores the obtained processing results in the storage device 13 as necessary.
[0018] The storage device 13 has a function as a storage medium for storing various information. The configuration of the storage device 13 is not particularly limited, but from the viewpoint of reducing the processing load on the user terminal 20, it may be configured to be able to store all the various information necessary for the control performed in the power supply system 1. Examples of such include HDDs and SSDs. However, the storage device for storing various information only needs to have a storage area in a state accessible by the server 10, and for example, it may be configured to have a dedicated storage area outside the server 10.
[0019] The user terminal 20 is managed by the user. The user here includes a telephone pointer and a called party who is the target of the call made by the telephone pointer. Examples of the user terminal 20 include, for example, a fixed telephone terminal, a mobile phone terminal, a smartphone, a personal computer, a tablet, and the like. The user terminal 20 may be the speaker-side terminal or the listener-side terminal described later.
[0020] The user terminal 20 may be connected to the communication network 30 and include hardware and software for executing various processes by communicating with the server 10. Each of the plurality of user terminals 20 may be configured to be able to communicate directly with each other without going through the server 10.
[0021] The user terminal 20 may include a processor 21, a memory 22, and a storage device 23. The processor 21 is, for example, a central processing unit such as a CPU (Central Processing Unit) that performs various operations and controls. Also, when the user terminal 20 includes a GPU (Graphics Processing Unit), a part of various operations and controls may be performed by the GPU. The user terminal 20 uses the data read into the memory 22 to execute various information processes by the processor 21, and stores the obtained processing results in the storage device 23 as needed. The storage device 23 has a function as a storage medium for storing various information.
[0022] FIG. 2 is a block diagram showing the configuration of a server corresponding to at least one embodiment of the present invention. The server 10 includes a calling unit 101, a receiver voice acquisition unit 102, a speaker voice generation unit 103, a speaker unit 104, and a speaker switching unit 105. The processor included in the server 10 refers to the connection program held in the storage device and executes the program to functionally realize the calling unit 101, the receiver voice acquisition unit 102, the speaker voice generation unit 103, the speaker unit 104, and the speaker switching unit 105.
[0023] The calling unit 101 has a function of making a call to the receiver-side terminal to be connected. The receiver voice acquisition unit 102 has a function of acquiring the receiver voice input to the receiver-side terminal. The speaker voice generation unit 103 has a function of generating a speaker voice for transmitting to the receiver-side terminal based on the content of the receiver voice. The speaker unit 104 has a function of transmitting the speaker voice to the receiver-side terminal. The speaker switching unit 105 has a function of switching the input channel of the speaker voice transmitted to the receiver-side terminal to the user of the speaker-side terminal when the receiver voice satisfies a predetermined condition.
[0024] The receiver terminal is the terminal that receives calls from the telephone appointment setter. For example, a company's receiver terminal may be a landline phone or a mobile phone. The receiver terminal is an example of the user terminal 20 described above. The user of the calling system 1 may specify the telephone number of the receiver terminal to be called. A list of telephone numbers of receiver terminals to be called is stored in memory 12 in advance, and the calling unit 101 may retrieve the telephone number of the receiver terminal to be called from the list.
[0025] The receiver's voice input to the receiver's terminal may be converted into digital data before it reaches the server 10.
[0026] [Generating spoken audio] When the target of a call is a person in charge or a manager at a company, the person in charge may not answer the phone. For example, a company receptionist may take the call temporarily and transfer it to the person in charge or a manager. Therefore, in the past, it took a long time for the telephone appointment setter to begin a conversation with the intended person. In the embodiment of this disclosure, in order to streamline operations, the call system 1, rather than a human telephone appointment setter, handles the interaction from the time the call is made to the target person (the person in charge or a manager mentioned above) until the call is connected to the intended person. Therefore, the call system 1 generates spoken audio to respond to the recipient until it transfers the call to the telephone appointment setter.
[0027] The speech generation unit 103 generates speech for transmission to the receiver terminal based on the content of the receiver's voice acquired from the receiver terminal.
[0028] The speech generation in the speech generation unit 103 may be performed using AI. For example, the speech generation unit 103 includes an LLM (Large-Scale Language Model). The speech generation unit 103 generates speech for transmission to the receiver terminal by inputting the receiver's voice or converted information obtained by converting the receiver's voice into the LLM.
[0029] The LLM may generate text data for transmission to the receiver terminal. The speech generation unit 103 converts the text data into speech. The LLM may also directly generate speech for transmission to the receiver terminal.
[0030] The speech generation unit 103 may generate speech using rule-based AI. In this case, the speech generation unit 103 analyzes the content of the receiver's voice obtained from the receiver's terminal. For example, it calculates the similarity between the text data obtained by transcribing the receiver's voice and the receiver's utterance scenarios that have been pre-registered in memory 12. For example, if the text data obtained by transcribing the receiver's voice is "This is XX Corporation," then this text data has a high similarity to the receiver's utterance scenario "Yes, this is XX Corporation" that has been registered in memory 12. The speech generation unit 103 extracts scenarios with a high similarity to the text data obtained by transcribing the receiver's voice from memory 12, and identifies the messages stored in memory 12 in association with the extracted scenarios as speech text. Such correspondences between receiver's utterance scenarios and messages as speech text may be pre-stored in memory 12 in the form of a table or the like.
[0031] The speech generation unit 103 may determine the content of the speech based on conditional branching implemented in the rule-based AI. For example, the speech generation unit 103 converts the receiver's speech into text and then determines whether or not a predetermined keyword is included in the text of the receiver's speech. Depending on whether or not the predetermined keyword is included, the content of the speech to be generated is changed.
[0032] Furthermore, the spoken voice does not necessarily have to be generated from scratch; it may also be generated by combining pre-recorded parts of the spoken voice according to the content of the received voice.
[0033] The speech unit 104 transmits the spoken voice to the receiver terminal. From the time the call is made to the target person until the call is connected to the target person (such as the person in charge or manager mentioned above), the speech unit 104 transmits the spoken voice generated by the speech voice generation unit 103 to the receiver terminal. On the other hand, after the call is connected to the target person, the speech unit 104 transmits the spoken voice uttered by the user of the speaker terminal to the receiver terminal. The speaker terminal is, for example, a terminal used by a human telephone appointment setter. The speaker switching unit 105 handles the switching of the speaker on this speaker terminal.
[0034] The speaker switching unit 105 has a function that, when the receiver's voice meets predetermined conditions, switches the input channel of the voice transmitted to the receiver terminal to the user of the speaker terminal.
[0035] The predetermined conditions for the speaker switching unit 105 to switch input channels may be, for example, the following conditions. (Condition 1) The person inputting the receiver's voice can be determined to be the person making the call. (Condition 2) It can be presumed that the person inputting the receiver's voice changes from someone who is not the intended caller to the intended caller.
[0036] Let's explain the above (Condition 1) in more detail. For the sake of explanation, let's refer to the person to be called as A. Person A is, for example, a company manager. Let's refer to someone who is not a target of the call as B. Person B is, for example, a company receptionist.
[0037] If the receiver's voice contains keywords that allow it to determine that the person currently receiving the call is the caller A, such as "Yes, this is A," the speaker switching unit 105 determines that the predetermined condition 1 is met and switches the input channel of the voice to the user on the speaker terminal. Alternatively, the voice characteristics of the caller may be pre-registered in the server 10's memory 12. The speaker switching unit 105 may also determine that the caller is inputting the receiver's voice by comparing the voice characteristics contained in the receiver's voice with the voice characteristics of the caller registered in memory 12. The speaker switching unit 105 may also determine that the caller is inputting the receiver's voice by combining keyword-based determination and voice characteristic comparison determination.
[0038] The above-mentioned (Condition 2) will be explained in more detail. If the receiver's voice contains a predetermined keyword that suggests the person answering the phone is changing from someone other than the intended caller to the intended caller, such as "Please wait a moment" or "I will transfer you to A," the speaker switching unit 105 considers that the predetermined condition 2 is met and switches the input channel of the voice to the user on the speaker's terminal. The predetermined keyword may be stored in memory 12 beforehand.
[0039] When the person answering the phone changes from person B (who is not the intended caller) to person A (who is the intended caller), hold music may be played during that time. Therefore, if the receiver's voice includes hold music, the speaker switching unit 105 determines that the predetermined condition 2 is met and switches the input channel of the voice to the user on the speaker's terminal.
[0040] Figure 3 is a flowchart showing an example of a power supply system process corresponding to at least one embodiment of the present invention.
[0041] The calling unit 101 initiates a call to the recipient terminal (St11). The recipient voice acquisition unit 102 acquires the recipient voice input to the recipient terminal (St12).
[0042] The speaker switching unit 105 determines whether the receiver's voice meets predetermined conditions (St13). If the receiver's voice meets the predetermined conditions (St13: YES), the process proceeds to step St16. If the receiver's voice does not meet the predetermined conditions (St13: NO), the process proceeds to step St14.
[0043] In step St14, the speech generation unit 103 generates speech for transmission to the receiver terminal based on the content of the receiver's voice. The speech generation unit 104 transmits the speech to the receiver terminal (St15). The process then returns to step St12, where the next receiver voice is acquired from the receiver terminal.
[0044] In step St16, the speaker switching unit 105 switches the input channel for the spoken audio to be transmitted to the receiver terminal to the user of the speaker terminal. Thereafter, the telephone appointment setter, who is the user of the speaker terminal, conducts a conversation with the receiver.
[0045] The following shows an example of a telephone conversation using the calling system 1 according to the embodiment of this disclosure. [Conversation Example 1] In Conversation Example 1, the intended caller A answered the phone at the receiver's terminal from the start. Therefore, the receiver's voice acquired in step St12 includes the phrase "Hello, this is A." Since the receiver's voice satisfies the above-mentioned (Condition 1), the process transitions from step St13 to step St16. The speaker switching unit 105 switches the input channel of the voice to be transmitted to the receiver's terminal to the user of the speaker's terminal. In other words, from the beginning, the speaker is a human user such as a telephone appointment setter.
[0046] [Conversation Example 2] In conversation example 2, initially, person B, who is not the intended caller, answers the phone on the receiver's terminal. The receiver's voice acquired in step St12 includes the voice saying, "Hello, this is XX Corporation." Since the receiver's voice does not satisfy either (condition 1) or (condition 2) above, the process transitions from step St13 to step St14. That is, the calling system 1 answers the phone call to person B, who is not the intended caller. For example, the speech generation unit 103 generates speech such as, "Hello, is Mr. / Ms. A, the sales department manager, available?" The speech unit 104 transmits the generated speech to the receiver's terminal.
[0047] Person B, who is not the intended recipient of the call, says on the phone, "This is A. I'll transfer you, please wait a moment," and puts the call into hold mode. Then, the next receiver voice acquired in step St12 satisfies the above-mentioned (condition 1) and (condition 2), so the process transitions from step St13 to step St16. The speaker switching unit 105 switches the input channel of the voice to be transmitted to the receiver terminal to the user of the speaker terminal. In other words, at the moment person B, who is not the intended recipient of the call, puts the call into hold mode, the entity handling the call on the speaker side switches from the calling system 1 to a human user such as a telephone appointment setter.
[0048] As described above, each embodiment of the present application solves one or more of the shortcomings. Note that the effects of each embodiment are non-limiting effects or examples of effects.
[0049] In each of the embodiments described above, the user terminal 20 and the server 10 execute the various processes described above in accordance with various control programs (for example, a call-making program) stored in their own storage devices. Furthermore, other computers, not limited to the user terminal 20 and the server 10, may also execute the various processes described above in accordance with various control programs (for example, a call-making program) stored in their own storage devices.
[0050] Furthermore, the configuration of the calling system 1 is not limited to the configuration described as an example of the embodiment above. For example, the server may perform some or all of the processes described as being performed by the user terminal, or the user terminal may perform some or all of the processes described as being performed by the server. Alternatively, the user terminal may be equipped with some or all of the storage unit (memory device) provided by the server. In other words, the calling system 1 may be configured such that one of the user terminals or the server provides some or all of the functions provided by the other.
[0051] Furthermore, the program may be configured to implement some or all of the functions described above as examples of each embodiment in a standalone device that does not include a communication network.
[0052] [Note] The above-described embodiments are written in such a way that at least the following invention can be put into practice by a person with ordinary skill in the art to which the invention pertains.
[0053] [1] In the processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient's terminal to the user of the speaker's terminal, A calling program that makes this possible.
[0054] According to the above calling program, it is possible to increase the efficiency of calling designated target individuals.
[0055] [2] The aforementioned predetermined conditions The ability to determine that the person inputting the recipient's voice is the person making the call. [1] The calling program described above. This allows the caller to switch from the automated calling system to a human operator the moment the intended recipient answers the phone.
[0056] [3] The aforementioned predetermined conditions It can be inferred that the person inputting the recipient's voice changes from someone who is not the intended caller to the intended caller. [1] The calling program described above.
[0057] [4] The aforementioned predetermined conditions The aforementioned receiver's voice includes hold music. The calling program described in [3].
[0058] According to these calling programs, if someone other than the intended recipient answers the phone first, the system can switch the caller's representative from the calling system to a human at the same time the receiving representative is transferred to the intended recipient.
[0059] [5] A power supply system comprising one or more processors, The aforementioned processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient's terminal to the user of the speaker's terminal, A power supply system that makes this possible.
[0060] According to the above calling system, the efficiency of calling designated target individuals can be increased.
[0061] [6] A method of powering a device equipped with a processor, The calling step involves making a call to the recipient's terminal, A receiver voice acquisition step, which involves acquiring the receiver's voice input to the receiver's terminal, A speech voice generation step, which generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech step of transmitting the aforementioned spoken audio to the recipient terminal, A speaker switching step, in which, when the recipient's voice meets predetermined conditions, the input channel of the spoken voice to be transmitted to the recipient terminal is switched to the user of the speaker terminal, A method of making an electric charge, having the following characteristics.
[0062] According to the above calling method, the efficiency of calling designated target individuals can be increased. [Industrial applicability]
[0063] According to one embodiment of the present invention, a calling program, calling system, and calling method are useful for improving the efficiency of calling predetermined target persons. [Explanation of Symbols]
[0064] 1. Power supply system 10 servers 11 processors 12 memory 13 Storage device 20, 20A, 20B User Terminals 21 processors 22 memory 23 Storage device 30 Communication Networks 101 Calling Unit 102 Receiver voice acquisition unit 103 Speech Generation Unit 104 Speech Unit 105 Speaker switching unit
Claims
1. In the processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient's terminal to the user of the speaker's terminal, A calling program that makes this possible.
2. The aforementioned predetermined conditions The ability to determine that the person inputting the recipient's voice is the person making the call. The calling program according to claim 1.
3. The aforementioned predetermined conditions It can be inferred that the person inputting the recipient's voice changes from someone who is not the intended caller to the intended caller. The calling program according to claim 1.
4. The aforementioned predetermined conditions The aforementioned receiver's voice includes hold music. The calling program described in claim 3.
5. A power supply system comprising one or more processors, The aforementioned processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient's terminal to the user of the speaker's terminal, A power supply system that makes this possible.
6. A method of powering a device equipped with a processor, The calling step involves making a call to the recipient's terminal, A receiver voice acquisition step, which involves acquiring the receiver's voice input to the receiver's terminal, A speech voice generation step, which generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech step of transmitting the aforementioned spoken audio to the recipient terminal, A speaker switching step, in which, when the recipient's voice meets predetermined conditions, the input channel of the spoken voice to be transmitted to the recipient terminal is switched to the user of the speaker terminal, A method of making an electric charge, having the following characteristics.
Citation Information
Patent Citations
Line connection changeover device
JP1990030269A
Operator's operation support system for call center
JP2007004000A
Information processing system, information processing method, and information processing program
JP2022032967A
Automatic call device and automatic call method
JP2021034757A