Calling program, calling system, and calling method

The calling system addresses inefficiencies in connecting with intended recipients by automating interactions and switching to human operators when conditions are met, improving call efficiency.

JP2026052688APending Publication Date: 2026-03-24DIAL SHIFT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-02
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing systems for making calls to predetermined target persons are inefficient, particularly when the intended recipient is unavailable, leading to prolonged delays in connecting with the intended person.

Method used

A calling system equipped with a processor that includes a calling function, recipient voice acquisition, speech voice generation, and a speaker switching function to seamlessly transition from automated interaction to a human operator when predetermined conditions are met.

Benefits of technology

Enhances the efficiency of calling designated targets by automating interactions until the intended recipient is reached and seamlessly switching to a human operator when conditions are met, reducing overall call setup time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026052688000001_ABST
    Figure 2026052688000001_ABST
Patent Text Reader

Abstract

To improve the efficiency of making calls to designated target individuals. [Solution] The calling program provides the processor with a calling function that initiates a call to the recipient terminal, a recipient voice acquisition function that acquires the recipient voice input to the recipient terminal, a speech voice generation function that generates speech voice to be sent to the recipient terminal based on the content of the recipient voice, a speaking function that sends the speech voice to the recipient terminal, and a speaker switching function that switches the input channel of the speech voice to be sent to the recipient terminal to the user of the speaker terminal when the recipient voice meets predetermined conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] At least one embodiment of the present invention relates to a power supply program, a power supply system, and a power supply method.

Background Art

[0002] Patent Document 1 describes an automatic power supply device that notifies a voice message. The automatic power supply device causes a computer device to execute functions of acquiring a power supply target list, selecting a power supply target from the power supply target list, making a call to the power supply target, providing a voice message when the power supply target is called, and selecting another power supply target from the power supply target list when a call to the power supply target is unavailable.

Prior Art Document

Patent Document

[0003]

Patent Document 1

[0006] An object of at least one embodiment of the present invention is to provide a calling program, a calling system, and a calling method that solve the above problems and improve the efficiency of calling predetermined target persons. [Means for solving the problem]

[0007] In a non-limiting view, a calling program according to one embodiment of the present invention provides the processor with a calling function for making a call to a recipient terminal to be called; a recipient voice acquisition function for acquiring recipient voice input to the recipient terminal; a speech voice generation function for generating speech voice to be sent to the recipient terminal based on the content of the recipient voice; a speaking function for sending the speech voice to the recipient terminal; and a speaker switching function for switching the input channel of the speech voice to be sent to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies predetermined conditions.

[0008] In a non-limiting view, a calling system according to one embodiment of the present invention is a calling system comprising one or more processors, wherein the processors implement: a calling function for making a call to a recipient terminal to be called; a recipient voice acquisition function for acquiring recipient voice input to the recipient terminal; a speech voice generation function for generating speech voice to be transmitted to the recipient terminal based on the content of the recipient voice; a speaking function for transmitting the speech voice to the recipient terminal; and a speaker switching function for switching the input channel of the speech voice to be transmitted to the recipient terminal to the user of the speaker terminal when the recipient voice satisfies predetermined conditions.

[0009] In a non-limiting view, a calling method according to one embodiment of the present invention is a calling method using a device equipped with a processor, comprising: a calling step of making a call to a receiver terminal to be called; a receiver voice acquisition step of acquiring receiver voice input to the receiver terminal; a speech voice generation step of generating speech voice to be transmitted to the receiver terminal based on the content of the receiver voice; a speaking step of transmitting the speech voice to the receiver terminal; and a speaker switching step of switching the input channel of the speech voice to be transmitted to the receiver terminal to the user of the speaker terminal when the receiver voice satisfies predetermined conditions. [Effects of the Invention]

[0010] Each embodiment of the present invention resolves one or more of the shortcomings. [Brief explanation of the drawing]

[0011] [Figure 1] A block diagram showing an example of the configuration of a power supply system corresponding to at least one embodiment of the present invention. [Figure 2] A block diagram showing the configuration of a server corresponding to at least one embodiment of the present invention. [Figure 3] A flowchart showing an example of a power supply system process corresponding to at least one embodiment of the present invention. [Modes for carrying out the invention]

[0012] Hereinafter, examples of embodiments of the present invention will be described with reference to the drawings. The various components in each embodiment described below can be combined as appropriate, provided that no inconsistencies arise. Furthermore, some aspects described in one embodiment may be omitted in other embodiments. Also, operations and processes unrelated to the characteristic features of each embodiment may be omitted. Moreover, the order of the various processes constituting the various flows and sequences described below is not necessarily in any particular order, provided that no inconsistencies arise in the processing content.

[0013] The following explanation uses a call-making program executed on a server, which is an example of a computer, as an example. However, the computer may be other devices such as a user terminal. Alternatively, the entire call-making system may execute the call-making program.

[0014] Figure 1 is a block diagram showing an example configuration of a call system corresponding to at least one embodiment of the present invention. The call system 1 comprises a server 10 and user terminals 20 used by users of the call system 1. User terminals 20A, 20B, and 20C are examples of user terminals 20. The configuration of the call system 1 is not limited thereto. For example, the call system 1 may be configured so that a single user terminal is used by multiple users. The call system 1 may also comprise multiple servers.

[0015] Server 10 and user terminal 20 are examples of devices. Server 10 and user terminal 20 are each connected to a communication network 30 in a communication-enabled manner. The connection between the communication network 30 and server 10, and the connection between the communication network 30 and user terminal 20, may be a wired or wireless connection. The communication network 30 may be an internet line, an analog telephone line, or a line combining these.

[0016] The telephone system 1, comprising a server 10 and a user terminal 20, realizes various functions for executing various processes in response to user operations.

[0017] Server 10 includes a processor 11, a memory 12, and a storage device 13. The processor 11 is, for example, a central processing unit such as a CPU (Central Processing Unit) that performs various operations and controls. Also, when the server 10 includes a GPU (Graphics Processing Unit), part of the various operations and controls may be performed by the GPU. The server 10 uses the data read into the memory 12 to execute various information processes by the processor 11, and stores the obtained processing results in the storage device 13 as necessary.

[0018] The storage device 13 has a function as a storage medium for storing various information. The configuration of the storage device 13 is not particularly limited, but from the viewpoint of reducing the processing load on the user terminal 20, it may be configured to be able to store all the various information necessary for the control performed in the power supply system 1. Examples of such include HDDs and SSDs. However, the storage device for storing various information only needs to have a storage area in a state accessible by the server 10, and for example, it may be configured to have a dedicated storage area outside the server 10.

[0019] The user terminal 20 is managed by the user. The user here includes a telephone pointer and the called party who is the target of the call made by the telephone pointer. Examples of the user terminal 20 include, for example, a fixed telephone terminal, a mobile phone terminal, a smartphone, a personal computer, a tablet, etc. The user terminal 20 may be the speaker-side terminal or the listener-side terminal described later.

[0020] The user terminal 20 may be connected to the communication network 30 and include hardware and software for executing various processes by communicating with the server 10. Each of the plurality of user terminals 20 may also be configured to be able to communicate directly with each other without going through the server 10.

[0021] The user terminal 20 may include a processor 21, a memory 22, and a storage device 23. The processor 21 is, for example, a central processing unit such as a CPU (Central Processing Unit) that performs various operations and controls. Also, when the user terminal 20 includes a GPU (Graphics Processing Unit), a part of various operations and controls may be performed by the GPU. The user terminal 20 uses the data read into the memory 22 to execute various information processes by the processor 21, and stores the obtained processing results in the storage device 23 as necessary. The storage device 23 has a function as a storage medium for storing various information.

[0022] FIG. 2 is a block diagram showing the configuration of a server corresponding to at least one embodiment of the present invention. The server 10 includes a calling unit 101, a recipient voice acquisition unit 102, a speaking voice generation unit 103, a speaking unit 104, and a speaker switching unit 105. The processor included in the server 10 refers to the connection program held in the storage device and executes the program to functionally realize the calling unit 101, the recipient voice acquisition unit 102, the speaking voice generation unit 103, the speaking unit 104, and the speaker switching unit 105.

[0023] The calling unit 101 has a function of making a call to the recipient-side terminal to be connected. The recipient voice acquisition unit 102 has a function of acquiring the recipient voice input to the recipient-side terminal. The speaking voice generation unit 103 has a function of generating a speaking voice for transmission to the recipient-side terminal based on the content of the recipient voice. The speaking unit 104 has a function of transmitting the speaking voice to the recipient-side terminal. The speaker switching unit 105 has a function of switching the input channel of the speaking voice transmitted to the recipient-side terminal to the user of the speaker-side terminal when the recipient voice satisfies a predetermined condition.

[0024] The receiver terminal is the terminal that receives calls from the telephone appointment setter. For example, a company's receiver terminal may be a landline phone or a mobile phone. The receiver terminal is an example of the user terminal 20 described above. The user of the calling system 1 may specify the telephone number of the receiver terminal to be called. A list of telephone numbers of receiver terminals to be called is stored in memory 12 in advance, and the calling unit 101 may retrieve the telephone number of the receiver terminal to be called from the list.

[0025] The receiver's voice input to the receiver's terminal may be converted into digital data before it reaches the server 10.

[0026] [Generating spoken audio] When the target of a call is a person in charge or a manager at a company, the person in charge may not answer the phone. For example, a company receptionist may take the call temporarily and transfer it to the person in charge or a manager. Therefore, in the past, it took a long time for the telephone appointment setter to begin a conversation with the intended person. In the embodiment of this disclosure, in order to streamline operations, the call system 1, rather than a human telephone appointment setter, handles the interaction from the time the call is made to the target person (the person in charge or a manager mentioned above) until the call is connected to the intended person. Therefore, the call system 1 generates spoken audio to respond to the recipient until it transfers the call to the telephone appointment setter.

[0027] The speech generation unit 103 generates speech for transmission to the receiver terminal based on the content of the receiver's voice acquired from the receiver terminal.

[0028] The speech generation in the speech generation unit 103 may be performed using AI. For example, the speech generation unit 103 includes an LLM (Large-Scale Language Model). The speech generation unit 103 generates speech for transmission to the receiver terminal by inputting the receiver's voice or converted information obtained by converting the receiver's voice into the LLM.

[0029] The LLM may generate text data for transmission to the receiver terminal. The speech generation unit 103 converts the text data into speech. The LLM may also directly generate speech for transmission to the receiver terminal.

[0030] The speech generation unit 103 may generate speech using rule-based AI. In this case, the speech generation unit 103 analyzes the content of the receiver's voice obtained from the receiver's terminal. For example, it calculates the similarity between the text data obtained by transcribing the receiver's voice and the receiver's utterance scenarios that have been pre-registered in memory 12. For example, if the text data obtained by transcribing the receiver's voice is "This is XX Corporation," then this text data has a high similarity to the receiver's utterance scenario "Yes, this is XX Corporation" that has been registered in memory 12. The speech generation unit 103 extracts scenarios with a high similarity to the text data obtained by transcribing the receiver's voice from memory 12, and identifies the messages stored in memory 12 in association with the extracted scenarios as speech text. Such correspondences between receiver's utterance scenarios and messages as speech text may be pre-stored in memory 12 in the form of a table or the like.

[0031] The speech generation unit 103 may determine the content of the speech based on conditional branching implemented in the rule-based AI. For example, the speech generation unit 103 converts the receiver's speech into text and then determines whether or not a predetermined keyword is included in the text of the receiver's speech. Depending on whether or not the predetermined keyword is included, the content of the speech to be generated is changed.

[0032] Furthermore, the spoken voice does not necessarily have to be generated from scratch; it may also be generated by combining pre-recorded parts of the spoken voice according to the content of the received voice.

[0033] The speech unit 104 transmits the spoken voice to the receiver terminal. From the time the call is made to the target person until the call is connected to the target person (such as the person in charge or manager mentioned above), the speech unit 104 transmits the spoken voice generated by the speech voice generation unit 103 to the receiver terminal. On the other hand, after the call is connected to the target person, the speech unit 104 transmits the spoken voice uttered by the user of the speaker terminal to the receiver terminal. The speaker terminal is, for example, a terminal used by a human telephone appointment setter. The speaker switching unit 105 handles the switching of the speaker on this speaker terminal.

[0034] The speaker switching unit 105 has a function that, when the receiver's voice meets predetermined conditions, switches the input channel of the voice transmitted to the receiver terminal to the user of the speaker terminal.

[0035] The predetermined conditions for the speaker switching unit 105 to switch input channels may be, for example, the following conditions. (Condition 1) The person inputting the receiver's voice can be determined to be the person making the call. (Condition 2) It can be presumed that the person inputting the receiver's voice changes from someone who is not the intended caller to the intended caller.

[0036] Let's explain the above (Condition 1) in more detail. For the sake of explanation, let's refer to the person to be called as A. Person A is, for example, a company department head. Let's refer to someone who is not a target of the call as B. Person B is, for example, a company receptionist.

[0037] If the receiver's voice contains keywords that allow it to determine that the person currently receiving the call is the caller A, such as "Yes, this is A," the speaker switching unit 105 determines that the predetermined condition 1 is met and switches the input channel of the voice to the user on the speaker terminal. Alternatively, the voice characteristics of the caller may be pre-registered in the server 10's memory 12. The speaker switching unit 105 may also determine that the caller is inputting the receiver's voice by comparing the voice characteristics contained in the receiver's voice with the voice characteristics of the caller registered in memory 12. The speaker switching unit 105 may also determine that the caller is inputting the receiver's voice by combining keyword-based determination and voice characteristic comparison determination.

[0038] The above-mentioned (Condition 2) will be explained in more detail. If the receiver's voice contains a predetermined keyword that suggests the person answering the phone is changing from someone other than the intended caller to the intended caller, such as "Please wait a moment" or "I will transfer you to A," the speaker switching unit 105 considers that the predetermined condition 2 is met and switches the input channel of the voice to the user on the speaker's terminal. The predetermined keyword may be stored in memory 12 beforehand.

[0039] When the person answering the phone changes from person B (who is not the intended caller) to person A (who is the intended caller), hold music may be played during that time. Therefore, if the receiver's voice includes hold music, the speaker switching unit 105 determines that the predetermined condition 2 is met and switches the input channel of the voice to the user on the speaker's terminal.

[0040] Figure 3 is a flowchart showing an example of a power supply system process corresponding to at least one embodiment of the present invention.

[0041] The calling unit 101 initiates a call to the recipient terminal (St11). The recipient voice acquisition unit 102 acquires the recipient voice input to the recipient terminal (St12).

[0042] The speaker switching unit 105 determines whether the receiver's voice meets predetermined conditions (St13). If the receiver's voice meets the predetermined conditions (St13: YES), the process proceeds to step St16. If the receiver's voice does not meet the predetermined conditions (St13: NO), the process proceeds to step St14.

[0043] In step St14, the speech generation unit 103 generates speech for transmission to the receiver terminal based on the content of the receiver's voice. The speech generation unit 104 transmits the speech to the receiver terminal (St15). The process then returns to step St12, where the next receiver voice is acquired from the receiver terminal.

[0044] In step St16, the speaker switching unit 105 switches the input channel for the spoken audio to be transmitted to the receiver terminal to the user of the speaker terminal. Thereafter, the telephone appointment setter, who is the user of the speaker terminal, conducts a conversation with the receiver.

[0045] The following shows an example of a telephone conversation using the calling system 1 according to the embodiment of this disclosure. [Conversation Example 1] In conversation example 1, the intended caller A answered the phone at the receiver's terminal from the start. Therefore, the receiver's voice acquired in step St12 includes the phrase "Hello, this is A." Since the receiver's voice satisfies the above-mentioned (condition 1), the process transitions from step St13 to step St16. The speaker switching unit 105 switches the input channel of the voice to be transmitted to the receiver's terminal to the user of the speaker's terminal. In other words, from the beginning, the speaker is a human user such as a telephone appointment setter.

[0046] [Conversation Example 2] In conversation example 2, initially, person B, who is not the intended caller, answers the phone on the receiver's terminal. The receiver's voice acquired in step St12 includes the voice saying, "Hello, this is XX Corporation." Since the receiver's voice does not satisfy either (condition 1) or (condition 2) above, the process transitions from step St13 to step St14. That is, the calling system 1 answers the phone call to person B, who is not the intended caller. For example, the speech generation unit 103 generates speech such as, "Hello, is Mr. / Ms. A, the sales department manager, available?" The speech unit 104 transmits the generated speech to the receiver's terminal.

[0047] Person B, who is not the intended recipient of the call, says on the phone, "This is A. I'll transfer you, please wait a moment," and puts the call into hold mode. Then, the next receiver voice acquired in step St12 satisfies the above-mentioned (condition 1) and (condition 2), so the process transitions from step St13 to step St16. The speaker switching unit 105 switches the input channel of the voice to be transmitted to the receiver terminal to the user of the speaker terminal. In other words, at the moment person B, who is not the intended recipient of the call, puts the call into hold mode, the entity handling the call on the speaker side switches from the calling system 1 to a human user such as a telephone appointment setter.

[0048] As described above, each embodiment of the present application solves one or more of the shortcomings. Note that the effects of each embodiment are non-limiting effects or examples of effects.

[0049] In each of the embodiments described above, the user terminal 20 and the server 10 execute the various processes described above in accordance with various control programs (for example, a call-making program) stored in their own storage devices. Furthermore, other computers, not limited to the user terminal 20 and the server 10, may also execute the various processes described above in accordance with various control programs (for example, a call-making program) stored in their own storage devices.

[0050] Furthermore, the configuration of the calling system 1 is not limited to the configuration described as an example of the embodiment above. For example, the server may perform some or all of the processes described as being performed by the user terminal, or the user terminal may perform some or all of the processes described as being performed by the server. Alternatively, the user terminal may be equipped with some or all of the storage unit (memory device) provided by the server. In other words, the calling system 1 may be configured such that one of the user terminals or the server provides some or all of the functions provided by the other.

[0051] Furthermore, the program may be configured to implement some or all of the functions described above as examples of each embodiment in a standalone device that does not include a communication network.

[0052] [Note] The above-described embodiments are written in such a way that at least the following invention can be put into practice by a person with ordinary skill in the art to which the invention pertains.

[0053] [1] In the processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient terminal to the user of the speaker terminal, A calling program that makes this possible.

[0054] According to the above calling program, it is possible to increase the efficiency of calling designated target individuals.

[0055] [2] The aforementioned predetermined conditions The ability to determine that the person inputting the recipient's voice is the person making the call. The calling program described in [1]. This allows the caller to switch from the automated calling system to a human operator the moment the intended recipient answers the phone.

[0056] [3] The aforementioned predetermined conditions It can be inferred that the person inputting the recipient's voice changes from someone who is not the intended caller to the intended caller. The calling program described in [1].

[0057] [4] The aforementioned predetermined conditions The aforementioned receiver's voice includes hold music. The calling program described in [3].

[0058] According to these calling programs, if someone other than the intended recipient answers the phone first, the system can switch the caller's representative from the calling system to a human at the same time the recipient's representative is transferred to the intended recipient.

[0059] [5] A power supply system comprising one or more processors, The aforementioned processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient terminal to the user of the speaker terminal, A power supply system that makes this possible.

[0060] According to the above calling system, the efficiency of calling designated target individuals can be increased.

[0061] [6] A method of powering a device equipped with a processor, The calling step involves making a call to the recipient's terminal, A receiver voice acquisition step, which involves acquiring the receiver's voice input to the receiver's terminal, A speech voice generation step, which generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech step of transmitting the aforementioned spoken audio to the recipient terminal, A speaker switching step, in which, when the recipient's voice meets predetermined conditions, the input channel of the spoken voice to be transmitted to the recipient terminal is switched to the user of the speaker terminal, A method of making electricity, having the following characteristics.

[0062] According to the above calling method, the efficiency of calling designated target individuals can be increased. [Industrial applicability]

[0063] According to one embodiment of the present invention, a calling program, calling system, and calling method are useful for improving the efficiency of calling predetermined target persons. [Explanation of Symbols]

[0064] 1. Power supply system 10 servers 11 processors 12 memory 13 Storage device 20, 20A, 20B User Terminals 21 processors 22 memory 23 Storage device 30 Communication Networks 101 Calling Unit 102 Receiver voice acquisition unit 103 Speech Generation Unit 104 Speech Unit 105 Speaker switching unit

Claims

1. In the processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient's terminal to the user of the speaker's terminal, A calling program that makes this possible.

2. The aforementioned predetermined conditions The ability to determine that the person inputting the recipient's voice is the person making the call. The calling program according to claim 1.

3. The aforementioned predetermined conditions It can be inferred that the person inputting the recipient's voice changes from someone who is not the intended caller to the intended caller. The calling program according to claim 1.

4. The aforementioned predetermined conditions The aforementioned receiver's voice includes hold music. The calling program described in claim 3.

5. A power supply system comprising one or more processors, The aforementioned processor, A calling function that initiates a call to the recipient's terminal, A receiver voice acquisition function that acquires the receiver's voice input to the receiver-side terminal, A speech voice generation function that generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech function that transmits the aforementioned spoken audio to the recipient terminal, A speaker switching function that, when the recipient's voice meets predetermined conditions, switches the input channel of the spoken voice transmitted to the recipient's terminal to the user of the speaker's terminal, A power supply system that makes this possible.

6. A method of powering a device equipped with a processor, The calling step involves making a call to the recipient's terminal, A receiver voice acquisition step, which involves acquiring the receiver's voice input to the receiver's terminal, A speech voice generation step, which generates speech voice for transmission to the recipient terminal based on the content of the recipient's voice, A speech step of transmitting the aforementioned spoken audio to the recipient terminal, A speaker switching step, in which, when the recipient's voice meets predetermined conditions, the input channel of the spoken voice to be transmitted to the recipient terminal is switched to the user of the speaker terminal, A method of making electricity, having the following characteristics.

Citation Information

Patent Citations

  • Automatic call device and automatic call method

    JP2021034757A