Response device, response method, and response program
The response device classifies callers and uses a large language model to generate tailored responses, addressing the challenge of inappropriate interactions by managing diverse caller types effectively.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NTT DOCOMO BUSINESS INC
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to provide appropriate responses to diverse callers, particularly when dealing with long conversations, difficult customers, scammers, or nuisance sales calls, failing to protect the recipient effectively.
A response device that classifies callers based on conversation characteristics and uses a large language model to generate tailored responses, including summaries, acknowledgments, or disconnecting conversations as needed.
Enables appropriate responses to various callers, effectively managing long conversations, complaints, beneficial sales calls, and fraudulent calls by providing targeted interactions or protective measures.
Smart Images

Figure 2026079507000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a response device, a response method, and a response program.
Background Art
[0002] There is a known technique for generating a response to a user who makes a call (hereinafter sometimes referred to as the "calling user") based on a large language model. For example, there is a known conventional technique for generating the response content of a dialogue agent based on the user's language information and the user's non-verbal information using a large language model (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the prior art, it may be difficult to give an appropriate response according to the calling party. For example, in the prior art, when the calling user is a person who continues a conversation for a long time, a monster caller, a fraudster, a person making a nuisance business call, etc., it is difficult to take an appropriate response to protect the called user (hereinafter referred to as the "responding user") from the calling user.
Means for Solving the Problems
[0005] Therefore, in order to solve the above-mentioned problems and achieve the objective, the response device of the present invention is characterized by comprising: a classification unit that classifies the user who initiated the conversation based on the natural language conversation input; a generation unit that generates the generated response by inputting a command to a large-scale language model, which is a prompt set according to the classification of the user who initiated the conversation, in accordance with the content of the natural language conversation by the user who initiated the conversation; and an output unit that outputs the generated response to the user who initiated the conversation. [Effects of the Invention]
[0006] The present invention has the effect of enabling an appropriate response according to the source of the power call. [Brief explanation of the drawing]
[0007] [Figure 1] Figure 1 is a diagram illustrating the overall structure of the response generation process according to the embodiment. [Figure 2] Figure 2 shows the configuration of the response device according to this embodiment. [Figure 3] Figure 3 is a table diagram showing an example of conversation destination user information according to the embodiment. [Figure 4] Figure 4 is a table diagram showing an example of classification conditions according to the embodiment. [Figure 5] Figure 5 is a table diagram showing an example of summary information according to the embodiment. [Figure 6] Figure 6 shows an example of the classification process according to the embodiment. [Figure 7] Figure 7 shows an example of the generation and output of a response to a first user according to the embodiment. [Figure 8] Figure 8 shows an example of the generation and output of a response to a second user according to the embodiment. [Figure 9] Figure 9 shows an example of the generation and output of a response to a third user according to the embodiment. [Figure 10] Figure 10 shows an example of the generation and output of a response to a third user according to the embodiment. [Figure 11] Figure 11 is a flowchart showing the processing by the response device according to the embodiment. [Figure 12] Figure 12 shows an example of a computer that implements the response device according to this embodiment. [Modes for carrying out the invention]
[0008] Hereinafter, embodiments for carrying out the present invention (hereinafter referred to as "embodiments") will be described with reference to the drawings. However, each embodiment is not limited to those described below.
[0009] <Overview> (background) Calls may be made from various sources with diverse requirements, and it is necessary to respond to these calls efficiently. Therefore, a reference technique is known that uses a large-scale language model to generate response content for the user who made the call.
[0010] However, the reference technology may have difficulty generating appropriate responses to the diverse requirements of the various callers mentioned above. For example, it is difficult to take appropriate responses to protect the recipient from the caller when the caller is someone who engages in long conversations, a difficult customer, a scammer, or someone making annoying sales calls.
[0011] (Processing by response device 100) Therefore, the response device 100 according to this embodiment generates a response according to the classification based on the conversation of the user making the call, and outputs an appropriate response to the user making the call.
[0012] Here, we will explain the overall process performed by the response device 100. Figure 1 is a diagram illustrating the overall process of generating a response according to this embodiment. The response device 100 shown in Figure 1 is an example of a computer that provides the technology to realize the information processing described below.
[0013] First, the response device 100 provides a predetermined large language model with the voice quality of the conversation partner user, conversation characteristics, etc. (Fig. 1(1-1)) as prior knowledge (Fig. 1(1-2)). The above "provision of prior knowledge" means inputting predetermined information into the large language model in advance, including learning and tuning of the large language model, input of information into the large language model, insertion or substitution of information related to the prior knowledge into the prompt input to the large language model, etc.
[0014] Next, the response device 100 classifies the conversation originator user who makes a call based on the conversation of the conversation originator user (Fig. 1(2)). Subsequently, the response device 100 inputs a generation command for a response corresponding to the conversation originator user generated according to the classification into the large language model using a prompt expressed in natural language text (Fig. 1(3-1)), and causes the large language model to generate a response corresponding to the conversation originator user (Fig. 1(3-2)).
[0015] Then, the response device 100 outputs the response generated according to the conversation originator user for each conversation originator user (Fig. 1(4-1)).
[0016] For example, when classified as user 10a who has a long conversation without getting to the point before stating the original requirements, the response device 100 outputs responses such as "echo", "conversation summary", "summary confirmation", etc. to the user 10a (Fig. 1(4-2)). Thereby, the response device 100 can appropriately echo the user 10a who has a long conversation without getting to the point before stating the original requirements, and appropriately ask about the original requirements based on the conversation summary.
[0017] Also, when classified as user 10b who makes a complaint to the conversation partner user, the response device 100 outputs responses such as "echo", "conversation summary", etc. to the user 10b (Fig. 1(4-3)). Thereby, the response device 100 can appropriately echo the user 10b who conveys a complaint while having an angry emotion, and summarize the content of the complaint to enable consideration of appropriate countermeasures, etc.
[0018] Furthermore, if the user 10c is classified as making a beneficial sales call, the answering device 100 outputs a response to that user 10c such as "a summary of the conversation to be presented to the user on the other end of the conversation" (Figure 1 (4-4)). This allows the answering device 100 to not simply block sales calls, but to extract information from the user 10c who is making a beneficial call and to transmit that useful information to the user on the other end of the conversation.
[0019] Furthermore, if the caller is classified as a user 10d making a fraudulent or unwanted sales call, the answering device 100 outputs a response such as "disconnecting the conversation" to the user 10d (Figure 1 (4-5)). This prevents the answering device 100 from connecting calls from users 10d making unwanted sales or fraudulent calls to the intended user, thereby protecting the intended user from unwanted sales or fraudulent calls.
[0020] In this way, the response device 100 according to this embodiment can provide an appropriate response by generating a response according to the classification result based on the conversation of the user initiating the conversation and outputting it for each user initiating the conversation. As a result, the response device 100 has the effect of enabling an appropriate response according to the caller.
[0021] <Description of response device 100> The configuration of the response device 100 according to this embodiment will now be described. Figure 2 is a diagram showing the configuration of the response device 100 according to this embodiment. As shown in Figure 2, the response device 100 has a communication unit 110, a storage unit 120, and a control unit 130.
[0022] Although not shown in Figure 2, the response device 100 may be equipped with an input unit such as a keyboard or mouse to receive input such as operations from an administrator. The response device 100 may also be equipped with a display for showing the administrator user information of the user receiving the conversation, the content of the conversation from the initiating user, the classification result of the initiating user, and the generated response.
[0023] (Communications Department 110) The communication unit 110 performs data communication related to input such as user information of the conversation recipient and the conversation of the conversation source. The communication unit 110 also performs data communication related to outputting information about the generated response.
[0024] The communication unit 110 is implemented using a NIC (Network Interface Card) or the like, and controls communication via telecommunication lines such as a LAN (Local Area Network) or the Internet. The communication unit 110 can be connected to the network via wired or wireless connection as needed, and can send and receive information bidirectionally with terminal devices such as mobile phones and smartphones.
[0025] (Storage unit 120) The storage unit 120 stores data and programs used for various processes by the control unit 130, as well as various data acquired through the operation of the control unit 130. The storage unit 120 is implemented using semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as hard disks and optical discs. As shown in Figure 2, the storage unit 120 also includes a conversation destination user information DB 121, a classification condition DB 122, a summary information DB 123, and a generation model DB 124.
[0026] (Conversation target user information DB121) The conversation user information DB121 is a database that stores information about the conversation user (conversation user information), including information such as the conversation user's voice quality and characteristics of their conversational style during a call. This information is used when the generation unit 135, described later, generates a response that mimics the conversation user's voice quality and the words they use. Specifically, the conversation user information DB121 stores information such as attribute information about the conversation user (attribute information), voice information about the conversation user's voice quality, and characteristics of the conversation user's conversational style (conversation characteristics).
[0027] Here, an example of conversation user information stored in the conversation user information DB121 will be explained using Figure 3. Figure 3 is a table diagram showing an example of conversation user information according to the embodiment. The conversation user information DB121 stores information related to each item, such as "No," which is information that identifies individual data included in the conversation user information, "conversation user identification information," "attribute information," "voice information," and "conversation features," in a table format as shown in Figure 3. The letters "A to D" written in each item of the table diagram shown in Figure 3 are legends for the information included in each item.
[0028] The "user identification information for the person you are talking to" mentioned above is information used to identify the user to whom the call is connected (the person you are talking to). This includes, for example, the name or title of the person you are talking to, a number that identifies the person you are talking to, and other information that combines text, numbers, symbols, etc. to identify the person you are talking to, but in a form that excludes information that could identify an individual related to that person you are talking to.
[0029] Furthermore, "attribute information" includes, for example, the identification information of the user being conversed with, gender and age, hobbies and preferences, etc., in a form that excludes information that could identify the individual user. For example, "attribute information" may include attribute information related to the user's family and friends, as well as attribute information related to the community or industry to which the user belongs.
[0030] "Voice information" refers to information about the voice quality of the conversation partner, including acoustic characteristics such as pitch and timbre of the voice spoken by the conversation partner, as well as information on recordings of the conversation partner's voice. "Conversation features" also include characteristic information representing the conversation partner's speaking style, language used, vocabulary, etc., as well as information obtained by converting the conversation spoken by the conversation partner into text, etc.
[0031] (Classification condition DB122) The classification condition DB122 is a database that stores information (classification conditions) related to the classification processing of the conversation source user by the classification unit 134, which will be described later. Specifically, the classification condition DB122 stores information that associates classification conditions with the classification names that apply to those classification conditions.
[0032] Here, an example of classification conditions stored in the classification condition DB122 will be explained using Figure 4. Figure 4 is a table diagram showing an example of classification conditions according to the embodiment. The classification condition DB122 stores information related to each item, such as "No," which is information that identifies individual data included in the classification conditions, "classification condition," and "classification name," in a table format as shown in Figure 4.
[0033] The "classification criteria" mentioned above are conditions established for classifying the user who initiated the conversation based on their conversation, and the system performs classification to a classification name that matches these criteria. The "classification name" is information that identifies the result of the classification based on the classification criteria, and includes information represented by a combination of text, numbers, symbols, etc.
[0034] For example, in the table diagram shown in Figure 4, if a conversation initiated by a user is determined to match the classification criteria "conversation lasts for more than one minute, purpose is unclear," identified by No. "1," that user is classified as "User 1." Similarly, if a conversation initiated by a user is determined to match the classification criteria "user's emotion is 'anger,' conversation includes complaints," identified by No. "2," that user is classified as "User 2." Furthermore, if a conversation initiated by a user is determined to match the classification criteria "registered as a spam call, conversation includes fraud," identified by No. "3," that user is classified as "User 3."
[0035] (Summary Information DB123) The summary information DB123 is a database that stores information (summary information) about the summary of a conversation by a conversation-starting user, which is generated by the generation unit 135 described later. Specifically, the summary information DB123 stores information such as a summary of a conversation by a conversation-starting user, associated with information that identifies the conversation-starting user (conversation-starting user identification information).
[0036] Here, an example of summary information stored in the summary information DB123 will be explained using Figure 5. Figure 5 is a table diagram showing an example of summary information according to the embodiment. The summary information DB123 stores information related to each item, such as "No," which is information that identifies individual data included in the summary information, "Conversation source user identification information," and "Summary," in a table format as shown in Figure 5. The letters "E and F" written in each item of the table diagram shown in Figure 5 are legends for the information included in each item.
[0037] The "Initiating User Identification Information" mentioned above is information used to identify the user who initiated the call (initiating user), and includes, for example, the name or title of the initiating user, a number that identifies the initiating user, and other information combining text, numbers, symbols, etc., to identify the initiating user, but in a form that excludes information that could identify the individual in question.
[0038] Furthermore, the "summary" is information summarized by the user who initiated the conversation, and includes, for example, the main points of the conversation, or information that summarizes the conversation content to a predetermined number of characters. In addition, the above-mentioned "summary" may include conversation logs of past sales calls or fraudulent calls.
[0039] (Generative model DB124) The generative model DB124 is a database that stores predetermined models used by the generation unit 135 (described later) to generate responses corresponding to the conversation of the user initiating the conversation. For example, the generative model DB124 can store large-scale language models as generative models.
[0040] Specifically, the response device 100 according to this embodiment can use at least one of the following as a large-scale language model: "ChatGPT®", a large-scale language model possessing general-purpose knowledge, and "tsuzumi®", a predetermined large-scale language model on which adapter tuning is performed (see, for example, References 1 and 2).
[0041] (Reference 1):ChatGPT(OpenAI),<URL:https: / / openai.com / chatgpt> ,<Searched on October 7, 2021> (Reference 2): NTT version of large-scale language model "tsuzumi",<URL:https: / / www.rd.ntt / research / LLM_tsuzumi.html> ,<Searched on October 7, 2021>
[0042] Furthermore, the response device 100 according to this embodiment can perform processing using a predetermined learning model that falls within the category of machine learning models, in addition to the aforementioned "ChatGPT®" and "tsuzumi®".
[0043] (Control unit 130) Now, let's return to Figure 2 and continue the explanation. The control unit 130 has an internal memory for temporarily storing programs and processing data that define various processing procedures of the response device 100, and is realized by electronic circuits such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit), and integrated circuits such as an ASIC (Application Specific Integrated Circuit) and an FPGA (Field Programmable Gate Array). As shown in Figure 2, the control unit 130 has an acquisition unit 131, a learning unit 132, a receiving unit 133, a classification unit 134, a generation unit 135, and an output unit 136.
[0044] (Acquisition part 131) The acquisition unit 131 acquires user information from an external information processing device, etc., via the communication unit 110, etc., which is used by the output unit 136 (described later) to output a response to the user initiating the conversation by mimicking the voice quality of the user the user is talking to.
[0045] For example, the acquisition unit 131 acquires information such as attribute information of the conversation destination user ("attribute information" in Figure 3), voice information of the conversation destination user ("voice information" in Figure 3), and conversation characteristics of the conversation destination user ("conversation characteristics" in Figure 3) from the information processing device that manages the conversation destination user information, and stores it in the conversation destination user information B121.
[0046] (Learning Section 132) The learning unit 132 inputs the attribute information of the conversational user, the voice information of the conversational user, and the conversational features of the conversational user into a large-scale language model using prompts, and trains the large-scale language model.
[0047] Specifically, the learning unit 132 learns a large-scale language model to generate a response that reproduces the "conversation method" of the user initiating the conversation, using attributes such as gender and age, sounds such as pitch and timbre, and conversational features such as the way the conversational user speaks and the words they use.
[0048] (Reception desk 133) The reception unit 133 receives the conversation from the user who initiated the conversation, via a terminal device such as a telephone or smartphone operated by the user. The reception unit 133 then converts the received conversation into text data, audio data, etc., and transmits it to the classification unit 134, which will be described later.
[0049] (Classification section 134) The classification unit 134 classifies the user initiating the conversation based on the input natural language conversation. Specifically, the classification unit 134 uses at least one of the following, extracted based on the input natural language conversation: the voice quality of the user initiating the conversation, the content of the conversation, and information identifying the user initiating the conversation, to classify the user engaging in the conversation into a first user, a second user, and a third user.
[0050] Here, an example of the classification process by the classification unit 134 will be explained using Figure 6. Figure 6 is a diagram showing an example of the classification process according to the embodiment. Figure 6 shows an example in which the classification unit 134 classifies the conversation source user using a large-scale language model as an example of the classification process. Note that the classification process using the large-scale language model shown in Figure 6 is merely an example, and classification processing of the conversation source user using other known technologies may also be performed.
[0051] The classification unit 134 inputs a command to the large-scale language model, using a prompt (Figure 6(1)), to perform a predetermined classification on the user in question, in response to the conversation of the user in question, causing the large-scale language model to perform the classification of the user in question. Specifically, as shown in Figure 6(1), the classification unit 134 uses prompts including "<role>", "<constraints>", "<command>", etc., to cause the large-scale language model to classify the user making the call.
[0052] The "<Role>" above is text that defines the role of a person who embodies a specific role in the large-scale language model and generates information for the large-scale language model. As described above, the classification unit 134 can use a prompt that includes "Role Definition" to accurately classify users who make calls to the large-scale language model based on the desired classification.
[0053] For example, the classification unit 134 can assign the large-scale language model the role of "classifying the initiating user who made the call based on predetermined conditions" by using a prompt that includes the "<Role>" shown in (1-1) of Figure 6. As a result, the classification unit 134 can perform more accurate classification processing by having the large-scale language model act as an expert in classifying users who make calls.
[0054] Furthermore, the "<Constraints>" mentioned above is text that describes predetermined constraints on the execution of processing when the large-scale language model generates information. As described above, the classification unit 134 can accurately classify users who make calls to the large-scale language model based on the desired classification by using prompts that include "processing instructions based on processing constraints".
[0055] For example, the classification unit 134 can use prompts containing the "<Constraints>" shown in Figure 6 (1-2) to impose constraints on the large-scale language model, such as "use the conversation content of the original user" and "classify the original user according to the classification conditions." As a result, the classification unit 134 can prevent irrelevant classifications from being performed when executing the original user classification process on the large-scale language model, and can perform a more accurate original user classification process.
[0056] Furthermore, the "<command>" mentioned above is a text containing instructions for causing the large-scale language model to perform a desired process. The classification unit 134 can cause the large-scale language model to perform a desired process by inputting a prompt containing the "<command>" that describes the desired process into the large-scale language model.
[0057] For example, the classification unit 134 can use a prompt containing the "<command>" shown in (1-3) of Figure 6 to cause the large-scale language model to perform processes such as "reading the conversation content (text) of the initiating user" and "classifying the initiating user according to the voice quality of the receiving user, the conversation content, information identifying the initiating user (telephone number, etc.), and the sentiment of the conversation estimated from the voice quality and conversation content." The "sentiment of the conversation" mentioned above may be estimated based on known sentiment estimation techniques using the content of the text included in the conversation content, the voice information (voice quality) of the initiating user, etc.
[0058] Then, based on the above prompt (Figure 6(1)), the classification unit 134 can classify the user who initiated the conversation as shown in Figure 6(2). For example, if the user 10a before classification falls under the category of "conversation continues for more than one minute, and the purpose is unclear," the classification unit 134 classifies the user 10a as "the first user."
[0059] Furthermore, if user 10b before classification falls under the category of "the user initiating the conversation is angry, and the conversation contains content related to complaints," the classification unit 134 classifies user 10b as a "second user." Also, if users 10c and 10d before classification fall under the category of "registered as spam calls, and the conversation contains content related to fraud," the classification unit 134 classifies user 10c or user 10d as a "third user."
[0060] (Generation unit 135) The generation unit 135 generates a response by inputting a command to the large-scale language model, which is set according to the classification of the conversation source user, to generate a response corresponding to the content of the conversation in natural language by the conversation source user.
[0061] Specifically, the generation unit 135 inputs a command to a large-scale language model using a prompt expressed in natural language text to generate a summary of the conversation content or a response including interjections, based on input from the user initiating the conversation, and generates the summary information or interjections as a response. Furthermore, the generation unit 135 generates a response using a voice quality identical or similar to that of a predetermined person. A specific example of the response generation process by the generation unit 135 will be explained in subsequent sections.
[0062] (Output section 136) The output unit 136 outputs the generated responses to the user who initiated the conversation. For example, the output unit 136 outputs responses such as acknowledgments, summaries, and confirmations of summaries generated by the generation unit 135 to the user who initiated the conversation and is classified as a first user. The output unit 136 also outputs responses such as acknowledgments and summaries generated by the generation unit 135 to the user who initiated the conversation and is classified as a second user. Furthermore, the output unit 136 outputs summaries generated by the generation unit 135 and responses such as disconnecting the conversation to the user who initiated the conversation and is classified as a third user. Specific examples of the response output processing by the output unit 136 will be explained in subsequent sections.
[0063] (An example of processing) From here, using Figures 7 to 10, an example of the response generation and output processing by the response device 100 will be explained. The first example shown below is an example in which the response device 100 outputs responses such as acknowledgments, summaries, and confirmation of summaries to a conversation starter user classified as a first user. The second example is an example in which the response device 100 outputs responses such as acknowledgments and summaries to a conversation starter user classified as a second user. The third example is an example in which the response device 100 outputs responses such as summaries and conversation termination to a conversation starter user classified as a third user.
[0064] The large-scale language models used in the following examples 1 through 3 are, for example, models that have been pre-trained by a learning unit, which provides attribute information of the conversational user, voice information of the conversational user, and conversational features of the conversational user.
[0065] (Example 1) First, as a first example, we will explain, using Figure 7, an example in which the response device 100 generates and outputs an appropriate response to a first user who is classified as a conversation source user who continues to talk for a long time without stating the original requirements. Figure 7 is a diagram showing an example of the generation and output of a response to a first user according to the embodiment.
[0066] In the first example, when the user engaging in the conversation is classified as a first user, the response device 100 (generation unit) inputs a command to the large-scale language model using a prompt to generate summary information or acknowledgments as a response, corresponding to the content of the conversation input by the first user, and generates summary information or acknowledgments as a response.
[0067] Furthermore, the response device 100 (generation unit) inputs a command to the large-scale language model using prompts to generate a determination result regarding whether the summary information entered by the first user and the summary information generated based on the first user's conversation are identical or similar in content, and generates the determination result.
[0068] For example, as shown in (1) of Figure 7, the response device 100 (generation unit) uses prompts including "<role>", "<constraints>", "<commands>", etc., to cause a large-scale language model to generate a response or judgment result.
[0069] Specifically, the response device 100 (generation unit) uses prompts with the "<Role>" shown in (1-1) of Figure 7 to assign the large-scale language model the role of "a telephone operator who provides appropriate responses to a user (first user) who speaks ramblingly for a long time."
[0070] Furthermore, the response device 100 (generation unit) uses prompts containing the "<Constraints>" shown in (1-2) of Figure 7 to impose constraints on the large-scale language model, such as "nodding along while the original user's conversation continues," "not interrupting the original user's conversation midway," and "responding by mimicking the voice quality and conversational characteristics of the other user."
[0071] Furthermore, the response device 100 (generation unit) uses prompts with the "<command>" shown in (1-3) of Figure 7 to cause the large-scale language model to perform tasks such as "acknowledging the conversation, summarizing the conversation content, outputting the summary, and determining whether the summary is correct."
[0072] The response device 100 (output unit) then outputs to the first user the result of determining whether the response generated in response to the first user or the summary content input by the user is correct or incorrect (Figure 7(2)).
[0073] For example, in response to the first user's conversation, "It's hot today, isn't it? By the way, yesterday... (Figure 7 (2-1))", the response device 100 (output unit) outputs an acknowledgment such as, "Yes, it's hot. I see. I understand... (Figure 7 (2-2))".
[0074] The response device 100 (output unit) outputs summary information, such as, "To summarize what you've just said, is this what you mean?... (Figure 7 (2-3))", at the moment the first user's conversation is interrupted. This allows the response device 100 to make the first user aware that they have been continuing a conversation that is unrelated to the original requirements.
[0075] The response device 100 (output unit) responds to the first user's conversation, "Yes, that's right. By the way, the reason I called today was because of the requirement of XX (Figure 7 (2-4))," by outputting a response such as, "Understood. XX is... Did you understand my explanation? (Figure 7 (2-5))." Note that this response and the question to the first user may be implemented by a known generation process based on a large-scale language model.
[0076] Here, the response device 100 receives the first user's response, "...Yes... (Figure 7 (2-6))". The conversation shown in Figure 7 (2-6) indicates that the first user is unsure whether they understand the response shown in Figure 7 (2-5).
[0077] The response device 100 (output unit) outputs the response, "Just to be sure, could you please repeat what you just told me to me again? (Figure 7 (2-7))" in response to the conversation of the first user shown in (2-6) of Figure 7.
[0078] Then, the response device 100 (output unit) outputs a response such as "Yes! That understanding is correct. (Figure 7 (2-9))" in response to the first user's conversation, "...(explains the content)...(Figure 7 (2-8))".
[0079] As described above, the response device 100 has the first user summarize the explanation, and determines whether the first user has correctly understood the explanation based on whether the summary matches the explanation.
[0080] (Second example) Next, as a second example, an example of how the response device 100 generates and outputs an appropriate response to a second user who is classified as a user who conveys a complaint or claim will be explained using Figure 8. Figure 8 is a diagram showing an example of the generation and output of a response to a second user according to the embodiment.
[0081] In the second example, when the user engaging in the conversation is classified as a second user, the response device 100 (generation unit) inputs a command to a large-scale language model using a prompt expressed in natural language text to generate summary information or acknowledgments as a response, corresponding to the content of the conversation input by the second user, and generates summary information or acknowledgments as a response.
[0082] For example, as shown in (1) of Figure 8, the response device 100 (generation unit) uses prompts including "<role>", "<constraints>", "<commands>", etc., to cause a large-scale language model to generate a response.
[0083] Specifically, the response device 100 (generation unit) uses a prompt with the "<Role>" shown in (1-1) of Figure 8 to assign the role of "a telephone operator who provides an appropriate response to a user (second user) who has a complaint."
[0084] Furthermore, the response device 100 (generation unit) uses prompts containing the "<Constraints>" shown in (1-2) of Figure 8 to impose constraints on the large-scale language model, including "nodding along while the original user's conversation is ongoing," "not interrupting the original user's conversation," and "responding by mimicking the voice quality and conversational characteristics of the other user," as well as "never negating what the original user says" and "communicating the response immediately if it is something that can be addressed immediately."
[0085] Furthermore, the response device 100 (generation unit) uses prompts with the "<command>" shown in (1-3) of Figure 8 to cause the large-scale language model to perform actions such as "acknowledging the conversation, summarizing the conversation, outputting the summary, requesting confirmation of the summary, and outputting a statement that the content has been received."
[0086] The response device 100 (output unit) then outputs the response generated in response to the second user to the second user (Figure 8 (2)).
[0087] For example, the response device 100 (output unit) outputs a response such as "We are very sorry. How can I get to you?" (Figure 8 (2-2)) in response to a second user's conversation (complaint) "What exactly is going on?" (Figure 8 (2-1)). This response may be implemented by a known generation process based on a large-scale language model.
[0088] The response device 100 (output unit) outputs summary information, such as, "To summarize what you've just said, is this correct? (Figure 8 (2-3))," at the moment the second user's conversation is interrupted. This allows the response device 100 to inform the second user that it has accurately received the content of the complaint.
[0089] The response device 100 (output unit) outputs a response such as "I am very sorry. First, I will take care of XX (Figure 8 (2-5))" in response to the second user's conversation, "Yes. What are you going to do? (Figure 8 (2-4))". The initial response proposal for the second user may be implemented by a known generation process based on a large-scale language model.
[0090] Here, the response device 100 outputs an acknowledgment such as "I understand (Figure 8 (2-6))" in response to the first user's conversation, "I understand (Figure 8 (2-7))."
[0091] As described above, the response device 100 outputs acknowledgments and summary information in a way that suppresses the second user's anger, thereby performing appropriate complaint handling.
[0092] (Third example) Next, as a third example, an example of how the response device 100 generates and outputs an appropriate response to a third user classified as a user making sales calls or fraudulent calls will be described using Figures 9 and 10. Figures 9 and 10 are diagrams showing an example of the generation and output of a response to a third user according to the embodiment.
[0093] Figure 9 shows an example in which the answering device 100 generates an appropriate response to a third user classified as a user making a beneficial sales call and outputs it to that third user. Figure 10 also shows an example in which the answering device 100 generates an appropriate response to a third user classified as a user making a fraudulent call and outputs it to that third user.
[0094] In the third example, when the user engaging in the conversation is classified as a third user, the response device 100 (generation unit) inputs a command to a large-scale language model using a prompt expressed in natural language text to generate summary information or acknowledgments as a response, corresponding to the content of the conversation input by the third user, and generates summary information or acknowledgments as a response.
[0095] The response device 100 (output unit) outputs the generated summary information as a response to the user the third user wishes to converse with, if the user engaging in the conversation is classified as a third user and the third user does not meet predetermined conditions. The response device 100 (output unit) also disconnects the conversation if the user engaging in the conversation is classified as a third user and the third user meets predetermined conditions.
[0096] For example, the response device 100 (generation unit), as shown in (1) of Figure 9, uses prompts including "<role>", "<constraints>", "<commands>", etc., to cause a large-scale language model to generate a response.
[0097] Specifically, the response device 100 (generation unit) uses a prompt with the "<Role>" shown in (1-1) of Figure 9 to assign the role of "a telephone operator who provides an appropriate response to a user (third user) who is presumed to be making sales calls or fraudulent calls" to the large-scale language model.
[0098] Furthermore, the response device 100 (generation unit) uses prompts containing the "<Constraints>" shown in (1-2) of Figure 9 to impose constraints on the large-scale language model, including "nodding while the original user's conversation is ongoing," "not interrupting the original user's conversation midway," and "responding by mimicking the voice quality and conversational characteristics of the other user," as well as "using conversation logs from past sales calls or scam calls."
[0099] Furthermore, the response device 100 (generation unit) uses prompts containing "<command>" as shown in (1-3) of Figure 9 to perform actions such as "acknowledging and summarizing the conversation content" in addition to "determining whether the user who initiated the conversation is making a nuisance or fraudulent call," "outputting a message to call back if it is not a nuisance sales call or fraudulent call," and "disconnecting the conversation if it is a nuisance sales call or fraudulent call."
[0100] The above-mentioned "determination of whether the user making the call is a user making spam or fraudulent calls" means, in addition to the classification of the user making the call (third user) by the classification unit 134, the determination of whether the content of the call made by the third user is "a beneficial sales call or a spam sales call" and "whether or not it is a fraudulent call".
[0101] Here, using Figure 9(2), an example of the output of a response to a third user classified as a user making a beneficial sales call will be explained. As shown in Figure 9(2), the response device 100 (output unit) outputs a response generated in accordance with the third user making a beneficial sales call to that third user.
[0102] For example, the response device 100 (output unit) outputs responses such as "Yes, I see..." (Figure 9 (2-2)) in response to a conversation by a third user making a profitable sales call, such as "My name is XX. I'd like to introduce you to our XX product..." (Figure 9 (2-1)).
[0103] The response device 100 (output unit) outputs summary information, such as, "To summarize what you've just said, is this correct? (Figure 9 (2-3))" at the moment the third user's conversation is interrupted.
[0104] This allows the answering device 100 to verify whether the summary information about the third user's conversation is accurate. Furthermore, the answering device 100 performs a determination of whether the user making the call is a spam or fraudulent caller based on the registered phone number of the user making the call and past conversation logs. As a result, the answering device 100 determines that the call from the third user shown in Figure 9 is a "beneficial sales call".
[0105] The response device 100 (output unit) outputs a response such as "Yes, that's right. How about it? (Figure 9 (2-4))" to the third user's conversation. The response to the third user may be implemented by a known generation process based on a large-scale language model.
[0106] As described above, when the third user's sales call is beneficial, the response device 100 appropriately transmits summary information to the user without unnecessarily disconnecting the conversation.
[0107] Next, using Figure 10, we will explain an example of the output of a response to a third user classified as a user making nuisance sales calls or fraudulent calls. As shown in Figure 10(2), the response device 100 (output unit) outputs a response generated in response to the third user making nuisance sales calls or fraudulent calls to that third user. Note that the prompt shown in Figure 10(1) is the same as that in Figure 9(1), so its explanation is omitted.
[0108] For example, the response device 100 (output unit) outputs responses such as "Yes, I see..." (Figure 10 (2-2)) in response to a conversation from a third user making a fraudulent call, such as "My name is XX. I have a very profitable investment product to offer..." (Figure 10 (2-1)).
[0109] The response device 100 (output unit) outputs summary information, such as, "To summarize what you've just said, is this correct? (Figure 10 (2-3))", at the moment the third user's conversation is interrupted.
[0110] This allows the answering device 100 to verify the accuracy of the summary information regarding the third user's conversation. Furthermore, the answering device 100 performs a determination of whether the original caller is a user making spam or fraudulent calls, based on past conversation logs, etc. As a result, the answering device 100 determines that the call made by the third user shown in Figure 10 is a fraudulent call.
[0111] The response device 100 (output unit) outputs a response such as "No thank you. Please don't call again. If you persist, we will take legal action. (Figure 10 (2-5))" in response to the third user's conversation, "Yes, that's right. How about it? (Figure 10 (2-4))". The response to the third user may be implemented by a known generation process based on a large-scale language model.
[0112] As described above, if the call from the third user is a fraudulent call, the response device 100 will disconnect the call without connecting it to the user on the other end of the line in order to protect the user on the other end of the line.
[0113] (Procedure for processing by the response device 100) Next, the processing procedure implemented by the response device 100 according to this embodiment will be explained using Figure 11. Figure 11 is a flowchart showing the processing performed by the response device 100 according to this embodiment.
[0114] The response device 100 waits to process until a call is made (No in S101). When a call is made (Yes in S101), the response device 100 starts processing.
[0115] The classification unit 134 classifies the user who initiated the conversation based on the conversation (S102). Next, the generation unit 135 inputs a command to generate a response to the user who initiated the conversation, using a prompt corresponding to the classification of the user who initiated the conversation, into the large-scale language model and generates the response (S103).
[0116] The output unit 136 outputs the generated response in a predetermined format (S104). Then, the response device 100 completes the process.
[0117] (effect) The effects of the response device 100 according to this embodiment will now be explained. Conventionally, it can be difficult to take an appropriate response to protect the receiving user from the initiating user when the initiating user is someone who continues the conversation for a long time, a difficult customer, a fraudster, or someone making annoying sales calls.
[0118] Therefore, the classification unit 134 of the response device 100 according to this embodiment classifies the user who initiated the conversation based on the natural language conversation that is input. The generation unit 135 of the response device 100 inputs a command to generate a response corresponding to the content of the natural language conversation by the user who initiated the conversation, and prompts set according to the classification of the user who initiated the conversation, into a large-scale language model, and generates a response. The output unit 136 of the response device 100 outputs the generated response to the user who initiated the conversation.
[0119] Therefore, the response device 100 according to this embodiment has the effect of enabling an appropriate response according to the source of the call. Furthermore, the response device 100 according to this embodiment achieves predetermined effects by performing the processes described below.
[0120] The generation unit 135 inputs a command to the large-scale language model using prompts to generate a summary of the conversation content or a response including interjections, based on the input from the user who initiated the conversation, and generates the summary information or interjections as a response.
[0121] Through the process described above, the response device 100 can generate interjections and summary information corresponding to the conversation of the user initiating the conversation, and output them to that user. As a result, the response device 100 has the effect of enabling appropriate responses according to the conversation of the user initiating the conversation.
[0122] The generation unit 135 generates a response using a voice quality identical or similar to that of a predetermined person. Through the above-described process, the response device 100 is able to provide a response that mimics the user on the other end of the line. As a result, the response device 100 enables the user who initiated the call to make the call without feeling any discomfort, and achieves an appropriate response for that user.
[0123] The classification unit 134 uses at least one of the following, extracted based on the input natural language conversation: the voice quality of the conversation source user, the content of the conversation, and information that identifies the conversation source user, to classify the conversation source user into a first user, a second user, or a third user.
[0124] Through the process described above, the response device 100 classifies the user initiating the conversation based on predetermined classification conditions, thereby enabling an appropriate response according to the classified user. As a result, the response device 100 has the effect of enabling an appropriate response in each situation, even in cases where the appropriate response differs significantly, such as "when the conversation is unnecessarily long," "when it is a complaint," "when it is a beneficial sales call," or "when it is a nuisance sales call or a scam call."
[0125] When the user engaging in the conversation is classified as a first user, the generation unit 135 prompts the large-scale language model to input a command to generate a determination result regarding whether the summary information entered by the first user and the summary information generated based on the first user's conversation are identical or similar in content, and generates the determination result.
[0126] Through the processing described above, the response device 100 is able to clarify the content that the user who initiated the conversation (the first user) who has been classified as a user who engages in meaningless, long conversations actually wanted to discuss, and then provide an appropriate response. As a result, the response device 100 has the effect of improving the first user's satisfaction with the response. Furthermore, the response device 100 has the effect of reducing the burden on telephone operators and others who were previously forced to listen to meaningless, long conversations.
[0127] When the user engaging in the conversation is classified as a second user, the generation unit 135 inputs a command to the large-scale language model using a prompt to generate an interjection as a response corresponding to the content of the conversation input by the second user, and generates an interjection as a response.
[0128] Through the processing described above, the response device 100 enables the conversational user (second user), who has been classified as a user making a complaint, to receive the complaint details and provide initial support without provoking the second user's emotions. As a result, the response device 100 has the effect of improving the second user's satisfaction with the complaint handling. Furthermore, the response device 100 has the effect of reducing the distress and burden on telephone operators who have to handle complaints, which can sometimes be distressing.
[0129] The output unit 136 outputs the generated summary information as a response to the user the third user wishes to interact with, if the user engaging in the conversation is classified as a third user and the third user does not meet predetermined conditions. The output unit 136 also disconnects the conversation if the user engaging in the conversation is classified as a third user and the third user meets predetermined conditions.
[0130] Through the processing described above, the response device 100 enables appropriate responses to the initiating user (third user) whose usefulness is unknown. Specifically, if the initiating user makes a "useful sales call," the response device 100 enables appropriate responses such as transferring the call to the appropriate person without dismissing it. Furthermore, if the initiating user makes a "nuisance sales call" or a "fraudulent call," the response device 100 effectively protects the initiating user by terminating the conversation without transferring it to the intended user.
[0131] <Variation> The following describes modifications implemented by the response device 100 according to this embodiment.
[0132] (Data, etc.) The terms used in the description of the above embodiment, such as the conversation recipient user, the conversation source user, the first to third users, interjections, conversation summarization, conversation termination, user classification, the names of the functional parts of the response device 100, steps, processes, and the names of steps or processes, are merely examples and can be changed at will.
[0133] For example, while it was explained that the conversation destination user information DB121 stores information related to each of the following items in a table format, such as "No," which identifies individual data included in the conversation destination user information, "User identification information," "Attribute information," "Voice information," and "Conversation features," the items and information within each item that are stored are not limited to those shown in Figure 3. Similarly, while it was explained that the classification condition DB122 stores information related to each of the following items in a table format, such as "No," which identifies individual data included in the classification conditions, "Condition," and "Classification name," the items and information within each item that are stored are not limited to those shown in Figure 4. Similarly, while it was explained that the summary information DB123 stores information related to each of the following items in a table format, such as "No," which identifies individual data included in the summary information, and "Summary," the items and information within each item that are stored are not limited to those shown in Figure 5.
[0134] (Regarding the use of generative models) In this embodiment, the model (large-scale language model) used by the response device 100 is described as being stored in the generation model DB 124 of the storage unit 120, but this is not limited to this. For example, the response device 100 can access an external information processing device (server, etc.) and use a predetermined model.
[0135] (Flowcharts, etc.) In flowcharts, each step may be rearranged as long as it does not create inconsistencies, and some steps may be omitted. Furthermore, conjunctions such as "next," "continue," "in addition," "at this time," and "on this occasion" in flowchart descriptions do not limit the order or timing of the processes in the flowchart.
[0136] <Hardware Configuration> Each component of the illustrated device is a functional concept and does not necessarily have to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. Furthermore, each processing function performed by each device can be implemented, all or any part of it, by a CPU and the program that is analyzed and executed by that CPU, or by hardware using wired logic.
[0137] Furthermore, among the processes described in this embodiment, all or part of those described as being performed automatically can be performed manually using known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the drawings can be arbitrarily changed unless otherwise specified.
[0138] <Program> In one embodiment, the various devices constituting the response device 100 can be implemented by installing a response program as packaged software or online software on a desired computer. For example, by having the above-mentioned response program executed on an information processing device, the various devices constituting the response device 100 can be made to function. The information processing device referred to here includes desktop or notebook personal computers. In addition, the information processing device also includes mobile communication terminals such as smartphones and mobile phones, and slate terminals such as PDAs (Personal Digital Assistants).
[0139] Figure 12 shows an example of a computer that implements the response device 100 according to the embodiment. The computer 1000 has, for example, memory 1010 and CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0140] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.
[0141] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the various devices constituting the response device 100 is implemented as a program module 1093 in which code executable by a computer is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for performing processes similar to the functional configuration of the various devices constituting the response device 100 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0142] Furthermore, the configuration data used in the processing of the embodiment described above is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes the processing of the embodiment described above.
[0143] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN, WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read from the other computer by the CPU 1020 via a network interface 1070.
[0144] <Other> Although this embodiment has been described above, this embodiment is not limited by the description and drawings that constitute part of the disclosure. That is, all other embodiments, examples, and operational techniques made by those skilled in the art based on this embodiment are included in the scope of this embodiment. [Explanation of Symbols]
[0145] 100 Response device 110 Communications Department 120 Storage section 121 Conversation Destination User Information Database 122 Classification condition DB 123 Summary Information Database 124 Generative Model DB 130 Control Unit 131 Acquisition Department 132 Learning Department 133 Reception Department 134 Classification Department 135 Generation part 136 Output section
Claims
1. A classification unit that classifies the user who initiated the conversation based on the input natural language conversation, A command to generate a response corresponding to the natural language conversation content by the user initiating the conversation inputs prompts set according to the classification of the user initiating the conversation to a large-scale language model, and generates the response. An output unit that outputs the generated response to the user who initiated the conversation, A response device characterized by having the following features.
2. The generating unit is A command to generate a response including summary information or acknowledgments of the conversation content input by the user initiating the conversation is input to the large-scale language model using the prompt expressed in natural language text, thereby generating the summary information or acknowledgments as the response. The response device according to feature 1.
3. The generating unit is The response is generated using a voice quality identical or similar to that of a specified person. The response device according to feature 2.
4. The aforementioned classification unit is Based on the input natural language conversation, the system classifies the user into a first user, a second user, or a third user using at least one of the following: the voice quality of the user who initiated the conversation, the content of the conversation, and information identifying the user who initiated the conversation. The response device according to any one of claims 1 to 3.
5. The generating unit is If the user engaging in the aforementioned conversation is classified as the first user, The system generates the determination result by inputting a command to the large-scale language model using a prompt expressed in natural language text, which determines whether the summary information entered by the first user and the summary information generated based on the conversation of the first user are identical or similar in content. The response device according to feature 4.
6. The generating unit is If the user engaging in the aforementioned conversation is classified as a second user, A command to generate an interjection as a response corresponding to the content of the conversation input by the second user is input to the large-scale language model using the prompt expressed in natural language text, and the interjection is generated as the response. The response device according to feature 4.
7. The output unit is, If the user engaging in the aforementioned conversation is classified as a third user, and the third user does not meet the predetermined conditions, the generated summary information is output as the response to the user to whom the third user wishes to engage in dialogue. The conversation is terminated if the user engaging in the aforementioned conversation is classified as a third user, and the third user meets the predetermined conditions. The response device according to feature 4.
8. The generating unit is As the aforementioned large-scale language model, at least one of the following is used: a large-scale language model possessing general knowledge, and a predetermined large-scale language model on which adapter tuning is performed. The response device according to any one of claims 1 to 3.
9. A response method to be executed by a response device, A classification process that classifies the user who initiated the conversation based on the input natural language conversation, A command to generate a response corresponding to the conversation content in natural language by the initiating user involves inputting prompts set according to the classification of the initiating user into a large-scale language model to generate the response, and a generation process is also included. Output step of outputting the generated response to the user who initiated the conversation, A response method characterized by including
10. A classification step that classifies the user who initiated the conversation based on the input natural language conversation, A command to generate a response corresponding to the natural language conversation content by the initiating user involves inputting prompts set according to the classification of the initiating user into a large-scale language model to generate the response, and An output step of outputting the generated response to the user who initiated the conversation, A response program that instructs a computer to execute a command.