Conversation provision system and device

The system allows users to comfortably end conversations with virtual agents by using AI to detect intent and adjust the agent's presentation, addressing the pressure felt when the agent expects speech.

JP7866106B1Active Publication Date: 2026-05-26NTT DOCOMO INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NTT DOCOMO INC
Filing Date
2025-04-14
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Users may feel pressured to continue a conversation with a virtual agent when the agent remains in a mode indicating it is expecting a speech, even when they want to end the conversation.

Method used

A conversation provision system and device that estimates when a user intends to end the conversation and presents the virtual agent in a manner indicating it does not expect further utterances, using artificial intelligence to analyze conversation history and user behavior.

Benefits of technology

Enables a smooth and comfortable termination of conversations with virtual agents by accurately detecting user intent and adjusting the agent's presentation to reflect non-expectation of user input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007866106000001_ABST
    Figure 0007866106000001_ABST
Patent Text Reader

Abstract

This system provides a conversation service that allows users to comfortably end conversations with virtual agents. [Solution] The conversation provision system comprises conversation units 54, 56, and 58 that instruct a task processing device 32, which processes tasks using artificial intelligence in response to task processing instructions, to generate conversation content for a virtual agent to the user, and acquire the conversation content generated by the task processing device 32 and output it to the user; and presentation units 46, 48, and 52 that present an embodiment agent, which is an embodiment of the virtual agent, to the user, and if it is presumed that the user intends to end the conversation with the virtual agent, the presentation units 46, 48, and 52 present the embodiment agent in a manner that indicates that they do not expect any utterances from the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a conversation providing system and apparatus that provides a conversation with a virtual agent to a user using artificial intelligence.

Background Art

[0002] For users, conversations with virtual agents are provided using artificial intelligence. Regarding such virtual agents, during negotiations with users, the user's emotions are analyzed, and in response to the user's emotions, aspects of the image showing the virtual agent presented to the user, such as expressions and gestures, are changed (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Regarding the aspect of the image showing the virtual agent presented to the user, it is in a mode indicating that it is expecting a speech from the user. However, even when the user wants to end the conversation with the virtual agent, if it remains in the mode indicating that it is expecting a speech from the user, the user may feel pressured to continue the conversation with the virtual agent.

[0005] An object of the present invention is to provide a conversation providing system and apparatus that can comfortably end a conversation with a virtual agent.

Means for Solving the Problems

[0006] One aspect of the present invention is a conversation provision system comprising: a conversation unit that instructs a task processing device, which processes tasks using artificial intelligence in response to task processing instructions, to generate conversation content for a virtual agent to a user, and acquires the conversation content generated by the task processing device and outputs it to the user; and a presentation unit that presents an embodiment agent, which is an embodiment of the virtual agent, to the user, and if it is estimated that the user intends to end the conversation with the virtual agent, the presentation unit presents the embodiment agent in a manner that indicates it does not expect any utterance from the user.

[0007] Another aspect of the present invention is a conversation providing device comprising: a conversation output unit that instructs a task processing device, which processes tasks using artificial intelligence in response to task processing instructions, to generate conversation content for a virtual agent to a user, and acquires and outputs the conversation content generated by the task processing device; and a presentation instruction unit that instructs the device to present an embodiment agent, which is an embodiment of the virtual agent, to the user, and if it is presumed that the user intends to end the conversation with the virtual agent, the presentation instruction unit that instructs the device to present the embodiment agent in a manner that indicates it does not expect any utterance from the user. [Effects of the Invention]

[0008] This invention makes it possible to smoothly end a conversation with a virtual agent. [Brief explanation of the drawing]

[0009] [Figure 1] A block diagram showing the hardware configuration of a conversation provisioning system according to one embodiment of the present invention. [Figure 2] A block diagram showing the functional configuration of a conversation provision system according to one embodiment of the present invention. [Figure 3A] A first sequence diagram illustrating a conversation delivery method according to one embodiment of the present invention. [Figure 3B] A second sequence diagram illustrating a conversational delivery method according to one embodiment of the present invention. [Figure 3C] A third sequence diagram illustrating a conversational delivery method according to one embodiment of the present invention. [Figure 4] A table / diagram illustrating a method for estimating a user's intention to terminate service in one embodiment of the present invention. [Figure 5] A table / diagram illustrating the characteristics of an agent image in one embodiment of the present invention. [Modes for carrying out the invention]

[0010] An embodiment of the present invention will be described with reference to Figures 1 to 5. In the following, a virtual agent that engages in conversation with a user using artificial intelligence will be referred to as an AI agent. While such AI agents may provide information to the user according to a specific purpose or assist the user in performing various procedures, AI agents that primarily engage in conversation with the user are expected to go beyond simply providing information and performing procedures according to a specific purpose, and instead build a close relationship with the user, gain the user's trust, and become a supportive presence for the user.

[0011] In this embodiment, even when a user wants to end a conversation with the AI ​​agent, if the agent image representing the AI ​​agent remains in a state that indicates it is expecting a response from the user, the user may feel pressured to continue the conversation with the AI ​​agent. Therefore, when it is estimated that the user intends to end the conversation with the AI ​​agent, the agent image is displayed in a state that indicates it is not expecting a response from the user. This allows the user to comfortably end the conversation with the AI ​​agent and enables the building of a better relationship between the user and the AI ​​agent.

[0012] As shown in Figures 1 and 2, the conversation provision system of this embodiment includes a conversation provision server 10, a generation server 30, and a user terminal 40.

[0013] Referring to Figure 1, the hardware configuration of the conversation provisioning system of this embodiment will be described. Since the hardware configurations of the conversation provider server 10 and the generation server 30 are the same, they will be described together as servers below.

[0014] As shown in Figure 1, servers 10 and 30 are physically configured as computers including a processor 211, memory 212, storage 213, communication device 214, input device 215, output device 216, and buses connecting them. Each of these devices operates on power supplied from a battery (not shown). In the following description, the term "device" can be read as a circuit, device, unit, etc. The hardware configuration of servers 10 and 30 may include one or more of the devices shown in the figure, or it may be configured without some of the devices. Alternatively, multiple devices with different enclosures may be connected via communication to constitute servers 10 and 30.

[0015] Each function in servers 10 and 30 is realized by loading predetermined software (programs) onto hardware such as the processor 211 and memory 212, which causes the processor 211 to perform calculations, control communication by the communication device 214, and control at least one of the reading and writing of data in the memory 212 and storage 213.

[0016] The processor 211 controls the entire computer, for example, by running the operating system. The processor 211 may consist of a central processing unit (CPU) that includes interfaces with peripheral devices, control units, arithmetic units, registers, etc. Alternatively, a baseband signal processing unit, a call processing unit, etc., may be implemented by the processor 211.

[0017] Processor 211 reads programs (program codes), software modules, data, etc. from at least one of storage 213 and communication device 214 into memory 212, and executes various processes according to them. As the program, a program for causing a computer to execute at least a part of the operations described later is used. The functional blocks of servers 10 and 30 may be stored in memory 212 and realized by a control program operating in processor 211. Various processes may be executed by one processor 211, or may be executed simultaneously or sequentially by two or more processors 211. Processor 211 may be implemented by one or more chips. Note that the program may be transmitted to servers 10 and 30 via a telecommunication line.

[0018] Memory 212 is a computer-readable recording medium, and may be constituted by at least one of, for example, ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 212 may be referred to as a register, a cache, a main memory (main storage device), etc. Memory 212 can store a program (program code), a software module, etc. executable for implementing the method according to the present embodiment.

[0019] Storage 213 is a computer-readable recording medium, and may be constituted by at least one of, for example, optical discs such as CD-ROM (Compact Disc ROM), hard disk drives, flexible disks, magneto-optical disks (e.g., compact discs, digital versatile discs, Blu-ray (registered trademark) discs), smart cards, flash memories (e.g., cards, sticks, key drives), floppy (registered trademark) disks, magnetic strips, etc. Storage 213 may be referred to as an auxiliary storage device.

[0020] The communication device 214 is hardware (a transceiver device) for performing communication between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, a communication module, etc.

[0021] Each device such as the processor 211 and the memory 212 is connected by a bus for communicating information. The bus may be configured using a single bus, or may be configured using different buses for each device.

[0022] The servers 10, 30 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc., and some or all of each functional block may be realized by the hardware. For example, the processor 211 may be implemented using at least one of these hardware.

[0023] The user terminal 40 is a computer such as a smartphone, a mobile phone, a tablet, or a wearable terminal. Physically, the user terminal 40 is configured as a computer device including a processor 261, a memory 262, a storage 263, a communication device 264, an input device 265, an output device 266, and a bus connecting these. The processor 261, the memory 262, and the storage 263 are the same hardware as the processor 211, the memory 212, and the storage 213 of the servers 10, 30.

[0024] The communication device 264 may include, for example, a high-frequency switch, duplexer, filter, frequency synthesizer, etc., to implement at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the transmitting and receiving antennas, amplifier section, transmitting and receiving section, transmission path interface, etc., may be implemented by the communication device 264. The transmitting and receiving section may be physically or logically separated into a transmitting section and a receiving section.

[0025] The input device 265 is an input device that accepts input from an external source (e.g., a key, microphone, switch, button, sensor, reader, scanner, etc.). The output device 266 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 265 and the output device 266 may be configured as an integrated unit (e.g., a touchscreen).

[0026] Referring to Figures 2 and 3, the functional configuration and conversation provision method of this embodiment will be described. Hereinafter, each functional configuration of the conversation provision system will be described in accordance with each step of the conversation provision method. The conversation provision program of this embodiment works in cooperation with the hardware configuration of the conversation provision system of this embodiment described above to realize each functional configuration of the conversation provision system of this embodiment and to execute each step of the conversation provision method.

[0027] Furthermore, in one embodiment described below, the user is provided with a conversation with the AI ​​agent through voice input and voice output, but the user may also be provided with a conversation with the AI ​​agent through text input and text output on the screen.

[0028] As shown in Figure 2, in the conversation provision system, the conversation provision server stores user database 14 (hereinafter also referred to as "user DB14"), which associates with user identification information that identifies the user, agent identification information that identifies the AI ​​agent that primarily converses with the user, conversation history information that shows the conversation history between the user and the AI ​​agent, and estimation criteria information that shows the criteria for estimating whether the user intends to end the conversation with the AI ​​agent. In addition, the agent database 16 (hereinafter also referred to as "agent DB16") stores agent characteristic information that shows the characteristics of each AI agent, which is associated with the agent identification information of a large number of AI agents. The agent characteristic information includes character characteristic information that shows the character characteristics of each AI agent, and voice information for generating speech.

[0029] As shown in Figures 2 and 3A, the conversation provision system and conversation provision method display the agent image in a speech expectation state, as described below, to indicate that the system is expecting a speech from the user.

[0030] Conversation application startup step S12 The user terminal 40 has a conversation application installed for engaging in conversation with the AI ​​agent. On the user terminal 40, the application launch unit 42 launches the conversation application in response to the user's operation to launch the conversation application.

[0031] Agent designation request step S14 On the user terminal 40, the agent selection unit 44 presents multiple AI agents prepared for the conversation application and allows the user to select the AI ​​agent with which they will primarily converse. Subsequently, the agent selection unit 44 sends an agent designation request to the conversation service server 10, along with user identification information to identify the user and agent identification information to identify the AI ​​agent, designating the AI ​​agent selected by the user as the AI ​​agent that will primarily converse with the user. The multiple AI agents prepared for the conversation application correspond to existing characters and original characters of the conversation application, and the user can select their favorite character as the AI ​​agent that will primarily converse with them.

[0032] Agent designation step S16 In the conversation server 10, the agent designation unit 12 stores the agent designation information in the user database 14 in association with the user designation information, in response to an agent designation request received from the user terminal 40 along with the user designation information and agent designation information.

[0033] Speech expectation pattern setting step S18 In the user terminal 40, when the conversation application is launched by the application launch unit 42, the manner setting unit 46 sets the manner of the agent image, which is an image representing the AI ​​agent and acts as a materialized agent, to the expected speech manner that anticipates a utterance from the user. The agent image representing the AI ​​agent includes not only an image representing the AI ​​agent itself, but also an image representing the background in which the AI ​​agent exists. Examples of expected speech manners for agent images will be described in detail later.

[0034] Speech expectation pattern display step S22 On the user terminal 40, the image generation unit 48 generates an agent image of the expected speech pattern set by the pattern setting unit 46, and the image display unit 52 displays the agent image of the expected speech pattern generated by the image generation unit 48 to the user. By seeing the agent image of the expected speech pattern that anticipates a speech from the user, the user will naturally feel the urge to speak to the AI ​​agent.

[0035] As shown in Figures 2 and 3B, the conversation provision system and conversation provision method, as described below, involve the AI ​​agent responding to the user's utterances.

[0036] Speech acquisition step S24 In the user terminal 40, the voice acquisition unit 54 acquires the voice spoken by the user to the AI ​​agent and obtains voice information indicating the spoken voice.

[0037] Answer generation request step S26 On the user terminal 40, the conversation execution unit 58 sends a response generation request to the conversation provision server 10, along with the utterance voice information and user identification information acquired by the voice acquisition unit 54, requesting the AI ​​agent to generate a response for the user. This response generation request asks the conversation generation AI to generate a response for the AI ​​agent for the user.

[0038] Speech analysis step S28 In the conversation provision server 10, the speech analysis unit 18 analyzes the speech audio information received from the user terminal 40 along with user identification information, generates speech audio text data by converting the speech audio into text, and outputs it to the conversation generation instruction unit 22 along with the user identification information.

[0039] Answer generation instruction step S32 In the conversation provision server 10, when the conversation generation instruction unit 22 receives spoken speech text data along with user identification information from the speech analysis unit 18, it extracts agent identification information and conversation history information associated with the user identification information from the user DB 14, and extracts character feature information of agent feature information associated with the agent identification information from the agent DB 16. The conversation history information shows the conversation history between the user and the AI ​​agent, and if the amount of conversation history information is large, only a certain amount of the most recent conversation history information is extracted. The character feature information shows the character characteristics of the AI ​​agent.

[0040] Next, the conversation generation instruction unit 22 generates a response generation prompt based on the utterance speech text data, conversation history information, and character feature information, and sends it to the generation server along with the user identification information. The response generation prompt instructs the conversation generation AI to generate response text data that indicates the AI ​​agent's response to the most recent utterance from the user indicated by the utterance speech text data. The instruction is to ensure that the AI ​​agent's response is natural in light of the conversation history and matches the AI ​​agent's character characteristics.

[0041] Answer generation step S34 In the generation server, the conversation generation unit 32, which acts as a task processing unit, is equipped with a generation AI for conversation generation that utilizes LLM (Large Language Models). It inputs the response generation prompt received from the conversation provision server 10 along with user identification information into the conversation generation AI, receives response text data indicating the AI ​​agent's response from the conversation generation AI, and sends the response text data along with the user identification information to the conversation provision server 10.

[0042] Conversation history update step S36 In the conversation provision server 10, the conversation generation instruction unit 22 updates the conversation history information in the user DB 14 by adding the utterance text data input from the speech analysis unit 18 along with the user identification information and the response text data received from the generation server 30 along with the user identification information to the conversation history information associated with the user identification information. The conversation generation instruction unit 22 also outputs the response text data along with the user identification information to the speech generation unit 24.

[0043] Answer voice generation step S38 In the conversation provision server 10, when response text data is input from the conversation generation instruction unit 22 along with user identification information, the voice generation unit 24 extracts agent identification information associated with the user identification information from the user DB 14, extracts voice information included in the agent feature information associated with the agent identification information from the agent DB 16, generates response voice data by converting the response text data into voice using the AI ​​agent's voice, and transmits it to the user terminal 40 identified by the user identification information. Thus, in this embodiment, the conversation output unit is formed by the conversation generation instruction unit 22 and the voice generation unit 24.

[0044] Answer voice output step S42 In the user terminal 40, the conversation execution unit 58 acquires the response voice data received from the conversation provision server 10 and outputs it to the voice output unit 56. The voice output unit 56 then outputs the AI ​​agent's response to the user as voice based on the response voice data input from the conversation execution unit 58. Thus, in this embodiment, the conversation unit is formed by the voice output unit 56 and the conversation execution unit 58.

[0045] As shown in Figures 2 and 3C, the conversation provision system and conversation provision method estimate whether the user intends to end the conversation with the AI ​​agent, as described below. If it is estimated that the user intends to end the conversation with the AI ​​agent, the agent image is displayed in a non-speech expectation mode, indicating that the system does not expect the user to speak. Hereinafter, the user's intention to end the conversation with the AI ​​agent will also be referred to as the user's termination intention.

[0046] Termination intention analysis request step S44 In the conversation provision server 10, the intention analysis instruction unit 26 extracts conversation history information associated with user identification information from the user DB 14 when conversation history information is updated in the user DB 14 by the conversation generation instruction unit 22. If the amount of conversation history information is large, it extracts a certain amount of the most recent conversation history information as conversation history information. Based on the extracted conversation history information, the intention analysis instruction unit 26 generates an termination intention analysis prompt and sends it to the generation server 30 along with the user identification information. The termination intention analysis prompt instructs the generation AI for termination intention analysis to analyze the user's termination intention from the conversation history.

[0047] Termination intention analysis step S46 In the generation server, the intention analysis unit 34 is equipped with a generation AI for termination intention analysis that uses LLM to analyze the user's intention to end the conversation from the conversation history. The intention analysis unit 34 inputs the termination intention analysis prompt received from the conversation provision server 10 along with user identification information to the generation AI for termination intention analysis, receives termination intention analysis information from the generation AI for termination intention analysis indicating whether the user intends to end the conversation with the AI ​​agent, and transmits the termination intention analysis information along with the user identification information to the conversation provision server 10.

[0048] Estimated related information gathering step S48 On the user terminal 40, the related information collection unit 62 uses various functions of the user terminal 40 to collect estimated related information for estimating the user's intention to terminate the conversation, and transmits it to the conversation provision server 10 along with the user identification information.

[0049] Termination intention estimation step S52 The intention estimation unit 28, acting as a presentation instruction unit, estimates the user's termination intention based on termination intention analysis information and estimation-related information received from the generation server and user terminal 40 along with user identification information. This information is referenced in the user DB 14, where it indicates the criteria for estimating the user's termination intention, and is associated with the user identification information. The intention estimation unit 28 may also update the estimation criteria information depending on whether the estimation of the user's termination intention is correct or incorrect. An example of the method for estimating the user's termination intention will be described in detail later.

[0050] Step S54: Instruction to set non-speech expectation behavior In the conversation provision server 10, if the intention estimation unit 28 estimates that the user intends to end the conversation with the AI ​​agent, it sends a non-utterance expectation mode setting instruction to the user terminal 40, which is identified by the user identification information, instructing the agent image to be set to a non-utterance expectation mode that indicates that the user does not expect to speak.

[0051] Step S56: Setting the expected non-speech pattern In the user terminal 40, when the mode setting unit 46 receives a non-utterance expectation mode setting instruction from the conversation provision server 10, it sets the mode of the agent image to a non-utterance expectation mode that indicates that the user does not expect to utter anything. An example of a non-utterance expectation mode for the agent image will be described in detail later.

[0052] Non-speech expectation indication step S58 In the user terminal 40, the image generation unit 48 generates an agent image of the non-speech expectation mode set by the mode setting unit 46, and the image display unit 52 displays the agent image of the non-speech expectation mode generated by the image generation unit 48 to the user. Thus, in this embodiment, the presentation unit is formed by the mode setting unit 46, the image generation unit 48, and the image display unit 52.

[0053] In this way, when a user is ready to end a conversation with an AI agent, the AI ​​agent naturally switches to a non-utterance expectation mode, not expecting any further utterances from the user. As a result, the user does not feel pressured to continue the conversation with the AI ​​agent, allowing them to comfortably end the conversation and build a better relationship with the AI ​​agent.

[0054] Refer to Figure 4 to explain the method for estimating the user's intention to terminate the service. As mentioned above, AI agents that engage in conversations with users are expected to be empathetic to the user. If the agent's image is set to a non-speaking expected state when the user does not intend to end the conversation with the AI ​​agent, the AI ​​agent will behave in a way that the user does not want, which is undesirable. Therefore, it is desirable to accurately estimate the user's intention to end the conversation.

[0055] As a method for estimating a user's intention to terminate the service, for example, the following estimation method can be used, as shown in Figure 4.

[0056] (1) Estimation based on the content of the conversation between the user and the AI ​​agent (1-1) Estimation based on the user's conversation termination statement Using a generative AI for termination intention analysis, the conversation history between the user and the AI ​​agent is analyzed. If the user makes a statement indicating they want to end the conversation towards the end of the conversation history, such as a farewell or a greeting indicating they want to end the conversation, or a declaration that they will take action other than conversation, such as "I'm going to do such-and-such" or "I'm going to such-and-such a place," it is estimated that the user wants to end the conversation with the AI ​​agent. The accuracy of this estimation can be improved by fine-tuning the generative AI for termination intention analysis itself and by applying RAG (Retrieval-Augmented Generation) to each user. Alternatively, instead of using a generative AI for termination intention analysis, the system may estimate the user's intention to end the conversation with the AI ​​agent by appropriately detecting a user's statement indicating they want to end the conversation from the user's most recent utterance.

[0057] (1-2) Estimation based on the context of the conversation Using a generative AI for termination intention analysis, the conversation history between the user and the AI ​​agent is analyzed, and if it is determined from the context of the conversation that the user has ended the conversation, it is estimated that the user intends to end the conversation with the AI ​​agent. The improvement in the accuracy of such estimations using generative AI is described in (1-1).

[0058] (2) Estimation based on user attitude (2-1) Estimation based on non-conversation time On the user terminal, the timing function provided in the conversation application measures the time elapsed since the last conversation between the user and the AI ​​agent. If the non-conversation time during which the user has not spoken with the AI ​​agent exceeds a predetermined standard non-conversation time, it is presumed that the user intends to end the conversation with the AI ​​agent. The standard non-conversation time is stored in the user DB14 as estimation standard information, associated with the user identification information. Furthermore, regarding the standard non-conversation time, if, after presuming that the user intends to end the conversation with the AI ​​agent, the user does not actually speak to the AI ​​agent and it is determined that the estimation of the user's intention to end the conversation was correct, the standard non-conversation time is maintained as is, or if it is set to a relatively long time, it is updated to a slightly shorter time. Conversely, if, despite presuming that the user intends to end the conversation with the AI ​​agent, the user actually speaks to the AI ​​agent relatively soon afterward and it is determined that the estimation was incorrect, the standard non-conversation time is updated to a slightly longer time.

[0059] (2-2) Estimation based on gaze deviation On the user terminal, the user's facial image acquired by the camera function is analyzed to obtain the direction of the user's gaze and determine whether the user's gaze is deviating from the agent image. If the user's gaze deviates from the agent image by a predetermined deviation angle or a predetermined deviation time, it is presumed that the user intends to end the conversation with the AI ​​agent. The deviation angle and deviation time are also stored as estimation reference information in the user DB14. If it is determined that the estimation of the user's intention to end the conversation was correct, these values ​​are maintained, or if they are set to a relatively large angle and long duration, they are updated to a slightly smaller angle and shorter duration. If it is determined that the estimation was incorrect, the values ​​are updated to a slightly larger angle and longer duration.

[0060] (2-3) Estimation based on worsening facial expression On the user terminal, the system analyzes the user's facial image obtained via the camera function to acquire the user's facial expression. If the user's expression worsens beyond a predetermined standard level of facial deterioration (e.g., extremely bored, sleepy), it is estimated that the user intends to end the conversation with the AI ​​agent. The standard level of facial deterioration is also stored in the user DB14 as estimation standard information. If it is determined that the estimation of the user's intention to end the conversation was correct, the value is maintained as is, or if it is relatively high, it is updated to be slightly lower. If it is determined that the estimation was incorrect, it is updated to be slightly higher.

[0061] (2-4) Estimation based on movement On the user terminal, various sensors detect the user's movement, such as walking. If the user starts moving and continues moving for longer than a predetermined standard travel time, it is presumed that the user intends to end the conversation with the AI ​​agent. The standard travel time is also stored in the user DB14 as estimated standard information. If it is determined that the estimation of the user's intention to end the conversation was correct, the standard is maintained as is, or if it is set to a relatively long time, it is updated to a slightly shorter time. If it is determined that the estimation was incorrect, it is updated to a slightly longer time.

[0062] (3) Estimation based on the user's external circumstances (3-1) Estimation based on deviation from normal conversation time On the user terminal, the timing function provided in the conversation application retrieves the time the user is conversing with the AI ​​agent. If the time deviates by a predetermined threshold from the normal conversation time when the user typically converses with the AI ​​agent, it is presumed that the user intends to end the conversation with the AI ​​agent. The normal conversation time is also stored in the user DB14 as estimation threshold information. Initially, it is set to exclude nighttime hours, and is updated as appropriate when a conversation takes place between the user and the AI ​​agent. The time outside the threshold is also stored in the user DB14 as estimation threshold information. If it is determined that the estimation of the user's intention to end the conversation was correct, it is maintained as is, or if it is set to a relatively long time, it is updated to a slightly shorter time. If it is determined that the estimation was incorrect, it is updated to a slightly longer time.

[0063] (3-2) Estimation based on the approach of the user's schedule On the user terminal, the scheduled time and location of the user's most recent schedule are obtained from the scheduling function, and the current time and current location are obtained from the timing function and current location detection function. When the current time and the user's current location are closer to the scheduled time and location of the schedule than predetermined reference approach time and reference approach distance, it is estimated that the user intends to end the conversation with the AI ​​agent. The reference approach time and reference approach distance are also stored as estimation reference information in the user DB14. If it is determined that the estimation of the user's intention to end the conversation was correct, these values ​​are maintained as they are, or if they are set to a relatively short time and short distance, they are updated to a slightly longer time and distance. If it is determined that the estimation was incorrect, they are updated to a slightly shorter time and distance.

[0064] (4) Estimation based on the operating status of other applications (4-1) Estimation based on the launch of other applications If a user launches an application other than the conversation application on their terminal, it is presumed that the user intends to end the conversation with the AI ​​agent. The other applications that are subject to this estimation of the user's intention to end the conversation are also stored in the user DB14 as estimation criteria information. If it is determined that the estimation of the user's intention to end the conversation was correct, the application is either maintained or added to the list of applications subject to estimation. If it is determined that the estimation was incorrect, the application is removed from the list of applications subject to estimation of the user's intention to end the conversation.

[0065] (4-2) Estimation based on notifications from other applications If a notification is received from an application other than the conversation application on the user's terminal, it is presumed that the user intends to end the conversation with the AI ​​agent. The other applications that are subject to the estimation of the user's intention to end the conversation are stored and updated in the same way as described in (4-1) above.

[0066] The methods for estimating a user's intention to terminate the game described above may be used individually or in combination. When used in combination, the degree of estimation of the user's intention to terminate the game may be determined by weighting, and an overall judgment may be made. When an overall judgment is made using weighting, each weight may also be stored in the user DB14 as estimation criterion information, and each weight may be updated as appropriate according to the accuracy of the estimation of the user's intention to terminate the game.

[0067] Furthermore, even if it is estimated that the user intends to end the conversation with the AI ​​agent, a condition may be added to prevent the agent image from being set to a non-speaking expected state.

[0068] For example, (2-2) Regarding estimation based on gaze deviation, even if the user's gaze deviates from the agent image by a predetermined deviation angle or a predetermined deviation time, if the user's movement is detected simultaneously by various sensors, the agent image configuration is not set to the non-speech expectation configuration, as it may be that the user is only temporarily looking away due to movement.

[0069] Furthermore, (4-2) Regarding estimations based on notifications from other applications, if the notifications from other applications are not particularly important and are merely notifications from entertainment applications that compete with the conversational application, the agent image may be kept in the speech expectation state rather than set to the non-speech expectation state, thereby encouraging the user to speak.

[0070] Refer to Figure 5 to explain the aspects of the agent image. As mentioned above, AI agents that engage in conversations with users are expected to be empathetic to the user, and it is desirable that the appearance of the agent image representing the AI ​​agent be similar to the attitude that a real person would actually take.

[0071] As examples of expected speech and non-expected speech patterns for agent images, the following patterns are used, as shown in Figure 5. In the following, the direction from the screen displaying the agent image to the user in the user terminal 40 is described as facing forward. Note that in (2) non-expected speech patterns, the degree to which the agent does not expect a speech from the user increases in the order of patterns (2-1) to (2-6).

[0072] (1) Expected manner of utterance The AI ​​agent is positioned in the background, slightly leaning forward, with its posture and gaze directed straight ahead.

[0073] (2) Non-utterance expectation (2-1) The AI ​​agent's gaze is averted from directly in front of it. (2-2) In addition to the above (2-1), the AI ​​agent's stance is not facing forward. (2-3) In addition to the above (2-2), the AI ​​agent has withdrawn somewhat. (2-4) In addition to (2-3) above, the AI ​​agent is positioned slightly behind the front in the background. For example, it is sitting on a sofa that is positioned slightly behind the front, rather than on a chair that is positioned in front. (2-5) In addition to (2-4) above, the AI ​​agent temporarily performs actions other than conversation. For example, it may operate a mobile phone, read a paperback book it is carrying, or eat or drink food and drinks that have been provided. (2-6) The AI ​​agent is performing an action that is entirely unrelated to conversation. For example, it is watching TV in the background or in a different room from the one where the conversation usually takes place, training with equipment, or sleeping in bed.

[0074] As mentioned above, it is undesirable to display the agent image in a non-speaking expectation state when the user does not intend to end the conversation with the AI ​​agent. This is especially true when a state that indicates a high degree of disinterest in user responses is selected as the non-speaking expectation state. Therefore, as described above, when combining methods for estimating the user's intention to end the conversation and making an overall judgment based on weighted estimations, it may be possible to select states that indicate a high degree of disinterest in user responses, in the order of (2-1) to (2-6), as the degree of disinterest in user responses increases.

[0075] In the embodiment described above, when it is estimated that the user wishes to end the conversation with the AI ​​agent, the agent image representing the AI ​​agent is set to indicate that it does not expect the user to speak, thus enabling a pleasant end to the conversation with the AI ​​agent. In particular, the estimation of whether the user wishes to end the conversation with the AI ​​agent is based on the content of the conversation between the user and the agent, the user's attitude, the user's external circumstances, and the operating status of applications other than the conversation application, thus enabling accurate estimation. Furthermore, regarding the agent image representing the AI ​​agent, in order to indicate that it does not expect the user to speak, the agent uses expressions such as its gaze averted from the front, its posture averted from the front, it is slightly withdrawn, it is positioned slightly behind the front in the background, it is temporarily performing an action other than conversation, or it is performing an action completely other than conversation, thus appropriately indicating to the user that the AI ​​agent does not expect the user to speak.

[0076] In the embodiment described above, speech audio information representing the user's utterance is converted into speech audio text data, a response generation prompt including the speech audio text data is generated and input to a conversation generation AI, the response text data is output from the conversation generation AI, and the response text data is converted into response audio information representing the AI ​​agent's response based on the AI ​​agent's voice information. Alternatively, a response generation prompt including speech audio information representing the user's utterance and the AI ​​agent's voice information may be generated and input to a conversation generation AI, and response audio information representing the AI ​​agent's response may be output from the conversation generation AI. In this case, the speech audio information and response audio information may be stored as conversation history information.

[0077] The key disclosures of this application can be summarized as follows: The first disclosure is a conversation-providing device comprising: a conversation unit that instructs a task processing device, which processes tasks using artificial intelligence in response to task processing instructions, to generate conversation content for a virtual agent to a user, and acquires the conversation content generated by the task processing device and outputs it to the user; and a presentation unit that presents an embodiment agent, which is an embodiment of the virtual agent, to the user, and if it is presumed that the user intends to end the conversation with the virtual agent, the presentation unit presents the embodiment agent in a manner that indicates it does not expect any utterance from the user.

[0078] In this disclosure, if it is presumed that the user wishes to end the conversation with the virtual agent, the manifested agent, which embodies the virtual agent, is presented in a manner that indicates it does not expect any further utterances from the user, thus enabling a pleasant termination of the conversation with the virtual agent.

[0079] The second disclosure is that the aforementioned estimation is based on the content of the conversation between the user and the virtual agent, and is a conversation-providing device as described in the first disclosure.

[0080] This disclosure makes it possible to appropriately estimate whether a user intends to end a conversation with a virtual agent, based on the content of the conversation between the user and the virtual agent.

[0081] The third disclosure is the conversational device of the first disclosure, which makes the aforementioned assumptions based on the user's attitude.

[0082] This disclosure allows us to estimate whether a user intends to end a conversation with the virtual agent based on their attitude, making it possible to appropriately estimate whether a user intends to end a conversation with the virtual agent.

[0083] The fourth disclosure is that the aforementioned estimation is based on the user's external circumstances and is a conversational device as described in the first disclosure.

[0084] This disclosure makes it possible to appropriately estimate whether a user intends to end a conversation with the virtual agent, based on the user's external circumstances.

[0085] The fifth disclosure is the conversation-providing device of the first disclosure, in which the presumption is made based on the operating state of functions other than the conversation section of the conversation-providing device.

[0086] This disclosure makes it possible to appropriately estimate whether a user intends to end a conversation with a virtual agent, based on the operating status of functions other than the conversation section of the conversation provider.

[0087] The sixth disclosure is a conversation-providing device of the first disclosure, wherein the presentation unit presents the materialized agent in a manner that does not face forward, in order to indicate that it does not expect an utterance from the user.

[0088] In this disclosure, the embodied agent is presented in a manner that does not face forward, as this indicates that the virtual agent does not expect any utterances from the user. Therefore, it is possible to appropriately show the user that the virtual agent does not expect any utterances from the user.

[0089] The seventh disclosure is a conversation-providing device of the first disclosure, wherein the presentation unit presents the embodiment agent in a manner that is not conversing, in order to indicate that it does not expect utterances from the user.

[0090] In this disclosure, the embodied agent is presented in a manner that indicates it does not expect user utterances, by performing actions other than conversation. This makes it possible to appropriately show the user that the virtual agent does not expect user utterances.

[0091] The eighth disclosure is that the presenting unit is a conversation-providing device as described in the first disclosure, which, even if it is presumed that the user intends to end the conversation with the virtual agent, does not present the manifestation of the agent in a manner that indicates it does not expect an utterance from the user, if the retention condition is met.

[0092] This disclosure ensures that even if it is presumed that the user wishes to end the conversation with the virtual agent, if the conditions for holding off are met, the manifested agent will not be presented in a manner that indicates it does not expect any utterance from the user, thereby encouraging the user to continue the conversation with the virtual agent.

[0093] The ninth disclosure concerns a conversation providing device comprising: a conversation output unit that instructs a task processing device, which processes tasks using artificial intelligence in response to task processing instructions, to generate conversation content for a virtual agent to a user, and acquires and outputs the conversation content generated by the task processing device; and a presentation instruction unit that instructs the device to present an embodiment agent, which is an embodiment of the virtual agent, to the user, and if it is presumed that the user intends to end the conversation with the virtual agent, the presentation instruction unit that instructs the device to present the embodiment agent in a manner that indicates it does not expect any utterance from the user. This disclosure has the same effect as the first disclosure. [Explanation of Symbols]

[0094] 10...Conversation server 12...Agent designation unit 14...User database 16…Agent DB 18…Voice Analysis Unit 22…Conversation Generation Instruction Unit 24…Voice Generation Unit 26…Intention Analysis Instruction Unit 28…Intention Estimation Unit 30…Generation Server 32…Conversation Generation Unit 34…Intention Analysis Unit 40…User Terminal 42…Application Launch Unit 44…Agent Selection Unit 46…Pattern Setting Unit 48…Image Generation Unit 52…Image Display Unit 54…Audio Acquisition Unit 56...Audio output unit 58...Conversation execution unit 62...Related information collection unit

Claims

1. A task processing device that processes tasks using artificial intelligence in response to task processing instructions, a conversation unit that instructs the task processing device to generate conversation content for a virtual agent to the user, and acquires the conversation content generated by the task processing device and outputs it to the user, A presentation unit that presents to the user an embodiment agent that embodies the virtual agent, and if it is presumed that the user intends to end the conversation consisting of utterances and responses between the user and the virtual agent after the user or the virtual agent has finished speaking or responding, the presentation unit presents the embodiment agent in a manner that indicates it does not expect any utterances from the user. A conversation-providing system equipped with the following features.

2. The aforementioned estimation is made based on the content of the conversation between the user and the virtual agent. The conversation provision system according to claim 1.

3. The aforementioned estimation is based on the user's attitude. The conversation provision system according to claim 1.

4. The aforementioned estimation is made based on the user's external circumstances. The conversation provision system according to claim 1.

5. The aforementioned estimation is made based on the operating status of functions other than the conversation unit of the conversation provision system. The conversation provision system according to claim 1.

6. The presentation unit presents the materialized agent in a manner that does not face forward, as an indication that it does not expect a utterance from the user. The conversation provision system according to claim 1.

7. The presentation unit, in an action that indicates it does not expect a utterance from the user, presents the embodiment agent in a manner that is not conversing. The conversation provision system according to claim 1.

8. Even if the presentation unit estimates that the user intends to end the conversation with the virtual agent, if the holding condition is met, it will not present the manifestation agent in a manner that indicates it does not expect an utterance from the user. The conversation provision system according to claim 1.

9. A task processing device that processes tasks using artificial intelligence in response to task processing instructions, a conversation output unit that instructs the task processing device to generate conversation content for a virtual agent to the user, and acquires and outputs the conversation content generated by the task processing device, A presentation instruction unit that instructs the user to present an embodiment agent that embodies the virtual agent, and if it is presumed that the user intends to end the conversation consisting of utterances and responses between the user and the virtual agent after the user or the virtual agent has finished speaking or responding, the presentation instruction unit instructs the embodiment agent to present in a manner that indicates it does not expect any utterances from the user. A conversation-providing device equipped with the following features.