Conversation system, control program, and control method

The integration of a large-scale language model and speech synthesis in conversation systems allows conversational agents to express emotions naturally, improving the richness and naturalness of human-agent interactions.

JP2026082410APending Publication Date: 2026-05-19ATR ADVANCED TELECOMM RES INST INT
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ATR ADVANCED TELECOMM RES INST INT
Filing Date
2024-11-07
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional conversation systems fail to enable conversational agents to express emotions naturally, limiting the richness and naturalness of human-agent interactions.

Method used

A conversation system utilizing a large-scale language model to generate response sentences and emotional parameters, combined with speech synthesis to output synthesized speech, enabling conversational agents to express emotions naturally and richly.

Benefits of technology

Conversational agents can engage in more natural and emotionally rich conversations with humans, reducing response time delays and enhancing user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026082410000001_ABST
    Figure 2026082410000001_ABST
Patent Text Reader

Abstract

This invention provides a conversational system, control program, and control method that enable conversational agents to converse naturally and emotionally. [Solution] The conversation system includes a conversation agent capable of conversing with humans, such as a communication robot, and comprises a conversation text acquisition means S1 that acquires conversation text from humans, a response text generation means S3 that inputs the conversation text into a large-scale language model and generates a response text and emotion parameters corresponding to the conversation text, a voice generation means S5 that generates synthesized speech using the response text and emotion parameters, and a voice output means S7 that causes the conversation agent to output the synthesized speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a conversation system, a control program, and a control method, and more particularly, to a conversation system, a control program, and a control method including a conversation agent such as a communication robot and a virtual character that can communicate with humans.

Background Art

[0002] An example of a conventional conversation system of this type is disclosed in Patent Document 1. The technology of Patent Document 1 relates to a response method proposal system that proposes a response method for a chatbot communicatively connected to a user terminal. This response method proposal system includes a knowledge database that associates and records keywords, response candidates, frequencies, and likability, a dictionary that extracts keywords from text data obtained by textifying an inquiry from a user terminal, and a search unit that extracts response candidates corresponding to the keywords extracted by the dictionary and having both the frequency and the likability above a threshold value and transmits them to the chatbot. Further, this response method proposal system includes an emotion determination unit that calculates the likability of a user based on voice data obtained by recording a call between the user terminal and an operator terminal, a received power database that associates and records text data obtained by textifying the call with the likability, and a learning unit that learns the keywords and response candidates of the call based on the text data recorded in the received power database and associates and records the frequency of use of the response candidates and the likability calculated by the emotion determination unit with the learned keywords and response candidates. [[ID=I4]]

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Disclosure of the Invention

Problems to be Solved by the Invention

[0004] The technology described in Patent Document 1 estimates the user's (human's) level of liking and selects an answer accordingly, allowing for conversations that do not offend the user's feelings. However, in order to achieve more natural conversations between humans and conversational agents such as communication robots and virtual characters, it is desirable that the conversational agents can also express emotions, but this point is not considered in the technology described in Patent Document 1.

[0005] Therefore, the primary objective of this invention is to provide a novel conversation system, control program, and control method.

[0006] Another object of this invention is to provide a conversational system, a control program, and a control method that enable a conversational agent to converse with humans in a natural and emotionally rich manner. [Means for solving the problem]

[0007] The first invention is a conversation system comprising a conversation agent capable of conversing with a human, the system comprising: a conversation acquisition means for acquiring conversation text from a human; a response text generation means for inputting the conversation text acquired by the conversation acquisition means into a large-scale language model and generating a response text and emotion parameters corresponding to the conversation text; a speech generation means for generating synthesized speech using the response text and emotion parameters generated by the response text generation means; and a speech output means for outputting the synthesized speech generated by the speech generation means to the conversation agent.

[0008] According to the first invention, a large-scale language model is used to generate response sentences and emotional parameters for the response sentences in response to human conversation, and synthesized speech that takes the emotional parameters into account is generated and output to the conversation agent, so that the conversation agent can converse with humans in a natural and emotionally rich manner.

[0009] The second invention is subordinate to the first invention and uses a communication robot as a conversational agent.

[0010] According to the second invention, a communication robot can converse with humans in a natural and emotionally rich manner.

[0011] The third invention is dependent on the first or second invention, and the response sentence generation means generates an emotion parameter for each sentence included in the response sentence.

[0012] According to the third invention, conversational agents can converse with humans in a more natural and emotionally rich manner.

[0013] The fourth invention is subordinate to the third invention, wherein the voice generation means sequentially generates synthesized speech each time the response sentence generation means generates a sentence included in the response sentence, and the voice output means sequentially outputs the synthesized speech to the conversation agent each time the voice generation means generates synthesized speech for a sentence included in the response sentence.

[0014] According to the fourth invention, the time between when a user speaks to a conversational agent and when the conversational agent responds can be shortened, allowing the conversational agent to converse with the user more naturally.

[0015] The fifth invention is a control program executed in a control device for a conversation system equipped with a conversation agent capable of conversing with a human, wherein the control program causes the processor of the control device to function as a conversation text acquisition means for acquiring conversation text from a human, a response text generation means for inputting the conversation text acquired by the conversation text acquisition means into a large-scale language model and generating a response text and emotion parameters corresponding to the conversation text, a speech generation means for generating synthesized speech using the response text and emotion parameters generated by the response text generation means, and a speech output means for causing the conversation agent to output the synthesized speech generated by the speech generation means.

[0016] The sixth invention is a control method in a control device of a conversation system including a conversation agent capable of communicating with a human. The processor of the control device acquires a conversation sentence from the human, inputs the conversation sentence into a large language model, generates a response sentence corresponding to the conversation sentence and an emotion parameter for the response sentence, generates a synthesized voice using the response sentence and the emotion parameter, and outputs the synthesized voice to the conversation agent.

[0017] In the fifth and sixth inventions as well, the same operational effects as those of the first invention are achieved.

Effect of the Invention

[0018] According to this invention, conversation agents such as communication robots and virtual characters can converse with humans more naturally and emotionally richly.

[0019] The above object, other objects, features, and advantages of this invention will become more apparent from the following detailed description of the embodiments made with reference to the drawings.

Brief Description of the Drawings

[0020] [Figure 1] It is a block diagram showing a conversation system of an embodiment of this invention. [Figure 2] It is a block diagram showing an example of the electrical configuration of a control device included in the conversation system. [[ID=**28]] [Figure 3] It is a diagram showing the appearance of a communication robot included in the conversation system. [Figure 4] It is a block diagram showing an example of the electrical configuration of the communication robot. [Figure 5] It is an illustrative diagram showing an example of the memory map of the RAM of the control device. [Figure 6] It is a flowchart showing an example of the conversation control process executed by the CPU of the control device.

Best Mode for Carrying Out the Invention

[0021] Referring to FIG. 1, a conversation system 10 according to an embodiment of the present invention includes a communication robot 16 (hereinafter simply referred to as "robot 16") that can communicate with a user (human). As will be described in detail later, in the conversation system 10, when the user enjoys a conversation with the robot 16 like a conversation between friends or when the robot 16 acts as an operator to answer inquiries from the user, the robot 16 appropriately expresses emotions, enabling a natural conversation with the user.

[0022] As shown in FIG. 1, the conversation system 10 includes a control device 12 that comprehensively controls the entire conversation system 10. The control device 12 is a computer (information processing device) such as a general-purpose personal computer or a workstation. The robot 16 and the cloud server 18 etc. are connected to the control device 12 via a network 14 such as a LAN and the Internet. These may be connected by wire or wirelessly. However, the control device 12 may be a computer built into the robot 16.

[0023] FIG. 2 is a block diagram showing an example of the electrical configuration of the control device 12. As shown in FIG. 2, the control device 12 includes a CPU 20. The CPU 20 is connected to a memory 22, a communication I / F 24 (communication interface), and an input / output I / F 26 (input / output interface) via an internal bus.

[0024] The CPU 20 is a processor that controls the entire control device 12 and, by extension, the entire conversation system 10. The CPU 20 functions as a conversation sentence acquisition means, a response sentence generation means, a voice generation means, a voice output means, etc. according to the present invention.

[0025] The memory 22 includes a RAM, a ROM, an HDD, etc. The CPU 20 can control the robot 16 by executing various programs described later stored in the memory 22. Each program is stored in advance in the ROM or HDD and is deployed and executed in the RAM as needed.

[0026] The communication interface 24 is an interface for sending and receiving control signals and data between the robot 16 and external devices such as the cloud server 18 via the network 14, under the control of the CPU 20.

[0027] Input / output interface 26 is connected to input devices 28 such as a keyboard and computer mouse, and output devices 30 such as a monitor and speakers. The input / output interface 26 outputs operation data received from the input devices 28 to the CPU 20. It also outputs image data and sound data generated by the CPU 20 to the output devices 30, so that the screen corresponding to the image data is displayed on the monitor, and the sound corresponding to the sound data is output from the speakers. Note that the configuration of the control device 12 shown in Figure 2 is just an example and is not limited to this configuration.

[0028] Referring to Figure 3 along with Figure 1, robot 16 is an example of a conversational agent capable of conversing with a user, and can communicate with the user using voice and body movements. In this embodiment, RoboBee®, a humanoid robot manufactured and sold by Vstone Corporation, is used. However, any robot with a different appearance and structure can be used as robot 16, as long as it is capable of conversing with a user. For example, Sota® and its puppet version (puppet robot) manufactured by Vstone Corporation can be suitably used. In this embodiment, the communication behavior of robot 16 is determined by the control device 12, and voice data (robot voice data) and operation commands for robot 16 to communicate with the user are transmitted from the control device 12 to robot 16. The configuration of robot 16 will be briefly described below.

[0029] As shown in Figure 3, the robot 16 includes a trolley 50. Multiple wheels 52 are provided on the underside of the trolley 50. Each of the multiple wheels 52 is independently driven by a wheel motor 54 (see Figure 4), allowing the trolley 50 (and thus the robot 16) to move in any direction, forward, backward, left, or right. A body 56 is mounted upright on top of the trolley 50. A distance sensor 58, such as an infrared distance sensor or an ultrasonic distance sensor, is provided on the upper center of the front of the body 56. This distance sensor 58 measures the distance to an object in front of the robot 16.

[0030] Upper arms 62R and 62L are provided at the upper ends of both sides of the torso 56 via shoulder joints 60R and 60L. Each of the shoulder joints 60R and 60L has three orthogonal degrees of freedom, allowing the angles of the upper arms 62R and 62L to be controlled around each of these three orthogonal axes. Forearms 66R and 66L are provided at the ends of the upper arms 62R and 62L via elbow joints 64R and 64L. Each of the elbow joints 64R and 64L has one degree of freedom, allowing the angles of the forearms 66R and 66L to be controlled around this axis. Furthermore, right hands 68R and left hands 68L, which have the functions of human fingers and palms, are provided at the ends of the forearms 66R and 66L.

[0031] A head 72 is provided at the upper end of the torso 56 via a neck joint 70. The neck joint 70 has three degrees of freedom, and the angle of the head 72 can be controlled around each of these three axes. A speaker 74 is provided at the bottom of the head 72, and microphones 76R and 76L are provided on both sides of the head 72. In addition, eyeball sections 78R and 78L are provided on both sides of the front of the head 72, and eye cameras 80R and 80L are provided on these eyeball sections 78R and 78L, respectively.

[0032] Figure 4 is a block diagram showing the electrical configuration of robot 16. As shown in Figure 4, robot 16 includes a CPU 82. The CPU 82 is a processor that oversees the overall control of robot 16 and controls the robot's movements based on commands (operation instructions) from the CPU 20 of the control device 12. The memory 84, motor control board 86, sensor input / output board 88, and audio input / output board 90 are connected to this CPU 82 via a bus.

[0033] Memory 84 includes RAM, ROM, and HDD. RAM is used as work memory and buffer memory for CPU 82. Control programs for controlling the robot 16's movements (executing tasks) according to instructions from control device 12 are stored in ROM and HDD. These control programs include a voice output program for outputting synthesized speech from speaker 74 and an action program for controlling each motor (described later) to make the robot 16 perform body movements. The control programs also include a detection program for detecting the output of each sensor (sensor information) and a communication program for sending and receiving necessary data and commands to and from an external computer such as control device 12.

[0034] The motor control board 86, for example, is composed of a DSP and controls the drive of each axis motor, such as those for the shoulder joints 60R and 60L, elbow joints 64R and 64L, and neck joint 70. Specifically, the motor control board 86 receives control data from the CPU 82 and controls the rotation angles of a total of four motors (collectively referred to as "right arm motor 92" in Figure 4): three motors that control the angles of the three orthogonal axes of the shoulder joint 60R and one motor that controls the angle of the elbow joint 64R. Similarly, the motor control board 86 receives control data from the CPU 82 and controls the rotation angles of a total of four motors (collectively referred to as "left arm motor 94" in Figure 4): three motors that control the angles of the three orthogonal axes of the shoulder joint 60L and one motor that controls the angle of the elbow joint 64L. In addition, the motor control board 86 receives control data from the CPU 82 and controls the rotation angles of three motors (collectively referred to as "head motor 96" in Figure 4) that control the angles of the three orthogonal axes of the neck joint 70. Furthermore, the motor control board 86 receives control data from the CPU 82 and controls the rotation angle of the two motors that drive the wheels 52 (collectively referred to as "wheel motors 54" in Figure 4).

[0035] The sensor input / output board 88, like the motor control board 86, is composed of a DSP and receives signals from each sensor and provides them to the CPU 82. Specifically, distance data from the distance sensor 58 (for example, reflection time data) is input to the CPU 82 through this sensor input / output board 88. Similarly, video signals from the eye cameras 80R and 80L are also input to the CPU 82. The audio input / output board 90 is also composed of a DSP and outputs voice or speech from the speaker 74 according to the speech synthesis data provided by the CPU 82. Audio input from microphones 76R and 76L is also provided to the CPU 82 via the audio input / output board 90.

[0036] Furthermore, a communication unit 100 is connected to the CPU 82 via a bus. The communication unit 100 transmits data received from the CPU 82 to an external computer such as the control unit 12, and also provides data and commands received from the external computer to the CPU 82. Note that the configuration of the robot 16 shown in Figures 3 and 4 is just an example and is not limited to this configuration.

[0037] Returning to Figure 1, the cloud server 18 includes a first cloud server that provides a speech recognition service and a second cloud server equipped with a Large Language Model (LLM). In this embodiment, the "Speech-to-Text" API service provided by Google is used as the speech recognition service that converts speech data into text data. The control device 12 can obtain text data (user conversation data) by inputting the speech data of the conversation spoken by the user (user voice data) acquired by the microphones 76R and 76L of the robot 16 into the speech recognition service. In this embodiment, the "Chat GPT (specifically GPT-4 Turbo)" provided by OpenAI is used as the large language model. The control device 12 can obtain text data of the robot 16's response sentences according to the content of the user's conversation by inputting the speech-recognized user conversation data into the large language model. However, the types of speech recognition service and large language model described above are just examples and are not limited to these.

[0038] In a conversation system 10 configured in this way, the robot 16 enjoys conversing with the user as if they were friends and responds to inquiries from the user. In this process, the robot 16 expresses emotions by changing its voice tone according to the situation (conversation content), enabling natural and emotionally rich conversations with the user. In this embodiment, "VOICEPEAK," provided by AHS, is used as the speech synthesis software to change the voice tone of the robot 16. "VOICEPEAK" is text-to-speech software that can synthesize speech by inputting preferred sentences as text, and can adjust the voice tone according to emotion parameters. However, this type of speech synthesis software is just an example and is not limited to it. In this embodiment, the emotion parameters of the robot 16's conversation sentences (response sentences) are generated using a large-scale language model. The control device 12 can obtain synthesized speech data (robot voice data) to be spoken by the robot 16 by inputting data (robot response sentence data) including the text data of the robot 16's response sentences generated by the large-scale language model and the emotion parameters to be attached to these response sentences into the speech synthesis software. The following provides a detailed explanation.

[0039] In this conversation system 10, when a user and a robot 16 converse, the conversation (utterances) made by the user are acquired as voice data from the microphones 76R and 76L of the robot 16. This user voice data is input from the control device 12 to the voice recognition service of the cloud server 18, converted into text data, and then output back to the control device 12. In other words, the control device 12 acquires the user's conversation as text data.

[0040] When the control device 12 acquires the user's conversation text, it is input from the control device 12 to the large-scale language model on the cloud server 18. The large-scale language model estimates the emotion of the input user's conversation text, generates text data of a response appropriate to the content and emotion of the conversation text, and generates emotion parameters to be attached to the response text. The generated response text and emotion parameter data are output to the control device 12. In other words, the control device 12 uses the large-scale language model to generate the robot 16's response text and emotion parameters. When generating emotion parameters using the large-scale language model, it is preferable to generate emotion parameters for each sentence in the response text (for example, each sentence separated by punctuation marks including periods, commas, periods, question marks, and exclamation marks). This enables more natural and emotionally rich conversation. In addition, the output emotion parameters should correspond to the specifications of the speech synthesis software. For example, numerical values ​​from 0 to 100 are output for five items: happiness, anger, sadness, warmth, and distant gaze.

[0041] The following is an example of prompts to cause a large-scale language model to output the response sentences and sentiment parameters of robot 16. "We will roleplay as a chatbot with simulated emotions, following the conditions below." In the following conversation, you will behave as if you possessed the following five emotional parameters. Each emotional parameter will change sentence by sentence throughout the conversation. Your response tone and phrasing will change to reflect the current emotional parameter values ​​with each sentence. In subsequent conversations, please first output the current emotion parameters, and then output the conversation. Also, please ensure that the value of each emotion parameter is a multiple of 10, and output one sentence at a time, separated by punctuation marks. The standard emotion value is 70 or higher for enjoyment. Please be sure to include a colon (:) at the end of each sentence. Examples of emotion parameter values ​​are as follows: / / / Example of emotional parameter values / / / ·standard Enjoyment: 70, Anger: 0, Sadness: 0, Warmth: 50, Distant gaze: 0 ·joy Enjoyment: 90, Anger: 0, Sadness: 0, Warmth: 50, Distant gaze: 0 ·sorrow Enjoyment: 10, Anger: 0, Sadness: 90, Warmth: 50, Distant gaze: 10 ·anger Enjoyment: 10, Anger: 90, Sadness: 10, Warmth: 50, Distant gaze: 10 / / / The output format will be as follows: / / / / / / Answer text / / / Understood! Is there anything you'd like me to do? Please feel free to ask me anything. / / / Output statement / / / Enjoyment: 0-100, Anger: 0-100, Sadness: 0-100, Gentle: 0-100, Distant gaze: 0-100 Understood! Enjoyment: 0-100, Anger: 0-100, Sadness: 0-100, Gentle: 0-100, Distant gaze: 0-100 Is there anything you would like me to do? Enjoyment: 0-100, Anger: 0-100, Sadness: 0-100, Gentle: 0-100, Distant gaze: 0-100. Please ask me anything.

[0042] When the control device 12 obtains the response sentence and emotion parameters from the robot 16, these are input into the speech synthesis software, and synthesized speech is generated from the input data. In other words, the control device 12 generates synthesized speech using the response sentence and emotion parameters. The generated synthesized speech data is transmitted from the control device 12 to the robot 16, and the synthesized speech is played back by the robot 16's speaker 74. In other words, the control device 12 causes the robot 16 to output synthesized speech.

[0043] Furthermore, when having the robot 16 output synthesized speech, it is preferable to use a so-called streaming method in which the large-scale language model sequentially generates synthesized speech each time it generates a sentence included in the response, and the robot 16 sequentially outputs the synthesized speech each time it is generated. This shortens the time from when the user speaks to the robot 16 until the robot 16 responds, enabling more natural conversation with the user. It is also preferable to have the robot 16 utter fillers such as "um," "uh," and "well" while it is synthesizing speech (i.e., between uttering one sentence and uttering the next). This makes it less likely for the user to perceive a delay in the robot 16's response, enabling more natural conversation with the user.

[0044] In this way, by linking a large-scale language model with speech synthesis software that can change the tone of voice, the robot 16 can express emotions according to the content of the conversation, enabling natural and emotionally rich conversations between the robot 16 and the user.

[0045] Figure 5 shows an example of the memory map 200 of the RAM built into the control device 12. As shown in Figure 5, the RAM of the control device 12 includes a program storage area 202 and a data storage area 204. The program storage area 202 stores the conversation control program executed by the control device 12. The conversation control program includes a main processing program 202a, a communication program 202b, a speech recognition program 202c, a response sentence generation program 202d, a speech synthesis program 202e, and a robot control program 202f, etc.

[0046] The main processing program 202a is a program for executing the main routine of the conversation control processing of the control device 12 in this embodiment. The communication program 202b is a program for communicating (sending and receiving data, etc.) with external devices such as the robot 16 and the cloud server 18.

[0047] The speech recognition program 202c is a program for converting user voice data acquired by the microphones 76R and 76L of the robot 16 into text data (i.e., obtaining the user's conversation). In this embodiment, the program uses Google's speech recognition service (Speech-to-Text) to convert the user's voice data into text data.

[0048] The response sentence generation program 202d is a program for generating response sentences and emotion parameters for robot 16 in response to user conversation sentences. In this embodiment, the program uses a large-scale language model (GPT-4 Turbo) provided by OpenAI to generate response sentences and emotion parameters for robot 16.

[0049] The speech synthesis program 202e is a program for generating synthesized speech to be produced by the robot 16 using response sentences and emotion parameters generated using a large-scale language model. In this embodiment, the speech synthesis software (VOICEPEAK) provided by AHS Corporation is used as the speech synthesis program 202e.

[0050] The robot control program 202f is a program for controlling the communication behavior of the robot 16, and includes synthesized speech data generated using the speech synthesis program 202e, and a program for selecting and sending action commands to cause the robot 16 to perform physical movements.

[0051] Although not shown in the diagram, the program storage area 202 also appropriately stores other application programs, etc., in addition to the conversation control program of this invention.

[0052] Meanwhile, the data storage area 204 stores user voice data 204a, user conversation data 204b, robot response data 204c, robot voice data 204d, and operation command data 204e, among others.

[0053] User voice data 204a is data of the user's voice detected by microphones 76R and 76L installed on the robot 16 and transmitted from the robot 16. User conversation data 204b is data obtained by converting user voice data 204a into text data by the speech recognition program 202c, that is, text data of the conversation uttered by the user. Robot response data 204c is text data of the robot 16's response to the user's conversation, and emotion parameter data, generated by the response generation program 202d (large-scale language model). Robot voice data 204d is synthesized voice data of the response to be spoken by the robot 16, generated by the speech synthesis program 202e using the response and emotion parameters. Action command data 204e is data of action commands that indicate patterns of physical movements (gestures) included in the robot 16's communication behavior.

[0054] Although not shown in the diagram, the data storage area 204 stores other data necessary for the control device 12 to perform conversation control processing, and also contains timers (counters) and flags necessary for performing conversation control processing.

[0055] Figure 6 is a flowchart showing an example of conversation control processing performed by the CPU 20 of the control device 12. As shown in Figure 6, when the conversation control processing starts, the CPU 20 first obtains the conversation text from the user in step S1. That is, it obtains the user's voice data detected by the microphones 76R and 76L of the robot 16 from the robot 16, and converts this user's voice data into text data using the speech recognition service of the cloud server 18.

[0056] In the next step, S3, a large-scale language model is used to generate response sentences and emotion parameters for the robot 16. Specifically, the user's conversation text obtained in step S1 is sent to the large-scale language model on the cloud server 18, which outputs response sentences and emotion parameters appropriate to the content and emotion of that conversation text.

[0057] In the following step S5, synthesized speech is generated using the response sentence and emotion parameters. That is, the conversation sentence and emotion parameters generated in step S3 are input into speech synthesis software to generate synthesized speech to be spoken by the robot 16.

[0058] In the following step S7, the synthesized voice is output from the speaker 74 of the robot 16. That is, the synthesized voice data generated in step S5 is sent to the robot 16, and the synthesized voice is played back by the speaker 74. At this time, if the robot 16's response (communication behavior) involves physical movement, an action command indicating the pattern of physical movement is also sent to the robot 16 at the same time as the synthesized voice data.

[0059] Then, in step S9, it is determined whether or not there is a termination instruction. For example, if the termination button is operated on the control device 12 or if the series of conversations ends, it is determined that there is a termination instruction. If the answer in step S9 is "NO", that is, if there is no termination instruction, the process returns to step S1. On the other hand, if the answer in step S9 is "YES", that is, if there is a termination instruction, this conversation control process is terminated.

[0060] Note that the processing steps in the flowchart shown in Figure 6 are merely examples, and the order of processing steps can be changed if similar results can be obtained.

[0061] As described above, according to this embodiment, a large-scale language model is used to generate response sentences and emotion parameters corresponding to human conversation sentences, and synthesized speech that takes the emotion parameters into account is generated and output to the robot 16, so that the robot 16 can converse with humans in a natural and emotionally rich manner.

[0062] In the above-described embodiment, a robot was used as the conversation agent. However, the conversation agent in the conversation system according to this invention may be a virtual character (also called a CG agent) that is rendered using computer graphics and displayed on a screen. Furthermore, the conversation agent according to this invention refers to a proxy entity that can converse with a user on behalf of another human being, and includes robots existing in real space and characters existing in virtual space.

[0063] Furthermore, in the above-described embodiment, the user's speech was obtained by converting the user's voice acquired by the microphone into text data (i.e., by speech recognition), but it is also possible to obtain the user's speech from data entered using a keyboard.

[0064] Furthermore, in the above-described embodiment, emotions were expressed by changing the tone of voice of the conversational agent, such as a robot. However, in addition to this speech-based emotional expression, emotions may also be expressed by changing the conversational agent's physical movements (for example, facial expressions or gestures) according to emotional parameters generated by a large-scale language model. By adding emotional expression through physical movements to speech-based emotional expression, the conversational agent can converse with humans more naturally and emotionally. [Explanation of symbols]

[0065] 10...Conversation System 12 ... control device 16. Communication robot (conversational agent) 18. Cloud servers (large-scale language models, speech recognition services) 20 ...CPU of the control unit

Claims

1. A conversational system equipped with a conversational agent capable of conversing with humans, Conversation text acquisition means for acquiring conversation text from the aforementioned human, A response text generation means inputs the conversation text acquired by the conversation text acquisition means into a large-scale language model and generates a response text corresponding to the conversation text and emotion parameters for the response text. A voice generation means that generates synthesized speech using the response sentence generated by the response sentence generation means and the emotion parameters, and A conversation system comprising a voice output means for causing the synthesized voice generated by the voice generation means to output to the conversation agent.

2. The conversation system according to claim 1, wherein a communication robot is used as the conversation agent.

3. The conversation system according to claim 1 or 2, wherein the response sentence generation means generates the emotion parameter for each sentence included in the response sentence.

4. The voice generation means sequentially generates synthesized speech each time the response sentence generation means generates a sentence included in the response sentence, The conversation system according to claim 3, wherein the voice output means causes the voice generation means to sequentially output the synthesized voice to the conversation agent each time it generates synthesized voice for a sentence included in the response sentence.

5. A control program executed in the control device of a conversational system equipped with a conversational agent capable of conversing with humans, The processor of the control device is Conversation text acquisition means for acquiring conversation text from the aforementioned human, A response text generation means inputs the conversation text acquired by the conversation text acquisition means into a large-scale language model and generates a response text corresponding to the conversation text and emotion parameters for the response text. A voice generation means that generates synthesized speech using the response sentence generated by the response sentence generation means and the emotion parameters, and A control program that functions as a voice output means for causing the synthesized voice generated by the voice generation means to output to the conversation agent.

6. A control method for a control device of a conversation system equipped with a conversational agent capable of conversing with humans, The processor of the control device is Obtain the conversation text from the aforementioned human, The aforementioned conversation is input into a large-scale language model to generate a response sentence corresponding to the conversation sentence and emotion parameters for the response sentence. A synthesized voice is generated using the aforementioned response and emotion parameters, and A control method for causing the conversation agent to output the synthesized voice.