Interaction device, interaction system, interaction method and program
The interaction device addresses smooth dialogue by using turn-holding and turn-yielding utterances to manage dialogue flow, improving user interaction and preventing dialogue breakdowns.
Patent Information
- Application Number
- JP2024057018
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-10
AI Technical Summary
Conventional dialogue systems fail to facilitate smooth interaction with users by not considering appropriate timing for outputting responses and backchannels, leading to hindered dialogue.
An interaction device equipped with an utterance detection unit and an utterance output unit that controls a dialogue agent to output turn-holding and turn-yielding utterances before generating responses, ensuring smooth dialogue flow.
Facilitates conversational engagement by promoting turn-taking and reducing dialogue breakdowns, enhancing realism and user interaction in applications like nursing care and business negotiations.
Smart Images

Figure 2025154161000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a dialogue device, a dialogue system, a dialogue method, and a program. [Background technology]
[0002] Dialogue systems in which a dialogue agent automatically responds to messages from a user are known. For example, Patent Literature 1 discloses a voice dialogue method that includes the steps of inputting a user utterance, extracting prosodic features of the input user utterance, and generating a backchannel response to the user utterance based on the extracted prosodic features, and that adjusts the prosody of the backchannel when generating the backchannel so that the prosodic features match the prosodic features of the user utterance. Summary of the Invention [Problem to be solved by the invention]
[0003] However, conventional technologies can sometimes hinder smooth dialogue with users. For example, conventional technologies do not take into consideration how to output responses and backchannels that facilitate dialogue, which can prevent the dialogue agent from engaging in dialogue at the appropriate time.
[0004] In view of the above technical problems, an embodiment of the present invention aims to smoothly interact with a user using a dialogue agent. [Means for solving the problem]
[0005] One embodiment of the present invention is an interaction device that responds to a user's utterances using a dialogue agent, and is equipped with an utterance detection unit that detects a first utterance by the user, and an utterance output unit that controls the dialogue agent to output a response to the content of the first utterance detected by the utterance detection unit, and is characterized in that, when predetermined conditions are met, the utterance output unit controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response. [Effects of the Invention]
[0006] According to one embodiment of the present invention, a conversational agent can be used to facilitate a conversation with a user. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a diagram illustrating an example of a system configuration of a dialogue system according to an embodiment. [Figure 2] FIG. 2 illustrates an example of a dialogue agent according to an embodiment. [Figure 3] FIG. 2 illustrates another example of a dialogue agent according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a computer according to an embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of a hardware configuration of a terminal device according to an embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of a functional configuration of a server device according to the first embodiment. [Figure 7] FIG. 1 is a sequence diagram showing an example of a dialogue flow according to the prior art. [Figure 8] FIG. 3 is a sequence diagram showing an example of a dialogue flow according to the first embodiment. [Figure 9] 4 is a flowchart showing an example of an interaction method according to the first embodiment. [Figure 10] 10 is a flowchart showing an example of an utterance detection process according to the second embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a functional configuration of a server device according to a third embodiment. [Figure 12] FIG. 11 is a diagram showing an example of emotion determination rules according to the third embodiment. [Figure 13] FIG. 11 is a diagram illustrating an example of behavior determination rules according to the third embodiment. [Figure 14] 11 is a flowchart showing an example of emotion recognition processing according to the third embodiment. [Figure 15]FIG. 10 is a diagram illustrating an example of a functional configuration of a server device according to a fourth embodiment. [Figure 16] FIG. 13 is a sequence diagram showing an example of a dialogue flow according to the fourth embodiment. [Figure 17] 13 is a flowchart showing an example of a motion detection process according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, components having the same functions are designated by the same reference numerals, and redundant description will be omitted.
[0009] [First embodiment] A first embodiment of the present invention is an information processing system that provides a dialogue service. Hereinafter, the information processing system according to this embodiment will be referred to as a "dialogue system." The dialogue service is an information communication service that allows a dialogue with a dialogue agent. In the dialogue service, the dialogue agent automatically responds to messages from the user, thereby progressing the dialogue between the user and the dialogue agent.
[0010] <System configuration> Fig. 1 is a diagram illustrating an example of a system configuration of an interactive system according to an embodiment. In the example of Fig. 1, the interactive system 1 includes a server device 100 and a terminal device 10 connected to a communication network N such as the Internet or a LAN (Local Area Network).
[0011] The server device 100 is, for example, an information processing device having a computer configuration, or a system configured by multiple computers. The server device 100 provides a dialogue service in which a dialogue agent automatically responds to messages from a user 11 using a terminal device 10, by causing the computer included in the server device 100 to execute a predetermined program. In other words, the server device 100 controls a dialogue with the user 11 using the dialogue agent. The server device 100 is an example of a dialogue device.
[0012] The terminal device 10 is an information terminal used by a user 11, such as a PC (Personal Computer), a tablet terminal, or a smartphone. The terminal device 10 is capable of communicating with the server device 100 via a communication network N. The user 11 can use the terminal device 10 to utilize an interactive service provided by the server device 100. In other words, the user 11 interacts with an interactive agent through the interactive service.
[0013] Preferably, the dialogue system 1 supports the performance of a predetermined task, such as business negotiations or caregiving, through a dialogue in which a dialogue agent automatically responds to a message from a user.
[0014] The system configuration of the dialogue system 1 shown in Fig. 1 is an example. The terminal device 10 is not limited to a general-purpose information terminal, but may be, for example, a dedicated terminal device or various electronic devices. The dialogue system 1 may be realized by, for example, a single information processing device having the configuration of a computer. Here, the following description will be given assuming that the dialogue system 1 has the system configuration shown in Fig. 1.
[0015] (Image of a conversational agent) A dialogue agent is a system that automatically responds to questions from users or customers using knowledge including registered information and knowledge, or AI (Artificial Intelligence).
[0016] Examples of use cases for the dialogue agent include web conferences, websites, smartphone apps, or unmanned AI avatars in the metaverse.
[0017] FIG. 2 shows an example of an image of a dialogue agent according to an embodiment. This figure shows an example of a dialogue screen 200 for business negotiations that the server device 100 displays on the terminal device 10. In the example of FIG. 2, a virtual human 201 generated by 3D (three-dimensional) modeling is displayed on the dialogue screen 200. The virtual human 201 is an example of a dialogue agent. For example, the server device 100 controls the virtual human 201 on the dialogue screen 200 to advance business negotiations while having a dialogue with the user 11.
[0018] As a suitable example, the business negotiation dialogue screen 200 displays a large display 202. The server device 100 can display, for example, a product proposed by a user on this display 202 and can also control the virtual human 201 to explain the product.
[0019] FIG. 3 shows another example of an image of a dialogue agent according to an embodiment. This figure shows an example of an interactive screen 300 for nursing care purposes that the server device 100 displays on the terminal device 10. In the example of FIG. 3, the interactive screen 300 displays another virtual human 301 generated by 3D modeling, similar to FIG. 2. The virtual human 301 is another example of a dialogue agent. The server device 100 controls the virtual human 301 on this interactive screen 300 so that the virtual human 301 communicates with, for example, elderly people living alone to prevent dementia.
[0020] As a preferred example, the interaction between the user 11 and the virtual human 301 can be a text-based interaction 302 in addition to (or instead of) voice, as shown in FIG.
[0021] In this way, the dialogue system 1 can change the dialogue scenario to change the dialogue content to suit various uses, such as business negotiations, caregiving, classes, or counseling.
[0022] <Hardware configuration> (Computer hardware configuration) The server device 100 has, for example, the hardware configuration of a computer 500 as shown in Fig. 4. Alternatively, the server device 100 is configured by a plurality of computers 500. Furthermore, the terminal device 10 may have, for example, the hardware configuration of a computer 500 as shown in Fig. 4.
[0023] Fig. 4 is a diagram showing an example of the hardware configuration of a computer according to an embodiment. As shown in Fig. 4, the computer 500 includes, for example, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, a RAM (Random Access Memory) 503, a HD (Hard Disk) 504, an HDD (Hard Disk Drive) controller 505, a display 506, an external device connection I / F (Interface) 507, a network I / F 508, a keyboard 509, a pointing device 510, a DVD-RW (Digital Versatile Disk Rewritable) drive 512, a media I / F 514, and a bus line 515.
[0024] Furthermore, when the computer 500 is the terminal device 10, the computer 500 further includes a microphone 521, a speaker 522, an audio input / output I / F 523, a CMOS (Complementary Metal Oxide Semiconductor) sensor 524, an image sensor I / F 525, and the like.
[0025] Of these, the CPU 501 controls the overall operation of the computer 500. The ROM 502 stores programs used to start up the computer 500, such as an IPL (Initial Program Loader). The RAM 503 is used, for example, as a work area for the CPU 501. The HD 504 stores programs such as an OS (Operating System), applications, and device drivers, as well as various data. The HDD controller 505 controls the reading and writing of various data from and to the HD 504, for example, under the control of the CPU 501. The HD 504 and the HDD controller 505 are examples of storage devices.
[0026] The display 506 displays various types of information such as a cursor, a menu, a window, characters, or an image. The display 506 may be provided outside the computer 500. The external device connection I / F 507 is an interface for connecting various external devices to the computer 500. The network I / F 508 is an interface for connecting the computer 500 to a communication network N and communicating with other devices.
[0027] The keyboard 509 is a type of input means having a plurality of keys for inputting characters, numbers, various instructions, etc. The pointing device 510 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, etc. The keyboard 509 and the pointing device 510 may be provided outside the computer 500.
[0028] The DVD-RW drive 512 controls reading and writing of various data from and to a DVD-RW 511, which is an example of a removable recording medium. The DVD-RW 511 is not limited to a DVD-RW, and may be another removable recording medium. The media I / F 514 controls reading and writing (storing) of data from and to a medium 513 such as a flash memory. The bus line 515 includes an address bus, a data bus, various control signals, and the like for electrically connecting the above components.
[0029] The microphone 521 is a built-in circuit that converts sound into an electrical signal. The speaker 522 is a built-in circuit that converts the electrical signal into physical vibrations to produce sound such as music or voice. The sound input / output I / F 523 is a circuit that processes input and output of sound signals between the microphone 521 and the speaker 522 under the control of the CPU 501.
[0030] The CMOS sensor 524 is a type of built-in imaging means that captures an image of a subject (e.g., a self-portrait) and obtains image data under the control of the CPU 501. Note that the computer 500 may have an imaging means such as a CCD (Charge Coupled Device) sensor instead of the CMOS sensor 524. The imaging element I / F 525 is a circuit that controls the driving of the CMOS sensor 524.
[0031] (Example of hardware configuration of terminal device) 5 is a diagram illustrating an example of the hardware configuration of a terminal device according to an embodiment. Here, an example of the hardware configuration of the terminal device 10 will be described in the case where the terminal device 10 is an information terminal such as a smartphone or a tablet terminal.
[0032] In the example of Figure 5, the terminal device 10 includes a CPU 601, a ROM 602, a RAM 603, a storage device 604, a CMOS sensor 605, an image sensor I / F 606, an acceleration / direction sensor 607, a media I / F 609, and a GPS (Global Positioning System) receiving unit 610.
[0033] Of these, the CPU 601 controls the overall operation of the terminal device 10 by executing a predetermined program. The ROM 602 stores a program used to start up the CPU 601, such as an IPL. The RAM 603 is used as a work area for the CPU 601. The storage device 604 is a large-capacity storage device that stores programs such as an OS and applications, various types of data, and the like, and is realized by, for example, an SSD (Solid State Drive), a flash ROM, or the like.
[0034] The CMOS sensor 605 is a type of built-in imaging means that captures an image of a subject (mainly a self-portrait) under the control of the CPU 601 to obtain image data. Note that the terminal device 10 may have an imaging means such as a CCD sensor instead of the CMOS sensor 605. The imaging element I / F 606 is a circuit that controls the driving of the CMOS sensor 605. The acceleration / azimuth sensor 607 is one of various sensors, such as an electronic magnetic compass or gyrocompass that detects geomagnetism, and an acceleration sensor. The media I / F 609 controls the reading or writing (storage) of data from a medium (storage medium) 608, such as a flash memory. The GPS receiver 610 receives GPS signals (positioning signals) from GPS satellites.
[0035] The terminal device 10 also includes a long-distance communication circuit 611, an antenna 611a of the long-distance communication circuit 611, a CMOS sensor 612, an image sensor I / F 613, a microphone 614, a speaker 615, an audio input / output I / F 616, a display 617, an external device connection I / F 618, a short-distance communication circuit 619, an antenna 619a of the short-distance communication circuit 619, and a touch panel 620.
[0036] Of these, the long-distance communication circuit 611 is a circuit that communicates with other devices via, for example, the communication network N. The CMOS sensor 612 is a type of built-in imaging means that captures an image of a subject and obtains image data under the control of the CPU 601. The image sensor I / F 613 is a circuit that controls the driving of the CMOS sensor 612. The microphone 614 is a built-in circuit that converts sound into an electrical signal. The speaker 615 is a built-in circuit that converts an electrical signal into physical vibrations to generate sounds such as music and voice. The sound input / output I / F 616 is a circuit that processes the input and output of sound wave signals between the microphone 614 and the speaker 615 under the control of the CPU 601.
[0037] The display 617 is a type of display means such as a liquid crystal display or organic electroluminescence (EL) display that displays an image of a subject, various icons, etc. The external device connection I / F 618 is an interface for connecting various external devices. The short-range communication circuit 619 includes a circuit for performing short-range wireless communication. The touch panel 620 is a type of input means that allows a user to operate the terminal device 10 by pressing the display 617.
[0038] The terminal device 10 also includes a bus line 621. The bus line 621 includes an address bus, a data bus, and the like for electrically connecting the components such as the CPU 601 shown in FIG.
[0039] 5 is an example of the hardware configuration of the terminal device 10. The terminal device 10 may have other hardware configurations as long as it has a computer configuration, a communication circuit, a display, a microphone, a speaker, and the like.
[0040] <Functional configuration> Fig. 6 is a diagram showing an example of the functional configuration of the server device according to the first embodiment. As shown in Fig. 6, the server device 100 includes an utterance detection unit 110, a feature extraction unit 120, a voice recognition unit 130, a response generation unit 140, a response storage unit 145, a state management unit 150, and an utterance output unit 160.
[0041] The speech detection unit 110, feature extraction unit 120, voice recognition unit 130, response generation unit 140, state management unit 150, and speech output unit 160 are realized, for example, by processing executed by the CPU 501 and network I / F 508 of a program expanded from ROM 502 to RAM 503 as shown in FIG.
[0042] The response storage unit 145 is realized using, for example, the HD 504 shown in Fig. 4. Reading or writing of data stored in the HD 504 is performed via, for example, the HDD controller 505.
[0043] The speech detection unit 110 detects speech by the user 11 who uses the terminal device 10. For example, the speech detection unit 110 detects a speech interval from a video (video image and audio) of the user 11 received from the terminal device 10, and acquires an acoustic signal indicating the speech uttered by the user 11. The speech interval may be detected using a technique such as VAD (Voice Activity Detection). Therefore, the speech detection unit 110 detects the start and end of an utterance by the user 11. Hereinafter, an utterance by the user 11 will be referred to as a "user utterance." The user utterance is an example of a first utterance.
[0044] The feature extraction unit 120 extracts acoustic features from the user utterance detected by the speech detection unit 110. Any acoustic features may be used as long as they are features that can be used for speech recognition.
[0045] The speech recognition unit 130 recognizes the user's speech based on the acoustic features extracted by the feature extraction unit 120. The speech recognition unit 130 outputs text data indicating the speech recognition result. The speech recognition unit 130 may use any speech recognition technology as long as it can generate text data based on the acoustic features.
[0046] The response generation unit 140 generates an utterance from the dialogue agent to respond to the user utterance based on the speech recognition result output by the speech recognition unit 130. The utterance from the dialogue agent may be voice only, or may be video including voice. The video may include non-verbal information of the dialogue agent. For example, the non-verbal information may include facial expressions, gestures, hand movements, and other physical actions. Hereinafter, an utterance to respond to a user utterance will be referred to as a "response utterance."
[0047] The response generation unit 140 may generate a response utterance in cooperation with an external device or system. Examples of the external device or system include various search engines, large language models (LLMs), image generation models, and text-to-speech (TTS) systems. Note that "external" means that the device or system is not included in the dialogue system 1. The server device 100 may communicate with the external device or system via a communication network N.
[0048] The response storage unit 145 stores utterances from the dialogue agent. The response storage unit 145 stores a plurality of pre-generated turn-holding utterances, a plurality of pre-generated turn-giving utterances, and response utterances generated by the response generation unit 140. The response storage unit 145 may store information indicating behavior when outputting a turn-holding utterance, a turn-giving utterance, and a response utterance from the dialogue agent. The turn-holding utterance and the turn-giving utterance are examples of utterances for facilitating a dialogue with a user.
[0049] A turn-maintaining utterance is an utterance made by a dialogue agent to maintain a turn. In other words, a turn-maintaining utterance is an utterance to indicate to the user 11 that the current turn state is the system's turn. A turn-maintaining utterance may include a first turn-maintaining utterance, a second turn-maintaining utterance, and a third turn-maintaining utterance, each of which has different output conditions. A turn-maintaining utterance is an example of a second utterance.
[0050] The first turn-maintaining utterance is a turn-maintaining utterance that is output before a response utterance to a user utterance and when the end of the user utterance is detected. The first turn-maintaining utterance may include a backchannel. The backchannel is a short utterance such as a backchannel or a filler. For example, the backchannel may be a short utterance such as "um," "ah," or "I see." If the dialogue agent is an entity capable of visual expression, such as a virtual human, the backchannel may include actions such as a nod, a gesture, or a facial expression.
[0051] The second turn-maintaining utterance is a turn-maintaining utterance that is output before a response utterance to the user utterance and after the first turn-maintaining utterance. The second turn-maintaining utterance may be generated by the response generation unit 140. The second turn-maintaining utterance may include an utterance showing empathy for the user 11 or mirroring. Mirroring is an utterance that repeats the content of the user utterance. Mirroring may be an utterance that summarizes the user utterance. Mirroring may be an utterance that asks the user utterance to repeat the content of the user utterance. Mirroring may be generated, for example, by summarizing the speech recognition results of the user utterance using a large-scale language model. If the dialogue agent is an entity that can be visually represented, such as a virtual human, mirroring may include, for example, an action that mimics the user's action.
[0052] The third turn-maintaining utterance is a turn-maintaining utterance that is output before a response utterance to a user utterance and when a certain processing time is required for the response generation unit 140 to generate an utterance. The third turn-maintaining utterance is an utterance for filling in gaps in a dialogue. The third turn-maintaining utterance may include a backchannel.
[0053] A turn-transferring utterance is an utterance made by the dialogue agent to transfer the turn to the user 11. In other words, a turn-transferring utterance is an utterance to indicate to the user 11 that the current turn state is the user's turn. A turn-transferring utterance is an utterance that is output before a response utterance to a user utterance and when the start of a user utterance is detected. As an example, a turn-transferring utterance may include an utterance that encourages the user 11 to speak. As an example, an utterance that encourages speaking may be an utterance that agrees with the user 11. As an example, an utterance that agrees with the user 11 may be an affirmative utterance such as "yes, yes." As an example, a turn-transferring utterance may be a question to the user 11. As an example, a question to the user 11 may be a short question such as "What's wrong?" The turn-transferring utterance is another example of a second utterance.
[0054] The state management unit 150 manages the turn state of the dialogue between the user 11 and the dialogue agent. The turn state may be information indicating whether the dialogue turn is the user 11's turn (hereinafter referred to as the "user's turn") or the dialogue agent's turn (hereinafter referred to as the "system's turn"). A turn means the right to speak. The speaker who has a turn takes turns as the dialogue progresses. Each speaker can speak regardless of whether or not it is their turn, but if each speaker speaks when they have a turn, the dialogue can proceed smoothly.
[0055] The state management unit 150 may transition the dialogue turn state based on the detection state of a user utterance. The detection state of a user utterance may be a state indicating whether or not a user utterance is being detected. The detection state of a user utterance may be a state in which a user utterance is being detected from the time the utterance detection unit 110 detects the start of a user utterance until the end of that user utterance. Furthermore, the detection state of a user utterance may be a state in which a user utterance is not being detected from the time the utterance detection unit 110 detects the end of a user utterance until the start of a new user utterance.
[0056] For example, when the state management unit 150 detects the start of a user utterance when the turn state is the system's turn, the state management unit 150 transitions the turn state to the user's turn. For example, when the state management unit 150 detects the end of a user utterance when the turn state is the user's turn, the state management unit 150 transitions the turn state to the system's turn.
[0057] The state management unit 150 may track the turn state of the dialogue based on the response state to the user utterance. The response state to the user utterance may be a state indicating whether or not a response to the user utterance has been made. For example, when the state management unit 150 finishes outputting a response utterance while the turn state is the system's turn, the state management unit 150 transitions the turn state to the user's turn.
[0058] The utterance output unit 160 outputs a response from the dialogue agent and an utterance that facilitates dialogue with the user. In this embodiment, the utterances from the dialogue agent may include a response utterance, a turn-holding utterance, and a turn-transfer utterance. The utterance output unit 160 may output an utterance stored in the response storage unit 145.
[0059] The utterance output unit 160 outputs a turn-maintaining utterance or a turn-transferring utterance when a predetermined condition is met. The predetermined condition may be a condition related to the turn state managed by the state management unit 150. For example, the utterance output unit 160 may output a first turn-maintaining utterance when it detects the end of a user utterance (i.e., when the turn state transitions to the system's turn). By having the dialogue agent output the first turn-maintaining utterance when the user finishes speaking, the user 11 can be made aware that the dialogue turn has been handed over to the dialogue agent.
[0060] For example, the utterance output unit 160 may output a second turn-maintaining utterance when the response generation unit 140 is generating a response utterance (i.e., when the turn state is the system's turn). By the dialogue agent outputting a second turn-maintaining utterance, the user 11 can recognize that the dialogue agent understands the user utterance, improving the sense of realism of the dialogue. Furthermore, when a response utterance is generated based on an external device or system, the response waiting time of the user 11 due to processing delays can be reduced. For example, when synthesizing the voice of a response utterance using an external speech synthesis system, processing may take a long time if the number of characters in the text to be synthesized is large.
[0061] For example, when the utterance output unit 160 detects the start of a user utterance when the turn state is the system's turn (i.e., when the turn state transitions to the user's turn), it may output a question to the user 11 as a turn-yielding utterance. If the user 11 interrupts and speaks when the dialogue agent has a turn and is speaking or is about to speak, it is considered that the user 11 is about to speak something important. In this case, if the turn is handed over to the user 11 and the user 11 is encouraged to speak, the dialogue can proceed smoothly.
[0062] <Dialogue flow> The dialogue flows of the prior art and this embodiment will be compared with each other with reference to FIGS.
[0063] FIG. 7 is a sequence diagram showing an example of a dialogue flow according to the prior art. First, assume that a user 11 makes a user utterance u1 such as "I'm from Takarazuka, Hyogo Prefecture." Acoustic signals representing the user utterance u1 are input to the server device 100 in the order of "I'm (u1-1)," "I'm from Hyogo Prefecture (u1-2)," and "I'm from Takarazuka (u1-3)." The server device 100 sequentially recognizes the input user utterances u1-1, u1-2, and u1-3, thereby sequentially generating speech recognition progresses t1 and t2 such as "I'm" and "I'm from Hyogo Prefecture," and finally obtaining a speech recognition result t3 such as "I'm from Takarazuka, Hyogo Prefecture."
[0064] When the end of the user utterance u1 is detected (p1), the server device 100 generates a response utterance to the user utterance u1 based on the speech recognition result t3 of the entire user utterance u1 (p2). In addition, the server device 100 may acquire external information along with generating the response utterance (p3).
[0065] However, generating a response utterance (p2) and acquiring external information (p3) require a certain amount of processing time. Therefore, in a conventional dialogue system, even if the end of the user utterance u1 is detected, a response cannot be made until the generation of the response utterance is completed. In this case, the user 11 cannot understand the state of the dialogue system and ends up making a user utterance u2 such as "Um?"
[0066] Once the generation of the response utterance (p2) is complete, the dialogue agent can output response utterances r1 and r2 such as "Is that Takarazuka in Hyogo?" and "Takarazuka is a wonderful place, isn't it?". The user 11 can attempt to continue the dialogue by responding to response utterances r1 and r2 with a user utterance u3 such as "Yes, xxx is famous...". However, once the acquisition of external information (p3) is complete, the dialogue agent outputs a response utterance r3 based on the external information such as "Takarazuka is famous for xxx, isn't it?". In this case, the dialogue between the user 11 and the dialogue agent will not mesh, and the dialogue will break down.
[0067] To address the problem shown in FIG. 7, it is conceivable to output a sound (e.g., a beep) indicating that the user's utterance has been accepted when the end of the user's utterance u1 is detected. However, outputting a beep or the like during a dialogue is undesirable because it diminishes the sense of realism of the dialogue. Furthermore, there are dialogue systems that output responses to the user's utterance u1, such as laughter or nodding, but these responses are intended to show empathy for the user 11 and do not have the function of making the user 11 aware of the turn.
[0068] 8 is a sequence diagram showing an example of the flow of a dialogue according to the first embodiment. In this embodiment, when the dialogue agent detects the start of a user utterance u1, it outputs a turn-giving utterance b1 such as "Yes, yes." Furthermore, when the dialogue agent detects the end of the user utterance u1 (p1), it outputs a first turn-holding utterance b2 such as "Ah."
[0069] When the user 11 makes a user utterance u1 and is responded to with a turn-giving utterance b1, the user 11 recognizes that the dialogue agent is waiting for the user utterance u1. In other words, the user 11 recognizes that the user 11 has a turn based on the first turn-giving utterance b1. Since it is the user's turn, the user 11 can continue speaking without worry.
[0070] When the user 11 makes a user utterance u1 and is responded to with the first turn-preserving utterance b2, the user 11 recognizes that the dialogue agent has accepted the user utterance u1 and is preparing a response. In other words, the user 11 recognizes from the first turn-preserving utterance b2 that the dialogue agent has a turn. The user 11 refrains from speaking until the dialogue agent responds to the user utterance u1, thereby avoiding a breakdown in the dialogue.
[0071] If the generation of the response utterance (p2) takes a certain amount of processing time, the dialogue agent may output a second turn-maintaining utterance b3 including a backchannel such as "I see" or a mirroring such as "Is that Takarazuka from Hyogo?". Also, if the acquisition of external information (p3) also takes a certain amount of processing time, the dialogue agent may output a third turn-maintaining utterance b4 such as "Umm."
[0072] The user 11 recognizes from the first turn-maintaining utterance b2 that the dialogue agent has a turn, and therefore recognizes from the second turn-maintaining utterance b3 and the third turn-maintaining utterance b4 that the dialogue agent's turn is continuing. Since the user 11 refrains from speaking until the dialogue agent outputs a subsequent response, even if it takes time to generate a response utterance (p2) or to acquire external information (p3), a breakdown in the dialogue can be avoided.
[0073] Although Figure 8 shows an example in which the dialogue agent outputs short backchannels b1, b2, b3, and b4, the dialogue agent may also output responses that include gestures, such as looking away from the user, moving a hand, or smiling.
[0074] As shown in Fig. 8, by returning a turn-holding utterance or a turn-yielding utterance to the user 11 while the dialogue system 1 processes the user utterance u1, turn-taking between the user 11 and the dialogue agent becomes smooth. This prevents the user 11 from interrupting the dialogue during processing by the dialogue system 1. In this case, an appropriate backchannel that does not annoy the user 11 may be returned by analyzing acoustic features in parallel depending on the speech recognition process or the length of the user utterance.
[0075] In this way, according to this embodiment, turn-taking between the user and the dialogue agent can be smoothly promoted. For example, in applications such as nursing care or counseling, the effect of attentive listening can be promoted. Also, in applications such as business negotiations, the customer can be encouraged to speak up.
[0076] <Processing Procedure> 9 is a flowchart showing an example of the dialogue method according to the first embodiment. In this embodiment, the dialogue method is a processing procedure from when a user speaks to when a dialogue agent responds. Therefore, the dialogue method is repeatedly executed while the dialogue between the user and the dialogue agent continues.
[0077] In step S1, the utterance detection unit 110 of the server device 100 detects the start of a user utterance. When the utterance detection unit 110 detects a user utterance, the state management unit 150 transitions the turn state to the user's turn.
[0078] In step S2, the utterance output unit 160 of the server device 100 reads out a turn-giving utterance from the response storage unit 145. The utterance output unit 160 may randomly select one turn-giving utterance from the turn-giving utterances stored in the response storage unit 145. The utterance output unit 160 may read out, together with the turn-giving utterance, information indicating a behavior when outputting the turn-giving utterance.
[0079] The utterance output unit 160 controls the output of the read turn-giving utterance from the dialogue agent. The utterance output unit 160 may output the turn-giving utterance only when a predetermined condition is satisfied. For example, the utterance output unit 160 may control the output of the turn-giving utterance only once every time a predetermined number of user utterances are detected.
[0080] In step S3, the utterance detection unit 110 of the server device 100 acquires an acoustic signal indicating a user utterance. The utterance detection unit 110 sends the acoustic signal indicating the user utterance to the feature extraction unit 120.
[0081] In step S4, the feature extraction unit 120 of the server device 100 receives the acoustic signal from the speech detection unit 110. The feature extraction unit 120 extracts acoustic features from the received acoustic signal. The feature extraction unit 120 sends the extracted acoustic features to the speech recognition unit 130.
[0082] In step S5, the speech recognition unit 130 of the server device 100 receives the acoustic features from the feature extraction unit 120. The speech recognition unit 130 performs speech recognition on the user utterance based on the received acoustic features. The speech recognition unit 130 sends the speech recognition result to the response generation unit 140.
[0083] In step S6, the utterance detection unit 110 of the server device 100 determines whether or not the end of the user's utterance has been detected. If the end of the user's utterance has been detected (YES), the utterance detection unit 110 proceeds to step S7. At this time, the utterance detection unit 110 notifies the state management unit 150 and the utterance output unit 160 of the end of the user's utterance.
[0084] On the other hand, if the end of the user's utterance has not been detected (NO), the utterance detection unit 110 returns the process to step S2. After returning to step S2, the utterance detection unit 110 acquires the next acoustic signal and sends the acquired acoustic signal to the feature extraction unit 120. The feature extraction unit 120 receives the next acoustic signal from the utterance detection unit 110 and extracts acoustic features from the received acoustic signal. In this way, the server device 100 repeatedly executes steps S2 to S6 until the end of the user's utterance is detected in step S6.
[0085] In step S7, the state management unit 150 of the server device 100 transitions the turn state to the system's turn in response to the notification from the utterance detection unit 110. In response to the transition of the turn state to the system's turn, the utterance output unit 160 reads out a first turn holding utterance from the response storage unit 145. The utterance output unit 160 may randomly select one first turn holding utterance from the first turn holding utterances stored in the response storage unit 145. The utterance output unit 160 may read out, together with the first turn holding utterance, information indicating behavior when outputting the first turn holding utterance. The utterance output unit 160 controls the output of the read out first turn holding utterance from the dialogue agent.
[0086] In step S8, the response generation unit 140 of the server device 100 receives the speech recognition result from the speech recognition unit 130. The response generation unit 140 starts generating a second turn-holding utterance and a response utterance based on the received speech recognition result.
[0087] In step S9, the response generation unit 140 of the server device 100 determines whether or not the generation of the response utterance has been completed. If the generation of the response utterance has been completed (YES), the response generation unit 140 stores the response utterance in the response storage unit 145 and proceeds to step S14. On the other hand, if the generation of the response utterance has not been completed (NO), the response generation unit 140 proceeds to step S10.
[0088] In step S10, the response generation unit 140 of the server device 100 determines whether the generation of the second turn-holding utterance has been completed. If the generation of the second turn-holding utterance has been completed (YES), the response generation unit 140 stores the second turn-holding utterance in the response storage unit 145, and proceeds to step S13. On the other hand, if the generation of the second turn-holding utterance has not been completed (NO), the response generation unit 140 sends the generation time of the response utterance to the utterance output unit 160, and proceeds to step S11. The generation time of the response utterance is the elapsed time since the generation of the response utterance started in step S8.
[0089] In step S11, the utterance output unit 160 of the server device 100 receives the generation time of the response utterance from the response generation unit 140. The utterance output unit 160 determines whether the generation time of the response utterance is equal to or greater than a predetermined threshold. If the generation time of the response utterance is equal to or greater than the threshold (YES), the utterance output unit 160 proceeds to step S12. On the other hand, if the generation time of the response utterance is less than the threshold (NO), the utterance output unit 160 skips step S12 and returns the process to step S9.
[0090] In step S12, the utterance output unit 160 of the server device 100 reads out a third turn holding utterance from the response storage unit 145. The utterance output unit 160 may randomly select one third turn holding utterance from the third turn holding utterances stored in the response storage unit 145. The utterance output unit 160 may read out, together with the third turn holding utterance, information indicating the behavior when outputting the third turn holding utterance. The utterance output unit 160 controls the output of the read third turn holding utterance from the dialogue agent. After outputting the third turn holding utterance, the utterance output unit 160 returns the process to step S9.
[0091] In step S13, the utterance output unit 160 of the server device 100 receives the second turn-holding utterance from the response generation unit 140. The utterance output unit 160 controls the dialogue agent to output the received second turn-holding utterance. After outputting the second turn-holding utterance, the utterance output unit 160 returns the process to step S9.
[0092] In step S14, the utterance output unit 160 of the server device 100 reads out the response utterance from the response storage unit 145. The utterance output unit 160 controls the output of the read response utterance from the dialogue agent. When the utterance output unit 160 finishes outputting the response utterance, it notifies the state management unit 150 of this fact. In response to the notification from the utterance output unit 160, the state management unit 150 transitions the turn state to the user's turn.
[0093] <Effects of the first embodiment> The server device 100 according to the first embodiment outputs a second utterance to facilitate a dialogue with the user before outputting a response utterance from the dialogue agent in response to the content of the first utterance by the user. The second utterance may include an utterance to hold the turn or an utterance to hand over the turn.
[0094] The second utterance for facilitating the dialogue is a non-verbal or verbal reaction or signal that the listener makes to the speaker in dialogue communication, for example, to show interest or understanding, or to encourage further conversation. By outputting the second utterance for facilitating the dialogue from the dialogue agent to the user, the dialogue between the dialogue agent and the user can proceed smoothly, and effective communication between them can be promoted. Therefore, in one aspect, according to this embodiment, a dialogue with the user can be smoothly conducted using the dialogue agent.
[0095] When the server device 100 detects the end of the user's utterance, it may transition the turn state to the system's turn, and output a second utterance when the turn state transitions to the system's turn. The user recognizes that it is the system's turn when the dialogue agent outputs the second utterance when the user finishes speaking. According to this embodiment, it is possible to clearly indicate to the user that it is the system's turn.
[0096] When the server device 100 detects a user utterance when the turn state is the system's turn, it may transition the turn state to the user's turn and output a turn-yielding utterance when the turn state transitions to the user's turn. The user recognizes that it is the user's turn when the turn-yielding utterance is output from the dialogue agent when the user starts speaking. According to this embodiment, it is possible to clearly indicate to the user that it is the user's turn.
[0097] The server device 100 may generate a response utterance to the user utterance based on the speech recognition result of recognizing the user utterance, and while generating the response utterance, may output a second turn-keeping utterance that summarizes the speech recognition result. The user recognizes that the dialogue agent understands the user utterance when the dialogue agent utters the summary of the user utterance. According to this embodiment, the sense of realism of the dialogue can be improved and the user's waiting time for a response can be reduced.
[0098] The server device 100 may output a third turn-preserving utterance when the generation time of the response utterance exceeds a threshold. When the third turn-preserving utterance is output from the dialogue agent when the response waiting time is long, the user recognizes that the dialogue agent is about to speak. According to this embodiment, the response waiting time of the user can be reduced.
[0099] [Second embodiment] In the first embodiment, a configuration has been described in which the turn state transitions to the user's turn when the utterance detection unit 110 detects a user utterance. For example, while the server device 100 is generating a response utterance, the user may forcibly interrupt the dialogue to correct the utterance, for example. This user utterance carries a high level of message and is important. Therefore, it is preferable for the server device 100 to transition to the user's turn even while generating a response utterance. However, if the server device 100 simply transitions to the user's turn upon detection of a user utterance, sounds that should be ignored, such as the user's throat clearing or noise, may be recognized as important messages. This may cause the dialogue to break down.
[0100] In the second embodiment, when the utterance detection unit 110 detects a user utterance, it determines whether to allow an interruption. Specifically, the utterance detection unit 110 determines whether to proceed to the user's turn based on the voice recognition result of the user utterance and the duration of the user utterance. For example, the utterance detection unit 110 may not proceed to the user's turn if the voice recognition result indicates an utterance that is not intended to change the speaker, such as a throat clearing or a filler. Furthermore, the utterance detection unit 110 may not proceed to the user's turn if the duration of the user utterance is equal to or less than a predetermined threshold.
[0101] On the other hand, for example, if the speech recognition result is not a throat clearing or a filler, and the duration of the user's utterance exceeds a predetermined threshold, the utterance detection unit 110 may transition to the user's turn. In this case, the utterance output unit 160 may output a turn-transfer utterance when the user's turn is transitioned. The user recognizes that the turn of the dialogue has returned to them, and can continue speaking with peace of mind.
[0102] <Utterance detection processing> 10 is a flowchart showing an example of the speech detection process according to the second embodiment. The speech detection process corresponds to step S1 in FIG.
[0103] In step S1-1, the speech detection unit 110 detects sound from the video of the user 11 received from the terminal device 10. The speech detection unit 110 may detect the sound, for example, based on the waveform of an acoustic signal included in the video of the user 11. The speech detection unit 110 acquires an acoustic signal including the detected sound.
[0104] In step S1-2, the speech detection unit 110 determines whether the audio signal acquired in step S1-1 is a voiced section. For example, by using a voice detection technique such as VAD, it is possible to determine whether the audio signal contains speech.
[0105] If the acoustic signal is in a voiced section (YES), the utterance detection unit 110 sends the acoustic signal in the voiced section to the feature extraction unit 120 and proceeds to step S1-3. On the other hand, if the acoustic signal is not in a voiced section (NO), the utterance detection unit 110 proceeds to step S1-8.
[0106] In step S1-3, the feature extraction unit 120 receives the acoustic signal for the voiced section from the speech detection unit 110. The feature extraction unit 120 extracts acoustic features from the acoustic signal for the voiced section. The feature extraction unit 120 sends the extracted acoustic features to the speech recognition unit .
[0107] In step S1-4, the speech recognition unit 130 receives acoustic features from the feature extraction unit 120. The speech recognition unit 130 performs speech recognition based on the received acoustic features. The speech recognition unit 130 sends the speech recognition result to the speech detection unit 110.
[0108] In step S1-5, the utterance detection unit 110 receives the speech recognition result from the speech recognition unit 130. The utterance detection unit 110 determines whether the received speech recognition result is a filler.
[0109] If it is a filler (YES), utterance detection unit 110 proceeds to step S1-8. On the other hand, if it is not a filler (NO), utterance detection unit 110 proceeds to step S1-6.
[0110] In step S1-6, the speech detection unit 110 determines whether the speech length of the voiced section is equal to or greater than a threshold. The speech length may be, for example, the number of characters in the speech recognition result or the speech time in the acoustic signal.
[0111] If the utterance length is equal to or greater than the threshold (YES), utterance detection unit 110 proceeds to step S1-7. On the other hand, if the utterance length is less than the threshold (NO), utterance detection unit 110 proceeds to step S1-8.
[0112] In step S1-7, the utterance detection unit 110 permits an interrupt. Specifically, the utterance detection unit 110 notifies the state management unit 150 that a user utterance has been detected. The state management unit 150 transitions the turn state to the user's turn.
[0113] In step S1-8, the utterance detection unit 110 does not permit interruption, in which case the utterance detection unit 110 discards the user utterance and ends the process.
[0114] <Effects of the second embodiment> The server device 100 according to the second embodiment determines whether a user utterance has been detected based on the speech recognition result of the user utterance and the duration of the user utterance. According to one aspect, even when a response utterance is being generated, the server device 100 can transition to the user's turn if the user makes an important utterance.
[0115] For example, in applications such as nursing care, business negotiations, and counseling, if it becomes possible to interrupt a dialogue, the user will no longer have to wait for the dialogue agent to speak, and will be able to smoothly take control of the dialogue. On the other hand, the turn will not be transferred to the user when the user clears their throat or makes a filler that should be ignored, and this will prevent the dialogue from breaking down.
[0116] [Third embodiment] In the first embodiment, a configuration in which the speech output unit 160 outputs a predetermined or random back channel has been described. In the third embodiment, a configuration in which a back channel is output in accordance with the user's emotion will be described. By outputting a back channel based on the user's emotion, a highly empathetic dialogue that is in tune with the user's emotion can be realized.
[0117] <Functional configuration> Fig. 11 is a diagram showing an example of the functional configuration of a server device according to the third embodiment. As shown in Fig. 11, the server device 100 according to this embodiment includes an utterance detection unit 110, a feature extraction unit 120, a voice recognition unit 130, a response generation unit 140, a response storage unit 145, a state management unit 150, an utterance output unit 160, and an emotion recognition unit 170. The server device 100 according to the third embodiment differs from the server device 100 according to the first embodiment in that it further includes the emotion recognition unit 170.
[0118] The emotion recognition unit 170 recognizes the user's emotion based on the user's utterance. The emotion recognition unit 170 may recognize the user's emotion based on a speech recognition result that recognizes the user's utterance. The emotion recognition unit 170 may recognize the user's emotion based on acoustic features extracted from the user's utterance. The emotion recognition unit 170 may recognize the user's emotion based on both the speech recognition result and the acoustic features. In this embodiment, an example will be described in which the user's emotion is recognized based on both the speech recognition result and the acoustic features.
[0119] For example, the emotion recognition unit 170 may classify acoustic features into one of a plurality of emotions based on a trained classification model. Alternatively, the emotion recognition unit 170 may classify speech recognition results into one of a plurality of emotions based on a large-scale language model or the like. As an example, emotions may be classified into three categories: positive, negative, and neutral. Positive refers to any positive emotion, such as "joy." Negative refers to any negative emotion, such as "sadness." Neutral refers to any emotion other than positive or negative.
[0120] The emotion recognition unit 170 may determine a final emotion recognition result based on a combination of the emotion recognition result based on the acoustic feature and the emotion recognition result based on the speech recognition result. The emotion recognition unit 170 may determine a final emotion recognition result based on predetermined emotion determination rules.
[0121] Fig. 12 is a diagram showing an example of emotion determination rules according to the third embodiment. As shown in Fig. 12, the emotion determination rules are rules that determine a final emotion recognition result for a combination of an emotion recognition result based on acoustic features and an emotion recognition result based on a speech recognition result.
[0122] For example, if both emotion recognition results are positive or negative, the final emotion recognition result will also be positive or negative. For example, if one emotion recognition result is positive or negative and the other emotion recognition result is neutral, the final emotion recognition result will be weak positive or weak negative. For example, if the emotion recognition results are a combination of positive and negative or both emotion recognition results are neutral, the final emotion recognition result will be "other."
[0123] In this embodiment, the utterance output unit 160 outputs the first turn-maintaining utterance based on the emotion recognition result obtained by the emotion recognition unit 170 recognizing the user's emotion. As an example, the utterance output unit 160 may output the first turn-maintaining utterance with a behavior according to the emotion recognition result. The utterance output unit 160 may determine the behavior when outputting the first turn-maintaining utterance based on predetermined behavior determination rules.
[0124] FIG. 13 is a diagram illustrating an example of a behavior determination rule according to the third embodiment. As illustrated in FIG. 13, the behavior determination rule is a rule that determines the behavior of a dialogue agent in response to an emotion recognition result. The behavior of the dialogue agent may include various modalities. The modalities may include, for example, facial expressions, gestures, nodding, voice pitch, speaking speed, etc.
[0125] For example, if the emotion recognition result is positive, weakly positive, or other, the facial expression of the dialogue agent may be a smiling expression. A smiling expression may be, for example, an expression in which the corners of the mouth are turned up and the corners of the eyes are turned down. On the other hand, if the emotion recognition result is negative or weakly negative, the facial expression of the dialogue agent may be a sympathetic expression. An sympathetic expression may be, for example, an expression of listening seriously. Note that, although a negative emotion is represented by "sadness," it does not have to be an expression that indicates sadness.
[0126] <Processing Procedure> 14 is a flowchart showing an example of emotion recognition processing according to the third embodiment. The emotion recognition processing may be executed before outputting the first turn holding utterance, and may be executed between step S6 and step S7, for example.
[0127] In step S11-1, the emotion recognition unit 170 recognizes the user's emotion based on the acoustic features extracted by the feature extraction unit 120. Specifically, the emotion recognition unit 170 classifies the speech recognition result as positive, negative, or neutral based on a trained classification model.
[0128] In step S11-2, the emotion recognition unit 170 recognizes the user's emotion based on the speech recognition result recognized by the speech recognition unit 130. Specifically, the emotion recognition unit 170 classifies the speech recognition result as positive, negative, or neutral based on a large-scale language model.
[0129] In step S11-3, the emotion recognition unit 170 determines a final emotion recognition result based on a combination of the emotion recognition result recognized in step S11-1 and the emotion recognition result recognized in step S11-2. Specifically, the emotion recognition unit 170 determines a final emotion recognition result from the combination of the emotion recognition results in accordance with emotion determination rules.
[0130] In step S11-4, the emotion recognition unit 170 sends the emotion recognition result determined in step S11-3 to the utterance output unit 160. The utterance output unit 160 determines the behavior of the dialogue agent based on the emotion recognition result received from the emotion recognition unit 170. Specifically, the emotion recognition unit 170 determines the behavior of the dialogue agent from the emotion recognition result in accordance with behavior determination rules.
[0131] <Effects of the third embodiment> The server device 100 according to the third embodiment recognizes the user's emotion based on the user's utterance, and outputs a first turn-maintaining utterance according to the emotion recognition result. According to one aspect, the present embodiment can output an utterance that is in tune with the user's emotion.
[0132] The server device 100 may recognize the user's emotion based on the acoustic features extracted from the user's utterance and the speech recognition result of recognizing the user's utterance. According to this embodiment, the emotion is recognized based on both the acoustic features and the content of the utterance, thereby enabling the user's emotion to be recognized in detail and accurately.
[0133] For example, in applications such as nursing care or counseling, if a conversation agent expresses emotions similar to those of the user, it can promote the effect of attentive listening. Also, in applications such as business negotiations, it can promote customer understanding.
[0134] [Fourth embodiment] In the first embodiment, a configuration has been described in which the utterance output unit 160 outputs a predetermined or random back channel. In the fourth embodiment, a configuration will be described in which, instead of a back channel, a dialogue agent operates in response to the user's actions.
[0135] In human-to-human conversations, for example, when a speaker laughs, the other person may laugh back even if the speaker is not amused. This is because laughing during a conversation has social implications, such as inducing laughter or livening up the conversation. The behavior of laughing back when the other person laughs is an important form of communication. Therefore, it is desirable to implement a similar behavior in a dialogue system. However, even if a dialogue system detects a user's laughter, it is not easy to achieve a dialogue system that can laugh back accordingly.
[0136] In this embodiment, speech recognition of user utterances is performed sequentially, and laughter is detected from the speech recognition results. When laughter is detected, the dialogue system outputs, for example, a short laugh, a smile, or a gesture to share the laughter. This creates a sense of realism as if the dialogue system is listening attentively to the user's utterance, and helps build a good relationship between the user and the dialogue system.
[0137] Although the explanation here has focused on laughter, this embodiment may be applied to any action by the user. Humans tend to feel a sense of affinity with others who behave in the same way as they do, so if the dialogue agent behaves in a way that is similar to the user's, it will help build a good relationship.
[0138] <Functional configuration> Fig. 15 is a diagram showing an example of the functional configuration of a server device according to the fourth embodiment. As shown in Fig. 15, the server device 100 according to this embodiment includes an utterance detection unit 110, a feature extraction unit 120, a voice recognition unit 130, a response generation unit 140, a response storage unit 145, a state management unit 150, an utterance output unit 160, and a motion detection unit 180. The server device 100 according to the fourth embodiment differs from the server device 100 according to the first embodiment in that it further includes a motion detection unit 180.
[0139] The movement detection unit 180 detects a movement of the user. The movement detection unit 180 may detect a movement of the user based on a speech recognition result of a user utterance. The movement detection unit 180 may detect a movement of the user by extracting a label indicating the movement of the user included in the speech recognition result.
[0140] The movement detection unit 180 may determine the degree of the detected movement. For example, the degree of laughter may be various types such as smiling, laughing out loud, laughing long with body shaking, etc. The movement detection unit 180 may determine the degree of movement by classifying the degree of movement based on a trained classification model, for example. The classification model may be trained to input acoustic features and output a label indicating the degree of movement. For example, the classification model may be a support vector machine (SVM) or the like.
[0141] <Dialogue flow> 16 is a sequence diagram showing an example of the flow of a dialogue according to the fourth embodiment. In this embodiment, when the user 11 makes a user utterance u1 such as "I'm from Takarazuka, Hyogo Prefecture," the user 11 laughs briefly between "I'm" and "Hyogo Prefecture." The server device 100 recognizes that the user 11 has laughed, and generates a speech recognition progress t2 such as "I'm laughing."
[0142] The server device 100 detects "laugh" included in the speech recognition progress t2 and outputs laughter b11. The user 11 thinks that the dialogue agent is laughing back at the user's laughter, and feels a sense of affinity with the dialogue agent. This builds a good relationship between the user 11 and the dialogue agent, making it easier for the user to talk to the dialogue agent and to express emotions more easily in the dialogue, thereby allowing the subsequent dialogue to proceed smoothly.
[0143] <Processing Procedure> 17 is a flowchart showing an example of the motion detection process according to the fourth embodiment. The motion detection process may be executed before detecting the end of the user's utterance (i.e., when the turn state is the user's turn), and may be executed between step S5 and step S6, for example.
[0144] In step S12-1, the action detection unit 180 detects a user's action. Specifically, the action detection unit 180 extracts a label indicating the action from the speech recognition result recognized by the speech recognition unit .
[0145] In step S12-2, the movement detection unit 180 determines whether or not a predetermined movement has been detected. The predetermined movement may be, for example, laughter. If the predetermined movement has been detected (YES), the movement detection unit 180 proceeds to step S12-3. On the other hand, if the predetermined movement has not been detected (NO), the movement detection unit 180 ends the movement detection process.
[0146] In step S12-3, the movement detection unit 180 determines the degree of the detected movement. Specifically, the acoustic feature of the speech section in which the predetermined movement is detected is input to a trained classification model to obtain a label indicating the degree of the movement.
[0147] In step S12-4, the action detection unit 180 notifies the action determined in step S12-3 to the utterance output unit 160. In response to the notification from the action detection unit 180, the utterance output unit 160 determines the action of the dialogue agent.
[0148] <Effects of the Fourth Embodiment> The server device 100 according to the fourth embodiment detects a user's action, and when the turn state is the user's turn, outputs an action from the dialogue agent according to the detection result of the user's action. In one aspect, according to this embodiment, a good relationship can be built between the user and the dialogue agent, and the dialogue can proceed smoothly.
[0149] For example, in applications such as nursing care or counseling, the sense of distance between the conversation agent and the user can be reduced.
[0150] [supplement] Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and conventional circuit modules designed to perform each of the above-described functions.
[0151] The devices described in the example are merely one of several computing environments for implementing the embodiments disclosed herein. In one embodiment, server apparatus 100 includes multiple computing devices, such as a server cluster, configured to communicate with each other via any type of communications link, including a network, shared memory, etc., and to perform the processes disclosed herein.
[0152] The functional components of the server device 100 may be integrated into one server device or may be divided among multiple devices. Furthermore, at least some of the functional components of the server device 100 may be included in the terminal device 10.
[0153] For example, aspects of the present invention are as follows.
[0154] (Appendix 1) A dialogue device that responds to a user's utterance using a dialogue agent, an utterance detection unit that detects a first utterance by the user; an utterance output unit that controls the dialogue agent to output a response to the content of the first utterance detected by the utterance detection unit; Equipped with and when a predetermined condition is met, the utterance output unit controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response. 1. An interactive device comprising: (Appendix 2) The second utterance includes a turn-holding utterance or a turn-yielding utterance. 2. The interactive device of claim 1. (Appendix 3) a state management unit that manages a turn state of the dialogue based on a detection state of the first utterance, the utterance output unit outputs the second utterance based on the condition related to the turn state. 3. The interactive device of claim 2. (Appendix 4) when the state management unit detects the end of the first utterance, it transitions the turn state to a turn of the dialogue agent; the utterance output unit outputs the second utterance that maintains the turn when the turn state transitions to the turn of the dialogue agent. 4. The interactive device of claim 3. (Appendix 5) the state management unit, when detecting the first utterance when the turn state is the dialogue agent's turn, transitions the turn state to the user's turn; the utterance output unit outputs the second utterance to hand over the turn when the turn state transitions to the user's turn. 4. The interactive device of claim 3. (Appendix 6) the utterance detection unit determines whether the first utterance has been detected based on a speech recognition result of recognizing the first utterance and a duration of the first utterance. 6. An interactive device according to any one of claims 1 to 5. (Appendix 7) further comprising an emotion recognition unit that recognizes an emotion of the user based on the first utterance; the utterance output unit outputs the second utterance according to the emotion recognition result. 7. An interactive device according to any one of claims 1 to 6. (Appendix 8) the emotion recognition unit recognizes the emotion based on acoustic features extracted from the first utterance and a speech recognition result of recognizing the first utterance. 8. The interactive device of claim 7. (Appendix 9) a response generation unit that generates a response utterance to the first utterance based on a speech recognition result that recognizes the first utterance, the utterance output unit, when generating the response utterance, outputs the second utterance summarizing the speech recognition result. 9. An interactive device according to any one of claims 1 to 8. (Appendix 10) the utterance output unit outputs the second utterance that holds the turn when a generation time of the response utterance exceeds a threshold. 10. The interactive device of claim 9. (Appendix 11) further comprising a motion detection unit that detects a first motion by the user; the utterance output unit outputs a second action from the dialogue agent in accordance with a result of detecting the first action; 11. An interactive device according to any one of claims 1 to 10. (Appendix 12) A dialogue system in which a terminal device operated by a user and a dialogue device that responds to utterances of the user using a dialogue agent can communicate with each other via a network, The dialogue device an utterance detection unit that detects a first utterance by the user; an utterance output unit that controls the dialogue agent to output a response to the content of the first utterance detected by the utterance detection unit; Equipped with and when a predetermined condition is met, the utterance output unit controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response. A dialogue system characterized by: (Appendix 13) A computer that responds to a user's utterances using a dialogue agent, detecting a first utterance by the user; a step of controlling the dialogue agent to output a response to the content of the first utterance detected in the step of detecting; Run the controlling step controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response when a predetermined condition is met; How to interact. (Appendix 14) A computer that responds to user utterances using a dialogue agent. detecting a first utterance by the user; a step of controlling the dialogue agent to output a response to the content of the first utterance detected in the step of detecting; Execute the controlling step controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response when a predetermined condition is met; program.
[0155] Although the embodiments of the present invention have been described above, the present invention is not limited to such specific embodiments, and various modifications and applications are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]
[0156] 1. Dialogue System 10 Terminal Equipment 100 Server device 110 Speech detection unit 120 Feature Extraction Unit 130 Voice Recognition Unit 140 Response Generation Unit 150 Status Management Unit 160 Speech output unit 170 Emotion Recognition Department 180 Motion detection unit 200, 300 dialogue screen 201, 301 Virtual Human (Dialogue Agent) 500 computers [Prior art documents] [Patent documents]
[0157] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-38501
Claims
1. A dialogue device that responds to a user's utterance using a dialogue agent, an utterance detection unit that detects a first utterance by the user; an utterance output unit that controls the dialogue agent to output a response to the content of the first utterance detected by the utterance detection unit; Equipped with and when a predetermined condition is met, the utterance output unit controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response.
1. An interactive device comprising:
2. The second utterance includes a turn-holding utterance or a turn-yielding utterance. The interactive device according to claim 1 .
3. a state management unit that manages a turn state of the dialogue based on a detection state of the first utterance, the utterance output unit outputs the second utterance based on the condition related to the turn state.
3. The interactive device according to claim 2.
4. the state management unit, when detecting the end of the first utterance, transitions the turn state to a turn of the dialogue agent; the utterance output unit outputs the second utterance that maintains the turn when the turn state transitions to the turn of the dialogue agent.
4. The interactive device according to claim 3.
5. the state management unit, when detecting the first utterance while the turn state is the dialogue agent's turn, transitions the turn state to the user's turn; the utterance output unit outputs the second utterance to hand over the turn when the turn state transitions to the user's turn.
4. The interactive device according to claim 3.
6. the utterance detection unit determines whether the first utterance has been detected based on a speech recognition result of recognizing the first utterance and a duration of the first utterance.
6. An interactive device according to any one of claims 1 to 5.
7. an emotion recognition unit that recognizes an emotion of the user based on the first utterance; the utterance output unit outputs the second utterance according to the emotion recognition result.
6. An interactive device according to any one of claims 1 to 5.
8. the emotion recognition unit recognizes the emotion based on acoustic features extracted from the first utterance and a speech recognition result of recognizing the first utterance.
8. An interactive device according to claim 7.
9. a response generation unit that generates a response utterance to the first utterance based on a speech recognition result that recognizes the first utterance, the utterance output unit, when generating the response utterance, outputs the second utterance summarizing the speech recognition result.
6. An interactive device according to any one of claims 1 to 5.
10. the utterance output unit outputs the second utterance that holds the turn when a generation time of the response utterance exceeds a threshold.
10. The interactive device of claim 9.
11. further comprising a motion detection unit that detects a first motion by the user; the utterance output unit outputs a second action from the dialogue agent in accordance with a detection result of the first action; 6. An interactive device according to any one of claims 1 to 5.
12. A dialogue system in which a terminal device operated by a user and a dialogue device that responds to utterances of the user using a dialogue agent can communicate with each other via a network, The dialogue device an utterance detection unit that detects a first utterance by the user; an utterance output unit that controls the dialogue agent to output a response to the content of the first utterance detected by the utterance detection unit; Equipped with and when a predetermined condition is met, the utterance output unit controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response. A dialogue system characterized by:
13. A computer that responds to a user's utterances using a dialogue agent, detecting a first utterance by the user; a step of controlling the dialogue agent to output a response to the content of the first utterance detected in the step of detecting; Run the controlling step controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response when a predetermined condition is met. How to interact.
14. A computer that responds to user utterances using a dialogue agent. detecting a first utterance by the user; a step of controlling the dialogue agent to output a response to the content of the first utterance detected in the step of detecting; Execute the controlling step controls the dialogue agent to output a second utterance that facilitates dialogue with the user before outputting the response when a predetermined condition is met. program.
Citation Information
Patent Citations
Voice interactive method and voice interactive system
JP2016038501A