Conversation system, conversation control method, and storage medium
By acquiring the user's language and non-language information to generate the dialogue agent's response content, the problem in the prior art that the dialogue agent cannot comprehensively consider language and non-language information is solved, and a more accurate dialogue response is achieved.
Patent Information
- Application Number
- CN202480010765.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-01-18
- Publication Date
- 2025-09-12
AI Technical Summary
In the prior art, the dialogue of the dialogue agent cannot generate appropriate response content based on the user's language information and non-language information.
The first acquisition unit acquires the user's language information, the second acquisition unit acquires the user's non-language information, the generation unit generates the dialogue agent's response content based on the two, and the control unit controls the dialogue agent's response.
The appropriateness of the response content of the dialogue agent in the dialogue system is improved, and it can more accurately understand the user's intention and respond.
Smart Images

Figure CN120641904A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a dialogue system, a dialogue control method, and a storage medium. Background Art
[0002] There are conversational systems in which conversational agents automatically respond to messages from users. Agent systems are also known that learn conversations with users and change attributes such as the tone and personality of the conversational agent (for example, Patent Document 1).
[0003] Citation List
[0004] Patent Literature
[0005] [Patent Document 1] Japanese Unexamined Patent Application Publication No. 2022-093479 Summary of the Invention
[0006] Technical issues
[0007] In the prior art, there is a problem that the dialogue agent's conversation cannot generate the dialogue agent's response content based on the user's language information and the user's non-language information.
[0008] An embodiment of the present disclosure is made in view of the above-mentioned problems. In a dialogue system that uses a dialogue agent to conduct dialogue with a user, the dialogue agent's response content can be generated based on the user's language information and the user's non-language information.
[0009] Solutions to the Problem
[0010] To address the above issues, a dialogue system according to one embodiment of the present disclosure uses a dialogue agent to conduct a dialogue with a user. The dialogue system includes a first acquisition unit, a second acquisition unit, a generation unit, and a control unit. The first acquisition unit acquires the user's language information from the dialogue. The second acquisition unit acquires the user's non-verbal information from the dialogue. The generation unit generates response content including a verbal response and a non-verbal response from the dialogue agent based on the user's language information and the user's non-verbal information. The control unit controls the dialogue agent based on the response content generated by the generation unit.
[0011] Effects of the present invention
[0012] According to an embodiment of the present disclosure, in a dialogue system that uses a dialogue agent to conduct a dialogue with a user, response content of the dialogue agent can be generated based on the user's language information and the user's non-language information. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] A more complete understanding of the embodiments of the present disclosure and many of its attendant advantages and features can be readily obtained and understood from the following detailed description taken in conjunction with the accompanying drawings.
[0014] [ Figure 1 ]
[0015] Figure 1 FIG. 1 is a diagram illustrating an example of a system structure of a dialogue system according to an embodiment of the present disclosure.
[0016] [ Figure 2 ]
[0017] Figure 2 FIG. 4 is a diagram illustrating an example of a conversational agent according to an embodiment of the present disclosure.
[0018] [ Figure 3 ]
[0019] Figure 3 FIG. 4 is a diagram illustrating another example of a conversational agent according to an embodiment of the present disclosure.
[0020] [ Figure 4 ]
[0021] Figure 4 This is a diagram showing an overview of a conversation process according to an embodiment of the present disclosure.
[0022] [ Figure 5 ]
[0023] Figure 5 This is a diagram showing an example of the hardware configuration of a computer according to an embodiment of the present disclosure.
[0024] [ Figure 6 ]
[0025] Figure 6 1 is a diagram illustrating an example of a hardware configuration of a terminal device according to an embodiment of the present disclosure.
[0026] [ Figure 7 ]
[0027] Figure 7 : is a diagram showing an example of a functional configuration of a dialogue system according to an embodiment of the present disclosure.
[0028] [ Figure 8 ]
[0029] Figure 8 This is a flowchart showing an overview of a conversation process according to an embodiment of the present disclosure.
[0030] [ Figure 9 ]
[0031] Figure 9is a diagram showing a functional configuration of a generation unit according to the first embodiment of the present disclosure.
[0032] [Figure 10]
[0033] Figure 10AA and Figure 10AB is a flowchart of a conversation process according to the first embodiment of the present disclosure, Figure 10BA and Figure 10BB is another flowchart of the dialogue process according to the first embodiment of the present disclosure.
[0034] [ Figure 11 ]
[0035] Figure 11 1 is a diagram showing an example of utilizing non-verbal information according to the first embodiment of the present disclosure.
[0036] [ Figure 12 ]
[0037] Figure 12 2 is a diagram showing the transition of a dialog scene according to the second embodiment of the present disclosure.
[0038] [ Figure 13 ]
[0039] Figure 13 is another diagram illustrating the transition of a dialogue scene according to the second embodiment of the present disclosure.
[0040] [ Figure 14 ]
[0041] Figure 14 2 is an example showing a conversation screen according to the third embodiment of the present disclosure.
[0042] [ Figure 15 ]
[0043] Figure 15 is a diagram showing a functional configuration of a dialogue system according to a third embodiment of the present disclosure.
[0044] [ Figure 16 ]
[0045] Figure 16 is a flowchart of a dialogue process according to the third embodiment of the present disclosure.
[0046] [ Figure 17 ]
[0047] Figure 17 is a diagram showing a functional configuration of a dialogue system according to a fourth embodiment of the present disclosure.
[0048] [Figure 18]
[0049] Figure 18A and Figure 18B is a diagram showing a conversation log according to a fourth embodiment of the present disclosure.
[0050] [ Figure 19 ]
[0051] Figure 19 is a diagram showing an example of a functional configuration of a dialogue system according to a fifth embodiment of the present disclosure.
[0052] [ Figure 20 ]
[0053] Figure 20 1 is a flowchart showing an example of a process of presenting a promotional slogan according to the fifth embodiment of the present disclosure.
[0054] [ Figure 21 ]
[0055] Figure 21 is a diagram showing an example of a functional configuration of a dialogue system according to a sixth embodiment of the present disclosure.
[0056] [ Figure 22 ]
[0057] Figure 22 2 is a diagram showing input and output information according to the sixth embodiment of the present disclosure.
[0058] [ Figure 23 ]
[0059] Figure 23 is a flowchart showing a dialogue process according to a sixth embodiment of the present disclosure.
[0060] [ Figure 24 ]
[0061] Figure 24 FIG. 1 is a diagram showing an example of a system configuration of a utilization scenario 1 according to an embodiment of the present disclosure.
[0062] [ Figure 25 ]
[0063] Figure 25 1 is a flowchart illustrating a conversation start process using scenario 1 according to an embodiment of the present disclosure.
[0064] [ Figure 26 ]
[0065] Figure 26 FIG. 1 is a diagram showing an example of a system configuration of a utilization scenario 2 according to an embodiment of the present disclosure.
[0066] [ Figure 27 ]
[0067] Figure 271 is a flowchart illustrating a conversation start process using scenario 2 according to an embodiment of the present disclosure.
[0068] [ Figure 28 ]
[0069] Figure 28 FIG. 1 is a diagram showing an example of a system configuration of a utilization scenario 3 according to an embodiment of the present disclosure.
[0070] [ Figure 29 ]
[0071] Figure 29 3 is a flowchart showing a conversation start process using scenario 3 according to an embodiment of the present disclosure.
[0072] The accompanying drawings are intended to depict embodiments of the present invention and should not be interpreted as limiting its scope. The accompanying drawings should not be considered to be drawn to scale unless explicitly noted. Likewise, the same or similar reference numerals represent the same or similar components in multiple views. DETAILED DESCRIPTION
[0073] In describing the embodiments shown in the drawings, specific terms are employed for the sake of clarity. However, the disclosure of this specification is not intended to be limited to the specific terms so selected, and it should be understood that each specific element includes all technical equivalents that have similar functions, operate in a similar manner, and achieve similar results.
[0074] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0075] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Figure 1 : is a diagram showing an example of a system structure of a dialogue system according to an embodiment of the present disclosure. Figure 1 In the example of FIG. 1 , the interactive system 1 includes a server device 100 and a terminal device 10 connected to a communication network N such as the Internet and a LAN (Local Area Network).
[0076] The server device 100 is, for example, an information processing device having a computer structure or a system composed of multiple computers. The server device 100 provides a conversation service in which a conversation agent automatically responds to messages from the user 11 using the terminal device 10 by executing a predetermined program on the computer included in the server device 100.
[0077] The terminal device 10 is an information terminal such as a PC (Personal Computer), a tablet terminal, or a smartphone, used by a user 11. The terminal device 10 can communicate with the server device 100 via a communication network N. The user 11 can use the interactive service provided by the server device 100 using the terminal device 10.
[0078] Preferably, the dialogue system 1 assists in performing a predetermined task such as business negotiation or nursing care by automatically responding to a message from a user through a dialogue agent.
[0079] Figure 1 The system structure of the dialogue system 1 shown is an example. In addition, the terminal device 10 is not limited to a general-purpose information terminal, and may be, for example, a dedicated terminal device or various electronic devices. In addition, the dialogue system 1 may also be implemented by an information processing device having a computer structure. Here, the dialogue system 1 is a Figure 1 The system having the system structure shown is described below.
[0080] A conversational agent is a system that automatically responds to inquiries from users or customers using information or knowledge registered in response to inquiries, such as knowledge from users or customers, or AI (Artificial Intelligence).
[0081] As a conversational agent, it can also be used as an online conference, website, smartphone application, or unmanned AI virtual image in the metaverse.
[0082] Figure 2 An example of an image of a conversation agent according to an embodiment is shown. This figure shows an example of a conversation screen 200 for business negotiation that the server device 100 causes the terminal device 10 to display. Figure 2 In the example shown in FIG. 2 , a virtual person 201 generated by 3D modeling is displayed on a dialogue screen 200. Virtual person 201 is an example of a dialogue agent. For example, server device 100 controls virtual person 201 in dialogue screen 200 so that the virtual person 201 can conduct a conversation with user 11 and advance a business negotiation.
[0083] As a preferred example, a large display 202 is displayed on the business negotiation dialogue screen 200. The server device 100 can also control the display 202 to display, for example, products and services proposed to the user and to cause the virtual person 201 to explain the products and services.
[0084] Figure 3Another example of an image of a dialogue agent according to an embodiment is shown. This figure shows an example of a dialogue screen 300 for care purposes that the server device 100 causes the terminal device 10 to display. Figure 3 In the example of Figure 2 Similarly, another virtual person 301 generated by 3D modeling is displayed. Virtual person 301 is another example of a dialogue agent. In this dialogue screen 300, server device 100 controls virtual person 301 so as to communicate with elderly people living alone, for example, to prevent dementia.
[0085] As a preferred example, Figure 3 As shown, the user 11 and the virtual person 301 can have a conversation 302 based on character strings in addition to (or instead of) voice.
[0086] As described above, the dialogue system 1 can change the dialogue scene to change the content of the dialogue according to various purposes such as business negotiation, nursing care, teaching, or consultation.
[0087] Figure 4 This diagram is used to explain an overview of conversation processing according to one embodiment. It shows an example of the relationship between user 11's verbal and non-verbal information and the conversational agent's verbal and non-verbal responses during a conversation between user 11 and the conversational agent, with the horizontal axis representing time.
[0088] exist Figure 4 In the example, when the user 11 starts the operation, at time t1, the server device 100 causes the dialogue agent to make a speech 401 such as a greeting or an opening as a verbal response, and to perform an opening 402 such as a bow or a smile as a non-verbal response.
[0089] Accordingly, at time t2, when the user speaks, server device 100 obtains the language information and non-language information of user 11. At this time, server device 100 may cause the dialogue agent to make a non-language response such as nodding 403, for example.
[0090] User 11's verbal information includes, for example, information representing the content of user 11's speech 411, converted into text using voice recognition technology. User 11's non-verbal information includes, for example, information other than verbal information, such as user 11's facial expressions, gaze, posture, or emotions, acquired through image recognition technology. This non-verbal information includes, for example, voice tone, speaking speed, pitch, intensity, coughs, sighs, laughter, or silence, acquired from the sound included in the user's video (paralanguage). As described above, non-verbal information such as images and sounds is utilized in a multimodal manner.
[0091] Language information is information that conveys the content of speech through language. For example, it conveys the meaning based on well-defined language rules and dictionaries, such as vocabulary, grammar, sentence structure, and context. Language information includes, for example, information representing the content of user 11's speech 411, which is converted into text data using voice recognition technology.
[0092] Non-verbal information is information conveyed through means other than language. Non-verbal information includes information other than language, such as user 11's facial expressions, gaze, posture, or emotions, acquired through image recognition technology. User 11's non-verbal information also includes, for example, audio information (paralanguage) acquired from the sound included in the image of user 11, other than language, such as voice tone, speaking speed, pitch, intensity, coughs, sighs, laughter, or silence. As described above, the dialogue system 1 according to this embodiment utilizes non-verbal information, such as images and sounds, in a multimodal manner.
[0093] Server device 100 interprets the intention of user 11's speech based on user 11's language information and taking into account non-language information of user 11. This allows server device 100 to improve the accuracy of intention interpretation compared to interpreting intention based on language information alone.
[0094] Server device 100 generates a conversational agent response corresponding to the intended meaning of user 11's speech. This response includes a verbal response representing the conversational agent's speech content, as well as a non-verbal response, such as the conversational agent's facial expressions or gestures. Preferably, server device 100 modifies the conversational agent's non-verbal response based on the acquired non-verbal information about user 11.
[0095] When time t3 arrives, server device 100 controls the dialogue agent based on the generated response. For example, server device 100 uses speech synthesis processing to convert the generated verbal response into speech, causing the dialogue agent to speak 404. Preferably, server device 100 moves the dialogue agent's mouth (lip synchronization) based on dialogue agent's speech 404. Furthermore, server device 100 causes the dialogue agent to perform non-verbal responses, such as facial expressions or gestures, based on the generated non-verbal response.
[0096] As described above, the dialogue system 1 according to this embodiment changes the response content (verbal and non-verbal) of the dialogue agent (virtual person 201, 301) based on the non-verbal information of the user 11. As a result, according to this embodiment, the dialogue system 1, which uses the dialogue agent to conduct a dialogue with the user, can provide a more appropriate response to the user 11.
[0097] The server device 100 has, for example, Figure 5Alternatively, the server device 100 may be composed of a plurality of computers 500. In addition, the terminal device 10 may also have, for example, Figure 5 The hardware structure of the computer 500 is shown.
[0098] Figure 5 FIG. 1 is a diagram showing an example of the hardware configuration of a computer according to an embodiment of the present invention. Figure 5 As shown, the computer 500 includes a CPU (central processing unit) 501, a ROM (read-only memory) 502, a RAM (random access memory) 503, an HD (hard disk) 504, an HDD (hard disk drive) controller 505, a display 506, an external device connection I / F (interface) 507, a network I / F 508, a keyboard 509, a pointing device 510, a DVD-RW (digital versatile disc rewritable) drive 512, a medium I / F 514, a bus 515, and the like.
[0099] When the computer 500 is the terminal device 10, the computer 500 further includes a microphone 521, a speaker 522, a sound input and output I / F 523, a CMOS (Complementary Metal Oxide Semiconductor) sensor 524, and an imaging element I / F 525.
[0100] The CPU 501 controls the overall operation of the computer 500. The ROM 502 stores programs for starting up the computer 500, such as an IPL (Initial Program Loader).
[0101] The RAM 503 is used as, for example, a work area for the CPU 501. The HD 504 stores, for example, an OS (operating system), application programs, programs such as device drivers, and various data. The HDD controller 505 controls, for example, reading various data from the HD 504 and writing various data to the HD 210 in accordance with the control of the CPU 501. The HD 504 and the HDD controller 505 are examples of storage devices.
[0102] The display 506 displays various information, such as a cursor, menus, windows, characters, and images. Alternatively, the display 506 may be provided externally to the computer 500. The external device connection I / F 507 is an interface for connecting various external devices to the computer 500. The network I / F 508 is an interface for connecting the computer 500 to the communication network 2 for communication with other devices.
[0103] The keyboard 509 is a type of input unit having multiple keys for inputting text, numerical values, various instructions, etc. The pointing device 510 is an input device that performs functions such as selecting and executing various commands, selecting processing targets, and moving a cursor. Alternatively, the keyboard 509 and pointing device 510 may be provided externally to the computer 500.
[0104] The DVD-RW drive 512 controls the reading and writing of various data from and to a DVD-RW 511, which is an example of a removable recording medium. The DVD-RW 511 is not limited to a DVD-RW and may also be another removable recording medium. The media I / F 514 controls the reading and writing (storage) of data from and to a medium 513, such as a flash memory. The bus 515 includes an address bus, a data bus, and various control signals for electrically connecting the aforementioned components.
[0105] Microphone 521 is a built-in circuit that converts sound into electrical signals. Speaker 522 is a built-in circuit that converts electrical signals into physical vibrations to produce sounds such as music and voices. Audio input / output I / F 523 is a circuit that processes the input and output of audio signals between microphone 521 and speaker 522 under the control of CPU 501.
[0106] The CMOS sensor 524 is a built-in imaging device that captures an image of a subject (e.g., its own image) under the control of the CPU 501 and obtains image data. The computer 500 may include an imaging device such as a CCD (charge coupled device) sensor instead of the CMOS sensor 524. The imaging element I / F 525 is a circuit that controls the driving of the CMOS sensor 524.
[0107] Figure 6 1 is a diagram showing an example of the hardware configuration of a terminal device according to an embodiment. Here, an example of the hardware configuration of the terminal device 10 will be described when the terminal device 10 is an information terminal such as a smartphone or a tablet terminal.
[0108] exist Figure 6 In the example, the terminal device 10 includes a CPU 601 , a ROM 602 , a RAM 603 , a storage device 604 , a CMOS sensor 605 , an imaging element I / F 606 , an acceleration / orientation sensor 607 , a medium I / F 609 , and a GPS (Global Positioning System) receiver 610 .
[0109] The CPU 601 controls the overall operation of the terminal device 10 by executing a predetermined program. The ROM 602 stores programs such as an IPL for starting the CPU 601. The RAM 603 serves as a work area for the CPU 601. The storage device 604 is a large-capacity storage device that stores programs such as the OS and applications, as well as various data, and is implemented, for example, as an SSD (Solid State Drive) or flash ROM.
[0110] The CMOS sensor 605 is a built-in camera device that captures an object (mainly its own image) under the control of the CPU 601 and obtains image data. The terminal device 10 may have a camera device such as a CCD sensor instead of the CMOS sensor 605. The camera element I / F 606 is a circuit that controls the drive of the CMOS sensor 605. The acceleration / azimuth sensor 607 is a variety of sensors such as an electromagnetic compass, a gyrocompass, and an acceleration sensor for detecting geomagnetism. The medium I / F 609 controls the reading of data from or the writing (storage) of data to a medium (storage medium) 608 such as a flash memory. The GPS receiver 610 receives GPS signals (positioning signals) from GPS satellites.
[0111] The terminal device 10 includes a long-distance communication circuit 611, an antenna 611a of the long-distance communication circuit 611, a CMOS sensor 612, an imaging element I / F 613, a microphone 614, a speaker 615, a sound input and output I / F 616, a display 617, an external device connection I / F 618, a short-distance communication circuit 619, an antenna 619a of the short-distance communication circuit 619, and a touch screen 620.
[0112] Long-distance communication circuit 611 is a circuit for communicating with other devices via communication network 2, for example. CMOS sensor 612 is a built-in imaging device that captures images of a subject and obtains image data under the control of CPU 601. Image sensor I / F 613 is a circuit that controls the driving of CMOS sensor 612. Microphone 614 is a built-in circuit that converts sound into electrical signals. Speaker 615 is a built-in circuit that converts electrical signals into physical vibrations to produce sounds such as music and voices. Sound input / output I / F 616 is a circuit that processes the input and output of sound wave signals between microphone 614 and speaker 615 under the control of CPU 601.
[0113] The display 617 is a display unit that displays images of a subject, various icons, and the like. The display 617 includes a liquid crystal display (LCD) and an organic electroluminescent (EL) display. The external device connection interface 618 is an interface for connecting to various external devices. The short-range communication circuit 619 includes circuitry for performing short-range wireless communication. The touch panel 620 is an input device that allows the user to operate the terminal device 10 by pressing the screen of the display 617.
[0114] The terminal device 10 includes a bus 621. The bus 621 includes a bus for electrically connecting Figure 6 The address bus, data bus, etc. of each component such as the CPU 601 shown.
[0115] Figure 6 The hardware configuration of the terminal device 10 shown is an example, and the terminal device 10 may have other hardware configurations as long as it has a computer structure, a communication circuit, a display, a microphone, a speaker, and the like.
[0116] Figure 7 This is a diagram showing an example of the functional configuration of a dialogue system according to an embodiment of the present disclosure.
[0117] The server device 100 executes a predetermined program stored in a storage medium by the computer 500 included in the server device 100, thereby realizing, for example, Figure 7 The functional configuration shown. Figure 7 In the example of , the server device 100 includes a communication unit 701, a first acquisition unit 702, a second acquisition unit 703, a generation unit 704, a speech synthesis unit 711, a drawing unit 712, and an output unit 713. At least part of the above functional units can be implemented by hardware.
[0118] The server device 100 implements the storage unit 710 by, for example, a storage device such as the HD 504 and the HDD controller 505. The storage unit 710 may also be implemented by, for example, a storage server provided outside the server device 100 or a cloud service.
[0119] The communication unit 701 connects the server device 100 to the communication network N using, for example, the network I / F 508 and executes communication processing for communicating with other devices such as the terminal device 10 .
[0120] First acquisition unit 702 performs a first acquisition process to acquire user 11's language information from a conversation with user 11 using terminal device 10. For example, first acquisition unit 702 detects audio segments based on the video (moving images and audio) of user 11 received from terminal device 10 by communication unit 701 using a technique such as VAD (Voice Activity Detection), thereby acquiring user 11's speech. Furthermore, first acquisition unit 702 performs voice recognition processing on the acquired speech of user 11, converting the speech of user 11 into text. Furthermore, first acquisition unit 702 acquires the texted version of user 11's speech as user 11's language information.
[0121] Second acquisition unit 703 performs a second acquisition process to obtain non-verbal information of user 11 based on a conversation with user 11 using terminal device 10. For example, second acquisition unit 703 obtains non-verbal information of user 11, such as facial expressions, gaze, or emotions, through image processing based on the image (moving image and sound) of user 11 received by communication unit 701 from terminal device 10. Second acquisition unit 703 obtains non-verbal information of user 11, such as the volume, intonation, or timbre of the voice, from the image (moving image and sound) of user 11 received by communication unit 701 from terminal device 10.
[0122] Generation unit 704 performs a generation process to generate response content, including the dialogue agent's verbal response (conversation content) and non-verbal responses (such as the dialogue agent's actions and paralinguistics), based on the verbal information of user 11 acquired by first acquisition unit 702 and the non-verbal information of user 11 acquired by second acquisition unit 703. For example, generation unit 704 includes dialogue control unit 705, intention interpretation unit 706, and response generation unit 707. The backend of response generation unit 707 stores a large amount of actual conversation information (such as audio and video), which is used to construct response generation unit 707. If response generation unit 707 is a machine learning model (described later), this conversation information is used as learning data to improve the accuracy of dialogue generation.
[0123] The dialogue control unit 705 performs a dialogue control process including a process of inputting verbal information and non-verbal information of the user 11 and a process of outputting verbal responses and non-verbal responses of the dialogue agent.
[0124] Intent interpretation unit 706 performs intent interpretation processing, which interprets the intent of user 11's speech based on user 11's language information and further considers user 11's non-verbal information. For example, if user 11 says "That's fine," it is sometimes difficult to determine whether user 11 intended "OK" or "No," based solely on user 11's language information (speech text). Therefore, intent interpretation unit 706 involved in this embodiment uses not only user 11's language information (speech text) but also user 11's non-verbal information to interpret the intent of user 11's speech. As a result, intent interpretation unit 706 can improve the accuracy of intent interpretation processing.
[0125] For example, the intention interpretation unit 706 can input the language information and non-language information of the user 11 into the machine learning model to interpret the language intention of the user 11. The machine learning model is pre-trained by inputting the language information and non-language information of multiple users to interpret the intention of the user 11.
[0126] In this disclosure, machine learning refers to technology used to enable computers to acquire human-like learning capabilities. Machine learning refers to a method in which a computer autonomously generates the algorithms required for identification and other judgments based on previously acquired training data, and applies these algorithms to new data for prediction. The learning methods used for machine learning can be any of supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and deep learning, and can also be combinations of these learning methods. There is no limitation to the learning methods used for machine learning.
[0127] Response generation unit 707 generates a response from the conversational agent that corresponds to the intended meaning of user 11's speech. This response includes a verbal response representing the verbal content of the conversational agent's speech, as well as non-verbal responses such as expressions or gestures used by the conversational agent. Preferably, server device 100 modifies the conversational agent's verbal and non-verbal responses based on the acquired non-verbal information about user 11.
[0128] For example, the response generation unit 707 changes the content of the behavior of the dialogue agent based on the non-verbal information of the user 11. In addition, the response generation unit 707 changes the timing of the behavior of the dialogue agent based on the non-verbal information of the user 11.
[0129] To generate responses, rule-based natural language processing or large-scale language models can be used. For example, a large-scale language model can be applied, such as the text generation language model called GPT-3 (Generative Pre-trained Transformer 3). In rule-based natural language processing, the dialogue agent generates responses based on rules that predefine the response content, in response to the user's speech intent.
[0130] In the process of answering content, there are situational type and slot filling type.
[0131] The speech synthesis unit 711 performs speech synthesis processing to vocalize the speech response generated by the generation unit 704 using speech synthesis technology.
[0132] The drawing unit 712 performs a drawing process for drawing a dialogue screen depicting the dialogue agent according to the non-verbal response generated by the generating unit 704. For example, the drawing unit 712 reflects facial expressions, gaze, posture, or emotions on the dialogue screen according to the non-verbal response. Figure 2 A virtual person (dialogue agent) 201 is shown.
[0133] Preferably, the drawing unit 712 also draws lip synchronization that moves the dialogue agent's mouth according to the dialogue agent's speech.
[0134] Output unit 713 performs processing to output an image including the voice of the conversational agent voiced by voice synthesis unit 711 and the conversation screen drawn by drawing unit 712. For example, output unit 713 transmits an image including the voice of the conversational agent voiced by voice synthesis unit 711 and the conversation screen drawn by drawing unit 712 to terminal device 10 via communication unit 701.
[0135] The voice synthesis unit 711 , the drawing unit 712 , and the output unit 713 serve as a control unit 714 that controls the dialogue agent based on the response content generated by the generation unit 704 .
[0136] The storage unit 710 stores various information, data, programs, and the like, such as machine learning models, rules, setting information, and conversation logs used by the server device 100 .
[0137] The terminal device 10 can access the server device 100 by using a web browser etc. provided in the terminal device 10, and display the Figure 2 The dialogue screen 200 shown here can be configured in any functional manner by sending the video of the user 11 .
[0138] Figure 7The system structure of the dialogue system 1 shown in FIG. 1 is an example. For example, the dialogue system 1 may also be composed of Figure 7 The functional configuration of the server device 100 shown in FIG. 1 is constituted by an information processing device. Furthermore, at least a portion of the functional configuration of the server device 100 may also be provided by the terminal device 10. For example, the terminal device 10 may include a first acquisition unit 702, a second acquisition unit 703, a voice synthesis unit 711, a drawing unit 712, an output unit 713, and the like. In this case, the terminal device 10 may also transmit language information and non-language information to the server device 100 and display a dialogue screen based on the language responses and non-language responses received from the server device 100.
[0139] Figure 8 This is a flowchart showing an overview of the dialogue process performed by the dialogue system according to one embodiment of the present disclosure. Figure 7 An example of the processing repeatedly performed by the dialogue system 1 having the functional configuration shown. Figure 8 The start time of the processing is assumed to be that a conversation has already taken place between the user 11 using the terminal device 10 and the conversation agent provided by the server device 100.
[0140] In step S801, first acquisition unit 702 acquires user language information from the conversation between user 11 and the dialogue agent. For example, first acquisition unit 702 acquires user 11's speech from the video of user 11 received by communication unit 701 from terminal device 10. Furthermore, first acquisition unit 702 performs voice recognition processing on the acquired speech of user 11, and acquires a text of user 11's speech (language information) by converting the speech of user 11 into text.
[0141] In step S802, second acquisition unit 703 acquires non-verbal information about user 11 from the conversation between user 11 and the dialogue agent in parallel with the processing in step S801. For example, second acquisition unit 703 performs image processing on the image of user 11 received by communication unit 701 from terminal device 10 to acquire non-verbal information such as user 11's facial expression, gaze, or emotion. Second acquisition unit 703 also performs audio processing on the image of user 11 received by communication unit 701 from terminal device 10 to acquire non-verbal information such as the volume, intonation, or timbre of the voice.
[0142] In step S803 , the generation unit 704 interprets the intention of the user 11 's speech based on the language information of the user 11 acquired by the first acquisition unit 702 and the non-language information of the user 11 acquired by the second acquisition unit 703 .
[0143] In step S804 , the generation unit 704 generates a verbal response and a non-verbal response corresponding to the intention of the user 11 's speech.
[0144] In step S805 , the voice synthesis unit 711 synthesizes the speech voice of the dialogue agent based on the language response generated by the generation unit 704 .
[0145] In step S806 , the drawing unit 712 draws the dialogue agent based on the non-verbal response generated by the generation unit 704 in parallel with the process of step S805 .
[0146] In step S807, the output unit 713 outputs the dialogue screen including the speech sound of the dialogue agent synthesized by the sound synthesizing unit 711 and the dialogue agent drawn by the drawing unit 712. For example, the output unit 713 transmits the dialogue screen to the terminal device 10 using the communication unit 701.
[0147] The dialogue system 1 repeatedly executes Figure 8 This processing can change not only the conversational agent's voice but also its non-verbal response based on the non-verbal information of user 11. As a result, according to this embodiment, in conversational system 1 that uses a conversational agent to conduct a conversation with user 11, a more appropriate response to user 11 can be provided.
[0148] First embodiment
[0149] The dialogue system 1 of this embodiment can cope with various applications by changing the dialogue scene. In the first embodiment, an example of dialogue processing corresponding to a business negotiation application is described.
[0150] The dialogue system 1 according to the first embodiment has, for example, Figure 7 The generating unit 704 according to the first embodiment has, for example, Figure 9 Functional configuration shown.
[0151] Figure 9 : is a diagram showing an example of the functional configuration of the generation unit according to the first embodiment. Figure 9 As shown, the dialog control unit 705 of the generation unit 704 includes, for example, an input filtering unit 901 , a dialog state management unit 902 , and an output filtering unit 903 .
[0152] The input filter unit 901 includes, for example, an input I / F function for accepting input of language information and non-language information from the user 11, a misrecognition handling function, and a function for detecting inappropriate input. The misrecognition handling function and the function for detecting inappropriate input are optional functions and are not essential.
[0153] The dialog state management unit 902 has functions such as recording input information, storing the current business negotiation stage, controlling the business negotiation stage, and recording output information. The business negotiation stage is an example of a numerical definition of the progress of the business negotiation.
[0154] The output filter unit 903 has, for example, a function of outputting language-compatible and non-language-compatible output I / Fs of the dialogue agent and a function of detecting inappropriate outputs. The function of detecting inappropriate outputs is optional and not essential.
[0155] Intent interpretation section 706 performs intent interpretation processing to interpret the intent of user 11's speech based on the user 11's language information and non-language information received by dialogue control section 705. While intent interpretation section 706 can also infer user 11's intent based on user 11's language information and context, the possibility of more accurately interpreting user 11's intent increases when non-language information is taken into account.
[0156] For example, while user 11's utterances such as "Isn't it?" are often used as negative responses, they can also be used when user 11 expresses joy in a positive way, exceeding expectations. In such cases, the intended intention interpretation unit 706 uses user 11's non-verbal information as clues to more accurately interpret user 11's intentions.
[0157] For example, if user 11's voice has a high pitch and a bright expression as non-verbal information, intention interpretation unit 706 may interpret user 11's utterance, such as "Isn't it?", as "positive (happy)." In this case, generation unit 704 may also set the dialogue agent's expression to a smiley face, maintaining the current conversational scene.
[0158] On the other hand, if the pitch of user 11's voice is low or the expression in the image of user 11 is dull, intention interpretation unit 706 may also determine that user 11's utterance, such as "Isn't it?", is "negative." In this case, generation unit 704 may, for example, reduce the dialogue agent's gestures and transition to a dialogue (product) scenario containing more detailed examples (or to a scenario featuring another product).
[0159] In business scenarios like business negotiations, it's difficult for the other person to express their emotions. However, nonverbal information perceived as negative can be crucial, affecting not only the progress of the negotiation but also the long-term emotional development of the next negotiation. Therefore, careful handling is required. For example, it is preferable to consider the depth of user 11's voice and the degree of gloom on their facial expression when determining whether to change the scene.
[0160] The response generation unit 707 includes multiple dialogue scenes 911-917 corresponding to multiple business negotiation stages 1-7, a product and service recommendation unit 918, and a judgment unit 919. The product and service recommendation unit 918 and the judgment unit 919 can be provided outside the response generation unit 707.
[0161] Dialogue scene 911, corresponding to the first stage, is used when starting a business negotiation, such as a greeting at the beginning of a business negotiation or searching for customer data. Dialogue scene 912, corresponding to the second stage, is used, for example, to exchange business cards or engage in small talk. Dialogue scene 913, corresponding to the third stage, is used, for example, to listen to business details or hear about device usage. Dialogue scene 914, corresponding to the fourth stage, is used, for example, to confirm specific needs or explore potential needs.
[0162] Dialogue scene 915 corresponding to the fifth stage includes, for example, prompting recommended products, providing sales copy to encourage purchases, determining whether to postpone or terminate a business negotiation, etc. Dialogue scene 916 corresponding to the sixth stage includes, for example, confirming delivery dates and guiding electronic contracts. Dialogue scene 917 corresponding to the seventh stage includes, for example, creating a daily report or preparing / distributing a questionnaire.
[0163] Dialogue state management unit 902 of dialogue control unit 705 selects a dialogue scenario to use from multiple dialogue scenarios 911-917 based on the current state of the business negotiation. For example, dialogue control unit 705 begins the business negotiation with dialogue scenario 911 corresponding to the first stage and advances the stage as the negotiation progresses. Alternatively, if user 11 declines the business negotiation, dialogue control unit 705 may lower the stage.
[0164] Thus, the generation unit 704 can change the response content of the dialogue agent according to the plurality of pre-set business negotiation stages. The business negotiation stage is an example of the plurality of pre-set dialogue stages.
[0165] For example, in the fifth stage, the product and service recommendation unit 918 performs product recommendation processing to select products recommended to the user 11 based on the content of the conversation in the first to fourth stages. For example, in the fifth stage, the judgment unit 919 performs judgment processing to determine whether to postpone or terminate the business negotiation based on the content of the conversation in the first to fifth stages.
[0166] Figure 9 The number of the plurality of business negotiation stages 1 to 7 shown is an example, and other numbers of 2 or more may be used. Figure 9The dialogue contents of the illustrated plurality of dialogue scenes 911 to 917 are merely examples, and other contents may be used.
[0167] Figure 10AA and Figure 10AB is a flowchart of an example of a dialogue process according to the first embodiment of the present disclosure. Figure 7 The functional configuration of the server device 100 shown in FIG. Figure 9 The functional configuration of the generation unit 704 shown is an example of the dialogue process performed by the dialogue system 1.
[0168] In step S1001, the dialogue system 1 starts a dialogue in the dialogue scene 911 corresponding to the first stage and determines whether there is customer data related to the user 11. If there is customer data, the dialogue system 1 moves the process to step S1002. On the other hand, if there is no customer data, the dialogue system 1 moves the process to step S1008.
[0169] The processes of steps S1002-S1005 and S1008-S1011 involve the same business negotiation phase, but utilize different conversation scenarios. For example, in steps S1002-S1005, since dialogue system 1 has customer data, it is preferable to use a conversation scenario that advances the business negotiation based on past business negotiation data. On the other hand, in steps S1008-S1011, since dialogue system 1 does not have customer data, it is preferable to use a conversation scenario that includes the information necessary to create the customer data and carefully listens to the conversation. This reduces the number of times the dialogue agent hears the same content from user 11.
[0170] If the process proceeds to step S1002, the dialogue system 1 conducts a dialogue in the dialogue scene 912 corresponding to the second stage, and determines whether a business card exchange or a small talk has been performed. If a business card exchange or a small talk has been performed, the dialogue system 1 moves the process to step S1004. On the other hand, if a business card exchange or a small talk has not been performed, the dialogue system 1 ends the process, for example. Figure 10AA Preferably, the dialogue system 1 ends the business negotiation if no business card exchange or small talk is carried out within a predetermined time period after the dialogue starts in the dialogue scene 912 corresponding to the second stage.
[0171] If the process proceeds to step S1003, the dialogue system 1 conducts a dialogue in the dialogue scene 913 corresponding to the third stage, and determines whether the situation such as the content of the business or the use of the equipment is heard. If the situation can be heard, the dialogue system 1 moves the process to step S1004. On the other hand, if the situation cannot be heard, the dialogue system 1 ends the process, for example. Figure 10AAPreferably, the dialogue system 1 ends the business negotiation if it fails to hear the situation after a predetermined time has passed since the dialogue started in the dialogue scene 913 corresponding to the third stage.
[0172] When the process proceeds to step S1004, the dialogue system 1 conducts a dialogue in dialogue scene 914 corresponding to the fourth stage and determines whether it was able to hear a request, such as a potential request or a predicted request. If the request was able to be heard, the dialogue system 1 proceeds to step S1005. On the other hand, if the request was not heard, the dialogue system 1 returns to step S1003.
[0173] When the process proceeds to step S1005, the dialogue system 1 conducts a dialogue in dialogue scene 915 corresponding to the fifth stage and determines whether a product or service can be proposed. If a product or service can be proposed, the dialogue system 1 proceeds to step S1006. On the other hand, if a product or service cannot be proposed, the dialogue system 1 returns to step S1004 or step S1005.
[0174] For example, based on the information acquired in steps S1003 and S1004, the dialogue system 1 uses the product and service recommendation unit 918 to select a product to be proposed to the user 11. However, if the acquired information is insufficient and the product and service recommendation unit 918 cannot select a product to be proposed to the user 11, the dialogue system 1 returns the process to step S1004 or step S1005.
[0175] When the process proceeds to step S1006, the dialogue system 1 conducts a dialogue in dialogue scene 916 corresponding to the sixth stage and determines whether the contract has been signed. If the contract has been signed, the dialogue system 1 proceeds to step S1007. On the other hand, if the contract has not been signed, the dialogue system 1 returns to step S1005, for example.
[0176] If the process moves to step S1007, the dialogue system 1 conducts a dialogue in the dialogue scene 917 corresponding to the seventh stage and determines whether the business negotiation can be sorted out. If the business negotiation can be sorted out, the dialogue system 1 ends. Figure 10AA and Figure 10AB dialogue processing.
[0177] On the other hand, if the process moves from step S1001 to step S1008, the dialogue system 1 conducts a dialogue in the dialogue scene 912 (for new customers) corresponding to the second stage, and determines whether a business card exchange or a small chat has been performed. If a business card exchange or a small chat has been performed, the dialogue system 1 moves the process to step S1009. On the other hand, if a business card exchange or a small chat has not been performed, the dialogue system 1 ends the process. Figure 10AA and Figure 10AB Preferably, the dialogue system 1 ends the business negotiation if no business card exchange or small talk is carried out within a predetermined time period after the dialogue starts with the dialogue scene 912 (for new customers) corresponding to the second stage.
[0178] If the process proceeds to step S1009, the dialogue system 1 conducts a dialogue in the dialogue scene 913 (for new customers) corresponding to the third stage, and determines whether the user has heard the business details or the status of the equipment being used. If the user has heard the status, the dialogue system 1 proceeds to step S1010. On the other hand, if the user has not heard the status, the dialogue system 1 ends the process. Figure 10AA and Figure 10AB Preferably, the dialogue system 1 ends the business negotiation if it fails to hear the situation after a predetermined time has passed since the dialogue started in the dialogue scene 913 (for new customers) corresponding to the third stage.
[0179] When the process proceeds to step S1010, the dialogue system 1 conducts the dialogue in dialogue scenario 914 (for new customers) corresponding to the fourth stage and determines whether it was able to hear a demand, such as a potential demand or a predicted demand. If the demand was able to be heard, the dialogue system 1 proceeds to step S1011. On the other hand, if the demand was not heard, the dialogue system 1 returns to step S1009.
[0180] When the process proceeds to step S1011, the dialogue system 1 conducts a dialogue in dialogue scene 915 (for new customers) corresponding to the fifth stage and determines whether a product suggestion can be made. If a product suggestion can be made, the dialogue system 1 proceeds to step S1006. On the other hand, if a product suggestion cannot be made, the dialogue system 1 returns to step S1010.
[0181] pass Figure 10AA and Figure 10AB By processing, the dialogue system 1 can change the response content of the dialogue agent according to a plurality of preset dialogue stages.
[0182] Figure 10AA and Figure 10AB For example, if the contract cannot be signed in step S1006, the dialogue system 1 may also execute Figure 10BB Processing of steps S1021 and S1022.
[0183] Figure 10BA and Figure 10BB 1006 is another flowchart of an example of the dialogue process according to the first embodiment of the present disclosure. In a case where the contract cannot be concluded in step S1006, the dialogue system 1 shifts the process to step S1021.
[0184] When the process proceeds to step S1021, the dialogue system 1 determines whether the emotion analysis of the user 11 is positive. If the emotion analysis is positive, the dialogue system 1 returns the process to step S1005. On the other hand, if the emotion analysis is not positive (if it is negative), the dialogue system 1 proceeds to step S1022.
[0185] If the process proceeds to step S1022, the dialogue system 1 performs a greeting to end (or postpone) the conversation, and ends the conversation. Figure 10BA and Figure 10BB For example, the dialogue system 1 may cause the dialogue agent to perform a greeting and bow to mark the end of the business negotiation.
[0186] Figure 11 This diagram illustrates an example of utilizing non-verbal information in the first embodiment. For example, dialogue system 1 obtains a direction vector 1101 indicating the direction of user 11's face from an image 1100 of user 11. Based on this direction vector 1101 and the position 112 of user 11's pupil, dialogue system 1 obtains line of sight information indicating user 11's line of sight.
[0187] For example, if user 11 is interested in a product presented by the conversational agent, user 11 tends to focus on the product displayed on the conversation screen, resulting in little eye movement (small divergence), as shown in lines 1103a and 1103b. On the other hand, if user 11 is not interested in a product presented by the conversational agent, their attention is reduced, resulting in a large eye movement (large divergence), as shown in line 1103c, for example.
[0188] Therefore, for example, after presenting a product to user 11, dialogue system 1 may obtain gaze information indicating user 11's gaze, and if the gaze is slightly distracted, determine that user 11's emotional analysis is positive (continuing the business negotiation). Alternatively, dialogue system 1 may obtain gaze information indicating user 11's gaze after presenting a product to user 11, and if the gaze is significantly distracted, determine that user 11's emotional analysis is negative (ending or postponing the business negotiation).
[0189] This method is not limited to determining the end (or extension) of a business negotiation, and can also be used to determine whether to move to a higher business negotiation stage or return to a lower business negotiation stage.
[0190] Second embodiment
[0191] In the first embodiment, an example of conversation processing for nursing care applications is described. In nursing care applications, a conversation scenario for reminiscence therapy can be used. Reminiscence therapy is a psychotherapy in which elderly people, for example, talk about their past to calm their minds and improve their cognitive function.
[0192] In the reminiscence method, conversations about nostalgic memories are considered to be the left brain's process of verbalizing images that emerge in the right brain. Conversations that follow a storyline (introduction, development, transition, and conclusion) are called "when, where, who, what, and why (5W) conversations," while conversations centered on the setting and how things happened are called "how (1H) conversations." Conversations that focus on the setting or the moment are said to be twice as enjoyable as those that follow a storyline.
[0193] In the second embodiment, the dialogue system 1 uses a recall method of dialogue scenarios. In order to deepen the dialogue as it progresses, a plurality of "how to (1H) dialogue" dialogue scenarios are set and maintained, and the dialogue agent's response content is generated based on the dialogue scenarios.
[0194] The functional configuration of the dialogue system 1 according to the second embodiment can be Figure 7 The functional configuration of the dialogue system 1 described in is the same.
[0195] Figure 12 and Figure 13 This figure shows an example of the transition of the dialogue scene according to the second embodiment. This figure shows an example of the transition of the dialogue scene of the recall method. The actual transition varies according to the speech of the user 11, so this figure shows the transition of the dialogue scene according to the second embodiment. Figure 12 、 13 This is an example of transition when user 11 speaks.
[0196] For example, in state 1201, the dialog agent says "Did you do any sports at school?", and in state 1202, assume that the user 11 says "I did sport A at school."
[0197] In this case, as a first stage, the dialogue system 1 causes the dialogue agent to utter an utterance reviewing general knowledge about sport A. For example, in state 1203 , the dialogue agent says, "Where is your location?" Furthermore, in state 1204 , it is assumed that the user 11 says, "I'm at location B."
[0198] In this case, as the second stage, the dialogue system 1 causes the dialogue agent to speak in depth about the topic of sport A. For example, the dialogue agent randomly selects a state from states 1205 , 1209 , 1213 , and 1215 and transitions to the selected state.
[0199] As an example, when transitioning to state 1205, the dialogue agent says "Have you ever participated in a competition?" In state 1206, assume that the user 11 says "I have participated in competitions many times."
[0200] In this case, as the third stage, the dialogue system 1 causes the dialogue agent to speak further into the topic of state 1205. For example, in state 1207, the dialogue agent says, "Have you ever won any prizes in a competition?" In state 1208, assume that the user 11 says, "I participated in a county-level sports competition." In this case, the dialogue system 1 transitions the state to state 1217.
[0201] As another example, when transitioning from state 1204 to state 1209, the dialog agent says “How often do you do exercise A?” In state 1210, assume that the user 11 says “I do exercise A more than three times a week.”
[0202] In this case, as the third stage, the dialogue system 1 causes the dialogue agent to speak further into the topic of state 1209. For example, in state 1211, the dialogue agent says, "What do you like about Sport A?" In state 1212, assume that the user 11 says, "I can play on a team." In this case, the dialogue system 1 causes the state to transition to state 1217.
[0203] As another example, when transitioning from state 1204 to state 1213, the dialogue agent says, "Do you like sport A?" In state 1214, assume that the user 11 says, "Yes, I do." In this case, the dialogue system 1 transitions the state to state 1211.
[0204] As another example, when transitioning from state 1204 to state 1215, the dialogue agent asks, "Have you watched Sport A?" In state 1216, assume that the user 11 says, "Yes, I have." In this case, as an example, the dialogue system 1 transitions the state to state 1217. As described above, the dialogue system 1 can also omit the third stage of drilling down.
[0205] When transitioning to state 1217, assume that the dialog agent says, "Thank you for telling me. I see you're enjoying sport A. That's great." In state 1218, assume that the user 11 says, for example, "You're welcome."
[0206] In this case, the dialogue system 1 may, for example, end the dialogue or further change the state to Figure 13 Status 1301.
[0207] When transitioning to state 1301, the dialog agent says, for example, “Do you have a favorite team?” In state 1302, assume that user 11 says, “I like team C.”
[0208] In this case, as the fourth stage, the dialogue system 1 causes the dialogue agent to conduct an in-depth discussion about the user's favorite team (or election) in sport A. For example, in state 1303, the dialogue agent says, "What do you like about Team C?" In state 1304, assume that the user 11 says, for example, "I like Team C because they're strong." In this case, the dialogue system 1 causes the dialogue agent to conduct a closing greeting. For example, in state 1305, the dialogue agent says, "I understand. Thank you for telling me. Thank you for taking the time to share this with us. End of conversation."
[0209] pass Figure 12 、 Figure 13 In the migration, the dialogue system 1 can use the dialogue scene of the recall method, and as the dialogue proceeds, in order to specifically deepen the dialogue, the dialogue agent has multiple "1H dialogues".
[0210] Third embodiment
[0211] For example, in Figure 3 The illustrated conversation screen 300 includes not only voice-based conversations and actions of the virtual person 301 but also auxiliary visual information, thereby facilitating in-depth conversations during business negotiations and nursing care.
[0212] Figure 14 FIG. 1 is a diagram showing an example of a dialogue screen according to the third embodiment. Figure 14 In the example shown in FIG. 14 , in addition to a virtual person (dialogue agent) 1401 and a text-based dialogue 302, an illustration 1403, which is an image generated based on the content of the dialogue, is displayed on dialogue screen 1400. This illustration 1403 allows user 11 to easily visualize cross-country skiing, which is the content of the dialogue. Furthermore, illustration 1403 may also include audio information, such as sound effects or sounds different from the content of the dialogue.
[0213] Figure 15 1 is a diagram showing an example of a functional configuration of a dialogue system according to a third embodiment. Figure 15 As shown, the server device 100 involved in the third embodiment has Figure 7 In addition to the functional configuration of the server device 100 described in , the server device 100 further includes an image generation unit 1501.
[0214] Image generation unit 1501, for example, included in generation unit 704, performs image generation processing to generate illustration 1403, which is an image generated based on the content of the conversation with user 11. For example, image generation unit 1501 can generate illustration 1403 using a learned machine learning model (e.g., DALL·E, DALL·E2, or StableDiffusion) that generates images based on text information. Alternatively, image generation unit 1501 can generate illustration 1403, an image related to the content of the conversation, based on at least one of user 11's verbal information and non-verbal information.
[0215] For example, image generation section 1501 may generate illustration 1403 when user 11's emotion analysis determines that user 11 is "positive" based on verbal information such as "cross-country skiing" spoken by user 11 and non-verbal information such as the "high pitch" of user 11's voice. This allows dialogue system 1 to further stimulate user 11's recall and facilitate effective dialogue.
[0216] Each function configuration other than the image generation unit 1501 can be configured with Figure 7 The functional configuration of the dialogue system 1 according to an embodiment of the present disclosure is the same as that described in .
[0217] Figure 16 is a flowchart showing an example of a dialogue process according to the third embodiment. Figure 15 The functional configuration shown is an example of the dialogue process executed by the dialogue system 1.
[0218] In step S1601, the first acquisition unit 702 acquires the speech of the user 11. In addition, in step S1602, the first acquisition unit 702 performs a speech recognition process on the acquired speech of the user 11. Thus, the first acquisition unit 702 outputs the language information of the user 11 by converting the speech of the user 11 into text. In addition, the processing of steps S1601 and S1602 can also be used, for example, Figure 8 The processing of step S801.
[0219] In step S1603, the image generation unit 1501 extracts a summary or keywords etc. from the speech voice of the user 11. Furthermore, in step S1604, the image generation unit 1501 generates, for example, Figure 14 Illustration 1403 and other images described in.
[0220] In step S1604, the generation unit 704 generates a voice that causes the dialogue agent to speak. Figure 8The processing of steps S1604 and S1605 is performed in the same or substantially the same manner as steps S803 and S804. The generation unit 704 may also generate the image in the image generation unit 1501. Figure 14 In the case of the cross-country skiing illustration 1403 shown, a voice is generated to cause the dialogue agent to speak about cross-country skiing.
[0221] In step S1606, generation section 704 outputs the image generated by image generation section 1501 and the sound generated by generation section 704 to dialogue screen 1400. At this time, dialogue system 1 may cause virtual person 1401 to perform actions to assist displayed illustration 1403 (e.g., pointing with a finger).
[0222] pass Figure 16 The dialog system 1 can, for example, Figure 4 As shown, an illustration 1403 that is an image related to the content of the conversation is displayed on the conversation screen 1400.
[0223] Fourth embodiment
[0224] Figure 17 1 is a diagram showing an example of the functional configuration of the dialogue system of the fourth embodiment. Figure 17 As shown, the server device 100 of the fourth embodiment has Figure 7 In addition to the functional configuration of the server device 100 described in , the server device 100 further includes a summarizing unit 1701.
[0225] The summarizing unit 1701 is included in the generating unit 704, for example, and the conversation control unit 705 performs summary processing such as creating a report or the like by summarizing the conversation record stored in the storage unit 710.
[0226] The dialogue control unit 705 of the dialogue system 1 generates, for example, Figure 18A and Figure 18B The conversation record 1800 shown is stored in the storage unit 710 .
[0227] exist Figure 18A and Figure 18B In the example, conversation log 1800 includes information such as "time stamp," "speaker," "speech text," and "document name" as items. "Time stamp" indicates the date and time when user 11 or the conversational agent made a speech. "Speaker" indicates whether the speech in "speech text" was made by the user or the conversational agent. "Speech text" is a textual representation of the speech made by user 11 or the conversational agent. "Document name" indicates the file name of the speech recorded by user 11.
[0228] like Figure 18A and Figure 18B As shown, the conversation record 1800 is a complete record of the conversation between the user 11 and the conversation agent, and therefore, it is desirable to summarize it, for example, if it is submitted as a report.
[0229] The summarizing unit 1701 may summarize the conversation record 1800 by applying a large-scale language model, or may summarize the conversation record 1800 by using a cloud service disclosed as article summarization artificial intelligence (AI).
[0230] Important information summarized includes, for example, when, where, who, what, why, and how (5W1H) information, such as date and time, location, user information (attributes, new customers, existing customers, etc.), user challenges or needs, product or service proposal information, action items, and next schedule information. Summarizing unit 1701 summarizes conversation log 1800 and generates a conversation report or meeting record containing this information.
[0231] Summarizing unit 1701 may also determine that user 11 is interested in the product presented by the dialogue agent based on verbal information such as "yes" uttered by user 11, non-verbal information such as the "high pitch" of user 11's voice, and the "bright expression" of user 11. In this case, summarizing unit 1701 preferably creates the summary text so that the description of the products and services does not include any omissions.
[0232] Fifth embodiment
[0233] Figure 19 1 is a diagram showing an example of the functional configuration of the dialogue system of the fifth embodiment. Figure 19 As shown, the server device 100 according to the fifth embodiment has Figure 7 In addition to the functional configuration of the server device 100 described in , it further includes a sales copy generation unit 1901.
[0234] The sales copy generating unit 1901 is for example Figure 9 In the dialogue scene 915 corresponding to the fifth stage shown, product recommendations are made and a sales copy generation process is executed to generate a sales copy presented to the user 11. The sales copy is an advertisement or promotional text that attracts people's attention. In this case, it is a character string used to promote the product proposed to the user 11.
[0235] As an example, assume that the summary of products proposed by the conversational agent to the user is a demand analysis service with the following content.
[0236] This technology supports the quality of responses and reduces response times at support centers and call centers in sectors such as retail / wholesale, food and beverage, manufacturing, information and communications, services, pharmaceuticals, cosmetics, and tourism. Furthermore, contextual analysis of the vast number of inquiries from customers contributes to the development of sales promotion measures and new products and services.
[0237] However, these sentences make it difficult for the user 11 to understand the features of goods and services provided to the user 11. Therefore, the sales copy generation unit 1901 can generate, for example, the following sales copy.
[0238] "Support from customer engagement to action planning! AI thoroughly analyzes customers."
[0239] Alternatively, the sales copy generating unit 1901 may generate, for example, the following sales copy.
[0240] "AI learns and analyzes accumulated customer feedback! It provides timely guidance on optimal solutions."
[0241] As another example, assume that the summary of goods and services suggested to user 11 by the conversational agent is a sales support service having the following content.
[0242] "Customer experience and sales know-how are accumulated by individuals and not shared within the team. This leads to inefficiencies such as the time required to search through disparate customer data during handovers. Sharing information about our products and services before sales reduces time-consuming searches, which are often a personalized process. For example, sharing reference information such as proposals for similar cases created by experienced individuals can help eliminate the problem of data creation that can be biased by skill, and develop documents that lead to successful business negotiations."
[0243] However, these sentences make it difficult for the user 11 to understand the features of goods and services. Therefore, the sales copy generation unit 1901 can generate, for example, the following sales copy.
[0244] "Instantly capture customer interest! AI supports successful business negotiations."
[0245] Alternatively, the sales copy generating unit 1901 may generate, for example, the following sales copy.
[0246] "AI learning relies on the business style of personalization process! AI recommends proposal documents based on the customer's interests."
[0247] Such a sales copy can be efficiently generated by using, for example, a large-scale language model.The sales copy generating unit 1901 can generate a sales copy using a sales copy generating service provided by an external cloud service.
[0248] Figure 20 This is a flowchart showing an example of the sales copy processing of the fifth embodiment. Figure 9 The illustrated dialogue scene 915 corresponding to the fifth stage is an example of a process of generating a sales copy corresponding to the product proposed to the user 11 .
[0249] In step S2001, Figure 9 The product and service recommendation unit 918 is based on Figure 10AA The content of the conversation conducted in steps S1003 to S1004 determines the products and services proposed to the user 11.
[0250] In step S2002, Figure 19 The sales copy generating unit 1901 obtains the information of the determined goods and services from the storage unit 710 .
[0251] In step S2003, the sales copy generation unit 1901 uses the acquired product and service information to generate sales copies for the products and services determined by the product and service recommendation unit 918. For example, the sales copy generation unit 1901 may utilize a sales copy generation service provided by an external cloud service to generate sales copies. For another example, the sales copy generation unit 1901 may generate sales copies using a large-scale language model.
[0252] In step S2004, the dialogue system 1 prompts the user 11 with the goods and services proposed to the user 11 and the sales copy of the goods and services. Figure 2 The display 202 of the illustrated dialogue screen 200 displays information on recommended products and services and a sales copy of the products and services.
[0253] Figure 20 The processing shown is merely an example. For example, the goods and services offered to user 11 may be a package of goods and services obtained by combining multiple goods and services. In this case, sales copy generation unit 1901 obtains information on the multiple goods and services in step S2002 and generates a sales copy using the information on the multiple goods and services in step S2003.
[0254] According to the fifth embodiment, the dialogue system 1 can directly convey the value of products and services to the user 11 in an easily understandable manner.
[0255] Sixth embodiment
[0256] Figure 21 1 is a diagram showing an example of a functional configuration of a dialogue system according to a sixth embodiment of the present disclosure. Figure 21 As shown, the server device 100 according to the sixth embodiment has Figure 7In addition to the functional configuration of the server device 100 described in the specification, the storage unit 710 also has (stores) a past history DB (Database) 2101 and input and output information of non-language information (hereinafter referred to as input and output information) 2102.
[0257] The past history DB 2101 is a database storing, for example, past conversation records, non-verbal information, and physical condition information of the user 11 .
[0258] In the input / output information 2102, for example, Figure 22 As shown in FIG. 1 , the non-verbal information obtained (input) from the image and voice of the user 11 is included to determine whether it is positive or negative. In the input / output information 2102, for example, Figure 22 As shown, it includes information on whether the non-verbal information represented by the image and voice of the dialogue agent is positive or negative.
[0259] Thus, intention interpretation unit 706 can easily determine whether the non-verbal information contained in the image and voice of user 11 is positive or negative using input and output information 2102. Response generation unit 707 can use input and output information 2102 to obtain examples of positive or negative non-verbal information from the dialogue agent.
[0260] When acquiring non-verbal information about user 11 from a conversation with user 11 utilizing terminal device 10, second acquisition unit 703 of the sixth embodiment acquires non-verbal information (emotional) and non-verbal information (personality). Non-verbal information (emotional) includes, for example, user 11's emotions, attitude, language (intensity, speed, or intonation), physiological characteristics, or body movements (eye contact, facial expressions), which can change over time. For example, intention interpretation unit 706 can determine whether user 11 is positive or negative based on the non-verbal information (emotional) acquired by second acquisition unit 703.
[0261] On the other hand, non-verbal information (personality) includes, for example, non-verbal information (attribute information) such as user 11's gender, age, physical features, or body shape, which does not change or changes only slightly depending on the circumstances. For example, response generation unit 707 can generate a verbal or non-verbal response corresponding to the user's attributes (e.g., gender, age, or body shape) based on the non-verbal information (personality) acquired by second acquisition unit 703. Non-verbal information (personality) is an example of non-verbal information representing attributes of user 11.
[0262] Other functional configurations of the dialogue system 1 according to the sixth embodiment can be Figure 7 The functional configuration of the dialogue system 1 described in is the same.
[0263] Figure 23 1 is a flowchart showing an example of a dialogue process according to the sixth embodiment of the present disclosure. This process shows that after starting a dialogue between the user 11 and the dialogue agent, Figure 21 The following description omits the processing performed by the dialog system 1. Figure 8 A detailed description of processing contents that are the same as or substantially the same as the overview of the dialogue processing involved in one embodiment described in .
[0264] In step S2301 , the first acquiring unit 702 acquires language information of the user 11 from the conversation between the user 11 and the dialogue agent.
[0265] In steps S2302 and S2303 , in parallel with the processing of step S2301 , the second acquiring unit 703 acquires the non-verbal information (emotional information) and non-verbal information (personality information) of the user 11 based on the conversation between the user 11 and the dialogue agent.
[0266] In step S2304 , the generation unit 704 interprets the intention of the user 11 's utterance based on the language information acquired by the first acquisition unit 702 and the non-language information (emotion) acquired by the second acquisition unit 703 .
[0267] In step S2305, generation unit 704 refers to the non-verbal information (personality) acquired by second acquisition unit 703 or past history DB 2101 to generate a language response (dialogue sentence) corresponding to the intended speech of user 11. For example, generation unit 704 determines user 11's gender, interests, body type, etc. based on past conversation history with user 11 in past history DB 2101, and generates a different language response (dialogue sentence) based on user 11's gender, interests, body type, etc.
[0268] If there is no past history of user 11, generation unit 704 may detect the facial area from the image of user 11 and use age / gender estimation AI (artificial intelligence) to infer user 11's gender or age, for example. Generation unit 704 may also use body estimation AI to infer user 11's body shape based on the image of user 11. Generation unit 704 may also determine user 11's interests, etc. based on user 11's language information. The generation unit stores the inferred gender, age, body shape, etc. of user 11 in past history DB 2101.
[0269] As a specific example, assume that during a business negotiation, generation unit 704 determines based on user 11's verbal and non-verbal information that user 11 is a woman in her 40s who is interested in cosmetics. In this case, generation unit 704 can determine that cosmetics and services for people in their 40s are worth introducing or recommending, and, for example, generate a verbal response introducing the specific products and services.
[0270] As another example, the generation unit 704 may estimate the body shape of the user 11 based on an image of the user 11 during a business negotiation, compare it with the user 11's past body shape history, and perform a transition or comparison of the user 11's body shape with past body shapes. Thus, the generation unit 704 may generate a verbal response, for example, for the user 11 who has recently gained weight, offering products or services such as low-sugar ingredients or weight management applications.
[0271] As another example, generation unit 704 may estimate user 11's clothing fashion sense based on an image of user 11 during a business negotiation and compare it with user 11's past clothing fashion sense. Thus, generation unit 704 may generate a verbal response introducing specific products and services to user 11 whose clothing-related products and services are determined to be worthy of priority introduction.
[0272] As another example, generation unit 704 may estimate the body shape of user 11 based on an image of user 11 during a business negotiation, and determine whether a physical examination of user 11 is necessary based on past medical history information. Thus, generation unit 704 may generate a verbal response confirming the current physical condition of user 11 determined to require a physical examination.
[0273] In step S2306, the generation unit 704 determines the paralanguage of the dialogue agent (e.g., voice pitch, speaking speed, voice pitch, voice intensity, cough, sigh, laugh, or silence, etc.) based on the generated language response and the non-language information of the user 11. For example, the generation unit 704 may also refer to Figure 22 If the user 11's emotion analysis is determined to be positive based on the input / output information 2102 shown in FIG. 2 , the positive non-verbal information (paralanguage) of the dialogue agent is obtained from the input / output information 2102. Similarly, the generation unit 704 may also refer to the Figure 22 When it is determined that the emotion analysis of the user 11 is negative based on the input / output information 2102 shown, the negative non-verbal information (paralanguage) of the dialogue agent is acquired from the input / output information 2102 .
[0274] Figure 22The input / output information 2102 shown is an example. In the input / output information 200, various positive and negative non-verbal information of the user 11 and positive and negative non-verbal information of the dialogue agent are registered in advance.
[0275] In step S2307 , the control unit 714 synthesizes the response voice of the dialogue agent based on the language response generated by the generation unit 704 and the sublanguage determined by the generation unit 704 .
[0276] The server device 100 executes the processes of steps S2308 and S2309 in parallel with the processes of steps S2306 and S2307 .
[0277] In step S2308, the generation unit 704 determines the facial expression, sight, and gesture of the dialogue agent based on the non-verbal information of the user 11. For example, the generation unit 704 determines the facial expression, sight, and gesture of the dialogue agent based on the non-verbal information of the user 11. Figure 22 If the user 11's emotion analysis is determined to be positive based on the input / output information 2102 shown, the positive non-verbal information (facial expression, sight, and gesture) of the dialogue agent is obtained from the input / output information 2102. Figure 22 If it is determined that the emotion analysis of the user 11 is negative based on the input / output information 2102 shown, the negative non-verbal information (facial expression, eye gaze, and hand gesture) of the dialogue agent is acquired from the input / output information 2102 .
[0278] In step S2309 , the generation unit 704 determines the action (movement) of the dialogue agent based on the determined facial expression, line of sight, and gesture of the dialogue agent.
[0279] As a specific example, if generation unit 704 determines that user 11's emotion analysis is positive during a business negotiation, generation unit 122 can, for example, cause the dialogue agent to smile and increase its hand gestures. If generation unit 704 determines that user 11's emotion analysis is negative, it can also, for example, cause the dialogue agent to have a silent face, nod, bow, etc. In addition to positive / negative judgments, the dialogue agent can also be caused to perform actions (movements) based on non-verbal information (personality). For example, in a positive situation, the dialogue agent can be caused to perform actions corresponding to (similar to) user 11's non-verbal information (personality), such as user 11's hand gestures, the shape of crossed arms, or the speed and rhythm of the conversation recorded in past history DB 2101.
[0280] In step S2310, control unit 714 draws the interactive agent based on the action of the interactive agent determined by generation unit 704 and outputs a dialogue screen including the drawn interactive agent and the synthesized response speech. For example, output unit 713 transmits the dialogue screen to terminal device 10 using communication unit 701.
[0281] The dialogue system 1 repeatedly executes Figure 8 The processing can provide a more appropriate response to user 11 based on the non-verbal information (personality) of user 11 or past history DB2101.
[0282] Next, an example of a usage scenario of the dialogue system 1 of this embodiment will be described.
[0283] Figure 24 This is a diagram showing an example of a system configuration of a usage scenario 1 according to an embodiment of the present disclosure. Figure 1 This is an example of a case where the terminal device 10 is a signage terminal 2400 of a digital signage. Figure 24 In the example of FIG, the signage terminal 2400 includes an input device 2401 such as a camera and a microphone, and a hardware structure of a computer.
[0284] Figure 25 This is a flowchart showing an example of a session start process in the utilization scenario 1 according to an embodiment of the present disclosure.
[0285] In step S2501, the interactive system 1 detects the face of the user 11 from an image captured by the input device 2401 included in the identification terminal 2400. As a specific example, the interactive system 1 extracts the facial image of the user 11 from the image captured by the input device 2401 and performs facial authentication on the extracted facial image. If the extracted facial image passes facial authentication, the interactive system 1 determines that the face of the user 11 has been detected.
[0286] In step S2502, the dialogue system 1 determines whether face detection has continued for a predetermined time. For example, the dialogue system 1 determines whether the state of detecting the user 11's face has continued for a predetermined time (e.g., 5 seconds). If face detection has continued for the predetermined time, the dialogue system 1 moves the process to step S2503. On the other hand, if face detection has not continued for the predetermined time, the dialogue system 1 returns the process to step S2501. The processes of steps S2501 and S2502 can be performed by either the identification terminal 2400 or the server device 100.
[0287] In step S2503, server device 100 determines whether there is a past history of user 11. For example, server device 100 refers to past history DB 2101 and, if past conversation records of user 11 exist, determines that there is a past history of user 11. If there is a past history, server device 100 proceeds to step S2504. On the other hand, if there is no past history, server device 100 proceeds to step S2505.
[0288] In step S2504, the server device 100 determines a scenario to be used in the dialogue process based on the past history (past dialogue records, etc.) of the user 11. This allows the dialogue system 1 to prevent the same question or speech from being repeated multiple times to the same user 11.
[0289] In step S2505, the server device 100 selects a predetermined scenario (for example, a scenario for a new customer) as a scenario to be used in the dialogue process.
[0290] In step S2506, the dialogue system 1 performs, for example, Figures 1 to 23 The dialog process described in . Figure 25 Through this process, interactive system 1 can provide interactive services to user 11 using signage terminal 2500. Interactive system 1 can modify the content of the conversation provided to user 11 based on user 11's past conversation history, etc. The processes of steps S2703 to S2705 are optional and not required. For example, interactive system 1 may also determine the scenario for the conversation during the conversation process of step S2506.
[0291] Figure 26 This is a diagram showing an example of a system configuration for a usage scenario 2 according to an embodiment of the present disclosure. Figure 1 The terminal device 10 in FIG. 2 is an example of a display terminal 2600 for a metaverse. The display terminal 2600 includes a metaverse display such as a head-mounted display or a space reproduction display, and a computer. The dialogue system 1 provides a dialogue service to the user 11 using a dialogue agent in a virtual space.
[0292] Figure 27 This is a flowchart showing an example of a session start process in the utilization scenario 2 according to an embodiment of the present disclosure.
[0293] In step S2701, the dialogue system 1 detects the approach of the avatar of the user 11 in the virtual space. For example, the dialogue system 1 detects whether the avatar of the user 11 is within a predetermined range (e.g., within 1 meter) based on the user 11's login information, the coordinates of the avatar of the user 11 in the virtual space, and the coordinates of the dialogue agent.
[0294] In step S2702, the interactive system 1 determines whether the avatar of user 11 has been within a predetermined range (e.g., within 1 meter) for a predetermined time (e.g., 5 seconds). If the avatar of user 11 has been within the predetermined range for a predetermined time, the interactive system 1 proceeds to step S2703. On the other hand, if the avatar of user 11 has not been within the predetermined range for a predetermined time, the interactive system 1 returns to step S2701.
[0295] In step S2703, server device 100 determines whether there is a past history record for user 11. For example, server device 100 refers to past history record DB 2101 and, if past conversation records for user 11 exist, determines that there is a past history record for user 11. If there is a past history record, server device 100 proceeds to step S2704. On the other hand, if there is no past history record, server device 100 proceeds to step S2705.
[0296] In step S2704, the server device 100 determines a scenario to be used in the conversation process based on the past history (past conversation records) of the user 11. In step S2705, the server device 100 selects a predetermined scenario (e.g., a new user scenario) as the scenario to be used in the conversation process.
[0297] In step S2706, the dialogue system 1 performs a reference operation on the virtual space, for example Figures 1 to 23 The dialog process described in . Figure 27 By processing, the dialogue system 1 can use the display terminal 2600 used in the metaverse to provide the user 11 with a dialogue service in the virtual space.
[0298] Figure 28 This diagram illustrates an example system configuration for Utilization Scenario 3 according to an embodiment of the present disclosure. Utilization Scenario 3 illustrates an example of a situation in which user 11, using terminal device 10, conducts a web conference with a conversational agent provided by server device 100. User 11 can participate in a web conference hosted by conference server 2810 outside the system, or server device 100 can host the web conference.
[0299] Figure 29 This is a flowchart showing an example of a session start process using scenario 3 according to an embodiment of the present disclosure.
[0300] In step S2901, user 11 uses terminal device 10 to participate in a web conference with a dialogue agent provided by dialogue system 1. For example, user 11 participates in the web conference by using terminal device 10 to access a link for participating in the web conference with the dialogue agent.
[0301] In step S2902, interactive system 1 determines whether a conversation start operation from user 11 has been accepted during the online conference. If the conversation start operation from user 11 has been accepted, interactive system 1 proceeds to step S2903. On the other hand, if the conversation start operation from user 11 has not been accepted, interactive system 1, for example, repeatedly executes the process of step S2902.
[0302] In step S2903, server device 100 determines whether there is a past history record for user 11. For example, server device 100 refers to past history record DB 2101 and, if past conversation records for user 11 exist, determines that there is a past history record for user 11. If there is a past history record, server device 100 proceeds to step S2904. On the other hand, if there is no past history record, server device 100 proceeds to step S2905.
[0303] In step S2904, the server device 100 determines a scenario to be used in the conversation process based on the past history (past conversation records) of the user 11. In step S2905, the server device 100 selects a predetermined scenario (e.g., a new user scenario) as the scenario to be used in the conversation process.
[0304] In step S2906, the dialogue system 1 performs, for example, Figures 1 to 23 The dialog process described in . Figure 29 By processing, the dialogue system 1 can provide dialogue services to the user 11 using network conferences.
[0305] As described above, according to the embodiments of the present invention, in the dialogue system 1 that conducts a dialogue with the user 11 using the dialogue agent, a more appropriate response to the user 11 can be made.
[0306] The functions of the above-described embodiments can be implemented by one or more processing circuits or circuit systems. In the embodiments of the present disclosure, "processing circuit" includes processors that execute various functions in a software manner, such as processors installed in electronic circuits, as well as devices such as ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), FPGAs (Field Programmable Gate Arrays), and existing circuit modules designed to perform the functions described above.
[0307] The groups of devices or apparatuses described in the above embodiments represent only one of multiple computing environments for implementing the embodiments of the present disclosure. In some embodiments, the server device 100 includes multiple computing devices, such as a server cluster. The multiple computing devices are configured to communicate with each other via any type of communication link (including a network, shared memory, etc.) to perform the processing of the present disclosure.
[0308] The various functional configurations of the server device 100 may be integrated into one server device or divided into multiple devices. At least a portion of the various functional configurations of the server device 100 may also be included in the terminal device 10 .
[0309] The following describes some aspects of the present disclosure. In this specification, a dialog system, a dialog control method, and a program according to several examples are disclosed.
[0310] Aspect 1
[0311] A dialogue system uses a dialogue agent to conduct a dialogue with a user. The dialogue system includes a first acquisition unit, a second acquisition unit, a generation unit, and a control unit. The first acquisition unit acquires the user's language information from the dialogue. The second acquisition unit acquires the user's non-verbal information from the dialogue. The generation unit generates response content including a language response and a non-verbal response of the dialogue agent based on the user's language information and the user's non-verbal information. The control unit controls the dialogue agent based on the response content generated by the generation unit.
[0312] Aspect 2
[0313] In the dialogue system according to aspect 1, the response content of the dialogue agent includes the non-verbal response of the dialogue agent. The generating unit is configured to generate the non-verbal response of the dialogue agent based on the non-verbal information of the user.
[0314] Aspect 3
[0315] In the dialogue system according to aspect 2, the generation unit is configured to change the content of the action of the dialogue agent according to the non-verbal information of the user.
[0316] Aspect 4
[0317] In the dialogue system according to aspect 2 or 3, the generation unit is configured to change the timing of the action of the dialogue agent according to the non-verbal information of the user.
[0318] Aspect 5
[0319] In the dialogue system according to any one of aspects 1 to 4, the non-verbal information of the user includes information on facial expression, sight, posture, or emotion, each of which is acquired from an image of the user.
[0320] Aspect 6
[0321] In the dialogue system according to any one of aspects 1 to 5, the non-verbal information of the user includes information on voice volume, voice intonation, or voice timbre obtained from the user's voice, respectively.
[0322] Aspect 7
[0323] In the dialogue system according to any one of aspects 1 to 6, the generation unit is configured to change the response content of the dialogue agent according to a scenario of the dialogue.
[0324] Aspect 8
[0325] In the dialogue system according to any one of aspects 1 to 7, the generation unit is configured to change the response content of the dialogue agent according to a plurality of preset dialogue stages.
[0326] Aspect 9
[0327] In the dialogue system according to aspect 8, the generation unit is configured to change the dialogue phase according to the line of sight information of the user.
[0328] Aspect 10
[0329] The dialogue system according to any one of aspects 1 to 9 further includes an image generation unit configured to generate an image related to the dialogue content based on at least one of the user's language information or the user's non-language information. The control unit also uses the image to conduct a dialogue with the user in addition to the dialogue agent.
[0330] Aspect 11
[0331] The dialog system according to any one of aspects 1 to 10 further includes a summarizing unit configured to summarize the dialog based on the dialog record of the dialog.
[0332] Aspect 12
[0333] In the dialogue system according to any one of aspects 1 to 11, the dialogue is a business negotiation with the user, and the control unit is configured to propose products and services based on the dialogue content of the business negotiation.
[0334] Aspect 13
[0335] In the dialogue system according to aspect 12, the control unit is configured to present sales copies of goods and services based on the dialogue content of the business negotiation.
[0336] Aspect 14
[0337] The dialogue system according to any one of aspects 1 to 13 further comprises a database configured to store past histories of the dialogues, and the generation unit configured to change the scenario of the dialogue based on the past histories of the dialogues.
[0338] Aspect 15
[0339] The dialogue system according to any one of aspects 1 to 14 further comprises a database configured to store past records of the dialogue, and the generation unit is configured to generate the language response of the dialogue agent with reference to the past records of the dialogue.
[0340] Aspect 16
[0341] In the dialogue system according to any one of aspects 1 to 15, the second acquisition unit is configured to acquire the non-verbal information indicating user attributes from the dialogue, and the generation unit is configured to generate the verbal response or the non-verbal response according to the user attributes.
[0342] Aspect 17
[0343] A conversation control method is performed by a conversation system that uses a conversation agent to conduct a conversation with a user. The conversation control method includes: obtaining the user's language information from the conversation; obtaining the user's non-verbal information from the conversation; generating response content including a verbal response and a non-verbal response from the conversation agent based on the user's language information and the user's non-verbal information; and controlling the conversation agent based on the generated response content.
[0344] Aspect 18
[0345] A program executed by a computer, the program being processed by a dialogue system that uses a dialogue agent to conduct a dialogue with a user, the processing comprising: obtaining language information of the user from the dialogue; obtaining non-verbal information of the user from the dialogue; generating response content including a language response and a non-verbal response of the dialogue agent based on the language information of the user and the non-verbal information of the user; and controlling the dialogue agent based on the generated response content.
[0346] The above embodiments are illustrative and do not limit the present invention. Therefore, many additional modifications and variations are possible based on the above teachings. For example, elements and / or features of different illustrative embodiments may be combined and / or replaced with each other within the scope of the present invention. Any of the above operations may be performed in various other ways, such as in an order different from that described above.
[0347] The above embodiments are illustrative and do not limit the present invention. Therefore, many additional modifications and variations are possible based on the above teachings. For example, elements and / or features of different illustrative embodiments may be combined with each other and / or replaced with each other within the scope of the present invention. Any of the above operations may be performed in various other ways, such as in an order different from that described above.
[0348] The present disclosure can be implemented in any convenient form, such as using dedicated hardware or a mixture of dedicated hardware and software. The present disclosure can be implemented as computer software applied by one or more networked processing devices. The processing facilities can be applied to any appropriately programmed device, such as a general-purpose computer, a personal digital assistant, a mobile phone (such as a WAP or 3G compatible phone), etc. Because the present disclosure can be applied as software, each aspect of the present disclosure includes computer software that can be executed on a programmable device. The computer software can be provided to a programmable device using any conventional carrier medium (carrier method). The carrier medium can accommodate transient carrier methods, such as electrical, optical, microwave, acoustic or radio frequency signals carrying computer code. An example of such a transient method is a TCP / IP signal that carries computer code on an IP network (such as the Internet). The carrier medium can also include a storage medium for storing processor-readable code, such as a floppy disk, a hard disk, a CD ROM, a magnetic tape device or a solid-state storage device.
[0349] The functions of the embodiments of the present disclosure can be implemented using circuits or processing circuits, which include general-purpose processors, special-purpose processors, integrated circuits, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), conventional circuits, and / or combinations thereof, configured or programmed to perform the disclosed functions. A processor is a processing circuit or circuits that include transistors and other circuits. In the present disclosure, a circuit, unit, or device is hardware that performs or is programmed to perform the functions described. The hardware can be any hardware known in the present disclosure or otherwise, programmed or configured to perform the functions described. When the hardware is a processor of a type of circuit, the circuit, device, or unit is a combination of hardware and software, with the software being used to configure the hardware and / or processor.
[0350] This patent application is based upon and claims the benefit of priority from Japanese Patent Application No. 2023-017067 filed with the Japan Patent Office on February 7, 2023, and Japanese Patent Application No. 2023-221852 filed with the Japan Patent Office on December 27, 2023, the disclosures of which are incorporated herein by reference in their entirety.
[0351] Reference Signs List
[0352] 1: Dialogue system
[0353] 10: Terminal device
[0354] 100: Server device
[0355] 200, 300: Dialogue screen
[0356] 201, 301, 1401: Virtual Human (Dialogue Agent)
[0357] 500: Computer
[0358] 702: First acquisition unit
[0359] 703: Second acquisition unit
[0360] 704:Generation Unit
[0361] 714: Control Unit
[0362] 1501: Image generation unit
[0363] 1701: Summary Unit
[0364] 1901: Sales copy generation unit
Claims
1. A dialogue system for conducting a dialogue with a user using a dialogue agent, the dialogue system comprising: a first acquiring unit, configured to acquire the language information of the user from the conversation; a second acquiring unit, configured to acquire non-verbal information of the user from the conversation; a generating unit configured to generate response content including a language response and a non-language response of the dialogue agent based on the language information of the user and the non-language information of the user; as well as A control unit is configured to control the conversation agent based on the response content generated by the generation unit.
2. The dialogue system according to claim 1, in, The response content of the dialogue agent includes the non-verbal response of the dialogue agent, and The generating unit is configured to generate the non-verbal response of the dialogue agent according to the non-verbal information of the user.
3. The dialogue system according to claim 2, in, The generating unit is configured to change the content of the action of the dialogue agent according to the non-verbal information of the user.
4. The dialogue system according to claim 2, in, The generating unit is configured to change the timing of the action of the conversational agent according to the non-verbal information of the user.
5. The dialogue system according to any one of claims 1 to 4, in, The non-verbal information of the user includes information about facial expressions, sight lines, postures, or emotions, each of which is acquired from an image of the user.
6. The dialogue system according to claim 5, in, The non-verbal information of the user includes information on voice volume, voice intonation, or voice timbre respectively obtained from the voice of the user.
7. The dialogue system according to claim 1, in, The generating unit is configured to change the response content of the dialogue agent according to a scenario of the dialogue.
8. The dialogue system according to claim 1, in, The generating unit is configured to change the response content of the dialogue agent according to a plurality of preset dialogue stages.
9. The dialogue system according to claim 8, in, The generating unit is configured to change the dialogue phase according to the sight line information of the user.
10. The dialogue system according to claim 1, further comprising an image generating unit configured to generate an image related to the conversation content based on at least one of the language information of the user or the non-language information of the user, in, The control unit performs a conversation with the user using the image in addition to the conversation agent.
11. The dialogue system according to claim 1, It further includes a summarizing unit configured to summarize the conversation based on the conversation record of the conversation.
12. The dialogue system according to claim 1, in, The conversation is a business negotiation with the user, and The control unit is configured to propose goods and services based on the conversation content of the business negotiation.
13. The dialogue system according to claim 12, in, The control unit is configured to present sales copies of goods and services based on the conversation content of the business negotiation.
14. The dialogue system according to claim 7, further comprising a database configured to store past histories of said conversations, and in, The generation unit is configured to change the scenario of the conversation based on the past history of the conversation.
15. The dialogue system according to claim 1, further comprising a database configured to store past history of said conversations, and in, The generating unit is configured to generate the language response of the dialogue agent with reference to the past history of the dialogue.
16. The dialogue system according to claim 1, in, The second acquiring unit is configured to acquire the non-language information representing user attributes from the conversation, and The generating unit is configured to generate the language response or the non-language response according to the user attributes.
17. A dialog control method, performed by a dialog system that uses a dialog agent to conduct a dialog with a user, the dialog control method comprising: Acquiring language information of the user from the conversation; obtaining non-verbal information of the user from the conversation; generating response content including a language response and a non-language response of the dialogue agent based on the user's language information and the user's non-language information; as well as The conversational agent is controlled based on the generated response content.
18. A storage medium storing computer-readable program code, which, when executed by a dialog system for conducting a dialog with a user using a dialog agent, causes the dialog system to: Acquiring language information of the user from the conversation; obtaining non-verbal information of the user from the conversation; generating, based on the user's language information and the user's non-language information, response content including a language response and a non-language response of the dialogue agent; as well as The conversational agent is controlled based on the generated response content.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP2022093479A
gaming machines
JP2023017067A