Dialogue system, dialogue method and dialogue program
The interactive system addresses the limitations of conventional dialogue systems by using intention estimation and product selection models to enable accurate product selection and dialogue, improving the reliability of automated procedures.
Patent Information
- Application Number
- JP2024047476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2044-03-25
AI Technical Summary
Conventional dialogue systems using language models struggle with generating erroneous information, making them unsuitable for automating procedures like reservations, and lack the ability to interactively select and order products.
An interactive system that includes an intention estimation unit, a product selection unit, and text generation units to facilitate product selection and dialogue through voice or text input, using models like GPT to infer user intentions and extract necessary information.
Enables users to select products through dialogue, ensuring accurate responses and preventing errors, thus enhancing the usability of dialogue systems for procedures like reservations.
Smart Images

Figure 2025147275000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a dialogue system, a dialogue method, and a dialogue program. [Background technology]
[0002] In recent years, dialogue systems using large language models, which are a type of natural language processing model, have been put into practical use. For example, Cited Document 1 discloses a response generation method performed in an electronic device, the response generation method including the steps of acquiring at least one utterance data, acquiring a first context corresponding to the utterance data from a context candidate set, generating one or more dialogue sets including the first context and the utterance data, receiving a second context from a user, and acquiring a response corresponding to the second context using a language model based on the one or more dialogue sets. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2023-73220 Summary of the Invention [Problem to be solved by the invention]
[0004] In the above-mentioned conventional technology, responses to users are generated using a language model. However, since language models can output erroneous information, it is difficult to use them in dialogue systems for automating procedures, such as reservation systems.
[0005] Furthermore, it was not possible to interactively select a product from among multiple products and make a reservation or order.
[0006] The present invention has been made in view of the above-mentioned problems, and has as its object to realize a dialogue system that allows a user to select a product from among a plurality of products through dialogue. [Means for solving the problem]
[0007] An interactive system according to one embodiment includes an intention estimation unit that uses an estimation model to estimate an intention from voice information or text information, a product selection unit that uses a product selection model to select a product corresponding to the voice information or text information if the intention is product selection, a first text generation unit that generates a first text based on the product selection result, and an output unit that outputs the first text. [Effects of the Invention]
[0008] According to one embodiment, it is possible to realize a dialogue system that allows a user to select a product from among a plurality of products through dialogue. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 illustrates an example of the configuration of a dialogue system 1000. [Figure 2] 1 is a diagram illustrating an example of a hardware configuration of a dialogue device 1. FIG. [Figure 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of a user terminal 2. [Figure 4] 2 is a diagram illustrating an example of a functional configuration of the dialogue device 1. FIG. [Figure 5] FIG. 10 is a diagram showing an example of procedure information 124. [Figure 6] FIG. 10 is a diagram illustrating an example of subject information. [Figure 7] FIG. 10 is a diagram showing an example of Q&A information. [Figure 8] 10 is a flowchart showing an example of processing executed by the dialogue system 1000. [Figure 9] 10 is a flowchart illustrating an example of a method for generating a first text. [Figure 10]10 is a flowchart illustrating an example of a method for generating a second text. [Figure 11] FIG. 10 is a schematic diagram showing a specific example of a dialogue. [Figure 12] FIG. 10 is a diagram showing an example of procedure information 124′. [Figure 13] 2 is a diagram illustrating an example of a functional configuration of the dialogue device 1. FIG. [Figure 14] FIG. 12 is a diagram showing an example of product information 1202. [Figure 15] FIG. 10 is a diagram showing an example of procedure information 124. [Figure 16] 10 is a flowchart showing an example of processing executed by the dialogue system 1000. [Figure 17] 10 is a flowchart illustrating an example of a method for generating a first text. [Figure 18] FIG. 10 is a schematic diagram showing a specific example of a dialogue. [Figure 19] FIG. 10 is a diagram showing an example of procedure information 124′. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, each embodiment of the present invention will be described with reference to the accompanying drawings. Note that, in the description of the specification and drawings relating to each embodiment, components having substantially the same functional configuration are designated by the same reference numerals, and redundant description will be omitted.
[0011] [First embodiment] <System configuration> First, an overview of the dialogue system 1000 according to this embodiment will be described. The dialogue system 1000 is a system that automatically conducts dialogue with a user by utilizing a natural language processing model. The dialogue system 1000 can be used as a dialogue system for a user to execute a procedure with a target person via text chat, voice chat, or call (telephone). Chat may be text input or voice input.
[0012] The user is a user of the dialogue system 1000 and a person who performs a procedure.
[0013] The target is a person who accepts procedures from a user. Examples of the target include, but are not limited to, service providers (such as accommodation facilities and restaurants), product or service sellers (such as stores), public institutions, administrative agencies, or those who have been authorized by these organizations to accept procedures on their behalf.
[0014] A procedure is something that a user performs on a target, such as, but not limited to, a reservation, an order, a purchase, an application, or a change or cancellation thereof.
[0015] Fig. 1 is a diagram showing an example of the configuration of a dialogue system 1000. As shown in Fig. 1, the dialogue system 1000 includes a dialogue device 1 and a user terminal 2, which are communicably connected to each other via a network N. The network N is, for example, a wired local area network (LAN), a wireless LAN, the Internet, a public line network, a mobile data communication network, or a combination of these. In the example of Fig. 1, the dialogue system 1000 includes one dialogue device 1 and one user terminal 2, but may include multiple of each.
[0016] The interactive device 1 is an information processing device that automatically responds to input from a user. The interactive device 1 is, for example, but not limited to, a PC (Personal Computer), a smartphone, a tablet terminal, a server device, or a microcomputer. In the example of FIG. 1, the interactive device 1 is one information processing device, but may also be realized as a system consisting of multiple information processing devices connected via a network N.
[0017] The user terminal 2 is an information processing device for a user to execute a procedure. The user terminal 2 is, for example, but not limited to, a PC, a smartphone, or a tablet terminal. The user inputs information for the procedure via the user terminal 2. The user inputs information such as, for example, but not limited to, voice, text, or button operation.
[0018] <Hardware configuration of the interactive device 1> Next, we will explain the hardware configuration of the dialogue device 1. Fig. 2 is a diagram showing an example of the hardware configuration of the dialogue device 1. As shown in Fig. 2, the dialogue device 1 includes a processor 101, a memory 102, a storage 103, a communication I / F 104, an input device 105, an output device 106, and a drive device 107, which are connected to each other via a bus B1.
[0019] The processor 101 controls each component of the interactive device 1 and realizes the functions of the interactive device 1 by expanding into the memory 102 and executing various programs including an OS (Operating System) and an interactive program stored in the storage 103. The processor 101 is, for example, but not limited to, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), or a DSP (Digital Signal Processor).
[0020] The memory 102 is, for example, a read-only memory (ROM), a random access memory (RAM), or a combination thereof. The ROM is, for example, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a combination thereof. The RAM is, for example, but not limited to, a dynamic random access memory (DRAM) or a static random access memory (SRAM).
[0021] The storage 103 stores various programs and data, including an OS and an interactive program. The storage 103 is, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a storage class memory (SCM), but is not limited to these.
[0022] The communication I / F 104 is an interface for connecting the interactive device 1 to an external device via the network N and controlling communication. The communication I / F 104 is, for example, Bluetooth (registered trademark), Wi-Fi (registered trademark), ZigBee (registered trademark), or Ethernet (registered trademark), but is not limited to these.
[0023] The input device 105 is a device for inputting information to the interactive apparatus 1. The input device 105 is, for example, a mouse, a keyboard, a touch panel, a microphone, a scanner, a photographing device (camera), various sensors, or an operation button, but is not limited to these.
[0024] The output device 106 is a device for outputting information from the interactive apparatus 1. The output device 106 is, for example, a display device, a projector, a printer, a speaker, or a vibrator, but is not limited to these.
[0025] The drive device 107 is a device that reads and writes data from and to the recording medium 108. The drive device 107 is, for example, but not limited to, a magnetic disk drive, an optical disk drive, a magneto-optical disk drive, or an SD card reader. The recording medium 108 is, for example, but not limited to, a CD (Compact Disc), a DVD (Digital Versatile Disc), an FD (Floppy Disk), an MO (Magneto-Optical disk), a BD (Blu-ray (registered trademark) Disc), a USB (registered trademark) memory, or an SD card.
[0026] In this embodiment, the dialogue program may be written into memory 102 or storage 103 during the manufacturing stage of dialogue device 1, or may be provided to dialogue device 1 via network N, or may be provided to dialogue device 1 via a non-transitory computer-readable recording medium such as recording medium 108.
[0027] <Hardware configuration of user terminal 2> Next, we will explain the hardware configuration of the user terminal 2. Fig. 3 is a diagram showing an example of the hardware configuration of the user terminal 2. As shown in Fig. 3, the user terminal 2 includes a processor 201, a memory 202, a storage 203, a communication I / F 204, an input device 205, and an output device 206, which are interconnected via a bus B2.
[0028] The processor 201 controls each component of the user terminal 2 and realizes the functions of the user terminal 2 by expanding into the memory 202 and executing various programs, including the OS and interactive programs, stored in the storage 203. The processor 201 is, for example, but is not limited to, a CPU, an MPU, a GPU, an ASIC, or a DSP.
[0029] The memory 202 is, for example, a ROM, a RAM, or a combination thereof. The ROM is, for example, a PROM, an EPROM, an EEPROM, or a combination thereof. The RAM is, for example, but not limited to, a DRAM or an SRAM.
[0030] The storage 203 stores various programs and data, including the OS and interactive programs. The storage 203 is, for example, a flash memory, an HDD, an SSD, or an SCM, but is not limited to these.
[0031] The communication I / F 204 is an interface for connecting the user terminal 2 to an external device via the network N and controlling communication. The communication I / F 204 is, for example, Bluetooth (registered trademark), Wi-Fi (registered trademark), ZigBee (registered trademark), or Ethernet (registered trademark), but is not limited to these.
[0032] The input device 205 is a device for inputting information to the user terminal 2. The input device 205 is, for example, a mouse, a keyboard, a touch panel, a microphone, a scanner, a photographing device (camera), various sensors, or an operation button, but is not limited to these.
[0033] The output device 206 is a device for outputting information from the user terminal 2. The output device 206 is, for example, a display device (display), a projector, a printer, a speaker, or a vibrator, but is not limited to these.
[0034] In this embodiment, the dialogue program may be written into memory 202 or storage 203 during the manufacturing stage of the user terminal 2, or may be provided to the user terminal 2 via the network N, or may be provided to the user terminal 2 via a non-transitory computer-readable recording medium such as the recording medium 208.
[0035] <Functional configuration of the dialogue device 1> Next, a description will be given of the functional configuration of the dialogue device 1. Fig. 4 is a diagram showing an example of the functional configuration of the dialogue device 1. As shown in Fig. 4, the dialogue device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0036] The communication unit 11 is realized by the communication I / F 104. The communication unit 11 transmits and receives information to and from the user terminal 2 via the network N. The communication unit 11 receives an input from the user terminal 2 and transmits a response to the user terminal 2.
[0037] The storage unit 12 is realized by a memory 102 and a storage 103. The storage unit 12 stores voice information 121, text information 122, an estimation model 123, procedure information 124, an extraction model 125, subject information 126, Q&A information 127, a first response rule 128, a second response rule 129, and procedure content information 130.
[0038] The voice information 121 is information about the voice spoken by the user.
[0039] The text information 122 is information indicating text input by the user or information obtained by converting the voice information 121 into text.
[0040] The inference model 123 is a natural language processing model that infers a user's intention (context) from the text information 122. The inference model 123 is, for example, but not limited to, a large-scale language model. When the text information 122 is input, the inference model 123 is trained to output an intention corresponding to the input text information 122 from among one or more predetermined intentions. The inference model 123 is, for example, but not limited to, a GPT (Generative Pretrained Transformer) that is fine-tuned by learning the relationship between the text information 122 and the intention.
[0041] The intent corresponds to the purpose of the text indicated by the text information 122, i.e., the intent of the user who input the utterance or text. The intent may be, for example, but not limited to, a procedure, a question, a positive or negative response. The intent may be, for example, but not limited to, a reservation, an order, a purchase, an application, or a change or cancellation thereof. The question may be, for example, but not limited to, a question about business hours, price, deadline, number of people, availability, cancellation conditions, product type, or whether a reservation can be made.
[0042] The estimation model 123 may be a natural language processing model that estimates a user's intention from the speech information 121. In this case, when the speech information 121 is input, the estimation model 123 is trained to output an intention corresponding to the input speech information 121 from among one or more predetermined intentions. The estimation model 123 is, for example, a multimodal LLM (Large Language Model) that is fine-tuned by learning the relationship between the speech information 121 and the intention, but is not limited to this.
[0043] The procedure information 124 is information indicating slot information set for each procedure (intention). The slot information is information set in advance as information necessary to execute a procedure. The procedure information 124 may be set in advance, or may be arbitrarily set by the subject.
[0044] Fig. 5 is a diagram showing an example of the procedure information 124. The procedure information 124 in Fig. 5 includes, as information items, a "PID", a "procedure", one or more "slots", and a "confirmation flag".
[0045] "PID" is identification information that uniquely identifies a procedure. Below, a procedure whose "PID" is "Pxxx" may be referred to as procedure Pxxx.
[0046] "Procedure" is the name of the procedure (intention).
[0047] "Slot" is the name of the slot information required to execute the procedure.
[0048] The "confirmation flag" is information indicating whether or not the slot information of the procedure information 124 has been confirmed. "On" indicates that it has been confirmed, and "off" indicates that it has not been confirmed.
[0049] 5, the "procedure" of procedure P001 is "reservation," "slot 1" is "date," "slot 2" is "time," "slot 3" is "name," "slot 4" is "number of people," and the "confirmation flag" is "off." This indicates that procedure P001 is a reservation procedure, and that the "date," "time," "name," and "number of people" are set as information necessary to execute the reservation, and the procedure information 124 has not been confirmed.
[0050] It should be noted that the information included in the procedure information 124 is not limited to the above examples. The procedure information 124 may not include some of the above information, or may include information other than the above.
[0051] The extraction model 125 is a natural language processing model that extracts slot information from the text information 122. The extraction model 125 is, for example, but not limited to, a large-scale language model. When the text information 122 is input, the extraction model 125 is trained to extract slot information corresponding to the input text information 122 from one or more pieces of preset slot information. The extraction model 125 is, for example, but not limited to, a GPT that is fine-tuned by learning the relationship between the text information 122, intention, and slot information.
[0052] The subject information 126 is information about the subject. The subject information 126 is stored for each subject.
[0053] Fig. 6 is a diagram showing an example of target person information 126. Target person information 126 in Fig. 6 includes, as information items, "TID", "Name", "Address", "Telephone number", and "Business hours".
[0054] "TID" is identification information that uniquely identifies a subject. Hereinafter, a subject whose "TID" is "Txxx" may be referred to as subject Txxx.
[0055] "Name" is the name of the subject.
[0056] "Address" is the address of the subject.
[0057] "Phone number" is the telephone number of the subject.
[0058] "Business hours" is the business hours of the target person.
[0059] It should be noted that the information included in the subject information 126 is not limited to the above examples. The subject information 126 may not include some of the above information, or may include information other than the above.
[0060] The Q&A information 127 is information indicating the correspondence between question information indicating a question and answer information indicating an answer. The Q&A information 127 may be set in advance or may be arbitrarily set by the subject.
[0061] Fig. 7 is a diagram showing an example of the Q&A information 127. The Q&A information 127 in Fig. 7 includes information items such as "QID," "question," and "answer."
[0062] "QID" is identification information that uniquely identifies a question (intention). In the following, a question with "QID" "Qxxx" may be referred to as question Qxxx.
[0063] "Question" is question information indicating the name and content of the question.
[0064] The "answer" is answer information indicating text that is an answer to a question. The "answer" may be set in advance or may be arbitrarily set by the subject. The "answer" may also refer to subject information 126.
[0065] In the example of Figure 7, the "Question" for question Q001 is "Seating," and the "Answer" is "You're asking about seating, right? We have table seating and private rooms available." This indicates that question Q001 is a question about seating, and the answer corresponding to this question is set to the text "You're asking about seating, right? We have table seating and private rooms available."
[0066] Furthermore, the "Question" of question Q002 is "Business hours," and the "Answer" is "You're asking about business hours, right? Our business hours are {Business hours}." This indicates that question Q002 is a question about business hours, and the answer corresponding to this question is set to the text "You're asking about business hours, right? Our business hours are {Business hours}." Here, {Business hours} refers to the "Business hours" in the target person information 126.
[0067] Furthermore, the "Question" of question Q003 is "Lunch course" and the "Answer" is "Lunch course includes Lunch A, Lunch B, and Lunch C." This indicates that question Q003 is a question about the type of lunch course (product), and the answer corresponding to this question is set to the text "Lunch course includes Lunch A, Lunch B, and Lunch C." Here, the type of lunch course may be set in advance, or may be set by referring to product information 1202, which will be described later.
[0068] It should be noted that the information included in the Q&A information 127 is not limited to the above examples. The Q&A information 127 may not include some of the above information, or may include information other than the above.
[0069] The first response rule 128 is a rule set in advance for generating a first text. The first text is a text that responds to the previous input from the user. The first text is, for example, a text indicating a backchannel or a text that confirms the content of the user's input, but is not limited to these. The first response rule 128 will be described in detail later.
[0070] The second response rule 129 is a preset rule for generating a second text. The second text is the next text generated in response to a series of inputs from the user. The second text is, for example, but not limited to, text requesting the user to input slot information, text requesting the user to confirm slot information, or text notifying the user that the procedure has been finalized. The second response rule 129 will be described in detail later.
[0071] The procedure content information 130 is information relating to the content of a procedure executed by the interactive device 1. The procedure content information 130 includes slot information.
[0072] The control unit 13 is realized by the processor 101 reading and executing a program from the memory 102 and working in cooperation with other hardware components. The control unit 13 includes a text conversion unit 131, an intention estimation unit 132, a slot extraction unit 133, an answer acquisition unit 134, a first text generation unit 135, a second text generation unit 136, an output unit 137, a procedure unit 138, and a notification unit 139.
[0073] The text conversion unit 131 converts the voice information 121 into text information 122 by voice recognition processing.
[0074] The intention estimation unit 132 estimates the user's intention from the text information 122 using the estimation model 123 .
[0075] The slot extraction unit 133 executes so-called slot filling. That is, when the user's intention is a procedure, the slot extraction unit 133 uses the extraction model 125 to extract slot information corresponding to the procedure from the text information 122, and temporarily stores the extracted slot information in the storage unit 12 as the slot information of the procedure.
[0076] If the user's intention is to ask a question, the answer acquisition unit 134 refers to the Q&A information 127 and acquires answer information corresponding to the question.
[0077] The first text generator 135 generates the first text based on at least one of the slot information and the answer information in accordance with the first response rule 128.
[0078] The second text generator 136 generates the second text based on the slot information in accordance with the second response rule 129 .
[0079] The output unit 137 outputs the first text and the second text. Specifically, when the dialogue with the user is performed by voice, the output unit 137 synthesizes voice corresponding to the first text and the second text and outputs the obtained voice information. The output voice information is transmitted to the user terminal 2 and output from the output device 206 (speaker). Furthermore, when the dialogue with the user is performed by text, the output unit 137 generates and outputs text information corresponding to the first text and the second text. The generated text information and the output text information are transmitted to the user terminal 2 and output to the output device 206 (display device).
[0080] The procedure unit 138 executes the procedure according to the contents of the procedure information 124. For example, if the procedure is a reservation, the procedure unit 138 executes the reservation according to the contents of the procedure information 124. Here, it is assumed that the procedure unit 138 functions as a procedure system (e.g., a reservation system) that executes the procedure, but a procedure system connected to the interactive device 1 via the network N may be provided separately from the interactive device 1. In this case, the procedure unit 138 instructs the procedure system to execute the procedure according to the contents of the procedure information 124.
[0081] The notification unit 139 notifies the user terminal 2 of the details of the executed procedure. The notification is performed by, for example, SMS, chat, email, voice, or push notification.
[0082] The functional configuration of the dialogue device 1 is not limited to the above example. For example, the dialogue device 1 may have some of the above functional configurations, with the user terminal 2 having the rest. The dialogue device 1 may also have functional configurations other than those described above. Each functional configuration of the dialogue device 1 may be realized by software, as described above, or by hardware such as an IC chip, a SoC (System on Chip), an LSI (Large Scale Integration), or a microcomputer.
[0083] <Learning Process Executed by the Dialogue System 1000> Next, a description will be given of the processing executed by the dialogue system 1000. Fig. 8 is a flowchart showing an example of the processing executed by the dialogue system 1000. The following description will be given taking as an example a case where a user performs a procedure over the phone. The dialogue system 1000 repeatedly executes the following processing until the call is ended.
[0084] (Step S101) The user terminal 2 acquires voice information of the voice uttered by the user through the input device 205 (microphone) (step S101).
[0085] (Step S102) The communication unit 21 of the user terminal 2 transmits the acquired voice information to the interactive device 1 (step S102).
[0086] (Step S103) The communication unit 11 of the interactive device 1 receives (acquires) the voice information from the user terminal 2 (step S103), and stores it in the storage unit 12 as voice information 121.
[0087] (Step S104) The text conversion unit 131 converts the voice information 121 into text information 122 by voice recognition processing (step S104), and stores it in the storage unit 12.
[0088] (Step S105) The intention estimation unit 132 estimates the user's intention from the text information 122 using the estimation model 123 (step S105). The intention estimation unit 132 estimates the user's intention as, for example, a procedure, a question, an affirmative, a negative, or a combination of these. For example, when the user utters, "I'd like to make a reservation for three people. Are there any tables available?", the intention estimation unit 132 estimates that the user's intention is a procedure and a question.
[0089] If the user's intention is a procedure, the process proceeds to step S106. If the user's intention is a question, the process proceeds to step S109. If the user's intention is a positive one, the process proceeds to step S118. If the user's intention is at least one combination of a procedure, a question, and a positive one, the process proceeds to the corresponding step.
[0090] (Step S106) The slot extraction unit 133 checks whether the procedure information 124 corresponding to the procedure estimated in step S105 has been set (step S106). If the procedure information 124 has not been set (step S106: NO), the process proceeds to step S107. If the procedure information 124 has been set (step S106: YES), the process proceeds to step S108.
[0091] (Step S107) The slot extraction unit 133 refers to the procedure information 124, and sets the procedure information 124 corresponding to the procedure estimated in step S105 as the procedure information 124 for this call (step S107). Hereinafter, the procedure information 124 set here will be referred to as procedure information 124'. At this point, the value of the slot information included in the set procedure information 124' is blank.
[0092] (Step S108) The slot extraction unit 133 uses the extraction model 125 to extract slot information corresponding to the procedure information 124' from the text information 122 (step S108), and adds the slot information to the procedure information 124'. Then, the process proceeds to step S110.
[0093] (Step S109) The answer acquisition unit 134 refers to the Q&A information 127 and acquires answer information corresponding to the question estimated in step S105 (step S109).
[0094] (Step S110) The first text generation unit 135 generates a first text based on at least one of the slot information extracted in step S108 and the answer information acquired in step S109 in accordance with the first response rule 128 (step S111). More specifically, if the user's intention is a procedure, the first text generation unit 135 generates the first text based on the slot information extracted in step S108. Furthermore, if the user's intention is a question, the first text generation unit 135 generates the first text based on the answer information acquired in step S109. Furthermore, if the user's intention is both a procedure and a question, the first text generation unit 135 generates the first text based on the slot information extracted in step S108 and the answer information acquired in step S109. The first text generation unit 135 may generate the first text based on preset priorities of procedures and questions. For example, if the priority of the procedure is higher than the priority of the question, the first text generation unit 135 generates text based on the slot information, and then generates text based on the answer information, and outputs the entirety of these as the first text.
[0095] (Step S111) The second text generator 136 generates second text based on the slot information of the procedure information 124' in accordance with the second response rule 129 (step S111). If the procedure information 124' is not set, the process proceeds to step S112.
[0096] (Step S112) The output unit 137 synthesizes speech corresponding to the first text and the second text (step S112).
[0097] (Step S113) The output unit 137 transmits the voice information obtained in step S112 to the user terminal 2 via the communication unit 11 (step S113).
[0098] (Step S114) The user terminal 2 outputs the voice information received from the interactive device 1 by the output device 206 (speaker) (step S114). Through the above processing, a voice interaction (call) between the user and the interactive device 1 is realized.
[0099] (Step S115) Meanwhile, the procedure unit 138 of the interactive device 1 refers to the procedure information 124' and determines whether the slot information has been confirmed (step S115). Specifically, if the "confirmation flag" of the procedure information 124' is "on", it is determined that the slot information has been confirmed (step S115: YES), and the process proceeds to step S116. On the other hand, if the "confirmation flag" of the procedure information 124' is "off" or if the procedure information 124' has not been set, it is determined that the slot information has not been confirmed (step S115: NO), and the process ends.
[0100] (Step S116) The procedure unit 138 executes the procedure based on the contents of the procedure information 124' (step S116). Through the above processing, the procedure is executed by voice (call).
[0101] (Step S117) The notification unit 139 notifies the user terminal 2 of the contents of the procedure executed in step S117. The notified contents include the slot information of the procedure information 124'.
[0102] When the user interacts with the dialogue device 1 through text, the user terminal 2 acquires the text information input by the user (step S101) and transmits it to the dialogue device 1 (step S102). The dialogue device 1 then skips steps S103 and S104 and estimates the user's intention from the text information received from the dialogue device 1 (step S105). After generating a first text and a second text (steps S110 and S111), the dialogue device 1 generates text information corresponding to the first text and the second text (step S113) and transmits the text information to the user terminal 2 (step S113). The user terminal 2 displays the text information received from the dialogue device 1 on the output device 206 (display device) (step S114).
[0103] (Step S118) The intention estimation unit 132 refers to the procedure information 124' and checks whether there is a vacant slot in the slot information (a "slot" with a blank value) (step S118). If there is a vacant slot in the slot information (step S118: YES), the process proceeds to step S110. If there is no vacant slot in the slot information (step S118: NO), the process proceeds to step S119.
[0104] (Step S119) The intention estimation unit 132 sets the "confirmation flag" of the procedure information 124' to "on" (step S119). This confirms the slot information of the procedure information 124'. Thereafter, the process proceeds to step S110.
[0105] <How to generate the first text> Next, a method for generating a first text (first response rule 128) will be described. Fig. 9 is a flowchart showing an example of a method for generating a first text determined by the first response rule 128. Fig. 9 corresponds to the internal processing of step S110 in Fig. 8.
[0106] (Step S201) If the user's intention estimated in step S105 is a question (step S201: YES), the process proceeds to step S202. If the user's intention estimated in step S105 is not a question (step S201: NO), the process proceeds to step S203.
[0107] (Step S202) The first text generation unit 135 generates a first text based on the answer information acquired in step S109 (step S202). The first text generation unit 135 may use the answer information as the first text as is, or may generate the first text by referring to the target person information 126. In this way, a text including a backchannel and an answer to the question is generated as the first text.
[0108] (Step S203) If the user's intention estimated in step S105 is a procedure (step S203: YES), the process proceeds to step S204. If the user's intention estimated in step S105 is not a procedure (step S203: NO), the process proceeds to step S208.
[0109] (Step S204) If it is the first time that the user's intention has been estimated as a procedure (step S204: YES), the process proceeds to step S205. If it is not the first time that the user's intention has been estimated as a procedure, that is, if the user's intention has already been estimated as a procedure in previous interactions (step S204: NO), the process proceeds to step S206.
[0110] (Step S205) The first text generation unit 135 generates the text "It's {procedure name}, isn't it?" as the first text. {procedure name} is the name of the procedure. For example, "Procedure" in the procedure information 124' is referenced as the {procedure name}. When the procedure is a reservation and the procedure information 124 in FIG. 5 is referenced, the text "It's a reservation, isn't it?" is generated. Furthermore, the procedure information 124 may be provided with an information item referenced by {procedure name} in addition to "Procedure." This allows for polite expressions such as "It's a reservation, isn't it?". This generates text indicating a backchannel as the first text.
[0111] (Step S206) If slot information is extracted in step S108 (step S206: YES), the process proceeds to step S207. If slot information is not extracted in step S108 (step S206: NO), the process ends.
[0112] (Step S207) The first text generation unit 135 references the procedure information 124' and generates text such as "{S} is {s}, isn't it?" as the first text. {S} is the name of the extracted slot information, and {s} is the value of the extracted slot information. In addition to "slot," the procedure information 124' may also be provided with information items referenced by {S} and {s}. This enables, for example, honorific expressions such as "Your name is Sato-sama, isn't it?". When multiple pieces of slot information are extracted, text is generated in which "{S} is {s}" is repeated for each piece of slot information. This generates, as the first text, text that confirms the user's input.
[0113] (Step S208) The first text generator 135 refers to the procedural information 124' and checks whether the "confirmation flag" is "on" (step S208). If the "confirmation flag" is "on" (step S208: YES), the process proceeds to step S209. If the "confirmation flag" is "off" (step S208: NO), the process ends.
[0114] (Step S209) The first text generation unit 135 generates text "{S} will confirm {procedure name} with {s}" as the first text (step S209). {S} is the name of the slot information in the procedure information 124'. {s} is the value of the slot information in the procedure information 124'. In addition to "slot" and "procedure", information items referenced by {S}, {s}, and {procedure name} may be provided in the procedure information 124'. This enables polite expressions such as "We will confirm your reservation with your name Sato." If multiple slot information is set, text in which "{S} will {s}" is repeated for each slot information is generated. This generates text as the first text notifying the user that the procedure has been confirmed.
[0115] In this way, the first text is generated in accordance with the first response rule 128. More specifically, the first text is generated by applying preset answer information and slot information to a predefined text format. By generating the first text in this way, it is possible to prevent erroneous information from being output as the first text (hallucination). Furthermore, because the first text is not generated by a natural language processing model, attacks such as prompt injection against the dialogue system 1000 can be prevented.
[0116] <How to generate the second text> Next, a method for generating the second text (second response rule 129) will be described. Fig. 10 is a flowchart showing an example of a method for generating the second text determined by the second response rule 129. Fig. 10 corresponds to the internal processing of step S111 in Fig. 8.
[0117] (Step S301) The second text generator 136 refers to the procedure information 124' and checks whether there is a vacant slot in the slot information (a "slot" with a blank value) (step S301). If there is a vacant slot in the slot information (step S301: YES), the process proceeds to step S302. If there is no vacant slot in the slot information (step S301: NO), the process proceeds to step S303.
[0118] (Step S302) The second text generator 136 references the procedure information 124' and generates text "Please tell me {S}" as the second text (step S302). {S} is the name of the slot information whose value is blank in the procedure information 124' ("slot"). An information item referenced by {S} may be provided in the procedure information 124', separate from "slot". This allows for polite expressions such as "Please tell me your name". Furthermore, if multiple pieces of slot information are blank, similar text is generated for each piece of slot information. This generates text as the second text, requesting the user to input slot information.
[0119] (Step S303) The second text generator 136 refers to the procedural information 124' and checks whether the "confirmation flag" is "on" (step S303). If the "confirmation flag" is "on" (step S303: YES), the process proceeds to step S304. If the "confirmation flag" is "off" (step S303: NO), the process ends.
[0120] (Step S304) The second text generator 136 references the procedure information 124' and generates text "Are you sure that {S} is {s}?" as the second text (step S304). {S} is the name of the slot information ("slot") in the procedure information 124'. {s} is the value of the slot information in the procedure information 124'. In addition to "slot", an information item referenced by {S} and {s} may be provided in the procedure information 124'. This enables polite expressions such as "Are you sure that your name is Sato-sama?" If multiple slot information is set, text is generated in which "{S} is {s}" is repeated for each slot information. This generates text as the second text requesting the user to confirm the slot information.
[0121] In this way, the second text is generated in accordance with the second response rule 129. More specifically, the second text is generated by applying preset answer information and slot information to a predefined text format. By generating the second text in this way, it is possible to prevent erroneous information from being output as the second text (hallucination). Furthermore, because the second text is not generated by a natural language processing model, attacks such as prompt injection against the dialogue system 1000 can be prevented.
[0122] <Examples of dialogue> Next, a specific example of a dialogue between a user and the dialogue device 1 will be described. Fig. 11 is a schematic diagram showing a specific example of a dialogue. In the following, an example will be described in which a user makes a reservation by telephone.
[0123] (Step S401) After a call between the interactive device 1 and the user terminal 2 starts, when the user utters "Excuse me, I'd like to make a reservation," the user terminal 2 acquires the voice information of the utterance (step S101) and transmits the voice information to the interactive device 1 (step S102).
[0124] (Step S402) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). At this point, the procedure information 124' for "reservation" has not been set (step S106: NO), so the slot extraction unit 133 refers to the procedure information 124 and sets the procedure information 124' for "reservation" (step S107).
[0125] 12 is a diagram showing an example of procedure information 124' for "reservation." In the example of FIG. 12, procedure information 124' has "date," "time," "name," and "number of people" as slot information. At the time of step S402, "date," "time," and "name" are blank, and the "confirmation flag" is "off."
[0126] After setting the procedure information 124', the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, no slot information is extracted.
[0127] Since this is the first time that the user's intention has been estimated as a procedure (step S201: NO, step S203: YES, step S204: YES) and slot information has not been extracted (step S206: NO), first text generation unit 135 generates the text "You've made a reservation" as the first text (step S205). In this way, the first text is generated (step S110).
[0128] Because there is a vacant slot information in the procedure information 124' (step S301: YES), the second text generator 136 generates the text "Please tell me the date" as the second text (step S302). In this way, the second text is generated (step S111). Note that although it is assumed here that slot information with blank values is checked in order from the top, slot information with blank values may also be checked all at once.
[0129] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "This is a reservation, right? Please tell me the date" (step S114).
[0130] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0131] (Step S403) When the user utters "August 10th, please," the user terminal 2 acquires the voice information of the utterance (step S101) and transmits the voice information to the dialogue apparatus 1 (step S102).
[0132] (Step S404) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, "August 10th" is extracted as the "date." The slot extraction unit 133 adds the extracted slot information to the procedure information 124'.
[0133] Since this is not the first time that the user's intention has been estimated to be a procedure (step S201: NO, step S203: YES, step S204: NO), and "August 10th" has been extracted as the "date" (step S206: YES), first text generation unit 135 generates the text "The date is August 10th, isn't it?" as the first text (step S207). In this way, the first text is generated (step S110).
[0134] Since there is an empty slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me the time” as the second text (step S302). In this way, the second text is generated (step S111).
[0135] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "The date is August 10th, isn't it? Please tell me the time" (step S114).
[0136] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0137] (Step S405) When the user utters (asks) "What time does it start?", the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the interactive apparatus 1 (step S102).
[0138] (Step S406) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be a question about "business hours" (step S105: question), and the answer acquisition unit 134 refers to the Q&A information 127 and acquires answer information corresponding to "business hours" (text such as "Your question is about business hours. Our business hours are {business hours}") (step S109).
[0139] Since the first text generation unit 135 estimates that the user's intention is a question (step S201: YES, step S203: NO), it generates text "Your question is about business hours. Business hours are from 9:00 to 17:00" as the first text based on the answer information (step S202). Here, since the answer information includes {business hours}, the first text generation unit 135 generates the first text by referring to the target person information 126. In this way, the first text is generated (step S110).
[0140] Since there is an empty slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me the time” as the second text (step S302). In this way, the second text is generated (step S111).
[0141] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "Your question is about business hours, right? Our business hours are from 9:00 to 18:00. Please tell us the hours" (step S114).
[0142] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0143] (Step S407) When the user utters "Three people please, starting at 5 pm," the user terminal 2 acquires the voice information of the utterance (step S101) and transmits the voice information to the dialogue apparatus 1 (step S102).
[0144] (Step S408) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, "5:00 PM" is extracted as the "time" and "3" as the number of people. The slot extraction unit 133 adds the extracted slot information to the procedure information 124'.
[0145] Since this is not the first time that the user's intention has been estimated to be a procedure (step S201: NO, step S203: YES, step S204: NO), and "17:00" has been extracted as the "time" and "3" as the number of people (step S206: YES), first text generation unit 135 generates the text "The time is 17:00 and there are three people" as the first text (step S207). In this way, the first text is generated (step S110).
[0146] Since there is a vacant slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me your name” as the second text (step S302). In this way, the second text is generated (step S111).
[0147] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "The time is 5 p.m. and there are three people. Please tell us your names" (step S114).
[0148] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0149] (Step S409) When the user utters "Sato desu," the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the dialogue apparatus 1 (step S102).
[0150] (Step S410) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, "Sato" is extracted as the "name". The slot extraction unit 133 adds the extracted slot information to the procedure information 124'.
[0151] Since this is not the first time that the user's intention has been estimated to be a procedure (step S201: NO, step S203: YES, step S204: NO), and "Sato" has been extracted as the "name" (step S206: YES), first text generation unit 135 generates the text "Your name is Sato-sama, isn't it?" as the first text (step S207). In this way, the first text is generated (step S110).
[0152] Since there is no vacancy in the slot information of procedure information 124' (step S301: NO) and the "confirmation flag" is "off" (step S303: NO), second text generation unit 136 generates text as the second text saying "Are you sure the date is August 10th, the time is 5:00 PM, the name is Sato, and the number of people is three?" (step S304). In this way, the second text is generated (step S111).
[0153] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "Your name is Sato, right? The date is August 10th, the time is 5:00 PM, your name is Sato, and the number of people is three. Is this correct?" (step S114).
[0154] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0155] (Step S411) When the user utters "yes," the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the interactive apparatus 1 (step S102).
[0156] (Step S412) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "affirmative" (step S105: YES). Because there is no vacant slot information in the procedure information 124' (step S118: NO), the intention estimation unit 132 sets the "confirmation flag" of the procedure information 124' to "on" (step S119).
[0157] Since the user's intention is neither a question nor a procedure (step S201: NO, step S203: NO) and the "confirm flag" is "on" (step S208: YES), first text generation unit 135 generates the following text as the first text: "We will confirm your reservation for the date August 10th, time 5:00 PM, name Sato, and number of people 3" (step S209). In this way, the first text is generated (step S110).
[0158] Since there is no free space in the slot information of the procedure information 124′ (step S301: NO) and the “confirmation flag” is “on” (step S303: YES), second text generation unit 136 ends the process without generating second text.
[0159] The output unit 137 synthesizes voice information corresponding to the first text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "We are confirming your reservation for August 10th, 5:00 PM, your name is Sato, and the number of people is three" (step S114).
[0160] Since the slot information of the procedure information 124' has been confirmed (step S115: YES), the procedure unit 138 of the interactive device 1 executes the reservation using the contents of the procedure information 124' (step S116) and saves the contents as procedure content information 130 in the storage unit 12. As a result, a reservation is made for three people named Sato at 5:00 PM on August 10th.
[0161] Thereafter, the notification unit 139 notifies the user terminal 2 of the procedure content information 130 by a method such as SMS, push notification, or email (step S117).
[0162] <Summary> As described above, according to this embodiment, a dialogue system 1000 is realized which includes an intention estimation unit 132 that estimates an intention from the text information 122 using the estimation model 123, a slot extraction unit 133 that, if the intention is a procedure, extracts slot information corresponding to the procedure from the text information 122 using the extraction model 125, an answer acquisition unit 134 that, if the intention is a question, acquires answer information corresponding to the question, a first text generation unit 135 that generates a first text based on at least one of the slot information and the answer information, a second text generation unit 136 that generates a second text based on the slot information, and an output unit 137 that outputs the first text and the second text.
[0163] According to the dialogue system 1000, the first text and the second text are generated in accordance with the first response rule 128 and the second response rule 129. More specifically, the first text and the second text are generated by applying preset answer information and slot information to a predefined text format. By generating the first text and the second text in this manner, it is possible to prevent erroneous information from being output as the first text and the second text (hallucination). Furthermore, because the first text and the second text are not generated by a natural language processing model, attacks such as prompt injection against the dialogue system 1000 can be prevented.
[0164] [Second embodiment] The dialogue system 1000 according to the second embodiment is a system that enables a user to select a product from among a plurality of products through dialogue and perform procedures such as reservation and ordering. Note that a description of the same configuration as in the first embodiment will be omitted.
[0165] <Functional configuration of the dialogue device 1> Next, a description will be given of the functional configuration of the dialogue device 1. Fig. 13 is a diagram showing an example of the functional configuration of the dialogue device 1. As shown in Fig. 13, the dialogue device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0166] The storage unit 12 is realized by a memory 102 and a storage 103. The storage unit 12 stores voice information 121, text information 122, an estimation model 123, procedure information 124, an extraction model 125, target person information 126, Q&A information 127, a first response rule 128, a second response rule 129, procedure content information 130, a product selection model 1201, and product information 1202.
[0167] The inference model 123 is a natural language processing model that infers a user's intention (context) from the text information 122. The inference model 123 is, for example, but not limited to, a large-scale language model. When the text information 122 is input, the inference model 123 is trained to output an intention corresponding to the input text information 122 from among one or more predetermined intentions. The inference model 123 is, for example, but not limited to, a GPT that is fine-tuned by learning the relationship between the text information 122 and the intention.
[0168] The intent corresponds to the purpose of the text indicated by the text information 122, i.e., the intent of the user who input the utterance or text. The intent may be, for example, but not limited to, a procedure, a question, a product selection, a positive or negative response. The intent may be, for example, but not limited to, a reservation, an order, a purchase, an application, or a change or cancellation thereof. The question may be, for example, but not limited to, a question about business hours, price, deadline, number of people, availability, cancellation conditions, product type, or whether a reservation can be made.
[0169] The estimation model 123 may be a natural language processing model that estimates a user's intention from the speech information 121. In this case, when the speech information 121 is input, the estimation model 123 is trained to output an intention corresponding to the input speech information 121 from among one or more predetermined intentions. The estimation model 123 is, for example, a multimodal LLM that is fine-tuned by learning the relationship between the speech information 121 and the intention, but is not limited to this.
[0170] The product selection model 1201 is a natural language processing model that estimates a product that a user is going to select from the text information 122. The product selection model 1201 is, for example, but not limited to, a large-scale language model. When the text information 122 is input, the product selection model 1201 is trained to output a product corresponding to the input text information 122 from one or more products registered in advance. The product selection model 1201 is, for example, but not limited to, a Generative Pretrained Transformer (GPT) that is fine-tuned by learning the relationship between the text information 122 and products. The product selection model 1201 may also be a natural language processing model that extracts product-related keywords from the text information 122.
[0171] A product is any good or service provided or sold by the subject. For example, if the subject is a service provider, the service provided by the subject is registered as the product. Specifically, if the subject is a restaurant, a meal course (menu) is registered as the product. Furthermore, if the subject is an accommodation facility, an accommodation plan is registered as the product. Furthermore, if the subject is a seller of products or services, for example, the service or product sold by the subject is registered as the product. Furthermore, if the subject is a public institution, administrative institution, or entity acting on their behalf to accept procedures, the procedure provided by the subject is registered as the product. Note that products are not limited to the above examples and can be registered arbitrarily depending on the subject.
[0172] The product selection model 1201 may be a natural language processing model that estimates the product that the user is going to select from the speech information 121. In this case, when the speech information 121 is input, the product selection model 1201 is trained to output a product corresponding to the input text information 122 from among one or more products registered in advance. The product selection model 1201 is, for example, a multimodal LLM that is fine-tuned by learning the relationship between the speech information 121 and products, but is not limited to this.
[0173] The product information 1202 is information about products that can be selected by the user and is registered for each product. The product information 1202 is registered in advance by the target person.
[0174] Fig. 14 is a diagram showing an example of product information 1202. The product information 1202 in Fig. 14 is product information for a restaurant, and includes information items such as "IID," "course name," and "price."
[0175] "IID" is identification information that uniquely identifies a product. In the following, a product whose "IID" is "Ixxx" may be referred to as product Ixxx.
[0176] "Course name" is the name of the product.
[0177] "Amount" is the amount of the product.
[0178] In the example of Fig. 14, the "course name" of product I001 is "mushroom hotpot course" and the "price" is "4000." This indicates that product I001 is a mushroom hotpot course that costs 4000 yen.
[0179] Note that the information included in product information 1202 is not limited to the above example. Product information 1202 may not include some of the above information, or may include information other than the above. For example, product information 1202 may include information regarding product options (whether all-you-can-drink is available, large servings, some menu changes, product color, model number, size, etc.). Furthermore, product information 1202 may be unstructured data in which a description of each product is written in text, rather than structured data as shown in FIG. 14 .
[0180] FIG. 15 is a diagram showing an example of procedure information 124 according to this embodiment. In the example of FIG. 15, "slot 5" of procedure P001 is "product," and "slot 6" is "quantity." This indicates that procedure P001 is a reservation procedure, and that the "product" and its "quantity" are set as information necessary to execute the reservation. When procedure P001 is executed, the product (e.g., a course) selected by the user is reserved for the target person (e.g., a restaurant) in the quantity specified by the user.
[0181] It should be noted that the information included in the procedure information 124 is not limited to the above examples. The procedure information 124 may not include some of the above information, or may include information other than the above.
[0182] The control unit 13 is realized by the processor 101 reading and executing a program from the memory 102 and working in cooperation with other hardware components. The control unit 13 includes a text conversion unit 131, an intention estimation unit 132, a slot extraction unit 133, an answer acquisition unit 134, a first text generation unit 135, a second text generation unit 136, an output unit 137, a procedure unit 138, a notification unit 139, and a product selection unit 140.
[0183] The product selection unit 140 uses the product selection model 1201 to select a product corresponding to the text information 122. The product selected here corresponds to the product that the user is about to select.
[0184] The functional configuration of the dialogue device 1 is not limited to the above example. For example, the dialogue device 1 may have some of the above functional configurations, and the user terminal 2 may have the rest. The dialogue device 1 may also have functional configurations other than those described above. Each functional configuration of the dialogue device 1 may be realized by software, as described above, or by hardware such as an IC chip, an SoC, an LSI, or a microcomputer.
[0185] <Learning Process Executed by the Dialogue System 1000> Next, a description will be given of the processing executed by the dialogue system 1000. Fig. 16 is a flowchart showing an example of the processing executed by the dialogue system 1000. The following description will be given taking as an example a case where a user performs a procedure over the phone. The dialogue system 1000 repeatedly executes the following processing until the call is ended.
[0186] (Step S101) The user terminal 2 acquires voice information of the voice uttered by the user through the input device 205 (microphone) (step S101).
[0187] (Step S102) The communication unit 21 of the user terminal 2 transmits the acquired voice information to the interactive device 1 (step S102).
[0188] (Step S103) The communication unit 11 of the interactive device 1 receives (acquires) the voice information from the user terminal 2 (step S103), and stores it in the storage unit 12 as voice information 121.
[0189] (Step S104) The text conversion unit 131 converts the voice information 121 into text information 122 by voice recognition processing (step S104), and stores it in the storage unit 12.
[0190] (Step S105) The intention estimation unit 132 estimates the user's intention from the text information 122 using the estimation model 123 (step S105). The intention estimation unit 132 estimates the user's intention as, for example, a procedure, a question, a product selection, an affirmative, a negative, or a combination of these. For example, when the user utters, "What kind of lunch courses do you have?", the intention estimation unit 132 estimates that the user's intention is a question. Furthermore, when the user utters, "We'd like to make a reservation for three people. Are there any tables available?", the intention estimation unit 132 estimates that the user's intention is a procedure and a question.
[0191] If the user's intention is a procedure or product selection, the process proceeds to step S106. If the user's intention is a question, the process proceeds to step S109. If the user's intention is a positive one, the process proceeds to step S118. If the user's intention is a combination of at least one of a procedure, a question, and a positive one, the process proceeds to the corresponding step.
[0192] (Step S106) The slot extraction unit 133 checks whether the procedure information 124 corresponding to the procedure estimated in step S105 has been set (step S106). If the procedure information 124 has not been set (step S106: NO), the process proceeds to step S107. If the procedure information 124 has been set (step S106: YES), the process proceeds to step S118.
[0193] (Step S107) The slot extraction unit 133 refers to the procedure information 124 and sets the procedure information 124 corresponding to the procedure estimated in step S105 as the procedure information 124 for this call (step S107). Then, the process proceeds to step S120. Hereinafter, the procedure information 124 set here will be referred to as procedure information 124'. At this point, the value of the slot information included in the set procedure information 124' is blank.
[0194] (Step S120) If the intention estimated in step S105 is a procedure (step S120: procedure), the process proceeds to step S108. If the intention estimated in step S105 is a product selection (step S120: product selection), the process proceeds to step S121.
[0195] (Step S108) The slot extraction unit 133 uses the extraction model 125 to extract slot information corresponding to the procedure information 124' from the text information 122 (step S108), and adds the slot information to the procedure information 124'. Then, the process proceeds to step S110.
[0196] (Step S121) The product selection unit 140 uses the product selection model 1201 to select one or more products corresponding to the text information 122 (step S121). Specifically, the product selection unit 140 inputs the text information 122 and the product information 1202 into the product selection model 1201, and acquires the products that the product selection model 1201 outputs as products corresponding to the text information 122 as products corresponding to the text information 122. Alternatively, the product selection unit 140 may input the text information 122 into the product selection model 1201, compare the product information 1202 with keywords extracted from the text information 122 by the product model 1201, and select products corresponding to the text information 122. If there is one product corresponding to the text information 122, the product selection unit 140 adds that product to the procedure information 124′. Thereafter, the process proceeds to step 110.
[0197] (Step S109) The answer acquisition unit 134 refers to the Q&A information 127 and acquires answer information corresponding to the question estimated in step S105 (step S109).
[0198] (Step S110) The first text generation unit 135 generates a first text in accordance with the first response rule 128 based on at least one of the slot information extracted in step S108, the answer information acquired in step S109, and the product information 1202 (product selection result) of the product selected in step S121 (step S111). More specifically, if the user's intention is a procedure, the first text generation unit 135 generates the first text based on the slot information extracted in step S108. If the user's intention is a question, the first text generation unit 135 generates the first text based on the answer information acquired in step S109. If the user's intention is both a procedure and a question, the first text generation unit 135 generates the first text based on the slot information extracted in step S108 and the answer information acquired in step S109. The first text generation unit 135 may generate the first text based on preset priorities of procedures and questions. For example, if the priority of the procedure is higher than the priority of the question, first text generation unit 135 generates text based on slot information, then generates text based on answer information, and outputs the whole of these as the first text. Also, if the user's intention is to select a product, first text generation unit 135 generates the first text based on product information 1202 of the product selected in step S120 (product selection result).
[0199] (Step S111) The second text generator 136 generates second text based on the slot information of the procedure information 124' in accordance with the second response rule 129 (step S111). If the procedure information 124' is not set, the process proceeds to step S112.
[0200] (Step S112) The output unit 137 synthesizes speech corresponding to the first text and the second text (step S112).
[0201] (Step S113) The output unit 137 transmits the voice information obtained in step S112 to the user terminal 2 via the communication unit 11 (step S113).
[0202] (Step S114) The user terminal 2 outputs the voice information received from the interactive device 1 by the output device 206 (speaker) (step S114). Through the above processing, a voice interaction (call) between the user and the interactive device 1 is realized.
[0203] (Step S115) Meanwhile, the procedure unit 138 of the interactive device 1 refers to the procedure information 124' and determines whether the slot information has been confirmed (step S115). Specifically, if the "confirmation flag" of the procedure information 124' is "on", it is determined that the slot information has been confirmed (step S115: YES), and the process proceeds to step S116. On the other hand, if the "confirmation flag" of the procedure information 124' is "off" or if the procedure information 124' has not been set, it is determined that the slot information has not been confirmed (step S115: NO), and the process ends.
[0204] (Step S116) The procedure unit 138 executes the procedure based on the contents of the procedure information 124' (step S116). Through the above processing, the procedure is executed by voice (call).
[0205] (Step S117) The notification unit 139 notifies the user terminal 2 of the contents of the procedure executed in step S117. The notified contents include the slot information of the procedure information 124'.
[0206] When the user interacts with the dialogue device 1 through text, the user terminal 2 acquires the text information input by the user (step S101) and transmits it to the dialogue device 1 (step S102). The dialogue device 1 then skips steps S103 and S104 and estimates the user's intention from the text information received from the dialogue device 1 (step S105). After generating a first text and a second text (steps S110 and S111), the dialogue device 1 generates text information corresponding to the first text and the second text (step S113) and transmits the text information to the user terminal 2 (step S113). The user terminal 2 displays the text information received from the dialogue device 1 on the output device 206 (display device) (step S114).
[0207] (Step S118) The intention estimation unit 132 refers to the procedure information 124' and checks whether there is a vacant slot in the slot information (a "slot" with a blank value) (step S118). If there is a vacant slot in the slot information (step S118: YES), the process proceeds to step S110. If there is no vacant slot in the slot information (step S118: NO), the process proceeds to step S119.
[0208] (Step S119) The intention estimation unit 132 sets the "confirmation flag" of the procedure information 124' to "on" (step S119). This confirms the slot information of the procedure information 124'. Thereafter, the process proceeds to step S110.
[0209] <How to generate the first text> Next, a method for generating a first text (first response rule 128) will be described. Fig. 17 is a flowchart showing an example of a method for generating a first text determined by the first response rule 128. Fig. 17 corresponds to the internal processing of step S110 in Fig. 16.
[0210] (Step S201) If the user's intention estimated in step S105 is a question (step S201: YES), the process proceeds to step S202. If the user's intention estimated in step S105 is not a question (step S201: NO), the process proceeds to step S203.
[0211] (Step S202) The first text generation unit 135 generates a first text based on the answer information acquired in step S109 (step S202). The first text generation unit 135 may use the answer information as the first text as is, or may generate the first text by referring to the target person information 126. In this way, a text including a backchannel and an answer to the question is generated as the first text.
[0212] (Step S203) If the user's intention estimated in step S105 is a procedure (step S203: YES), the process proceeds to step S204. If the user's intention estimated in step S105 is not a procedure (step S203: NO), the process proceeds to step S208.
[0213] (Step S204) If it is the first time that the user's intention has been estimated as a procedure (step S204: YES), the process proceeds to step S205. If it is not the first time that the user's intention has been estimated as a procedure, that is, if the user's intention has already been estimated as a procedure in previous interactions (step S204: NO), the process proceeds to step S206.
[0214] (Step S205) The first text generation unit 135 generates the text "It's {procedure name}, isn't it?" as the first text. {procedure name} is the name of the procedure. For example, "Procedure" in the procedure information 124' is referenced as the {procedure name}. When the procedure is a reservation and the procedure information 124 in FIG. 5 is referenced, the text "It's a reservation, isn't it?" is generated. Furthermore, the procedure information 124 may be provided with an information item referenced by {procedure name} in addition to "Procedure." This allows for polite expressions such as "It's a reservation, isn't it?". This generates text indicating a backchannel as the first text.
[0215] (Step S206) If slot information is extracted in step S108 (step S206: YES), the process proceeds to step S207. If slot information is not extracted in step S108 (step S206: NO), the process proceeds to step S210.
[0216] (Step S207) The first text generation unit 135 references the procedure information 124' and generates text such as "{S} is {s}, isn't it?" as the first text. {S} is the name of the extracted slot information, and {s} is the value of the extracted slot information. In addition to "slot," the procedure information 124' may also be provided with information items referenced by {S} and {s}. This enables, for example, honorific expressions such as "Your name is Sato-sama, isn't it?". When multiple pieces of slot information are extracted, text is generated in which "{S} is {s}" is repeated for each piece of slot information. This generates, as the first text, text that confirms the user's input.
[0217] (Step S208) The first text generator 135 refers to the procedure information 124' and checks whether the "confirmation flag" is "on" (step S208). If the "confirmation flag" is "on" (step S208: YES), the process proceeds to step S209. If the "confirmation flag" is "off" (step S208: NO), the process proceeds to step S210.
[0218] (Step S209) The first text generator 135 generates the text "{S} will confirm {procedure name} with {s}" as the first text (step S209). {S} is the name of the slot information in the procedure information 124'. {s} is the value of the slot information in the procedure information 124'. In addition to "slot" and "procedure," the procedure information 124' may also include information items referenced by {S}, {s}, and {procedure name}. This enables, for example, honorific expressions such as "We will confirm your reservation for Mr. Sato." If multiple slot information items are set, text is generated in which "{S} will {s}" is repeated for each slot information item. This generates, as the first text, text notifying the user that the procedure has been confirmed. Then, processing proceeds to step S210.
[0219] In this way, the first text is generated in accordance with the first response rule 128. More specifically, the first text is generated by applying preset answer information and slot information to a predefined text format. By generating the first text in this way, it is possible to prevent erroneous information from being output as the first text (hallucination). Furthermore, because the first text is not generated by a natural language processing model, attacks such as prompt injection against the dialogue system 1000 can be prevented.
[0220] (Step S210) If the user's intention estimated in step S105 is to select a product (step S210: YES), the process proceeds to step S211. If the user's intention estimated in step S105 is not to select a product (step S210: NO), the process ends.
[0221] (Step S211) The first text generator 135 checks whether one product has been selected in step S121 (step S211). If one product has been selected (step S211: YES), the process proceeds to step S212. If multiple products have been selected (step S211: NO), the process proceeds to step S213.
[0222] (Step S212) The first text generation unit 135 generates the text "It's {product name}, isn't it?" as the first text. {product name} is the name of one product selected in step S121 as the product corresponding to the text information 122. For example, the "course name" in the product information 1202 is referenced as the {product name}. When the selected product is product I001 and the product information 1202 in FIG. 14 is referenced, the text "It's the mushroom hotpot course, isn't it?" is generated as the first text. This generates text indicating a backchannel, that is, text indicating one product selected from the text information 122 and confirming whether that product is correct. The process then ends.
[0223] (Step S213) The first text generation unit 135 generates text such as "We have {product name}, {product name}" as the first text. The {product name} is the name of multiple products selected in step S121 as products corresponding to the text information 122. The name of each of the selected multiple products is input into the {product name}. The number of {product name}s added to the first text is equal to the number of products selected in step S121. For example, if three products are selected in step S121, a first message such as "We have {product name}, {product name}, {product name}" is generated. For example, the "course name" in the product information 1202 is referenced as the {product name}. If the selected products are products I001 and I002 and the product information 1202 in FIG. 14 is referenced, a first text such as "We have a mushroom hotpot course and a mushroom hotpot course with an aperitif" is generated. As a result, text indicating a backchannel, i.e., text indicating multiple products selected from the text information 122, is generated as the first text. The processing then ends.
[0224] <Examples of dialogue> Next, a specific example of a dialogue between a user and the dialogue device 1 will be described. Fig. 18 is a schematic diagram showing a specific example of a dialogue. In the following, an example will be described in which a user makes a reservation by telephone.
[0225] (Step S401) After a call between the interactive device 1 and the user terminal 2 starts, when the user utters "Excuse me, I'd like to make a reservation," the user terminal 2 acquires the voice information of the utterance (step S101) and transmits the voice information to the interactive device 1 (step S102).
[0226] (Step S402) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). At this point, the procedure information 124' for "reservation" has not been set (step S106: NO), so the slot extraction unit 133 refers to the procedure information 124 and sets the procedure information 124' for "reservation" (step S107).
[0227] 19 is a diagram showing an example of procedure information 124' for "reservation." In the example of FIG. 19, procedure information 124' has "date," "time," "name," and "number of people" as slot information. At the time of step S402, "date," "time," and "name" are blank, and the "confirmation flag" is "off."
[0228] After setting the procedure information 124', the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, no slot information is extracted.
[0229] Since this is the first time that the user's intention has been estimated as a procedure (step S201: NO, step S203: YES, step S204: YES) and slot information has not been extracted (step S206: NO), first text generation unit 135 generates the text "You've made a reservation" as the first text (step S205). In this way, the first text is generated (step S110).
[0230] Because there is a vacant slot information in the procedure information 124' (step S301: YES), the second text generator 136 generates the text "Please tell me the date" as the second text (step S302). In this way, the second text is generated (step S111). Note that although it is assumed here that slot information with blank values is checked in order from the top, slot information with blank values may also be checked all at once.
[0231] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "This is a reservation, right? Please tell me the date" (step S114).
[0232] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0233] (Step S403) When the user utters "August 10th, please," the user terminal 2 acquires the voice information of the utterance (step S101) and transmits the voice information to the dialogue apparatus 1 (step S102).
[0234] (Step S404) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, "August 10th" is extracted as the "date." The slot extraction unit 133 adds the extracted slot information to the procedure information 124'.
[0235] Since this is not the first time that the user's intention has been estimated to be a procedure (step S201: NO, step S203: YES, step S204: NO), and "August 10th" has been extracted as the "date" (step S206: YES), first text generation unit 135 generates the text "The date is August 10th, isn't it?" as the first text (step S207). In this way, the first text is generated (step S110).
[0236] Since there is an empty slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me the time” as the second text (step S302). In this way, the second text is generated (step S111).
[0237] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "The date is August 10th, isn't it? Please tell me the time" (step S114).
[0238] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0239] (Step S405) When the user utters (asks) "What time does it start?", the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the interactive apparatus 1 (step S102).
[0240] (Step S406) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be a question about "business hours" (step S105: question), and the answer acquisition unit 134 refers to the Q&A information 127 and acquires answer information corresponding to "business hours" (text such as "Your question is about business hours. Our business hours are {business hours}") (step S109).
[0241] Since the first text generation unit 135 estimates that the user's intention is a question (step S201: YES, step S203: NO), it generates text "Your question is about business hours. Business hours are from 9:00 to 17:00" as the first text based on the answer information (step S202). Here, since the answer information includes {business hours}, the first text generation unit 135 generates the first text by referring to the target person information 126. In this way, the first text is generated (step S110).
[0242] Since there is an empty slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me the time” as the second text (step S302). In this way, the second text is generated (step S111).
[0243] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "Your question is about business hours, right? Our business hours are from 9:00 to 18:00. Please tell us the hours" (step S114).
[0244] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0245] (Step S407) When the user utters "Three people please, starting at 5 pm," the user terminal 2 acquires the voice information of the utterance (step S101) and transmits the voice information to the dialogue apparatus 1 (step S102).
[0246] (Step S408) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, "5:00 PM" is extracted as the "time" and "3" as the number of people. The slot extraction unit 133 adds the extracted slot information to the procedure information 124'.
[0247] Since this is not the first time that the user's intention has been estimated to be a procedure (step S201: NO, step S203: YES, step S204: NO), and "17:00" has been extracted as the "time" and "3" as the number of people (step S206: YES), first text generation unit 135 generates the text "The time is 17:00 and there are three people" as the first text (step S207). In this way, the first text is generated (step S110).
[0248] Since there is a vacant slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me your name” as the second text (step S302). In this way, the second text is generated (step S111).
[0249] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "The time is 5 p.m. and there are three people. Please tell us your names" (step S114).
[0250] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0251] (Step S409) When the user utters "Sato desu," the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the dialogue apparatus 1 (step S102).
[0252] (Step S413) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, "Sato" is extracted as the "name". The slot extraction unit 133 adds the extracted slot information to the procedure information 124'.
[0253] Since this is not the first time that the user's intention has been estimated to be a procedure (step S201: NO, step S203: YES, step S204: NO), and "Sato" has been extracted as the "name" (step S206: YES), first text generation unit 135 generates the text "Your name is Sato-sama, isn't it?" as the first text (step S207). In this way, the first text is generated (step S110).
[0254] Since there is a vacant slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me the course” as the second text (step S302). In this way, the second text is generated (step S111).
[0255] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "Your name is Sato, isn't it? Please tell me the course" (step S114).
[0256] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0257] (Step S414) When the user utters "I'd like the mushroom hotpot course, please," the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the dialogue apparatus 1 (step S102).
[0258] (Step S415) When the dialogue device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "product selection" (step S105: procedure / product selection, step S120: product selection). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the product selection unit 140 uses the product selection model 1201 and refers to the product information 1202 to select a product corresponding to the text information 122 (step S121). Here, three "courses (products)" are selected: "mushroom hotpot course," "mushroom hotpot course with aperitif," and "mushroom hotpot course with all-you-can-drink."
[0259] Because the user's intention is to select a product (step S210: YES) and multiple courses have been selected as "course (product)" (step S211: NO), first text generation unit 135 generates the text "We have a mushroom hotpot course, a mushroom hotpot course with an aperitif, and a mushroom hotpot course with all-you-can-drink" as the first text (step S213). In this way, the first text is generated (step S110).
[0260] Because there is an empty slot in the procedure information 124' (step S301: YES), the second text generation unit 136 generates text "Please tell me the course" as the second text (step S302). In this way, the second text is generated (step S111). Note that when multiple products are selected from the text information 122, the second text generation unit 136 may generate text requesting the selection of one product from the multiple products, such as "Which one would you like?" as the second text.
[0261] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "We have a mushroom hotpot course, a mushroom hotpot course with an aperitif, and a mushroom hotpot course with all-you-can-drink. Please tell me which course you want" (step S114).
[0262] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0263] (Step S416) When the user utters "I'd like an all-you-can-drink option, please," the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the interactive apparatus 1 (step S102).
[0264] (Step S417) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "product selection" (step S105: procedure / product selection, step S120: product selection). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the product selection unit 140 uses the product selection model 1201 and refers to the product information 1202 to select a product corresponding to the text information 122 (step S121). Here, "mushroom hotpot course with all-you-can-drink" is selected as the "course (product)." The product selection unit 140 adds the selected slot information to the procedure information 124'.
[0265] Because the user's intention is to select a product (step S210: YES) and one course is selected as the "course (product)" (step S211: YES), first text generation unit 135 generates the text "The mushroom hotpot course comes with all-you-can-drink" as the first text (step S212). In this way, the first text is generated (step S110).
[0266] Since there is a vacant slot in the procedure information 124′ (step S301: YES), the second text generator 136 generates the text “Please tell me the number” as the second text (step S302). In this way, the second text is generated (step S111).
[0267] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "The mushroom hotpot course includes all-you-can-drink. Please tell me the number" (step S114).
[0268] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0269] (Step S418) When the user utters "There are three," the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the interactive apparatus 1 (step S102).
[0270] (Step S419) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "reservation" (step S105: procedure). Because the procedure information 124' for "reservation" has already been set (step S106: YES), the slot extraction unit 133 extracts slot information included in the procedure information 124' from the text information 122 (step S108). Here, "3" is extracted as the "number." The slot extraction unit 133 adds the extracted slot information to the procedure information 124'.
[0271] Since this is not the first time that the user's intention has been estimated to be a procedure (step S201: NO, step S203: YES, step S204: NO), and "3" has been extracted as the "number" (step S206: YES), first text generation unit 135 generates the text "The number is three, isn't it?" as the first text (step S207). In this way, the first text is generated (step S110).
[0272] Because there is no vacancy in the slot information of the procedure information 124' (step S301: NO) and the "confirmation flag" is "off" (step S303: NO), the second text generation unit 136 generates the following text as the second text (step S304): "The date is August 10th, the time is 5:00 PM, the name is Sato, the number of people is 3, and the course is three mushroom hotpot courses with all-you-can-drink, are you sure?" In this way, the second text is generated (step S111).
[0273] The output unit 137 synthesizes voice information corresponding to the first text and the second text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "The course is a mushroom hotpot course with all-you-can-drink, correct? The date is August 10th, the time is 5:00 PM, the name is Sato, the number of people is three, and the course is three mushroom hotpot courses with all-you-can-drink, correct?" (step S114).
[0274] Since the slot information of the procedure information 124' has not been determined (step S115: NO), the procedure unit 138 of the interactive device 1 does not execute the reservation and ends the process.
[0275] (Step S420) When the user utters "yes," the user terminal 2 acquires voice information of the utterance (step S101) and transmits the voice information to the interactive apparatus 1 (step S102).
[0276] (Step S421) When the interactive device 1 acquires the voice information 121 from the user terminal 2 (step S103), the text conversion unit 131 converts the voice information 121 into text information 122 (step S104), and the intention estimation unit 132 estimates the user's intention using the estimation model 123 (step S105). Here, the user's intention is estimated to be "affirmative" (step S105: YES). Because there is no vacant slot information in the procedure information 124' (step S118: NO), the intention estimation unit 132 sets the "confirmation flag" of the procedure information 124' to "on" (step S119).
[0277] Because the user's intention is not a question, procedure, or product selection (step S201: NO, step S203: NO, step S210: NO), and the "confirm flag" is "on" (step S208: YES), first text generation unit 135 generates the following text as the first text: "We will confirm your reservation for the date August 10th, time 5:00 PM, name Sato, number of people 3, and three courses including the mushroom hotpot course with all-you-can-drink" (step S209). Thus, the first text is generated (step S110).
[0278] Since there is no vacancy in the slot information of the procedure information 124′ (step S301: NO) and the “confirmation flag” is “on” (step S303: YES), second text generation unit 136 ends the process without generating second text.
[0279] The output unit 137 synthesizes voice information corresponding to the first text (step S112), and the communication unit 11 transmits the obtained voice information to the user terminal 2 (step S113). Upon receiving the voice information, the user terminal 2 outputs a voice message saying, "Your reservation is confirmed for the date August 10th, the time 5:00 PM, your name is Sato, the number of people is three, and the course is a mushroom hotpot course with all-you-can-drink for three" (step S114).
[0280] Since the slot information of the procedure information 124' has been confirmed (step S115: YES), the procedure unit 138 of the interactive device 1 executes the reservation using the contents of the procedure information 124' (step S116) and saves the contents as procedure content information 130 in the storage unit 12. As a result, three reservations are made for three people named Sato at 5:00 pm on August 10th, with a mushroom hotpot course and all-you-can-drink.
[0281] Thereafter, the notification unit 139 notifies the user terminal 2 of the procedure content information 130 by a method such as SMS, push notification, or email (step S117).
[0282] <Summary> As described above, according to this embodiment, a dialogue system 1000 is realized which includes an intention estimation unit 132 that estimates an intention from the text information 122 using the estimation model 123, a slot extraction unit 133 that, if the intention is a procedure, extracts slot information corresponding to the procedure from the text information 122 using the extraction model 125, an answer acquisition unit 134 that, if the intention is a question, acquires answer information corresponding to the question, a first text generation unit 135 that generates a first text based on at least one of the slot information and the answer information, a second text generation unit 136 that generates a second text based on the slot information, and an output unit 137 that outputs the first text and the second text.
[0283] According to the dialogue system 1000, the first text and the second text are generated in accordance with the first response rule 128 and the second response rule 129. More specifically, the first text and the second text are generated by applying preset answer information and slot information to a predefined text format. By generating the first text and the second text in this manner, it is possible to prevent erroneous information from being output as the first text and the second text (hallucination). Furthermore, because the first text and the second text are not generated by a natural language processing model, attacks such as prompt injection against the dialogue system 1000 can be prevented.
[0284] Also, a dialogue system 1000 is realized which includes an intention estimation unit 132 which estimates an intention from text information 122 using an estimation model 123, a product selection unit 140 which, if the intention is product selection, selects a product corresponding to the text information 122 using a product selection model 1201, a first text generation unit 135 which generates a first text based on the product selection result, and an output unit 137 which outputs the first text.
[0285] According to the dialogue system 1000, even if there are multiple products that the user can reserve or order, the user can select a desired product from the multiple products through dialogue and make a reservation or order.
[0286] <Additional Notes> The present embodiment includes the following disclosure.
[0287] (Appendix 1) an intention estimation unit that estimates an intention from speech information or text information using an estimation model; a product selection unit that selects a product corresponding to the voice information or text information by using a product selection model when the intention is to select a product; a first text generation unit that generates a first text based on the product selection result; an output unit that outputs the first text; A dialogue system comprising:
[0288] (Appendix 2) When a plurality of products corresponding to the audio information or the text information are selected, the first text generation unit generates a first text indicating the plurality of products. Dialogue system according to appendix 1.
[0289] (Appendix 3) When one product corresponding to the audio information or the text information is selected, the first text generation unit generates a first text indicating the one product. Dialogue system according to appendix 1.
[0290] (Appendix 4) The product selection unit refers to product information in which information about a plurality of products is registered, and selects a product corresponding to the text information. Dialogue system according to appendix 1.
[0291] (Appendix 5) The product selection model is a natural language processing model. Dialogue system according to appendix 1.
[0292] (Appendix 6) The first text is generated according to a first preset response rule. Dialogue system according to appendix 1.
[0293] (Appendix 7) The system further includes a text conversion unit that converts audio information into the text information. Dialogue system according to appendix 1.
[0294] (Appendix 8) The output unit outputs audio information corresponding to the first text. Dialogue system according to appendix 1.
[0295] (Appendix 9) The output unit outputs text information corresponding to the first text. Dialogue system according to appendix 1.
[0296] (Appendix 10) A dialogue method executed by a dialogue system, comprising: an intention estimation step of estimating an intention from speech information or text information using an estimation model; a product selection step of selecting a product corresponding to the voice information or text information by using a product selection model when the intention is to select a product; a first text generation step of generating a first text based on the product selection result; an output step of outputting the first text; A method of interaction comprising:
[0297] (Appendix 11) For dialogue systems, an intention estimation step of estimating an intention from speech information or text information using an estimation model; a product selection step of selecting a product corresponding to the voice information or text information by using a product selection model when the intention is to select a product; a first text generation step of generating a first text based on the product selection result; an output step of outputting the first text; An interactive program for executing an interactive method comprising the steps of:
[0298] (Appendix 101) an intention estimation unit that estimates an intention from speech information or text information using an estimation model; a slot extraction unit that extracts slot information corresponding to the procedure from the speech information or text information using an extraction model when the intention is a procedure; an answer acquisition unit that acquires answer information corresponding to the question when the intention is a question; a first text generator that generates a first text based on at least one of the slot information and the answer information; a second text generator that generates a second text based on the slot information; an output unit that outputs the first text and the second text; A dialogue system comprising:
[0299] (Appendix 102) The first text is generated according to a first preset response rule. 102. The dialogue system of claim 101.
[0300] (Appendix 103) The second text is generated according to a second preset response rule. 102. The dialogue system of claim 101.
[0301] (Appendix 104) The answer information is set in advance. 102. The dialogue system of claim 101.
[0302] (Appendix 105) The estimation model is a natural language processing model. 102. The dialogue system of claim 101.
[0303] (Appendix 106) The extraction model is a natural language processing model. 102. The dialogue system of claim 101.
[0304] (Appendix 107) The voice information is further converted into text information by a text conversion unit. 102. The dialogue system of claim 101.
[0305] (Appendix 108) The procedures include reservations, orders, purchases, applications, or changes or cancellations thereof. 102. The dialogue system of claim 101.
[0306] (Appendix 109) The procedure involves presetting one or more slot information. 102. The dialogue system of claim 101.
[0307] (Appendix 110) The output unit outputs audio information corresponding to the first text and the second text. 102. The dialogue system of claim 101.
[0308] (Appendix 111) The output unit outputs text information corresponding to the first text and the second text. 102. The dialogue system of claim 101.
[0309] (Appendix 112) a procedure unit for executing the procedure 102. The dialogue system of claim 101.
[0310] (Appendix 113) The system further includes a notification unit that notifies the contents of the procedure executed by the procedure unit. 102. The dialogue system of claim 101.
[0311] (Appendix 114) A dialogue method executed by a dialogue system, comprising: an intention estimation step of estimating the intention of the speech information or the text information using the estimation model; a slot extraction step of extracting slot information corresponding to the procedure from the speech information or text information using an extraction model if the intention is a procedure; an answer acquisition step of acquiring answer information corresponding to the question when the intention is a question; a first text generating step of generating a first text based on at least one of the slot information and the answer information; a second text generating step of generating second text based on the slot information; an output step of outputting the first text and the second text; A method of interaction comprising:
[0312] (Appendix 115) For dialogue systems, an intention estimation step of estimating the intention of the speech information or the text information using the estimation model; a slot extraction step of extracting slot information corresponding to the procedure from the speech information or text information using an extraction model if the intention is a procedure; an answer acquisition step of acquiring answer information corresponding to the question when the intention is a question; a first text generating step of generating a first text based on at least one of the slot information and the answer information; a second text generating step of generating second text based on the slot information; an output step of outputting the first text and the second text; An interactive program for executing an interactive method comprising the steps of:
[0313] The embodiments disclosed herein are illustrative in all respects and should not be considered limiting. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims. Furthermore, the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of symbols]
[0314] 1: Interactive device 2: User terminal 11: Communications Department 12: Storage part 13: Control unit 121: Audio information 122:Text information 123: Estimation model 124: Procedural information 125: Extraction model 126: Target information 127: Q&A information 128: First response rule 129: Second response rule 130: Procedure information 1201: Product selection model 1202:Product information 131: Text conversion section 132: Intention estimation unit 133: Slot extraction unit 134: Answer acquisition part 135: First text generation unit 136: Second text generation unit 137: Output section 138: Procedure Division 139: Notification Department 140: Product selection section
Claims
1. an intention estimation unit that estimates an intention from speech information or text information using an estimation model; a product selection unit that selects a product corresponding to the voice information or text information by using a product selection model when the intention is to select a product; a first text generation unit that generates a first text based on a product selection result; an output unit that outputs the first text; A dialogue system comprising:
2. When a plurality of products corresponding to the audio information or the text information are selected, the first text generation unit generates a first text indicating the plurality of products. The dialogue system according to claim 1 .
3. When one product corresponding to the audio information or the text information is selected, the first text generation unit generates a first text indicating the one product. The dialogue system according to claim 1 .
4. The product selection unit refers to product information in which information about a plurality of products is registered, and selects a product corresponding to the text information. The dialogue system according to claim 1 .
5. The product selection model is a natural language processing model. The dialogue system according to claim 1 .
6. The first text is generated according to a first preset response rule. The dialogue system according to claim 1 .
7. The system further includes a text conversion unit that converts audio information into the text information. The dialogue system according to claim 1 .
8. The output unit outputs audio information corresponding to the first text. The dialogue system according to claim 1 .
9. The output unit outputs text information corresponding to the first text. The dialogue system according to claim 1 .
10. A dialogue method executed by a dialogue system, comprising: an intention estimation step of estimating an intention from speech information or text information using an estimation model; a product selection step of selecting a product corresponding to the voice information or text information by using a product selection model when the intention is to select a product; a first text generation step of generating a first text based on the product selection result; an output step of outputting the first text; A method of interaction comprising:
11. For dialogue systems, an intention estimation step of estimating an intention from speech information or text information using an estimation model; a product selection step of selecting a product corresponding to the voice information or text information by using a product selection model when the intention is to select a product; a first text generation step of generating a first text based on the product selection result; an output step of outputting the first text; An interactive program for executing an interactive method comprising the steps of:
Citation Information
Patent Citations
Voice recognition control system, voice recognition control method, and voice recognition control program
JP2013007917A
Search device and program
JP2020197945A
Co-clustering system, method, and program
WO2017159402A1
Information processing device, information processing method, and program
WO2017195388A1
Method of generating response using utterance and apparatus therefor
JP2023073220A
Cited By
Automatic voice response method, automatic voice response program, and automatic voice response device
JP7862054B1