Information processing system, information processing method, and program

The information processing system addresses the issue of misinterpreted user intentions and feelings in dialogue systems by determining and presenting required confirmation items, thereby reducing user discomfort and improving interaction accuracy.

JP2025073476AActive Publication Date: 2025-05-13STARLEY CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2023184318
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2025-05-13
Estimated Expiration
2043-10-26

AI Technical Summary

Technical Problem

Existing dialogue systems often fail to accurately convey the user's intentions and feelings, leading to discomfort in subsequent dialogue interactions.

Method used

An information processing system that acquires user statement data, determines required confirmation items based on relationship information between general speech characteristics and necessary confirmation items, and presents these items to the user to ensure accurate understanding and response.

Benefits of technology

This approach reduces discomfort in dialogue by ensuring that the user's intentions and feelings are properly conveyed, leading to more accurate and user-friendly interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025073476000001_ABST
    Figure 2025073476000001_ABST
Patent Text Reader

Abstract

To provide an information processing system or the like that can reduce awkwardness in a dialogue.SOLUTION: According to one aspect of the present invention, an information processing system including a processor is provided. In this information processing system, the processor acquires, in an acquisition step, user utterance data indicating a user's utterance in a dialogue. In a determination step, the processor determines a confirmation requiring matter in a dialogue in which an utterance indicated by the acquired user utterance data was made based on relationship information indicating a relationship between a characteristic of general utterance data indicating an utterance in a general dialogue and a confirmation requiring matter to be confirmed in the dialogue. In a presentation step, the determined confirmation requiring matter is presented to the user. In a reply step, a reply to the user is output based on the content of the utterance indicated by the acquired user utterance data. If there is a reply to the presented confirmation requiring matter, a reply reflecting the content of the reply is output.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] Patent Document 1 discloses a technology that performs morphological analysis on an input sentence, extracts words whose parts of speech in the morphologically analyzed sentence satisfy certain conditions, calculates the sum of co-occurrences with other words for each extracted word, determines words whose sum of co-occurrences is lower than a predetermined threshold value as being erroneous input words, and creates a correction sentence if any word is determined to be erroneous. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2008-180801 A Summary of the Invention [Problem to be solved by the invention]

[0004] There are known dialogue functions that automatically respond to user comments. When a user is interacting with such a dialogue function, if the user's intention or feelings are not properly conveyed, the content of the subsequent dialogue may differ from what the user wants, causing discomfort to the user.

[0005] In view of the above circumstances, the present invention provides an information processing system and the like that can reduce the sense of awkwardness felt during dialogue. [Means for solving the problem]

[0006] According to one aspect of the present invention, an information processing system including a processor is provided. In this information processing system, the processor acquires user utterance data indicating user utterances in a dialogue in an acquisition step. In a determination step, the processor determines matters requiring confirmation in a dialogue in which a utterance indicated by the acquired user utterance data was made based on characteristics of general utterance data indicating utterances in a general dialogue and relationship information indicating a relationship between the utterances in the dialogue and matters requiring confirmation to be confirmed. In a presentation step, the determined matters requiring confirmation are presented to the user. In a reply step, a reply to the user is output based on the content of the utterance indicated by the acquired user utterance data. If there is a reply to the presented matters requiring confirmation, a reply reflecting the content of the reply is output.

[0007] According to this embodiment, it is possible to reduce the sense of awkwardness felt during a conversation. [Brief description of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram showing an example of the overall configuration of an automated dialogue system 1. [Diagram 2] 2 is a diagram illustrating an example of a hardware configuration of a server device 10. FIG. [Diagram 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of a user terminal 20. [Figure 4] FIG. 1 is an activity diagram showing an example of an interaction process. [Diagram 5] FIG. 13 is a diagram showing an example of a displayed automatic dialogue screen. [Figure 6] FIG. 13 is a diagram illustrating an example of a dialogue. [Figure 7] FIG. 13 illustrates an example of a confirmation table. [Figure 8] FIG. 11 is a diagram showing an example of output confirmation information. [Figure 9] FIG. 13 is a diagram showing an example of a correction input field. [Figure 10] FIG. 13 is a diagram showing an example of an output reply. [Figure 11] FIG. 13 is a diagram showing another example of an output reply. [Figure 12]FIG. 11 is a diagram showing another example of output confirmation information. [Figure 13] FIG. 11 is a diagram showing another example of a correction input field. [Figure 14] FIG. 13 is a diagram showing another example of an output reply. [Figure 15] FIG. 11 is a diagram showing an example of output confirmation information. [Figure 16] FIG. 13 is a diagram showing an example of an operation image for correction. [Figure 17] FIG. 13 is a diagram showing an example of an output reply. [Figure 18] FIG. 13 is a diagram showing an example of a mode table. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described with reference to the drawings. Various characteristic features shown in the following embodiments can be combined with each other.

[0010] Incidentally, the program for realizing the software appearing in this embodiment may be provided as a non-transitory computer-readable recording medium, or may be provided so as to be downloadable from an external server, or may be provided so that the program is started on an external computer and its functions are realized on a client terminal (so-called cloud computing).

[0011] In addition, in this embodiment, the term "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In addition, in this embodiment, various information is handled, and this information is represented, for example, by physical values ​​of signal values ​​representing voltage and current, high and low signal values ​​as a binary bit group consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculation can be performed on the circuit in the broad sense.

[0012] In addition, a circuit in the broad sense is a circuit realized by at least appropriately combining a circuit, circuitry, a processor, a memory, etc. In other words, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0013] <Embodiment 1> 1. System Configuration The system configuration according to the first embodiment will be described below. Fig. 1 is a diagram showing an example of the overall configuration of an automated dialogue system 1. Fig. 1 shows an overview of each device provided in the automated dialogue system 1 and a user who uses those devices. Each overview will be explained from time to time with reference to other figures. The automated dialogue system 1 is an information processing system that executes dialogue processing in which a machine automatically dialogues with a user without user operation.

[0014] The automated dialogue system 1 includes a communication line 2, a server device 10, and a user terminal 20. The communication line 2 is not particularly limited, but may be configured, for example, by the Internet network. The communication line 2 may also include a local area network, a mobile communication network, and a VPN (Virtual Private Network), etc. The communication line 2 mediates data exchange between devices connected to the line. In the example of FIG. 1, the server device 10 is connected to the communication line 2 by wire, and the user terminal 20 is connected wirelessly. The connection of each device to the communication line 2 may be wired or wireless.

[0015] The server device 10 is an information processing device that executes dialogue processing. The server device 10 includes an AI module 3. The AI ​​module 3 is a module that is adjusted (tuned) using AI (Artificial Intelligence) technology to realize a dialogue function that, for example, when a statement is input from a person, outputs a reply corresponding to the statement. The AI ​​module 3 has a natural language processing model with improved accuracy by machine learning using a large-scale data set called LLM (Large Language Models).

[0016] By performing machine learning with LLM, the AI ​​module 3 can realize a variety of dialogues. The dialogue function realized by the AI ​​module 3 does not simply answer questions from the user, but can, for example, provide backchannel responses to user comments, develop the current topic, change the topic, provide new topics, and touch on past topics.

[0017] In addition to the dialogue function, the AI ​​module 3 is also adjusted to realize a task execution function for executing a specific task instructed by a user in dialogue, as is realized in a smart speaker, etc. The specific task is, for example, searching the Internet, writing a document, sending an email, and using a service (such as a reservation service) provided on a specific site. The AI ​​module 3 can be adjusted to realize functions other than the dialogue function and the task execution function.

[0018] The user terminal 20 is a terminal used by a user, such as a smartphone, a tablet terminal, or a personal computer. The user terminal 20 performs operations such as displaying images in dialogue processing and accepting operations by the user. More specifically, the user terminal 20 accepts input of voice or text indicating a statement by the user. The user terminal 20 also outputs a response from the server device 10 in the form of voice or text.

[0019] 2. Hardware Configuration The hardware configuration according to the first embodiment will be described below. 2 is a diagram showing an example of a hardware configuration of server device 10. Server device 10 includes a control unit 11, a storage unit 12, a communication unit 13, and a bus 14. Bus 14 electrically connects each unit included in server device 10.

[0020] (Control unit 11) The control unit 11 has at least one processor. The at least one processor may be configured, for example, by a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), one or more integrated circuits, one or more discrete circuits, or a combination of these (not shown).

[0021] The control unit 11 is a computer that realizes various functions related to the automated dialogue system 1 by reading out a predetermined program stored in the storage unit 12. That is, information processing by software stored in the storage unit 12 is specifically realized by the control unit 11, which is an example of hardware, and can be executed as each functional unit included in the control unit 11. Note that the control unit 11 is not limited to being single, and may be implemented with multiple control units 11 for each function. Also, a combination of these may be used.

[0022] (Storage unit 12) The storage unit 12 stores various pieces of information defined by the above description. This can be implemented, for example, as a storage device such as a solid state drive (SSD) or a hard disk drive (HDD) that stores various programs and the like related to the automated dialogue system 1 executed by the control unit 11, or as a memory such as a random access memory (RAM) that stores temporarily required information (arguments, arrays, etc.) related to the program calculations. The storage unit 12 stores various programs, variables, etc. related to the automated dialogue system 1 executed by the control unit 11.

[0023] (Communications Department 13) The communication unit 13 is configured by a communication module. The communication module may be a wireless communication module conforming to standards such as IEEE802.11a / b / g / n / ac / ax, LTE, 5G, and 6G, or may be a wired communication module conforming to standards such as IEEE802.3. The communication unit 13 is configured to be capable of transmitting various electrical signals from the server device 10 to external components. The communication unit 13 is also configured to be capable of receiving various electrical signals from the external components to the server device 10. More preferably, the communication unit 13 has a network communication function, and may be implemented so that various information can be communicated between the server device 10 and an external device via the communication line 2.

[0024] Fig. 3 is a diagram showing an example of a hardware configuration of the user terminal 20. The user terminal 20 includes a control unit 21, a storage unit 22, a communication unit 23, an input unit 24, an output unit 25, and a bus 26. The bus 26 electrically connects the various units included in the user terminal 20. The control unit 21, the storage unit 22, and the communication unit 23 are similar hardware to the control unit 11, the storage unit 12, and the communication unit 13 shown in Fig. 2, although the specifications, model, etc. may be different.

[0025] (Input section 24) The input unit 24 has keys, buttons, a touch screen, a mouse, etc., and receives input from the user. The input unit 24 also has a microphone, and receives voice input from the user.

[0026] (Output section 25) The output unit 25 has a display, a speaker, and the like, and displays visual information generated in a manner that is visible to the user, such as screens, images, icons, text, and the like, on the display surface of the display, and outputs sounds including voice.

[0027] 3. Information Processing The above-mentioned dialogue processing will be described below as the information processing according to the embodiment. In the following description, the server device 10 and the user terminal 20 are described as the subjects of each dialogue processing, but the information processing is executed by a processor possessed by the control unit of each device. Among the dialogue processing, the AI ​​module 3 is used for the process of generating a reply from a user's utterance, but for other dialogue processing, the AI ​​module 3 may be used or other modules (such as a display module) may be used.

[0028] Fig. 4 is an activity diagram showing an example of dialogue processing. The dialogue processing shown in Fig. 4 is started when the user of the user terminal 20 performs an operation to display a screen for providing the automatic dialogue service provided by the server device 10. First, the server device 10 generates screen data showing an automatic dialogue screen at A11 and transmits the generated screen data to the user terminal 20. The user terminal 20 displays the automatic dialogue screen shown by the transmitted screen data at A12.

[0029] Fig. 5 is a diagram showing an example of a displayed automatic dialogue screen. In the example of Fig. 5, the server device 10 displays an automatic dialogue screen C1 on the display screen of the user terminal 20. A dialogue partner image F11 and a dialogue end button B11 are displayed on the automatic dialogue screen C1. The dialogue end button B11 is an operation image that is operated when ending the dialogue.

[0030] The dialogue partner image F11 is an image that evokes an image of a dialogue partner, and is displayed during a dialogue. The dialogue partner imaged by the dialogue partner image F11 is hereinafter also referred to as a "dialogue AI." Note that the dialogue partner image F11 is an image that imitates a human face, but an image of an animal or a figure may also be used as the dialogue partner image.

[0031] Next, in A13, the server device 10 judges whether the user who has started using the automated dialogue service is a user who is having a dialogue using the automated dialogue service for the first time. If the server device 10 judges in A13 that this is not the first dialogue (NO), the server device 10 proceeds to A25, which will be described later. A25 will be described later. If the server device 10 judges in A13 that this is the first dialogue (YES), the server device 10 generates a response for the initial dialogue and transmits it to the user terminal 20 in A14.

[0032] In A15, the user terminal 20 outputs the transmitted response for the initial dialogue. The response may be output as voice or text. In the example of FIG. 5, the user terminal 20 outputs voice V11 saying, "Nice to meet you. My name is Tomodachi." as a response for the initial dialogue. Voice V11 represents a statement by the dialogue AI to introduce itself. Having heard voice V11, the user makes a self-introduction statement including his or her own name to the dialogue AI, for example.

[0033] FIG. 6 is a diagram showing an example of a dialogue. In the example of FIG. 6, the user utters a voice V12 indicating a statement, "Nice to meet you. My name is Tom." The user terminal 20 accepts this voice V12 as an input of a user's utterance in A21, and transmits voice data indicating the input utterance to the server device 10. The server device 10 acquires the transmitted voice data in A22 as user utterance data indicating the user's utterance to the dialogue AI.

[0034] Next, the server device 10 determines the items to be checked in A23 based on the acquired user utterance data. The items to be checked are items that are predetermined in the automated dialogue system 1 as items to be checked in the dialogue. There are two main methods for determining the items to be checked. The first method is to use the AI ​​module 3 shown in FIG. 1, and the second method is to use a check table to be described later.

[0035] The first method uses an AI module 3 that is adjusted to generate a learning model trained on a data set containing a large amount of dialogue exchanges and output items requiring confirmation as a judgment result when user utterance data is input based on the generated learning model. The first method will be explained in more detail later. Meanwhile, the second method uses a confirmation table that defines what items requiring confirmation should be judged when what type of user utterance data is acquired.

[0036] FIG. 7 is a diagram showing an example of a confirmation table. In the confirmation table TB1 shown in FIG. 7, the type of matter requiring confirmation, the matter requiring confirmation, and the characteristics of general utterance data are associated with each other. The "general utterance data" shown in the confirmation table TB1 is not limited to data indicating utterances in a dialogue between a user and a dialogue AI, but is data indicating utterances collected from general dialogues. The types of matters requiring confirmation include "confirmation of the user's intention" and "confirmation of the user's feelings."

[0037] The "user's intention" here includes, for example, the language that the user wants the dialogue AI to use in the dialogue. This language is, for example, the way the dialogue AI wants the user to call the user, or the way the dialogue AI wants the user to talk about personal information (such as occupation, age, or hometown) of the user. In addition, when the AI ​​module 3 realizes the above-mentioned task execution function, the content of the task instructions given by the user in the dialogue with the dialogue AI is included in the "user's intention."

[0038] If the "user's intent" is misinterpreted during a conversation, for example, the robot may address the user by a name that the user does not want to be called, or may ask for the user's personal information in a way that the user does not want to be said (for example, asking the user to say "I'm from Kanto" instead of "I'm from ~ prefecture"), or may execute a task different from what the user instructed, causing the user to feel uncomfortable.

[0039] Also, if the "user's feelings" are misunderstood in a dialogue, even if the dialogue function responds in accordance with the feelings, it may respond in a way that does not match the user's feelings, such as "It's fun" when the user is feeling depressed, or "It's hard" when the user is feeling happy, which may cause the user to feel uncomfortable. In this way, the "user's intention" and the "user's feelings" are matters that must be confirmed in order to continue the dialogue smoothly.

[0040] Items to be confirmed as "user intent," i.e., items to be checked, include, for example, the above-mentioned "name" of the user, the user's "personal information," "task content," and "task deadline." "Name" is an item to be checked to confirm how the user wants to be called, and "personal information" is an item to be checked to confirm how the user wants to express their personal information. "Task content" and "task deadline" are items to be checked to confirm whether the task is being carried out as instructed by the user.

[0041] Furthermore, items to be checked as "user's feelings", i.e., items to be checked, include "fatigue", "stress", "emotions (joy, anger, sadness, and happiness)" and "activity level". "Fatigue" is an item to be checked to check whether the user is currently in a tired state, and "stress" is an item to be checked to check whether the user is currently in a stressed state. Furthermore, "emotions (joy, anger, sadness, and happiness)" is an item to be checked to check the user's current emotions. Note that these items to be checked regarding "intentions" and "feelings" are merely examples, and are not limited to these.

[0042] Next, the characteristics of the general utterance data associated with each confirmation item will be described. The characteristics of the general utterance data include "context", "phrase", and "speech feature". "Context" and "phrase" refer to the characteristics of the context in a sentence forming a dialogue including the utterance indicated by the general utterance data, and the characteristics of the phrases included in the user's utterance in the dialogue.

[0043] For example, a sentence containing phrases used when meeting someone for the first time, such as "Nice to meet you" or "My name is," has the context of "self-introduction." A sentence containing phrases indicating instructions, such as "please do...," "please do...," or "please...," has the context of "task instructions." A sentence containing work-related phrases, such as "commute," "business trip," or "overtime," has the context of "work-related." A sentence containing phrases used to address a third party, such as "Mr. / Ms. / Mr. / Mr. / Director," has the context of "human relationships." In this way, the context can be determined based on the phrases contained in the sentence.

[0044] Furthermore, the type of phrase included in the user's utterance can be determined by, for example, having the server device 10 store dictionary data that associates each phrase with its type.

[0045] In confirmation table TB1, the item to be confirmed, "how to address," is associated with the context of "self-introduction" and the phrase "person's name" as features of general utterance data. The item to be confirmed, "personal information," is associated with the context of "self-introduction" and the phrases, "occupation," "generation," and "place of birth," as features of general utterance data. The item to be confirmed, "task content," is associated with the context of "task instructions" and the phrases, "search," "send," and "reserve," as features of general utterance data. The item to be confirmed, "task deadline," is associated with the context of "task instructions" and the phrases, "date," and "time," as features of general utterance data.

[0046] In A23, the server device 10 determines the context based on the characteristics of the utterance by the dialogue AI indicated by the voice V11 shown in FIG. 5 and the utterance by the user indicated by the voice V12 shown in FIG. 6. In this case, since both contain the phrase "Nice to meet you," the server device 10 determines that the user utterance data acquired in A22 has a characteristic indicating the context of "self-introduction." In this way, the server device 10 determines the context by including not only the utterance by the user indicated by the user utterance data acquired in A22 but also the utterance by the dialogue AI. Note that, when the dialogue progresses, the server device 10 may determine the context by including both past utterances (including utterances two or more times ago), such as the utterance by the user one time ago, the utterance by the dialogue AI one time ago, and the utterance by the user two times ago.

[0047] Next, since the user's utterance indicated by the voice V12 includes the name "Tom," the server device 10 determines that the user's utterance data acquired in A22 has a feature indicating the phrase "person's name." The server device 10 then determines that there is an item to be confirmed, "name," which is associated in the confirmation table TB1 with the context "self-introduction" and the phrase "person's name." If there is no feature indicated in the confirmation table TB1, the server device 10 determines that there is no item to be confirmed.

[0048] In A24, the server device 10 judges whether or not the judgment in A23 was that there is an item to be checked. If the judgment in A24 was that there is no item to be checked (NO), the server device 10 generates a response to the utterance indicated by the user utterance data acquired in A22, and transmits the generated response to the user terminal 20.

[0049] Note that in A25, the server device 10 generates a response to the user even if it determines in A13 that this is not the first dialogue (NO). In this case, the server device 10 may generate a response based on the previous dialogue, or may generate a response that is often used when resuming a dialogue.

[0050] In A26, the user terminal 20 outputs the transmitted reply. The reply may be output by voice or text, as in A15. After A26, the process returns to A21, where the user's utterance is input again.

[0051] If the server device 10 determines in A24 that there is an item to be checked (YES), then in A31 it generates confirmation information for checking the item to be checked determined in A23, and transmits the generated confirmation information to the user terminal 20. In A32, the user terminal 20 outputs the transmitted confirmation information.

[0052] FIG. 8 is a diagram showing an example of the output confirmation information. In the example of FIG. 8, the user terminal 20 displays a confirmation item D11 called "Tom" and emits a voice V13 indicating a statement "Is it okay to call you that?" The confirmation item D11 and the voice V13 are both text and voice indicated by the confirmation information. When displaying the confirmation item D11, the user terminal 20 also displays a correction button B22. The user inputs a response to the output confirmation information. The user terminal 20 accepts the user's response to the confirmation item at A33.

[0053] If the user is happy with the name displayed in the confirmation item D11, the user inputs a response by replying with an affirmative response such as "Yes" or "Okay." In this case, the user terminal 20 accepts the affirmative response input by voice and transmits answer data indicating the response to the server device 10.

[0054] Furthermore, if the user wishes to change the name displayed in the item to be confirmed D11, the user performs an operation of pressing the edit button B22. When the edit button B22 is pressed, the user terminal 20 displays an input field for correcting the item to be confirmed D11. Note that, instead of pressing the edit button B22, an operation of tapping the dialogue partner image F11 or an operation of tapping anywhere on the screen may be used as the operation of correction.

[0055] FIG. 9 is a diagram showing an example of a correction input field. In the example of FIG. 9, the user terminal 20 displays a name correction input field E23, a software keyboard B23 for inputting a name, and a confirmation button B24. The user operates the software keyboard B23 to input a name that the user wants the dialogue AI to call them, and when the user inputs the name in the correction input field E23, the user presses the confirmation button B24 to confirm the answer. When the user terminal 20 accepts the name thus input as the user's answer to the item to be checked, it transmits answer data indicating that the item to be checked has been corrected to the server device 10.

[0056] The user terminal 20 may accept a voice input operation such as "No, it's not Tom, it's Tommy" as an input operation of the name that the user wants to correct. In that case, the display of the correction input field E23 etc. may not be necessary. However, since voice input may fail, the user terminal 20 may be configured such that, when voice input has failed a predetermined number of times, it does not accept any further voice input and displays the correction input field E23 etc. and accepts only input operations therein.

[0057] When the answer data is transmitted from the user terminal 20, the server device 10 acquires the answer data in A34. Next, the server device 10 generates a response reflecting the answer indicated by the acquired answer data in A35. If the answer data indicates a positive answer to the matter to be confirmed, the server device 10 has determined that the answer data correctly captures the user's intention or feelings, and therefore generates a response reflecting that the matter to be confirmed was an affirmative answer. The server device 10 transmits the generated response to the user terminal 20, and the user terminal 20 outputs the transmitted response in A36.

[0058] Fig. 10 is a diagram showing an example of an output reply. In the example of Fig. 10, the user terminal 20 emits voice V14 indicating a statement, "Then I'll call you Tom!" Voice V14 indicates a reply reflecting an answer that it is okay to call the item to be confirmed D11, "Tom." When the reply data indicates a correction to the item to be confirmed, the server device 10 has found that the item to be confirmed differs from the user's intention or feeling, and therefore generates a reply reflecting the corrected item to be confirmed.

[0059] Fig. 11 is a diagram showing another example of an output reply. In the example of Fig. 11, the user terminal 20 emits voice V15 indicating a statement, "Okay, then I'll call you Tommy from now on!" Voice V15 indicates a reply that reflects the user's intention to be called "Tommy" as entered in the input field of Fig. 9, rather than "Tom", which is the item to be confirmed D11. After A36, the process returns to A21, and the user's statement is input again.

[0060] In Fig. 8 and other figures, a case where "name" is an item that needs to be checked has been described. Next, a case where "task content" and "task deadline" are items that need to be checked will be described. In A23, for example, if the user's voice contains phrases such as "please do...", "by..." or "please...", the server device 10 determines that the user utterance data acquired in A22 has a feature indicating the context of "task instruction".

[0061] Next, if the user's utterance contains, for example, a word "reservation," the server device 10 determines that the user's utterance data acquired in A22 has the characteristic of the word "reservation." Also, if the user's utterance contains a word indicating "date," the server device 10 determines that the user's utterance data acquired in A22 has the characteristic of the word indicating "date."

[0062] Then, the server device 10 determines that there is an item to be confirmed, "task content," which is associated with the context "task instruction" and the word "reservation" in the confirmation table TB1. The server device 10 also determines that there is an item to be confirmed, "task deadline," which is associated with the context "task instruction" and the word "date" in the confirmation table TB1. After this, the processes of A31 and A32 are executed, and confirmation information for confirming the items to be confirmed is output.

[0063] Fig. 12 is a diagram showing another example of the output confirmation information. The example in Fig. 12 shows confirmation information output when a user utters a voice V31 saying "Please make a reservation at the ADC Hotel on November 15th" and the server device 10 mistakenly recognizes "ADC Hotel" as "ABC Hotel". In this case, the user terminal 20 displays, as the "task content", a confirmation item D31 of "ABC Hotel" recognized as the reservation target, as the "task deadline", a confirmation item D32 of "November 15th" which is the date by which the reservation should be made, and a correction button B32 on the user terminal 20.

[0064] Furthermore, the user terminal 20 emits a voice V32 saying, "Is this reservation correct?" Since the displayed confirmation item D31 (ABC Hotel) differs from the user's instruction (ADC Hotel), the user operates the correction button B32 while selecting the confirmation item D31 that the user wants to correct (selection is shown in FIG. 13 by underlining the confirmation item D31), thereby displaying an input field for correcting the confirmation item.

[0065] Fig. 13 is a diagram showing another example of a correction input field. In the example of Fig. 13, the user terminal 20 displays a reservation input field E33, a software keyboard B33 for inputting a name, and a confirmation button B34. When the user terminal 20 accepts the input reservation correction as a user's answer to the confirmation item, it transmits answer data indicating that the confirmation item has been corrected to the server device 10. Since the answer data indicates a correction to the confirmation item, the server device 10 generates a reply reflecting the corrected confirmation item and outputs the reply to the user terminal 20.

[0066] Fig. 14 is a diagram showing another example of an output reply. In the example of Fig. 14, the user terminal 20 emits voice V33 indicating a statement, "Excuse me. I see that it is the ADC Hotel and not the ABC Hotel. I will make a reservation there then." Voice V33 indicates a reply that reflects a reply that correctly indicates the user's intention to make a reservation at the "ADC Hotel" entered in the input field of Fig. 13, since the confirmation item D31 indicating that the reservation target is the "ABC Hotel" is different from the user's intention.

[0067] Next, a case where the type of the item to be checked is "user's feelings" will be described. For example, for "fatigue" and "stress", the server device 10 determines the item to be checked by using the words and phrases contained in the context and the utterance as the features of the user utterance data, in the same manner as the determination described with reference to FIG. 8 etc. Also, for "emotions (joy, anger, sadness, and happiness)" and "activity level", the server device 10 determines the item to be checked by using the features of the voice as the features of the user utterance data.

[0068] In A23, the server device 10 uses well-known technology to determine the user's emotion and activity level as matters requiring confirmation from the feature quantities such as the pitch, frequency components, and volume of the user's voice indicated by the acquired user utterance data. The user's emotion is, for example, joy, anger, sorrow, and happiness, but is not limited thereto and may include positive, negative, satisfaction, dissatisfaction, normal, calm, surprise, fear, and other emotions.

[0069] When the server device 10 determines the items to be checked in A23, the server device 10 generates confirmation information for checking the determined items to be checked in A31, and transmits the generated confirmation information to the user terminal 20. The user terminal 20 outputs the transmitted confirmation information in A32.

[0070] Fig. 15 is a diagram showing an example of output confirmation information. In the example of Fig. 15, the emotion "fun" is determined based on user utterance data of voice V41 saying "Good morning!" uttered by the user, and the user terminal 20 displays a check item D41 saying "Feeling happy" and issues voice V42 indicating a reply saying "Good morning. Did something fun happen?" The check item D41 is text indicated by the confirmation information, and the voice V42 is voice indicated by the confirmation information.

[0071] The user terminal 20 displays a correction button B42 together with the item to be checked D41. If the user thinks that the emotion displayed in the item to be checked D41 is different from the user's own emotion, the user performs an operation of pressing the correction button B42. When the correction button B42 is operated, the user terminal 20 displays an operation image for correcting the item to be checked D41.

[0072] FIG. 16 is a diagram showing an example of an operation image for correction. In the example of FIG. 16, the user terminal 20 displays a plurality of selection buttons B43 showing options for emotions, and a confirmation button B44. The selection buttons B43 show character strings "joy", "anger", "sadness", "fun", "energetic", and "depressed". The user operates the selection button B43 that matches his / her emotion, and presses the confirmation button B44 to confirm the answer. Note that these selection buttons B43 are merely examples, and selection buttons showing other emotions may be displayed. Also, a software keyboard as shown in FIG. 9 may be displayed to allow the user to input text showing the emotion.

[0073] The user terminal 20 transmits answer data indicating that the item to be confirmed has been corrected to the server device 10. In A35, the server device 10 generates a response that reflects the answer indicated by the acquired answer data, and transmits the generated response to the user terminal 20. In A36, the user terminal 20 outputs the transmitted response.

[0074] Fig. 17 is a diagram showing an example of an output reply. In the example of Fig. 17, it is assumed that the selection button B43 "feeling energetic" is operated on the screen of Fig. 16. In this case, the user terminal 20 emits a voice V43 indicating a statement "You certainly seem to be in good spirits. Let's keep it up and work hard." The voice V43 indicates a reply that reflects the user's feelings of wanting a dialogue that matches the level of activity of "feeling energetic," rather than the confirmation item D41 of "feeling happy."

[0075] The above is an explanation of the second method of determining items requiring confirmation (method using a confirmation table). Here, we will supplement the above-mentioned first method of determination, i.e., a method of determining items requiring confirmation using the AI ​​module 3. In the first method of determination, when user utterance data is input to the AI ​​module 3, the AI ​​module 3 outputs items requiring confirmation as a determination result based on a learning model. This learning model is trained using a large-scale data set including general conversations in which it is clear that the conversation partner has given a response that differs from the user's intention or feelings. This data set includes data showing utterances collected from general conversations, i.e., general utterance data similar to that described in FIG. 7.

[0076] Examples of such dialogues include dialogues in which the person requests the way to address himself or herself or the way to express personal information (such as occupation, age, or hometown) when introducing himself or herself, dialogues in which the task content or deadline is revised when a task is requested, or dialogues in which the person points out that the other person's response differs from the person's own feelings, which are dialogues in situations assumed in the confirmation table TB1. The more such dialogues that are learned from a dataset that includes more dialogues, the more accurate the AI ​​module 3 becomes in determining items requiring confirmation.

[0077] In addition, once the server device 10 obtains a response in A34, the server device 10 generates a response that reflects the response in the subsequent dialogue, even if the response is to a statement that is not determined to be a matter requiring confirmation. For example, in the example of Fig. 11, the response is to call the user "Tommy," so when the server device 10 calls the user in the subsequent dialogue, the server device 10 generates a response that reflects the response, such as "Tommy-san...", in both A25 and A35.

[0078] As described above, the server device 10 executes, for example, an acquisition step of acquiring user utterance data indicating a user's utterance in a dialogue at A22 shown in Fig. 4. The server device 10 also executes, for example, a determination step of judging an item requiring confirmation in a dialogue in which a utterance indicated by the user utterance data acquired at the acquisition step was made at A23 shown in Fig. 4. This determination is made based on relationship information indicating the relationship between the characteristics of general utterance data indicating an utterance in a dialogue and the item requiring confirmation to be confirmed in that dialogue. The confirmation table TB1 shown in Fig. 7 and the learning model possessed by the AI ​​module 3 are both examples of relationship information.

[0079] Furthermore, the server device 10 executes a presentation step in which the items to be checked determined in the determination step are presented to the user at A32 shown in Fig. 4. Items to be checked D11, D31, D32, and D41 shown in Fig. 8, Fig. 12, and Fig. 15 are each an example of an item to be checked presented to the user. Note that in the example of Fig. 8 etc., the server device 10 presents the items to be checked to the user by displaying the items to be checked, but this is not limiting and the items to be checked may be presented to the user by outputting them as voice.

[0080] In addition, the server device 10 executes a reply step in A25 of outputting a reply to the user based on the content of the utterance indicated by the user utterance data acquired in the acquisition step. Voice V14 shown in Fig. 10 is an example of a reply output in the reply step. If there is an answer to the item to be confirmed presented in the presentation step, the server device 10 executes a reply step in A35 of outputting a reply reflecting the content of the answer. Voices V15, V33, and V43 shown in Figs. 11, 14, and 17 are examples of replies reflecting the content of the answer.

[0081] According to this aspect, when a conversation occurs that includes a matter that needs to be confirmed, the user is asked to confirm the matter and a reply reflecting the answer is provided, thereby reducing the sense of awkwardness in the conversation compared to when the matter that needs to be confirmed is not confirmed.

[0082] In the determination step, for example, if a particular type of phrase is included in the utterance indicated by the user utterance data acquired in the acquisition step, the server device 10 determines whether or not it is appropriate to use the phrase in the dialogue as a matter requiring confirmation. Then, in the presentation step, the server device 10 presents the phrase to the user as a matter requiring confirmation. For example, in the examples of Figs. 6 and 8, when a particular type of phrase, "person's name," is included in the utterance, the person's name ("Tom") is presented to the user as a matter requiring confirmation. According to this embodiment, the user can correctly use the phrase he or she wants to use in the dialogue.

[0083] In addition, in a determination step, if the content of the utterance indicated by the user utterance data acquired in the acquisition step includes a specific task, the server device 10 determines whether or not it is correct to execute the task as a matter to be confirmed. Then, in a presentation step, the server device 10 presents the content of the task to the user as a matter to be confirmed.

[0084] 12, when a specific task of reserving a hotel is included in a comment, the task content and the task deadline are presented to the user as items to be confirmed. According to this embodiment, the content of the task is confirmed by the user before proceeding with the task, so that it is possible to prevent the task from proceeding with incorrect content compared to the case where the items to be confirmed are not presented to the user.

[0085] Furthermore, in the determination step, when the user's feelings are identified from the user utterance data acquired in the acquisition step, the server device 10 determines whether or not it is appropriate to have a dialogue that matches the identified feelings as a matter to be checked. Then, in the presentation step, the server device 10 presents the content of the feelings to the user as a matter to be checked. For example, in the example of FIG. 15, when the feeling of "fun" is identified from the voice features, the feeling is presented to the user as a matter to be checked. According to this embodiment, it is possible to continue a dialogue that matches the user's feelings.

[0086] <Variation: Reflecting answers in judgment> In the determination step, the server device 10 may reflect the answer to the confirmation item presented in the presentation step in the determination of the subsequent confirmation items. In the embodiment, the server device 10, for example, determines the user's emotion and activity level as confirmation items from the feature amounts such as the pitch, frequency components, and volume of the user's voice indicated by the user utterance data by using a well-known technique.

[0087] For example, suppose that the server device 10 determines that the item to be checked is "fun" in the case of a certain feature. This is because the feature generally represents the characteristics of a voice that is emitted when "fun". However, if the answer to the item to be checked is "sad", then for that user, the feature represents the characteristics of a voice that is emitted when "sad" rather than "fun". Therefore, when user utterance data showing a voice with the same feature is obtained for that user, the server device 10 determines that the item to be checked is "sad" rather than "fun".

[0088] Also, for example, the server device 10 determines that "fatigue" is a matter requiring confirmation because the context is "work-related" and includes a phrase indicating "busy," but the answer to the matter requiring confirmation is a response indicating the emotion of "stress." In this case, the server device 10 edits the confirmation table TB1 for that user so that the context "work-related" and the phrase "busy" are associated with the matter requiring confirmation of "stress."

[0089] In this case, when the server device 10 acquires user utterance data indicating utterances containing the same words and phrases in the same context for that user, it determines that the matter to be checked is "stress" rather than "fatigue." When making a determination using the AI ​​module 3, the answers to the matters to be checked can be added to the teacher data of the learning model of the AI ​​module 3 and learned, so that the answers to the matters to be checked can be reflected in the determination of the matters to be checked thereafter. According to this embodiment, it is possible to improve the accuracy of the determination of the matters to be checked compared to the case where these answers are not reflected.

[0090] <Variation: Not indicating items to be confirmed> In the presenting step, if the same determination result as the check-required item determined in the determining step has been determined a predetermined number of times or more in the past, the server device 10 may not present the check-required item for that determination result. The predetermined number of times is, for example, once, but may be two or more times. The smaller the predetermined number of times, the fewer the number of times the same check-required item is presented to the user.

[0091] When the server device 10 judges that there is an item to be checked in A24 (YES), it stores the number of times the same item to be checked has been checked in association with the user. When the server device 10 judges that there is an item to be checked in the next A24 (YES), it reads out the number of times the item to be checked has been checked, and if the number of times, including the current judgment, is equal to or greater than a predetermined number, it proceeds to A25 rather than A31 to generate a reply, that is, it only responds without presenting the item to be checked. According to this embodiment, it is possible to avoid checking the same item to be checked multiple times.

[0092] <Variation: Change in presentation mode> When the server device 10 determines that there is a check-required item related to the task instruction, the server device 10 may change the manner in which the check-required item is presented in accordance with the length of the time until the task is executed in the presentation step. The server device 10 uses, for example, a manner table in which the time until the task is executed is associated with the presentation manner of the check-required item.

[0093] Fig. 18 is a diagram showing an example of a mode table. In the example of Fig. 18, a mode table TB2 is shown in which the time periods until task execution, "less than Th1", "between Th1 and Th2", and "between Th2 and above", are associated with the presentation modes, "red characters + highlighted frame", "blue characters + normal frame", and "black characters + no frame". When the server device 10 determines in A23 that the "task deadline" of the task is an item requiring confirmation, in A31, the server device 10 calculates the time until the task deadline as the time until task execution.

[0094] The server device 10 generates, as confirmation information, information indicating the matters to be checked in a presentation format associated with the calculated period in the format table TB2. For example, when the period until the task is executed is "Th2 or more," the server device 10 simply presents the matters to be checked in black text as in the example shown in Fig. 8, etc., but when the period until the task is executed is "Th1 or more and less than Th2," the server device 10 presents an image in which the text indicating the matters to be checked is in blue text and surrounded by a frame, and when the period until the task is executed is "less than Th1," the server device 10 presents an image in which the text indicating the matters to be checked is in red text and surrounded by a highlighted frame.

[0095] According to such an embodiment, the more emphasized the item to be checked is, the shorter the period until the task is executed, so that the user who is presented with the item to be checked can intuitively understand the length of the period until the task is executed. Note that the embodiment shown in FIG. 18 is an example and is not limited to this. For example, when the item to be checked is presented to the user by voice, the volume of the voice may be increased as the period until the task is executed is shorter. In short, any embodiment may be used as long as the user can understand the correspondence between the length of the period until the task is executed and the presentation mode of the item to be checked.

[0096] <Configuration variations> The configuration (overall configuration, hardware configuration, functional configuration, etc.) shown in FIG. 1 and the like is an example, and other configurations may be used as long as there is no inconvenience in implementation. For example, the server device 10 may be distributed across two or more devices, and may be provided in the form of SaaS (Software as a Service) or a cloud computing system. Furthermore, the information processing performed by the server device 10 may be collectively executed by the user terminal 20. In short, as long as the necessary information processing is executed in the entire automated dialogue system 1, the devices that execute the information processing may have any configuration.

[0097] Furthermore, the artificial intelligence module (AI module 3) may be an internal or external component of the server device 10, or may be an internal or external component of the automated dialogue system 1. Furthermore, the function realized by one artificial intelligence module may be distributed and realized by two or more artificial intelligence modules, or the function realized by two or more artificial intelligence modules may be integrated and realized by one artificial intelligence module.

[0098] The output destination of information or data (hereinafter referred to as "information, etc.") may be another device, a display, a memory unit (including a built-in memory unit and an external memory unit), etc. Acquisition of information, etc. includes acquiring information, etc. generated by the device itself, in addition to acquiring information, etc. transmitted from another device. Tables, etc. (tables, databases, etc.) that associate parameters are not limited to the tables, etc. shown in the figures, and the number of parameters may be reduced or increased. Furthermore, information, etc. corresponding to parameters may be obtained by a formula, a conditional formula, etc., without using a table, etc.

[0099] The above-described aspects of the embodiment are information processing devices such as the server device 10 and the user terminal 20, and information processing systems such as the automated dialogue system 1 including the server device 10 and the user terminal 20, but may also be information processing methods. The information processing methods include the same steps as those executed by the information processing system. The above-described aspects of the embodiment may also be programs. The programs cause a computer to execute the same steps as those executed by the information processing system.

[0100] <Additional Notes> Furthermore, it may be provided in the following aspects:

[0101] (1) An information processing system having a processor, wherein the processor, in an acquisition step, acquires user utterance data indicating a user's utterance in a dialogue, in a determination step, determines the matter requiring confirmation in a dialogue in which a utterance indicated by the acquired user utterance data was made based on relationship information indicating a relationship between characteristics of general utterance data indicating utterances in a general dialogue and matters requiring confirmation to be confirmed in the dialogue, in a presentation step, presents the determined matter requiring confirmation to the user, and in a reply step, outputs a reply to the user based on the content of the utterance indicated by the acquired user utterance data, and if there is a reply to the presented matter requiring confirmation, outputs a reply reflecting the content of the reply.

[0102] According to this embodiment, it is possible to reduce the sense of awkwardness felt during a conversation.

[0103] (2) In the information processing system described in (1) above, in the determination step, when a specific type of phrase is included in the utterance indicated by the acquired user utterance data, the processor determines whether or not it is appropriate to use the phrase in the dialogue as the item requiring confirmation, and in the presentation step, presents the phrase to the user as the item requiring confirmation.

[0104] According to this embodiment, the user can correctly use the words that he or she wants to use in the dialogue.

[0105] (3) In the information processing system described in (1) or (2) above, in the determination step, when the content of the utterance indicated by the acquired user utterance data includes a specific task, the processor determines whether it is appropriate to execute the task as the item requiring confirmation, and in the presentation step, presents the content of the task to the user as the item requiring confirmation.

[0106] According to this aspect, it is possible to prevent the task from proceeding with erroneous content.

[0107] (4) In the information processing system described above in (3), in the presenting step, the processor changes the manner in which the items to be confirmed are presented depending on the length of time until the task is executed.

[0108] According to this embodiment, it is possible to intuitively grasp the length of time remaining until the execution of a task.

[0109] (5) In the information processing system described in any one of (1) to (4) above, in the determination step, when the user's feelings are identified from the acquired user utterance data, the processor determines whether it is appropriate to have a dialogue that matches the identified feelings as the item to be confirmed, and in the presentation step, presents the content of the feelings to the user as the item to be confirmed.

[0110] According to this embodiment, the user can continue the dialogue in accordance with his / her mood.

[0111] (6) In the information processing system according to any one of (1) to (5) above, in the judgment step, the processor reflects the presented answer to the item to be confirmed in a subsequent judgment of the item to be confirmed.

[0112] According to this aspect, it is possible to improve the accuracy of determining items requiring confirmation.

[0113] (7) An information processing system according to any one of (1) to (6) above, wherein in the presentation step, if a judgment result identical to the judged item requiring confirmation has occurred a predetermined number of times in the past, the processor does not present the item requiring confirmation for that judgment result.

[0114] According to this embodiment, it is possible to avoid having to confirm the same intention or sentiment multiple times.

[0115] (8) An information processing method comprising the steps of the information processing system according to any one of (1) to (7) above.

[0116] According to this embodiment, it is possible to reduce the sense of awkwardness felt during a conversation.

[0117] (9) A program for causing a computer to execute each step of the information processing system according to any one of (1) to (7) above.

[0118] According to this embodiment, it is possible to reduce the sense of awkwardness felt during a conversation. Of course, this is not the case. Furthermore, the above-described embodiments and modifications may be combined in any desired manner.

[0119] Finally, although various embodiments of the present invention have been described, these are presented as examples and are not intended to limit the scope of the invention. The new embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The embodiments and their modifications are within the scope and spirit of the invention, and are included in the scope of the invention and its equivalents described in the claims. [Explanation of symbols]

[0120] 1: Automated dialogue system 2: Communication lines 3: AI module 10: Server device 11: Control section 20: User terminal 21: Control section

Claims

1. An information processing system including a processor, The processor, In the acquisition step, user utterance data indicating utterances of the user in the dialogue is acquired; In the determination step, the matter requiring confirmation in the dialogue in which the utterance indicated by the acquired user utterance data was made is determined based on relationship information indicating a relationship between a feature of general utterance data indicating an utterance in a general dialogue and a matter requiring confirmation to be confirmed in the dialogue; In the presenting step, the determined confirmation-required item is presented to the user; In the response step, outputting a response to the user based on the content of the utterance indicated by the acquired user utterance data; If an answer is given to the presented items to be confirmed, a reply reflecting the content of the answer is output. Information processing system.

2. 2. The information processing system according to claim 1, The processor, In the determination step, when a specific type of phrase is included in the utterance indicated by the acquired user utterance data, it is determined whether or not it is appropriate to use the specific phrase in the dialogue as the item to be confirmed; In the presenting step, the phrase is presented to the user as the item to be confirmed. Information processing system.

3. 2. The information processing system according to claim 1, The processor, In the determination step, when a specific task is included in the content of the utterance indicated by the acquired user utterance data, it is determined whether or not it is appropriate to execute the task as the item to be confirmed; In the presenting step, contents of the task are presented to the user as the items to be confirmed. Information processing system.

4. 4. The information processing system according to claim 3, The processor, In the presenting step, a manner in which the confirmation-required item is presented is changed depending on a length of time until the execution of the task. Information processing system.

5. 2. The information processing system according to claim 1, The processor, In the determination step, when the user's feelings are identified from the acquired user utterance data, it is determined whether or not it is appropriate to have a dialogue that matches the identified feelings, as the matter to be confirmed; In the presenting step, the content of the sentiment is presented to the user as the item to be confirmed. Information processing system.

6. 2. The information processing system according to claim 1, The processor, In the determination step, the answer to the presented confirmation item is reflected in a determination of the subsequent confirmation item. Information processing system.

7. 2. The information processing system according to claim 1, The processor, In the presenting step, if a determination result identical to the determined check-required item has been obtained a predetermined number of times or more in the past, the check-required item is not presented for the determination result. Information processing system.

8. 1. An information processing method, comprising: The information processing system according to any one of claims 1 to 7, Information processing methods.

9. A program, A computer is caused to execute each step of the information processing system according to any one of claims 1 to 7. program.

Citation Information

Patent Citations

  • Method, device, and program for command processing

    JP2002287793A

  • Device, method and program for dialog with user

    JP2008217444A

  • Schedule management system and schedule management method

    JP2012164285A

  • Task management device

    JP2016206841A

  • Voice interactive device and voice interactive method

    JP2017215468A