Dialogue device, dialogue method, and dialogue program

JP2026141931APending Publication Date: 2026-09-07HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025028697
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2026-09-07

Smart Images

  • Figure 2026141931000001_ABST
    Figure 2026141931000001_ABST
Patent Text Reader

Abstract

Seek prior confirmation of opinions for and against the proposal. [Solution] The dialogue device performs the following: a setup process to set up multiple AI agents that simulate people and generate sentences using a language model; a first input process to receive the user's topic of discussion; a dialogue process to select a first AI agent and a second AI agent from the multiple AI agents and to perform a dialogue between the first AI agent and the second AI agent using the language model on the topic entered in the input process; and a dialogue text generation process to determine the AI ​​agent to which the user will have a dialogue based on the dialogue results from the dialogue process and to generate a dialogue text to the user from the AI ​​agent to which the user will have a dialogue based on the language model.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present invention relates to a dialog device, a dialog method, and a dialog program for executing dialog. [[Background Art]]

[0002] Non-Patent Document 1 discloses the collective predictive coding hypothesis: symbol emergence as decentralized Bayesian inference. [[Prior Art Documents]] [[Non-Patent Documents]]

[0003] [[Non-Patent Document 1]] Taniguchi, Tadahiro. "Collective predictive coding hypothesis: Symbol emergence as decentralized bayesian inference." Frontiers in Robotics and AI 11 (2024): 1353870. [[Summary of the Invention]] [[Problem to be Solved by the Invention]]

[0004] Non-Patent Document 1 discloses a method in which correct proposals are generated through dialog between language models, but this method accepts proposals from an interlocutor language model as correct. Therefore, when different proposals are presented to a certain language model from other language models, no technique is disclosed regarding to what extent the opinions of the other language models should be accepted, and a language model capable of appropriately selecting from a plurality of opinions cannot be generated.

[0005] An object of the present invention is to confirm pros and cons opinions in advance. [[Means for Solving the Problem]]

[0006] An interactive device that represents one aspect of the invention disclosed in this application is an interactive device having a processor that executes a program and a storage device that stores the program, wherein the processor performs a setting process to set up a plurality of AI agents that simulate people and generate sentences using a language model; a first input process to receive input of a topic from the user; a dialogue process to select a first AI agent and a second AI agent from the plurality of AI agents and to perform a dialogue between the first AI agent and the second AI agent using the language model with respect to the topic entered by the input process; and a dialogue sentence generation process to determine the AI ​​agent to which the user will have a dialogue based on the dialogue result from the dialogue process and to generate a dialogue sentence to the user from the AI ​​agent to which the user will have a dialogue based on the language model. [Effects of the Invention]

[0007] According to a typical embodiment of the present invention, it is possible to obtain prior confirmation of opinions for and against the invention. Problems, configurations, and effects other than those mentioned above will be clarified by the following description of the embodiments. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is an explanatory diagram showing four types of pro- and con states. [Figure 2] Figure 2 is an explanatory diagram showing an example of the system configuration of a dialogue system. [Figure 3] Figure 3 is a block diagram showing an example of a computer hardware configuration. [Figure 4] Figure 4 is an explanatory diagram showing an example of multi-AI agent configuration information. [Figure 5] Figure 5 is an explanatory diagram showing an example of user persona information. [Figure 6] Figure 6 is an explanatory diagram showing an example of dialogue information. [Figure 7] Figure 7 is an explanatory diagram showing an example of AI agent confidence information. [Figure 8]Figure 8 is a flowchart showing an example of a dialogue processing procedure by a dialogue device. [Figure 9] Figure 9 is an explanatory diagram showing an example of a settings screen. [Figure 10] Figure 10 is an explanatory diagram showing an example of a dialogue display screen. [Figure 11] Figure 11 is a flowchart showing a detailed example of the AI ​​agent-to-agent dialogue process (step S803) shown in Figure 8. [Figure 12] Figure 12 is a flowchart showing a detailed example of the AI ​​agent group update process (step S808) shown in Figure 8. [Figure 13] Figure 13 is an explanatory diagram illustrating one usage scenario for the dialogue system. [Figure 14] Figure 14 is an explanatory diagram illustrating usage scenario 2 of the dialogue system. [Figure 15] Figure 15 is an explanatory diagram showing an example of a screen displaying the results of a dialogue between AI agents. [Modes for carrying out the invention]

[0009] <Figure 1: Four types of support / opposition> Figure 1 is an explanatory diagram illustrating four types of agreement / disagreement states. These four types of agreement / disagreement states refer to four possible combinations in a two-person dialogue: one person's opinion on a given topic, the other person's opinion (agreement or disagreement), and one's own opinion on the other person's opinion (agreement or disagreement). An agenda item is a subject for discussion, and includes topics of conversation not only in meetings but also in any setting.

[0010] In other words, the four types of pro / con states are, with respect to a given topic, (1) a state in which speaker X agrees with speaker Y's opinion, and speaker Y also agrees with speaker X's opinion; (2) a state in which speaker X agrees with speaker Y's opinion, but speaker Y disagrees with speaker X's opinion; (3) a state in which speaker X disagrees with speaker Y's opinion, but speaker Y agrees with speaker X's opinion; and (4) a state in which speaker X disagrees with speaker Y's opinion, and speaker Y also disagrees with speaker X's opinion.

[0011] Here, speaker X is a user of the dialogue system or an AI (Artificial Intelligence) agent that simulates a user implemented in the dialogue system. Speaker Y is an AI agent that simulates another user implemented in the dialogue system.

[0012] The dialogue system according to the present embodiment performs dialogue between speakers X and Y, and increases or decreases the reliability of speaker Y who utters an utterance that reaches a specific approval-disapproval state among the four types of approval-disapproval states described above. Finally, the dialogue system presents the utterance content from speaker Y with the highest reliability to the user.

[0013] If the specific approval-disapproval state is the approval-disapproval state of (1) above, opinions favorable for consensus building or opinions sympathizing with the opinion of speaker X can be obtained from the dialogue system.

[0014] If the specific approval-disapproval state is the approval-disapproval state of (2) above, this means that speaker Y knows important matters unknown to speaker X, and an opinion that expects a change in the value of speaker X can be obtained from the dialogue system.

[0015] If the specific approval-disapproval state is the approval-disapproval state of (3) above, this means that speaker Y does not know the important matters of speaker X, and an opinion that expects a change in the value of speaker Y can be obtained from the dialogue system.

[0016] If the specific approval-disapproval state is the approval-disapproval state of (4) above, the discussion diverges like brainstorming, and various opinions of speaker Y can be obtained from the dialogue system.

[0017] <Figure 2 Example of system configuration of dialogue system> Figure 2 is an explanatory diagram showing an example of the system configuration of a dialogue system. The dialogue system 200 includes a dialogue device 201 and a terminal 202. The dialogue device 201 and the terminal 202 are communicatively connected to each other via a network 203 such as the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network).

[0018] The dialogue device 201 has a language model 210, multi-AI agent configuration information 211, user persona information 212, dialogue information 213, and AI agent reliability information 214, and functions as a dialogue AI that interacts with the user 220.

[0019] Language model 210 is a type of probabilistic model used in natural language processing, and it is a model that probabilistically predicts how likely a given word or sentence is to occur in natural language. Specifically, language model 210 is a mathematical model used in the field of natural language processing to generate and understand natural language by deep learning language patterns and grammatical rules using a large amount of training datasets of conversations.

[0020] For example, when predicting the next word or sentence, the language model 210 automatically generates the most likely word or sentence based on the context by calculating the probability of occurrence of a given word sequence or sentence, or by comparing the probability of occurrence of multiple word sequences or sentences. In other words, because the language model 210 has been deeply trained using a large dataset of conversations, it can perform various tasks such as responding to questions, proofreading and summarizing texts, translating texts, and generating texts.

[0021] In this way, when the language model 210 receives a query called a prompt, it outputs a response to that query. The language model 210 may be implemented using, for example, open source software. Alternatively, the language model 210 may be a large-scale language model implemented on an external computer that can communicate with the dialogue device 201 via the network 203, such as BERT or ChatGPT.

[0022] The multi-AI agent configuration information 211 is information for configuring multiple AI agents that interact within the dialogue device 201, and will be described later in Figure 4. The user persona information 212 is user information, and will be described later in Figure 5. The dialogue information 213 is the content of the conversation between the user and the AI ​​agent, and will be described later in Figure 6. The AI ​​agent trust level information 214 is information for managing the trust level between AI agents, and will be described later in Figure 7.

[0023] The dialogue device 201 generates multiple AI agents within the dialogue device 201 using the language model 210 and multi-AI agent configuration information 211. The dialogue device 201 generates an AI agent within the dialogue device 201 that behaves like the user 220, using the language model 210, multi-AI agent configuration information 211, and user persona information 212.

[0024] Terminal 202 is a computer used by user 220. There can be one or more terminals 202. If there are multiple terminals 202, each terminal 202 is used by a different user 220. User 220 participates in conversations such as meetings and chats through terminal 202, for example. Terminal 202 performs voice input / output, text input / output, and speech recognition.

[0025] An AI agent is an instance that simulates the utterances of user 220 by converting person information and speech history into prompts and inputting them into the language model 210. The AI ​​agent acts as a facilitator or participant in conversations with user 220 or with other AI agents within the conversational AI. When the AI ​​agent acts as a facilitator, it simulates the utterances of user 220 and acts as a virtual moderator of the conversation, making utterances that support consensus building within the conversation. When the AI ​​agent acts as a facilitator, it simulates the utterances of user 220.

[0026] Furthermore, because the language model 210 is highly versatile, the AI ​​agent can simulate user 220's speech by providing user 220's persona data from user persona information 212 as input information, without requiring retraining for each AI agent. In this case, the AI ​​agent simulates user 220's speech by appropriately changing the input information (prompt) equivalent to a command to the language model 210, without retraining. For example, by inputting a prompt such as "You are an engineering student. What kind of career do you aspire to in the future?" into the language model 210, the language model 210 generates a response from user persona information 212 containing a string of user 220's name enclosed in double quotes (in this case, a profession), such as "Engineering student".

[0027] An instance is the actual program that generates prompts for the language model 210 from user persona information 212, inputs them into the language model 210, and generates response sentences, as in this AI agent. Since the language model 210 itself does not have state, it can be shared by all AI agents. Therefore, the instance does not include the language model 210 itself. Furthermore, because the language model 210 consumes a large amount of memory, it is not instantiated for each AI agent. Processing of the language model 210 is instructed to a shared LLM instance.

[0028] <Figure 3: Example of hardware configuration of a computer (interaction device 201, terminal 202)> Figure 3 is a block diagram showing an example of the hardware configuration of a computer. Computer 300 includes a processor 301, a storage device 302, an input device 303, an output device 304, and a communication interface (communication IF) 305. The processor 301, storage device 302, input device 303, output device 304, and communication IF 305 are connected by a bus 306. The processor 301 controls computer 300. The storage device 302 serves as the work area for the processor 301. The storage device 302 is a non-temporary or temporary recording medium that stores various programs and data. Examples of storage devices 302 include ROM (Read Only Memory), RAM (Random Access Memory), HDD (Hard Disk Drive), and flash memory. The input device 303 inputs data. Examples of input devices 303 include a keyboard, mouse, touch panel, numeric keypad, scanner, microphone, and sensor. The output device 304 outputs data. Output devices 304 include, for example, displays, printers, and speakers. The communication IF 305 connects to the network 203 and sends and receives data.

[0029] <Figure 4 Multi-AI Agent Configuration Information 211> Figure 4 is an explanatory diagram showing an example of multi-AI agent configuration information 211. The multi-AI agent configuration information 211 includes the number of multi-AI agents 401, internal dialogue upper limit information 402, user ingestion conditions 403, opinion ingestion conditions 404, and confidence level operation parameters 405.

[0030] The multi-AI agent count 401 is the number of AI agents that make up the conversational AI, i.e., the maximum number of AI agents generated within the conversational device 201. The internal dialogue limit information 402 is the upper limit for dialogue between AI agents within the conversational AI. This can be specified by the number of times or by time.

[0031] User acquisition condition 403 is a condition for acquiring user 220's persona information into the dialogue device 201 and generating an AI agent that behaves as user 220. In the example in Figure 4, it is set to "User = Agree, Dialogue AI = Disagree" and "User = Disagree, Dialogue AI = Agree," so when the A / B states shown in (2) and (3) in Figure 1 occur, user 220's persona information is acquired into the dialogue device 201, and an AI agent that simulates user 220 is generated.

[0032] Opinion Incorporation Condition 404 is the condition for an AI agent to incorporate the opinion of the other AI agent as its own opinion. In the example in Figure 4, since it is set to "AI Agent 1 = Opposed, AI Agent 2 = Agreed", when the A / B state shown in (3) in Figure 1 is reached, the user 220's opinion is incorporated as the opinion of the conversational AI.

[0033] The confidence level manipulation parameter 405 is used to increase or decrease the confidence level managed by the AI ​​agent confidence level information 214. The confidence level manipulation parameter 405 is set for each of the (1) to (4) affirmative / negative states shown in Figure 1.

[0034] In the example in Figure 4, "AI Agent 1" corresponds to speaker X in Figure 1, and "AI Agent 2" corresponds to speaker Y in Figure 1. Therefore, "AI Agent 1 = Agree, AI Agent 2 = Agree" represents the A / B state shown in (1) in Figure 1. "AI Agent 1 = Agree, AI Agent 2 = Disagree" represents the A / B state shown in (2) in Figure 1. "AI Agent 1 = Disagree, AI Agent 2 = Agree" represents the A / B state shown in (3) in Figure 1. "AI Agent 1 = Disagree, AI Agent 2 = Disagree" represents the A / B state shown in (4) in Figure 1.

[0035] Additionally, a value of "0.0" indicates no change in confidence level, "1.0" indicates a 1.0 increase in confidence level, and "-0.1" indicates a 0.1 decrease in confidence level.

[0036] <Figure 5 User Persona Information 212> Figure 5 is an explanatory diagram showing an example of user persona information 212. User persona information 212 has the fields ID 501, age 502, occupation 503, personality 504, and interests 505. The combination of values ​​for each field in the same row defines persona data that characterizes the profile of a single user.

[0037] ID 501 is identification information that uniquely identifies user 220. Age 502 is the number of years elapsed since the user's date of birth. Occupation 503 is the job that user 220 is engaged in, or their position in that job. Personality 504 is the tendency of user 220's emotions and will. Interests 505 are the things that user 220 is interested in.

[0038] Note that age 502, occupation 503, personality 504, and interests 505 are examples of user persona information 212, and are not limited to these. In addition, other items that define a person's profile, such as preferences and values, may also be defined as user persona information 212.

[0039] Thus, the AI ​​agent is generated within the dialogue device 201 in either the case where the language model 210 is used, or the case where the language model 210 of an external service is used.

[0040] In other words, person information and speech history indicating the state of the AI ​​agent are input to the language model 210 as prompts each time. Therefore, there is only one language model 210, and it can be implemented either inside or outside the dialogue device 201. When using the language model 210 of an external service, the state of the language model 210 is not stored in the external service's computer, but is input from the dialogue device 201 to the external service's computer as a prompt.

[0041] In this embodiment, by providing the language model 210 with persona data of a user 220 as a prompt, the dialogue device 201 behaves as an AI agent simulating the user 220, as if the user 220 were participating in the conversation.

[0042] <Figure 6 Dialogue Information 213> Figure 6 is an explanatory diagram showing an example of dialogue information 213. Dialogue information 213 includes an agenda item 601, a user's opinion 602, a dialogue AI's opinion 603, the user's opinion 604 (for or against the dialogue AI's opinion), and the dialogue AI's opinion 605 (for or against the user's opinion). Dialogue information 213 exists for each user 220.

[0043] Agenda item 601 is a string of characters indicating a topic for discussion with the conversational AI. Agenda item 601 is input from terminal 202 to the conversational device 201. User comment 602 is a string of characters indicating user 220's opinion on agenda item 601. User comment 602 is input from terminal 202 to the conversational device 201.

[0044] The conversational AI opinion 603 is a string indicating the opinion output by the conversational AI. The conversational AI is a group of AI agents, and the conversational AI opinion 603 is, for example, the opinion from the AI ​​agent with the highest confidence level. User 220's opinion 604 on conversational AI opinion 603 is a string indicating whether user 220 agrees or disagrees with conversational AI opinion 603 and the reasons for that agreement or disagreement. The conversational AI's opinion 605 on user opinion 602 is a string indicating whether the conversational AI agrees or disagrees with user opinion 602 and the reasons for that agreement or disagreement.

[0045] If user 220's opinion 604 for the conversational AI opinion 603 is "in favor" and the conversational AI's opinion 605 for the user opinion 602 is "in favor", then the (1) for / against status shown in Figure 1 is displayed.

[0046] If user 220's opinion 604 for the conversational AI opinion 603 is "in favor" and the conversational AI's opinion 605 for the user opinion 602 is "against", then the (2) for / against status shown in Figure 1 will be displayed.

[0047] If user 220's opinion 604 for the conversational AI opinion 603 is "disagreement" and the conversational AI's opinion 605 for user opinion 602 is "agreement", then the agreement / disagreement status shown in (3) in Figure 1 will be displayed.

[0048] If user 220's opinion 604 for or against conversational AI opinion 603 is "against" and the conversational AI's opinion 605 for or against user opinion 602 is "against", then the agreement / disagreement status shown in (4) in Figure 1 will be displayed.

[0049] <Figure 7 AI Agent Trust Level Information 214> Figure 7 is an explanatory diagram showing an example of AI agent trust information 214. AI agent trust information 214 has a source of trust 701 and trusted information 702. The source of trust 701 defines the AI ​​agents that constitute the conversational AI. In Figure 7, four AI agents A1 to A4 are defined as an example.

[0050] Trustee information 702 includes the trust level 721 of the source of trust 701 to other AI agents, and the number of interactions 722 between the source of trust 701 and other AI agents.

[0051] The trust level of 721 for other AI agents from source 701 indicates the degree to which the AI ​​agent from source 701 trusts other AI agents. For example, if source 701 is AI agent A1, the trust levels of each of the other AI agents A2-A4 are stored. Specifically, for example, if source 701 is AI agent 1, the trust level for AI agent A2 is "0.2", the trust level for AI agent A3 is "0.7", and the trust level for AI agent A4 is "0.1". Therefore, it is shown that AI agent A1 trusts AI agent A3 the most.

[0052] The trust scores of 721 for other AI agents, based on trust source 701, are normalized so that the sum is "1.0" for each trust source 701. For example, if trust source 701 is AI agent A1, the trust scores for AI agent A2 are "0.2", for AI agent A3 are "0.7", and for AI agent A3 are "0.1", so the sum is "1.0".

[0053] The number of interactions between Trust Source 701 and other AI agents, 722, represents the total number of times Trust Source 701 interacted with other AI agents. For example, if Trust Source 701 is AI Agent A1, this indicates that the total number of times AI Agent A1 interacted with AI Agents A2 through A4 is "3".

[0054] <Figure 8: Example of dialogue processing procedure by dialogue device 201> Figure 8 is a flowchart illustrating an example of a dialogue processing procedure by the dialogue device 201. Prior to the execution of this flowchart, the dialogue device 201 is set to a state where it can interact with the terminal 202 of user 220 who will be interacting with the dialogue device 201. That is, the dialogue device 201 is assumed to recognize its dialogue partner by the ID 501 of user 220 who will be interacting with the dialogue device 201.

[0055] (Step S801) The dialogue device 201 performs initial setup. Specifically, for example, the dialogue device 201 accepts input of user 220's persona data from terminal 202 and accepts the selection of an AI agent from terminal 202. Since the dialogue device 201 recognizes the ID 501 of user 220, which is the dialogue partner, it reads the persona data associated with that ID 501. The number of selected AI agents may exceed the multi-AI agent count of 401.

[0056] (Step S802) Returning to Figure 8, the dialogue device 201 receives input from the terminal 202 of the agenda item 601 that user 220 requests a dialogue on, and user comments 602 on that agenda item.

[0057] [Figure 9 Settings screen] Figure 9 is an explanatory diagram showing an example of a settings screen. The settings screen 901 is displayed on the display device 900 of terminal 202, which is an example of an output device 304. For example, when user persona information 212 is transmitted from the dialogue device 201 to terminal 202, the settings screen 901 is displayed on the display device 900 of terminal 202.

[0058] The settings screen 901 includes an AI agent selection area 902 and an agenda information input area 903.

[0059] The AI ​​agent selection area 902 displays radio buttons 921 associated with persona data such as age 502, occupation 503, personality 504, and interests 505 for each user 220. The radio buttons 921 are selected by the user 220 on terminal 202. In Figure 9, white circles indicate an unselected state, and black circles indicate a selected state.

[0060] The import confirmation button 922 is a user interface for importing the persona data of user 220, selected by radio button 921, into the dialogue device 201 upon user 220's press. Pressing the import confirmation button 922 imports the persona data of user 220 selected by radio button 921 into the dialogue device 201, and an AI agent simulating the selected user 220 is set up.

[0061] The agenda information input area 903 includes an agenda input area 931, a user comment input area 932, and a send button 933. The agenda input area 931 is an area that receives input of an agenda item 601 requested by user 220 and displays the entered string. The user comment input area 932 for agenda item 601 is an area that receives input of user comment 602 for agenda item 601 and displays the entered string.

[0062] The send button 933 is a user interface for sending a string entered in the agenda input area 931 as agenda item 601 and a string entered in the user comment input area 932 as user comment 602 to agenda item 601 to the dialogue device 201 when pressed by the user 220.

[0063] (Step S803) Returning to Figure 8, the dialogue device 201 executes AI agent-to-AI dialogue processing. AI agent-to-AI dialogue processing (step S803) is the process of executing dialogue between the AI ​​agents set up in step S801 about the topic 601 entered in step S802. At this time, user comments 602 on topic 601 entered in step S802 are also used. Details of AI agent-to-AI dialogue processing (step S803) will be described later in Figure 11.

[0064] (Step S804) The dialogue device 201 generates a dialogue text using the AI ​​agent with the highest confidence level based on the dialogue results in the AI ​​agent-to-AI agent dialogue processing (step S803) and sends it to the user 220's terminal 202. The AI ​​agent that generates the dialogue text may be selected probabilistically according to confidence level information. In this case, AI agents with higher confidence levels are selected with a higher probability. In addition, the number of dialogues with other AI agents (722) may be used as a weight in the probabilistic selection based on confidence level. In cases such as brainstorming aimed at diverging discussions, the number of dialogues with other AI agents (722) is used as a weight so that AI agents with fewer dialogues with other AI agents (722) are selected with a higher probability. In cases such as consensus building aimed at converging discussions, the number of dialogues with other AI agents (722) is used as a weight so that AI agents with more dialogues with other AI agents (722) are selected with a higher probability.

[0065] The dialogue text includes the dialogue AI's opinion 603 on agenda item 601 and the user's opinion 605 for or against the dialogue AI's opinion 602.

[0066] Specifically, for example, the most reliable AI agent includes its persona data and speech history (held in step S1105), generates a prompt requesting a dialogue AI opinion 603 on agenda item 601, inputs it to the language model 210, and retrieves the dialogue AI opinion 603 on agenda item 601 output from the language model 210.

[0067] Similarly, the most reliable AI agent, including its persona data and speech history (held in step S1105), generates a prompt requesting a pro / con opinion 605 on the user's opinion 602 (step S802) regarding agenda item 601, inputs it to the language model 210, and retrieves the pro / con opinion 605 on the user's opinion 602 output from the language model 210. The dialogue device 201 then sends the dialogue text thus retrieved to the user's terminal 202.

[0068] Through such interactions between AI agents, user 220 can review the AI's dialogue opinion 603 on the reliable agenda 601 and the user's opinion 602 (both positive and negative) before interacting with other users 220, who are represented by the most reliable AI agent.

[0069] (Step S805) The dialogue device 201 receives input from the terminal 202 of user 220's opinion 604 for or against the dialogue AI opinion 603.

[0070] [Figure 10 Interactive display screen] Figure 10 is an explanatory diagram showing an example of a dialogue display screen. The dialogue display screen 1000 is displayed on the display device 900 of terminal 202, which is an example of an output device 304. For example, in step S802, when terminal 202 receives input of agenda item 601 and user comments 602 on the agenda item requested by user 220, the display device 900 displays agenda item 601 and user comments 602 on the dialogue display screen 1000.

[0071] Then, in step S804, the dialogue text of the most reliable AI agent (the dialogue AI's opinion 603 on agenda item 601, and the dialogue AI's opinion for or against the user's opinion 605) is sent from the dialogue device 201 to the terminal 202 and displayed on the dialogue display screen 1000.

[0072] Next, in step S805, when terminal 202 receives input of user 220's opinion 604 for or against the conversational AI opinion 603, user 220's opinion 604 for or against the conversational AI opinion 603 is displayed on the dialogue display screen 1000. In this way, dialogue information 213 is displayed on the dialogue display screen 1000.

[0073] (Step S806) Returning to Figure 8, the dialogue device 201 determines whether the approval / rejection status between the user 220 and the dialogue AI corresponds to the user acquisition condition 403. If the approval / rejection status between the user 220 and the dialogue AI corresponds to the user acquisition condition 403 (step S806: Yes), the device proceeds to step S807. On the other hand, if it does not correspond to the condition (step S806: No), the device proceeds to step S809.

[0074] For example, in the case of the dialogue information 213 in Figure 6 displayed on the dialogue display screen 1000 in Figure 10, the user 220's opinion 604 for or against the dialogue AI's opinion 603 is "for," and the dialogue AI's opinion 605 for or against the user's opinion 602 is "against." Therefore, the user acquisition condition 403 of "User = For, Dialogue AI = Against" is met. Consequently, the process proceeds to step S807.

[0075] (Step S807) The dialogue device 201 generates an AI agent that simulates the user 220 based on the persona data of the user 220 acquired in step S801. Specifically, for example, if an AI agent simulating the user 220 has not yet been generated, the dialogue device 201 generates such an AI agent.

[0076] (Step S808) The dialogue device 201 performs an AI agent group update process. Specifically, for example, the dialogue device 201 adjusts the number of AI agents in the AI ​​agent group so that the number of multi-AI agents is 401 or less. The AI ​​agent group update process (step S808) will be described later in Figure 12.

[0077] (Step S809) The dialogue device 201 determines whether or not the termination condition for ending the dialogue with the user 220 is met. If the termination condition is not met (step S809: No), the process returns to step S802. If the termination condition is met (step S809: Yes), the dialogue processing by the dialogue device 201 ends.

[0078] For example, the dialogue device 201 determines that the termination conditions are met when the number of repetitions of steps S802 to S808 exceeds a predetermined number, when a predetermined time has elapsed for the dialogue resulting from the repetition of steps S802 to S808, or when it receives a dialogue termination instruction from the user 220's terminal 202 (step S809: Yes).

[0079] <Figure 11 AI agent-to-agent dialogue processing (step S803)> Figure 11 is a flowchart showing a detailed example of the AI ​​agent-to-agent dialogue process (step S803) shown in Figure 8.

[0080] (Step S1101) The dialogue device 201 selects a first AI agent and a second AI agent from a group of AI agents that constitute the dialogue AI, according to selection criteria. The group of AI agents is a group of AI agents with a multi-AI agent count of 401 or less, generated by the selection in step S801. The selection criteria may be, for example, the unselected pair whose combined mutual trust score (which may be the average) is the highest, the AI ​​agents in order of the number of interactions between them being smallest, or random.

[0081] The mutual trust between the two AI agents is expressed as AI agent i's trust in AI agent j, Tij, and AI agent j's trust in AI agent i, Tji. The two selected AI agents are the first AI agent and the second AI agent. In the trust operation parameter 405 in Figure 4, "AI agent 1" refers to the first AI agent, and "AI agent 2" refers to the second AI agent.

[0082] (Step S1102) The dialogue device 201 causes the first AI agent and the second AI agent to generate opinions on agenda item 601 using the language model 210. Specifically, for example, assuming the first AI agent is a user 220 and the second AI agent is a dialogue AI, the first AI agent generates a prompt that includes its persona data and speech history (held in step S1105) and requests user opinion 602 on agenda item 601, inputs it to the language model 210, and obtains the user opinion 602 (opinion of the first AI agent) on agenda item 601 output from the language model 210.

[0083] Similarly, assuming the second AI agent is user 220 and the first AI agent is a conversational AI, the second AI agent includes its persona data and speech history (held in step S1105), generates a prompt to request user opinion 602 on agenda item 601, inputs it to the language model 210, and obtains the user opinion 602 (the second AI agent's opinion) on agenda item 601 output from the language model 210.

[0084] (Step S1103) The dialogue device 201 instructs the first AI agent and the second AI agent to use the language model 210 to generate opinions for or against the user opinion 602 of the dialogue partner obtained in step S1102.

[0085] Specifically, for example, assuming the first AI agent is user 220 and the second AI agent is a conversational AI, the first AI agent includes its persona data and speech history (held in step S1105), generates a prompt to ask for approval or disapproval of user opinion 602 (opinion of the second AI agent), inputs it to the language model 210, and obtains the approval or disapproval of user opinion 602 (opinion of the second AI agent) output from the language model 210.

[0086] Similarly, assuming the second AI agent is user 220 and the first AI agent is a conversational AI, the second AI agent includes its persona data and speech history (held in step S1105), generates a prompt to ask for approval or disapproval of user opinion 602 (opinion of the first AI agent), inputs it to the language model 210, and obtains the approval or disapproval of user opinion 602 (opinion of the first AI agent) output from the language model 210.

[0087] (Step S1104) The dialogue device 201 updates the trust destination information 702 of the AI ​​agent trust information 214 according to the approval / disapproval status based on the approval / disapproval opinions of the first AI agent and the second AI agent in step S1103. Specifically, for example, the dialogue device 201 identifies which of the trust operation parameters 405 the approval / disapproval statuses of the first AI agent and the second AI agent correspond to, and updates the trust level 721 of the first AI agent toward the second AI agent and the trust level 721 of the second AI agent toward the first AI agent according to the identified approval / disapproval status.

[0088] For example, if the approval / disapproval status of the first AI agent and the second AI agent is "AI agent 1 = in favor, AI agent 2 = against," then only the first AI agent's trust level in the second AI agent (which is 721) will increase by 0.1.

[0089] Furthermore, if the confidence level 721 of the dialogue device 201 increases or decreases, the row-wise sum of the confidence levels 721 of the original AI agent will be normalized to 1. To prevent the confidence level from becoming negative, the lower limit of the confidence level is set to 0.

[0090] Furthermore, the dialogue device 201 also increases the number of interactions 722 with other AI agents by 1. For example, suppose the first AI agent is AI agent A2 and the second AI agent is AI agent A3. In this case, the dialogue device 201 updates the number of interactions 722 with other AI agents whose requestor 701 is AI agent A2 from "4" to "5", and updates the number of interactions 722 with other AI agents whose requestor 701 is AI agent A3 from "10" to "11".

[0091] (Step S1105) The dialogue device 201 stores the other party's opinion as its own if the approval / disapproval status of the first AI agent and the second AI agent matches the approval / disapproval status specified in the opinion acquisition condition 404. That is, the first AI agent retains the approval / disapproval opinion of the second AI agent as part of its speech history, and the second AI agent retains the approval / disapproval opinion of the first AI agent as part of its speech history.

[0092] Therefore, if the AI ​​agent selected as the first or second AI agent re-executes steps S1102 and S1103, and if speech history has been accumulated in step S1105, it generates a prompt including the speech history and inputs it to the language model 210. As a result, the language model 210 outputs a response that takes the speech history into account, improving the accuracy of the response.

[0093] (Step S1106) The dialogue device 201 determines whether the AI ​​agent-to-AI dialogue process (step S803) corresponds to the internal dialogue limit information 402. If it does not correspond to the internal dialogue limit information 402 (step S1106: No), it returns to step S1101. If it does correspond to the internal dialogue limit information 402 (step S1106: Yes), the dialogue device 201 terminates the AI ​​agent-to-AI dialogue process (step S803) and proceeds to step S804.

[0094] <Figure 12 AI agent group update process (step S808)> Figure 12 is a flowchart showing a detailed example of the AI ​​agent group update process (step S808) shown in Figure 8.

[0095] (Step S1201) The dialogue device 201 adds an AI agent to the AI ​​agent group. The AI ​​agent to be added is the AI ​​agent generated in step S807.

[0096] (Step S1202) The dialogue device 201 adjusts the number of AI agents in the AI ​​agent group and proceeds to step S809. Specifically, for example, if the number of AI agents in the AI ​​agent group exceeds the multi-AI agent count of 401, the dialogue device 201 removes AI agents from the AI ​​agent group so that the number of AI agents in the AI ​​agent group becomes 401 or less. The AI ​​agents to be removed are determined, for example, starting with the AI ​​agent with the lowest total (or average) trust level to other AI agents.

[0097] For example, in the AI ​​agent confidence information 214 shown in Figure 7, the total confidence score of 721 for AI agent A1 is "0.3", the total confidence score of 721 for AI agent A2 is "0.7", the total confidence score of 721 for AI agent A3 is "1.8", and the total confidence score of 721 for AI agent A4 is "0.9". Therefore, AI agent A1 is determined to be removed and is expelled from the AI ​​agent group. In this case, the updated AI agent group consists of AI agents A2 to A4, and the four AI agents added in step S1201.

[0098] Furthermore, the dialogue device 201 only needs to have 2 or more AI agents in the AI ​​agent group after adjustment in step S1202, and 401 or fewer multi-AI agents.

[0099] Furthermore, if the number of AI agents in the AI ​​agent group after adjustment in step S1202 falls below the multi-AI agent count of 401, the dialogue device 201 may generate AI agents based on the persona data of other unselected users 220 and add them to the AI ​​agent group, or it may reinstate AI agents that were previously removed, so as not to exceed the multi-AI agent count of 401.

[0100] <Figure 13: Usage Scenario 1 of Dialogue System 200> Figure 13 is an explanatory diagram illustrating usage scenario 1 of the dialogue system 200. Usage scenario 1 shows an example of use in an online meeting. Multiple participants P1 to P4, who are users 220 (four in Figure 13 as an example), are participating in an online meeting on the dialogue device 201 via terminal 202.

[0101] In usage scenario 1, for example, the dialogue device 201 generates AI agents A1 to A4 that simulate participants P1 to P4 based on the persona data of participants P1 to P4 during the initial setup (step S801).

[0102] Furthermore, the dialogue device 201 generates a group of AI agents (AI agents A1 to A4) as a dialogue AI, as well as a facilitator AI agent 1300. Like the AI ​​agents, the facilitator AI agent 1300 is an AI agent that simulates a certain user 220 based on the user's persona data. However, unlike AI agents A1 to A4, which simulate meeting participants Pa to Pd, the facilitator AI agent 1300 is configured to take on the role of managing the meeting proceedings. The facilitator AI agent 1300 is displayed on the display device 900.

[0103] In each loop of steps S802 to S809, the dialogue device 201, with the help of the facilitator AI agent 1300, determines a speaker from among the participants P1 to P4 and receives input of the agenda 601 and user opinion 602 from the speaker. The dialogue device 201 also receives input of the user 220's opinion 604 for or against the dialogue AI opinion 603 from the determined speaker (step S805).

[0104] The dialogue device 201 performs AI agent-to-AI agent dialogue processing (step S803) using AI agents A1 to A4 that simulate participants Pa to Pd.

[0105] <Figure 14: Usage Scenario 2 of Dialogue System 200> Figure 14 is an explanatory diagram illustrating usage scenario 2 of the dialogue system 200. Usage scenario 2 shows an example of a user 220 using the dialogue system 200 in a face-to-face interaction with a dialogue AI. Multiple participants P1 to P4, who are users 220 (four in Figure 14 as an example), are gathered in conference room 1400. A display device 1401, which is an example of an output device 304 of the dialogue device 201, is installed in conference room 1400.

[0106] In the example shown in Figure 14, the settings screen 901 shown in Figure 9 and the dialogue display screen 1000 shown in Figure 10 are displayed on the display device 1401. The dialogue device 201 can identify which participant P1 to P4 corresponds to which user 220 by facial recognition based on image data from a camera (not shown). The voices of participants P1 to P4 can also be input to the dialogue device 201 via a directional microphone (not shown). Therefore, for example, agenda items 601, user opinions 602, and opinions for or against the dialogue AI 604 are input to the dialogue device 201 via this microphone. Alternatively, participants P1 to P4 may input agenda items 601, user opinions 602, and opinions for or against the dialogue AI 604 as text from their terminal 202 and send them to the dialogue device 201.

[0107] In usage scenario 2, as in usage scenario 1, the dialogue device 201 generates a group of AI agents as well as a facilitator AI agent 1300 as a dialogue AI. The facilitator AI agent 1300 is displayed on the display device 1401.

[0108] <Figure 15: Screenshot showing the results of the AI ​​agent-to-AI dialogue> Figure 15 is an explanatory diagram showing an example of the AI ​​agent dialogue result screen. In usage scenario 1 of Figure 13, the AI ​​agent dialogue result screen 1500 is displayed on the display device 900, and in usage scenario 2 of Figure 14, it is displayed on the display device 1500.

[0109] The AI ​​agent dialogue results screen 1500 displays AI agent confidence information 214 as the result of the AI ​​agent dialogue. In the AI ​​agent confidence information 214, AI agents A1 to A4 correspond to participants P1 to P4, respectively.

[0110] Furthermore, the AI ​​agent dialogue result screen 1500 may also display the speaker selection area 1501. The speaker selection area 1501 is a user interface for the user 220 to select from among participants P1 to P4 which speaker will input the agenda item 601 and user opinion 602.

[0111] In this case, there is no facilitator AI agent 1300, and one of the participants P1 to P4 becomes the facilitator. In usage scenario 1 of Figure 13, the dialogue device 201 may display the speaker determination area 1501 only on the terminal 202 of user 220, who is the facilitator.

[0112] The speaker determination area 1501 includes a state selection radio button 1511 and a state participant pair display unit 1512. The state selection radio button 1511 is a user interface that accepts the selection of a pro / no state by the user 220. White circles indicate an unselected state, and black circles indicate a selected state. In the example in Figure 15, the pro / no state (2), where speaker X is in favor and speaker Y is against, is selected.

[0113] The status participant pair display unit 1512 is an area that displays participant pairs corresponding to the AI ​​agent pair with a pro / con status selected by the status selection radio button 1511. The facilitator refers to the status participant pair display unit 1512 to determine the next speaker and prompts that next speaker to input the agenda item 601 and user opinion 602.

[0114] As explained above, according to this embodiment, through dialogue between AI agents that generate sentences by simulating other users 220, user 220 can confirm the dialogue AI opinion 603 on the reliable agenda 601 and the pro / con opinion 605 on the user opinion 602 before interacting with the other user 220, who is simulated by the most reliable AI agent.

[0115] Therefore, it is possible to encourage user 220 to converge on user opinion 602 (pro / con state in (1)), change user opinion 602 (pro / con state in (2)), change conversational AI opinion 603 (pro / con state in (3)), and diverge from user opinion 602 (pro / con state in (4)) through simulated conversations with other users 220 regarding agenda item 601. This is particularly effective when agenda item 601 does not have an objectively correct answer.

[0116] Furthermore, in the AI ​​agent dialogue processing (step S803), the reliability of the dialogue text spoken by the AI ​​agent in step S804 can be improved by increasing or decreasing the trust level of each AI agent (721) in the AI ​​agent based on their affirmative / negative status.

[0117] It should be noted that the present invention is not limited to the embodiments described above, but includes various modifications and equivalent configurations within the spirit of the attached claims. For example, the embodiments described above are described in detail to make the present invention easier to understand, and the present invention is not necessarily limited to having all of the described configurations. Furthermore, some of the configurations of one embodiment may be replaced with those of another embodiment. Furthermore, some of the configurations of one embodiment may be added to those of another embodiment. Furthermore, some of the configurations of each embodiment may be added, deleted, or replaced with other configurations.

[0118] Furthermore, each of the aforementioned configurations, functions, processing units, and processing means may be implemented in hardware, for example, by designing them as integrated circuits, or they may be implemented in software by having a processor interpret and execute programs that realize each function.

[0119] Information such as programs, tables, and files that implement each function can be stored in memory, hard disks, SSDs (Solid State Drives), or on recording media such as IC (Integrated Circuit) cards, SD cards, and DVDs (Digital Versatile Discs).

[0120] Furthermore, the control lines and information lines shown are those deemed necessary for explanatory purposes and do not necessarily represent all control lines and information lines required for implementation. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]

[0121] 200 Dialogue Systems 201 Dialogue device 202 terminals 210 Language Models 211 Agent Configuration Information 212 User Persona Information 213 Dialogue Information 214 Agent Trust Information 220 users 300 calculator 301 Processor 302 Storage Devices

Claims

1. An interactive device having a processor for executing a program and a storage device for storing the program, The aforementioned processor, A configuration process for setting up multiple AI agents that use language models to simulate human speech and generate sentences, A first input process that accepts the user's agenda, A dialogue process that selects a first AI agent and a second AI agent from the plurality of AI agents, and uses the language model to perform a dialogue between the first AI agent and the second AI agent on the topic input by the first input process, Based on the dialogue results from the dialogue processing, the AI ​​agent that will be the user's dialogue partner is determined, and a dialogue text generation process is performed to generate dialogue text from the AI ​​agent that will be the user's dialogue partner to the user using the language model. A dialogue device characterized by performing the following actions.

2. The dialogue device according to claim 1, In the dialogue processing, the processor uses the language model to perform the following actions: the generation of a first opinion on the topic by the first AI agent; the generation of a second opinion on the topic by the second AI agent; the generation of a first opinion for or against the second opinion by the first AI agent; and the generation of a second opinion for or against the first opinion by the second AI agent. A dialogue device characterized by the following features.

3. The dialogue device according to claim 2, Both the first and second opinions for and against the aforementioned opinion support the opposing side's opinion. A dialogue device characterized by the following features.

4. The dialogue device according to claim 2, Of the first opinion and the second opinion, one is an opinion in favor of the other party's opinion, and the other is an opinion against the other party's opinion. A dialogue device characterized by the following features.

5. The dialogue device according to claim 2, Both the first and second opinions for and against the aforementioned opinion are opposing the other party's opinion. A dialogue device characterized by the following features.

6. The dialogue device according to claim 2, The aforementioned processor, The above dialogue process is repeatedly executed until the predetermined conditions are met. A dialogue device characterized by the following features.

7. The dialogue device according to claim 6, In the first dialogue process, which is one of the dialogue processes that is repeatedly executed, the processor stores the second opinion as the dialogue history of the AI ​​agent selected by the first AI agent, and stores the first opinion as the dialogue history of the AI ​​agent selected by the second AI agent, if the first opinion and the second opinion are in a specific agreement / disagreement state. In a second dialogue process, which is executed after the first dialogue process as one of the dialogue processes that is repeatedly executed, the processor uses the language model to generate a first opinion using the dialogue history if the AI ​​agent selected for the first AI agent has a dialogue history, generate a second opinion using the dialogue history if the AI ​​agent selected for the second AI agent has a dialogue history, generate a first pro / con opinion using the dialogue history if the AI ​​agent selected for the first AI agent has a dialogue history, and generate a second pro / con opinion using the dialogue history if the AI ​​agent selected for the second AI agent has a dialogue history. A dialogue device characterized by the following features.

8. The dialogue device according to claim 2, The memory device stores trust information for each AI agent, which represents the level of trust in other AI agents. In the first input processing, the processor receives input from the user regarding the agenda, In the aforementioned dialogue process, the processor updates the trust level of the first AI agent towards the second AI agent and the trust level of the second AI agent towards the first AI agent based on the approval / disapproval status of the first and second approval / disapproval opinions. In the dialogue generation process, the processor determines the AI ​​agent to interact with the user based on the updated confidence information from the dialogue process, and uses the language model to generate dialogue from the AI ​​agent to the user. A dialogue device characterized by the following features.

9. The dialogue device according to claim 8, In the dialogue generation process, the processor determines the AI ​​agent that has the highest sum or average confidence score from all other AI agents to be the AI ​​agent that interacts with the user, and uses the language model to generate a dialogue from the AI ​​agent that interacts with the user to the user. A dialogue device characterized by the following features.

10. The dialogue device according to claim 1, In the first input processing, the processor receives input from the user regarding the agenda, In the dialogue generation process, the processor generates the following as the dialogue: the AI ​​agent's dialogue partner's opinion on the topic, and the AI ​​agent's opinion for or against the user's opinion input by the first input process. The aforementioned processor, A second input process that receives the user's opinion of approval or disapproval of the dialogue AI opinion generated by the dialogue text generation process, An update process to update the plurality of AI agents based on the user's approval or disapproval of the conversational AI opinion input by the second input process and the approval or disapproval status of the user's opinion generated by the dialogue text generation process, A dialogue device characterized by performing the following actions.

11. The dialogue device according to claim 10, In the update process, if the approval / disapproval status is a specific approval / disapproval status, the multiple AI agents are updated. A dialogue device characterized by the following features.

12. The dialogue device according to claim 11, The aforementioned processor, If the aforementioned approval / disapproval status is a specific approval / disapproval status, an agent generation process is executed to generate an AI agent that simulates the user and generates sentences. In the update process, the processor adds the AI ​​agent generated by the agent generation process to the plurality of AI agents. A dialogue device characterized by the following features.

13. The dialogue device according to claim 12, The memory device stores trust information for each AI agent, which represents the level of trust in other AI agents. In the update process, the processor updates the plurality of AI agents based on the confidence level. A dialogue device characterized by the following features.

14. An interaction method performed by an interaction device having a processor that executes a program and a storage device that stores the program, The aforementioned processor, A configuration process for setting up multiple AI agents that use language models to simulate human speech and generate sentences, A first input process that accepts the user's agenda, A dialogue process that selects a first AI agent and a second AI agent from the plurality of AI agents, and uses the language model to perform a dialogue between the first AI agent and the second AI agent on the topic input by the first input process, Based on the dialogue results from the dialogue processing, the AI ​​agent that will be the user's dialogue partner is determined, and a dialogue text generation process is performed to generate dialogue text from the AI ​​agent that will be the user's dialogue partner to the user using the language model. A dialogue method characterized by performing the following:

15. In the processor, A configuration process for setting up multiple AI agents that use language models to simulate human speech and generate sentences, A first input process that accepts the user's agenda, A dialogue process that selects a first AI agent and a second AI agent from the plurality of AI agents, and uses the language model to perform a dialogue between the first AI agent and the second AI agent on the topic input by the first input process, Based on the dialogue results from the dialogue processing, the AI ​​agent that will be the user's dialogue partner is determined, and a dialogue text generation process is performed to generate dialogue text from the AI ​​agent that will be the user's dialogue partner to the user using the language model. An interactive program characterized by causing the execution of [a certain action].