Information processing apparatus, information processing method, and program

The information processing device enhances conversation flow by using AI to analyze and suggest actions based on utterances, addressing the challenge of engagement in discussions and meetings.

JP2026003608APending Publication Date: 2026-01-13SAYSAY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025105181
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-24
Filing Date
2025-06-20
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing systems fail to adequately encourage participants to engage in conversations, discussions, or meetings, leading to insufficient promotion of interaction and idea generation.

Method used

An information processing device that utilizes a generation AI to analyze utterances and suggest actions for next utterances, including personality information, behaviors, and attitudes to enhance conversation flow.

Benefits of technology

The system promotes better progress in conversations by suggesting appropriate actions and behaviors, encouraging participation and improving discussion quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026003608000001_ABST
    Figure 2026003608000001_ABST
Patent Text Reader

Abstract

To provide a device, a method and a program for promoting progress in conversation or examination.SOLUTION: An information processor capable of exchanging information with a generation AI device for executing generation AI capable of outputting information corresponding to an input acquires the utterance contents of one or a plurality of persons composed of a person and / or an agent, and transmits the utterance contents to the generation AI device. A second request procedure of requesting the generation and AI device to perform an action related to a next speech based on an AI result based on information obtained from the generation and AI device as a result of the request in the first request procedure and the speech content; and an outputting procedure of outputting information using the action related to the next speech obtained from the generation and AI device as a result of the request in the second request procedure.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Conventionally, people have been the facilitators of meetings (for example, small meetings such as internal meetings or large meetings such as business conferences), hosts of streaming programs, facilitators of games (for example, real games or computer games), and commentators for sports broadcasts. Also, people have been involved in discussions such as idea generation among multiple people, and sometimes people generate ideas alone. Face-to-face group conversations (for example, face-to-face drinking parties) and conversations using online conferencing tools (for example, drinking parties) are also conducted. For example, meeting support systems that assist in the progress of meetings have been developed. For example, Patent Document 1 discloses a meeting support system equipped with a facilitation means that assists in the progress of the meeting depending on the content of comments made in the meeting. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-163405 Summary of the Invention [Problem to be solved by the invention]

[0004] However, when a person is in charge of facilitating conversations in, for example, these meetings, streaming programs, games, or live sports broadcasts, they may not be able to adequately encourage participants to speak up. Furthermore, when one or more people are discussing ideas, for example, the idea generation may not go smoothly. Furthermore, for example, in face-to-face group conversations or conversations using online conferencing tools, the conversation may not flow well. Thus, there is a problem of insufficient promotion of conversation or discussions between one or more people.

[0005] The present invention has been made in view of the above problems, and has an object to provide an information processing device, an information processing method, and a program that can promote progress in conversation or discussion. [Means for solving the problem]

[0006] An information processing device according to a first aspect of the present invention is an information processing device capable of exchanging information with a generation AI device that executes a generation AI that can output information according to an input, at least one processor; The at least one processor: an acquisition step for acquiring utterances of one or more people and / or agents; a first request step of sending the utterance content to the generation AI device and requesting the generation AI device to output, based on the utterance content, at least one of a matter relating to the utterance content, a classification of at least one utterance content included in the utterance content, or the emotion and / or state of the speaker in at least one utterance included in the utterance content; a second request procedure for requesting the generation AI device to output an action regarding the next utterance based on the analysis results and utterance content based on the information obtained from the generation AI device as a result of the request made in the first request procedure; an output step of outputting output information using an action related to the next utterance obtained from the generation AI device as a result of the request made in the second request step; Execute.

[0007] According to this configuration, the output information suggests an action for the next utterance, thereby promoting progress in the conversation or discussion.

[0008] An information processing device according to a second aspect of the present invention is an information processing device according to the first aspect, wherein the at least one processor, before making a request to the generation AI device in the second request step, requests the generation AI device for a next behavior and / or attitude to be taken based on the classification of at least one utterance obtained from the generation AI device and the emotion and / or state of the utterer; The analysis results based on the information obtained from the generating AI device include the requested results, the next behavior and / or attitude obtained from the generating AI device.

[0009] According to this configuration, the next action to be taken regarding the next utterance can be determined based on the next behavior and / or attitude, thereby enabling better progress.

[0010] An information processing device according to a third aspect of the present invention is the information processing device according to the second aspect, wherein the at least one processor, when requesting the next behavior and / or attitude to be taken, further requests a priority for each of the next behavior and / or attitude to be taken from the generation AI device; The analysis results based on the information obtained from the generating AI device further include a priority for each next behavior and / or attitude requested and obtained from the generating AI device.

[0011] According to this configuration, the action regarding the next utterance can be determined based on the priority of each behavior and / or attitude to be taken next, thereby enabling better progress.

[0012] An information processing device according to a fourth aspect of the present invention is the information processing device according to any one of the first to third aspects, wherein the items related to the utterance content obtained from the generation AI device include a topic of the utterance; Before making a request to the generation AI device in the second request step, the at least one processor requests the generation AI device to output personality information indicating what kind of person is suitable for the target conversation or discussion based on the topic of the utterance and / or the content of the utterance obtained from the generation AI device; The analysis results based on the information obtained from the generation AI device include personality information indicating what kind of person would be suitable to add to the target conversation, obtained by requesting the generation AI device.

[0013] According to this configuration, appropriate character information to add to the target conversation can be obtained, and the action regarding the next utterance can be determined based on this character information, thereby enabling better progress.

[0014] An information processing device according to a fifth aspect of the present invention is an information processing device according to the fourth aspect, wherein the person profile information indicating what kind of person is suitable to be added to the target conversation is the person's background, the person's characteristics, or the role within the conversation or consideration, and when requesting the generation AI device to estimate the person's background, the person's characteristics, or the role, the at least one processor further requests from the generation AI device a score representing the degree of relevance to the topic for each of the person's background, the person's characteristics, or the role; The analysis results based on the information obtained from the generating AI device further include the score obtained from the generating AI device in response to the request.

[0015] This configuration allows for better progress because the next action to be taken can be determined based on the score for each person's background, the score for each person's characteristics, or the score for each role in the conversation that promotes the statement.

[0016] An information processing device according to a sixth aspect of the present invention is the information processing device according to the fourth or fifth aspect, wherein the action related to the next comment includes a next commenter and, if there is no person who matches the next commenter, a message indicating that there is no person who matches the next commenter; If, as a result of the request made in the second request procedure, the at least one processor receives from the generation AI device that there is no suitable person for the next speaker, it creates an agent that fulfills the job type or role and performs a process of adding the created agent to the conversation or consideration.

[0017] This configuration allows an appropriate agent to participate in the conversation or discussion, thereby facilitating the progress of the conversation or discussion.

[0018] An information processing device according to a seventh aspect of the present invention is an information processing device according to any one of the first to sixth aspects, wherein the action related to the next utterance includes at least one of the following: the next speaker to be encouraged to speak by a facilitator agent who acts as a facilitator in a conversation or discussion; the content of the facilitator agent's utterance; the necessity of the participation of a guest agent who has a profession with knowledge of the topic or a role in the conversation or discussion that encourages utterances; or the tension or nuance of the facilitator agent's utterance.

[0019] With this configuration, at least one of the following can be obtained as an action related to the next statement: the next speaker whom the facilitator encourages to speak, the content of the facilitator's statement, whether an expert needs to participate, or the tone or nuance of the facilitator's statement, allowing for better progress.

[0020] An information processing device according to an eighth aspect of the present invention is the information processing device according to any one of the first to seventh aspects, wherein the at least one processor transmits to the generation AI device a character image setting prompt requesting setting of a character image of an agent to be added to a conversation; and creating an agent using at least the character profile settings obtained from the generation AI device in response to the transmission of the character profile setting prompt.

[0021] According to this configuration, an agent can be created according to the settings of the character image.

[0022] An information processing device according to a ninth aspect of the present invention is the information processing device according to any one of the first to eighth aspects, wherein the at least one processor sends to the generation AI device an expert data collection prompt requesting a search keyword for collecting expert data of an agent to be added to a conversation; collecting URLs by performing an internet search using search keywords obtained from the generating AI device in response to sending the specialized data collection prompt; collecting data from the web pages of the collected URLs; and creating an agent using at least the collected data.

[0023] This configuration allows an agent to be created that has appropriate specialized data.

[0024] An information processing device according to a tenth aspect of the present invention is the information processing device according to the ninth aspect, wherein the at least one processor sends a determination prompt to the generation AI device requesting the generation AI device to determine whether the collected data is sufficient; If the judgment result obtained in response to sending the judgment prompt indicates that the result is sufficient, a procedure for collecting the URL is further executed, and if the judgment result indicates that the result is not sufficient, a procedure for creating the agent is executed.

[0025] According to this configuration, sufficient data can be collected as the knowledge and know-how assumed from the career history set for the agent.

[0026] An information processing device according to an eleventh aspect of the present invention is an information processing device according to any one of the first to tenth aspects, wherein the at least one processor transmits a summary prompt to the generation AI device requesting a summary of the utterance content; and creating an agent to add to the conversation using at least the summary obtained from the generation AI device in response to the summary prompt.

[0027] This configuration makes it possible to create an agent that can generate a response in response to the content of previous comments.

[0028] An information processing method according to a twelfth aspect of the present invention includes an acquisition step in which an acquisition means acquires utterance contents of one or more people and / or agents; a first request procedure in which a first request means transmits the utterance content to a generation AI device that executes a generation AI capable of outputting information according to input, and requests the generation AI device to output, based on the utterance content, at least one of a matter relating to the utterance content, a classification of at least one utterance content included in the utterance content, or the emotion and / or state of the speaker in at least one utterance included in the utterance content; a second request procedure in which a second request means requests the generation AI device to take an action regarding the next utterance based on the analysis results and utterance content based on the information obtained from the generation AI device as a result of the request made in the first request procedure; an output step in which an output means outputs output information using an action related to the next utterance obtained from the generation AI device as a result of the request made in the second request step; It has.

[0029] According to this configuration, the output information suggests an action for the next utterance, thereby promoting progress in the conversation or discussion.

[0030] A program according to a thirteenth aspect of the present invention includes the steps of: an acquisition step for acquiring utterances of one or more people and / or agents; a first request step of sending the utterance content to a generation AI device that executes a generation AI capable of outputting information according to input, and requesting the generation AI device to output, based on the utterance content, at least one of a matter relating to the utterance content, a classification of at least one utterance content included in the utterance content, or the emotion and / or state of the speaker in at least one utterance included in the utterance content; a second request procedure for requesting the generation AI device to take action regarding the next utterance based on the analysis results and utterance content based on the information obtained from the generation AI device as a result of the request in the first request procedure; an output step of outputting output information using an action related to the next utterance obtained from the generation AI device as a result of the request made in the second request step; This is a program for executing the above.

[0031] According to this configuration, the output information suggests an action for the next utterance, thereby promoting progress in the conversation or discussion. [Effects of the Invention]

[0032] According to one aspect of the present invention, this output information can facilitate progress in a conversation or discussion by suggesting actions for the next utterance. [Brief explanation of the drawings]

[0033] [Figure 1] 1 is a schematic configuration diagram of an information processing system according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an example of a schematic configuration of an information processing apparatus according to an embodiment of the present invention. [Figure 3] FIG. 2 is a schematic diagram of participants in a conversation in the first embodiment. [Figure 4] FIG. 10 is a sequence diagram showing the flow of processing at the initial stage of a conversation in the first embodiment. [Figure 5] FIG. 10 is a schematic diagram showing an example of the flow of a process for determining an action regarding the next utterance of a facilitator agent. [Figure 6]1 is an example of the content of a conversation in the first embodiment. [Figure 7A] This is an example of a prompt that asks you to analyze the content of a conversation. [Figure 7B] This is an example of the output result when the generating AI device processes the prompt of Figure 7A. [Figure 8A] 10 is an example of a prompt requesting information about additional potential guest agents to facilitate the conversation. [Figure 8B] 8B is an example of an output result obtained when the generating AI device processes the prompt of FIG. 8A. [Figure 9A] This is an example of a prompt that requests the user to score the relevance of the content of a statement for each predetermined category. [Figure 9B] This is an example of the output result when the generating AI device processes the prompt of Figure 9A. [Figure 10A] This is an example of a prompt that asks you to analyze the speaker's emotion or state. [Figure 10B] 10B is an example of the output of the generating AI device processing the prompt of FIG. 10A. [Figure 11A] Here is an example of a prompt requesting next speaker behavior: [Figure 11B] 11B is an example of an output result obtained when the generating AI device processes the prompt of FIG. 11A. [Figure 12A] This is an example of a prompt requesting an action regarding the next utterance. [Figure 12B] This is an example of the output that would result if the generating AI device processed the prompt in Figure 12A and did not create an agent. [Figure 12C] This is an example of the output result when the generating AI device processes the prompt in Figure 12A and creates an agent. [Figure 13] FIG. 10 is a schematic diagram of participants in a conversation in a second embodiment. [Figure 14] 10 is a flowchart showing an example of a conversation flow in the second embodiment. [Figure 15] This is a continuation of the flowchart in FIG. 14. [Figure 16]16 is a continuation of the flowchart in FIG. 15. [Figure 17] 10 is an example of the content of a conversation in a first modified example of the first embodiment. [Figure 18] 6 is an example of an output result in step S30 of FIG. 5 in the first modification of the first embodiment. [Figure 19] 6 is an example of an output result in step S40 of FIG. 5 in the first modification of the first embodiment. [Figure 20] 6 is an example of an output result in step S70 of FIG. 5 in the first modification of the first embodiment. [Figure 21A] 6 is an example of a prompt in step S80 of FIG. 5 in the first modification of the first embodiment. [Figure 21B] 6 is an example of an output result in step S80 of FIG. 5 in the first modification of the first embodiment. [Figure 22] 10 is an example of the content of a conversation in a second modification of the first embodiment. [Figure 23] 6 is an example of an output result in step S30 of FIG. 5 in the second modification of the first embodiment. [Figure 24] 6 is an example of an output result in step S40 of FIG. 5 in the second modification of the first embodiment. [Figure 25] 6 is an example of an output result in step S70 of FIG. 5 in the second modification of the first embodiment. [Figure 26A] 6 is an example of a prompt in step S80 of FIG. 5 in the second modification of the first embodiment. [Figure 26B] 6 is an example of an output result in step S80 of FIG. 5 in the second modification of the first embodiment. [Figure 27] FIG. 10 is a sequence diagram illustrating an example of a guest agent creation process. [Figure 28] 1 is an example of a summary prompt. [Figure 29] 10 is an example of a character profile setting prompt. [Figure 30] An example of a specialized data collection prompt. [Figure 31] 10 is an example of a decision prompt. [Figure 32]28 is a flowchart showing an example of details of the process in step S745 of FIG. 27. [Figure 33] 28 is a flowchart showing an example of details of the process in step S765 of FIG. 27. [Figure 34] This is an example of the contents of a Markdown file. [Figure 35] 10 is an example of instruction text. [Figure 36] 34 is a flowchart showing an example of details of the process in step S935 of FIG. 33. DETAILED DESCRIPTION OF THE INVENTION

[0034] The present embodiment will be described below with reference to the drawings. However, more detailed explanations than necessary may be omitted. For example, detailed explanations of well-known matters or redundant explanations of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following explanation and to facilitate understanding by those skilled in the art. In this embodiment, the subject is the progress of conversations in meetings, streaming programs, games, sports broadcasts, etc., or the progress of discussions such as idea generation among one or more people. Here, conversations in this embodiment include not only the exchange of opinions via voice but also the exchange of opinions via text, such as chat. Here, chat refers to multiple people exchanging conversations by typing text in real time over a computer network. Furthermore, "encouraging speech" during discussions also includes stopping speech for a certain period of time to allow time for discussion. In this embodiment, a conversation in a meeting will be described as an example.

[0035] Fig. 1 is a schematic diagram of an information processing system according to this embodiment. As shown in Fig. 1, the information processing system S according to this embodiment includes, as an example, terminals 1-1 to 1-N (N is a natural number), an information processing device 2 connected to each of the terminals 1-1 to 1-N via a communication network CN, and a generation AI device 3 connected to the information processing device 2 via the communication network CN.

[0036] The terminals 1-1 to 1-N are used by different users, for example, and may be mobile phones such as multi-function mobile phones (so-called smartphones), tablets, notebook computers, or desktop computers. Hereinafter, the terminals 1-1 to 1-N will also be collectively referred to as terminals 1.

[0037] The information processing device 2 is used by, for example, an administrator who manages the information processing system S according to this embodiment, and is, for example, a server. The information processing device 2 provides information in response to a request from a terminal 1-i (i is an index and is an integer from 1 to N). The information processing device 2 may be a single computer or may be composed of multiple computers.

[0038] The generation AI device 3 generates and outputs information (such as text) according to the input by executing a generation AI (Artificial Intelligence) capable of outputting text according to the input. The generation AI device 3 is, for example, a generation AI server, specifically, ChatGPT (registered trademark). In the following, the generation AI device 3 will be described as a generation AI server as an example.

[0039] Fig. 2 is an example of a schematic configuration diagram of an information processing device in this embodiment. As shown in Fig. 2, the information processing device 2 includes, as an example, an input interface 21, a communication module 22, a storage device 23, a memory 24, an output interface 25, and a processor 26. Note that, in this embodiment, the information processing device 2 is described as including one processor 26, but is not limited to this, and may include multiple processors, as long as it includes at least one processor.

[0040] The input interface 21 receives an input from, for example, an administrator (for example, an employee of a management organization) of the information processing device 2, and outputs an input signal corresponding to the received input to the processor . The communication module 22 is connected to the communication network CN and communicates with the terminals 1-1 to 1-N and the generating AI device 3. This communication may be wired or wireless, but the following description will be given assuming that it is wired.

[0041] The storage device 23 stores programs and various data that the processor 26 reads and executes. The memory 24 temporarily stores data and programs and is a volatile memory, such as a RAM (Random Access Memory). The output interface 25 can be connected to an external device and can output a signal to the external device.

[0042] The processor 26 loads the program from the storage device 23 into the memory 24 and executes a series of instructions contained in the program. Each embodiment will be described below using a conversation as an example.

[0043] Example 1 First, a first embodiment will be described. FIG. 3 is a schematic diagram of participants in a conversation in the first embodiment. In the first embodiment, as shown in FIG. 3, a conversation is explained as proceeding between three users, User A, User B, and User C, and a facilitator agent (e.g., an AI agent) that serves as a facilitator and is realized by an information processing device 2 using a generation AI device 3. Here, the facilitator agent is a facilitator in a conversation or discussion, and organizes the content of the participants' talk and / or encourages them to speak up to actively express their opinions. The number of users participating in a conversation may be two or less, or may be four or more. Furthermore, one or more agents may participate in a conversation.

[0044] 4 is a sequence diagram showing an example of the flow of processing at the initial stage of a conversation in Example 1. In the following sequence diagram, the processing of the processor will be described, but the processor will be omitted.

[0045] (Step S111) The terminal 1-i receives input of initial conversation settings from the user. Here, the initial conversation settings include, for example, an agenda (theme, purpose). The initial settings may also include advance information or participant information.

[0046] (Step S112) The terminal 1-i accepts a selection of a conversation method and / or platform from the user. Here, the conversation method and / or platform may be, for example, a text chat (e.g., Slack (registered trademark), Discord (registered trademark), etc.), a video or voice call (e.g., Zoom (registered trademark), Meet (registered trademark), LINE (registered trademark), etc.), or a browser.

[0047] (Step S113) The terminal 1-i transmits the selected initial settings, conversation method and / or platform to the information processing device 2.

[0048] (Step S211) For example, when the information processing device 2 receives the initial settings, the conversation method, and / or the platform, it generates an introductory greeting for the facilitator agent who will act as the facilitator in the conversation or discussion. Here, the greeting may include, for example, a self-introduction of the facilitator agent and / or an explanation of the initially set agenda.

[0049] (Step S212) Next, the information processing device 2 transmits the introductory greeting information, which is the generated greeting data, to the terminal 1-i.

[0050] (Step S113) When the terminal 1-i receives the introductory greeting information, it presents the introductory greeting of the facilitator agent (for example, by displaying it in text or outputting it as voice).

[0051] FIG. 5 is a schematic diagram showing an example of the process flow for determining an action regarding the next utterance of a facilitator agent. FIG. 6 shows an example of the content of a conversation in Example 1. In FIG. 6 and subsequent figures, A represents user A, B represents user B, and F represents the facilitator agent. FIG. 7A shows an example of a prompt requesting an analysis of the content of a conversation. This prompt includes, as an example, an instruction, an answer format, and the content of the conversation, which is a transcript of the audio data of the conversation. The instruction includes, as an example, a sentence requesting an inference about the content of the conversation based on each item and a score for each item regarding the state of the conversation. Examples of the answer format include "theme of the overall conversation," "topic of the last utterance," "appropriate utterance content for the next utterance," "last speaker," "to whom the last utterance was directed," "proportion of participants' utterances," and "analysis results regarding the state of the conversation." This answer format, as an example, outputs a score for each item listed as "analysis results regarding the state of the conversation." This score is, for example, a number between 0.0 and 1.0, with the closer to 1, the higher the score. Figure 7B is an example of the output result of the generative AI device processing the prompt in Figure 7A. The output result in Figure 7B is an answer based on the answer format in Figure 7A, and a score is shown for each item listed as "analysis results of the conversation state" in Figure 7B. Figure 8A is an example of a prompt requesting information about potential guest agents to promote the conversation. This prompt includes, for example, an instruction, an answer format, the content of the conversation transcribed from the audio data, and the analysis results of the conversation. For example, the instruction includes a sentence requesting a profile of a candidate person with certain attributes or characteristics to join the conversation to make it more lively, based on the topic of the last remark and the flow of the conversation. The answer format includes, for example, gender, age, personality, attributes, personality characteristics, appearance characteristics, speaking characteristics, comments, and a score.Here, this comment is a description of what is expected of that person in this conversation, and the score is a number ranging from 0.0 to 1.0 that indicates the level of relevance. Figure 8B is an example of the output result when the generative AI device processes the prompt in Figure 8A. Here, the output result in Figure 8B is an answer that follows the answer format in Figure 8A.

[0052] FIG. 9A is an example of a prompt requesting that the relevance of a utterance be scored for each predetermined category. FIG. 9B is an example of an output result obtained when the generative AI device processes the prompt of FIG. 9A. FIG. 10A is an example of a prompt requesting that the speaker's emotions or state be analyzed. FIG. 10B is an example of an output result obtained when the generative AI device processes the prompt of FIG. 10A.

[0053] Figure 11A is an example of a prompt requesting the behavior of the next speaker. Figure 11B is an example of the output result when the generative AI device processes the prompt of Figure 11A. Figure 12A is an example of a prompt requesting an action regarding the next utterance. Figure 12B is an example of the output result when the generative AI device processes the prompt of Figure 12A and does not create an agent. Figure 12C is an example of the output result when the generative AI device processes the prompt of Figure 12A and creates an agent.

[0054] Hereinafter, the processing of the processor 26 of the information processing device 2 will be described along FIG. 5 with reference to FIGS. 6 to 12C. (Step S10) The processor 26 acquires the utterances of each user from each of the terminals 1-i. Here, as an example, the description will be given assuming that the contents of the conversation shown in FIG.

[0055] (Step S20) Next, processor 26 identifies the speaker. Specifically, for example, processor 26 identifies who made the comment by performing any of the following: voice recognition, chat sender ID recognition, or video recognition.

[0056] Next, processor 26 transmits the utterance content to generation AI device 3 and executes a first request procedure to request generation AI device 3 to output, based on the utterance content, at least one of a matter related to the utterance content, a classification of at least one utterance content included in the utterance content, or the emotion and / or state of the speaker in at least one utterance included in the utterance content. In Figure 5, as an example, processor 26 requests generation AI device 3 to output, based on the utterance content, the topic of the utterance (see step S30) as an example of a matter related to the utterance content, the classification of the last utterance content (see step S50) as a classification of at least one utterance content included in the utterance content, and the emotion and / or state (see step S60) of the speaker in at least one utterance included in the utterance content (here, the last utterance content as an example). (Step S30) Next, processor 26 analyzes the topic using generation AI device 3. Specifically, for example, processor 26 sends a prompt (e.g., the prompt in Figure 7A) to generation AI device 3 requesting that the content of the current conversation be analyzed based on the history of the conversation since the previous topic was grouped, and receives an output result (e.g., the example output result in Figure 7B) output by generation AI device 3 after processing the prompt.

[0057] (Step S40) Next, processor 26 uses generation AI device 3 to output personality information indicating what kind of person would be suitable to add to the target conversation based on the content of the utterance and / or the topic. Specifically, for example, processor 26 requests generation AI device 3 to output personality information indicating what kind of person would be suitable to add to the target conversation based on the content of the utterance and / or the topic. This provides personality information that is suitable to add to the target conversation, and based on this personality information, it is possible to determine an action regarding the next utterance, thereby enabling better progress. Here, this person profile information includes, for example, a person's background, a person's characteristics, a person's attributes (e.g., a woman in her 20s), or a role in a conversation (or discussion). A person's background is an element that indicates how the person arrived at their current state, and may include the person's experience, knowledge, educational history, work history, current occupation, cultural background, etc. On the other hand, a person's characteristics are an element that indicates the person's current characteristics and abilities, and may include character, personality, behavioral patterns, skills, interests, preferences, etc. As an example, processor 26 uses generation AI device 3 to select, based on the content and / or topic of the utterance, an occupation (e.g., a professional) with knowledge of the relevant topic as an example of the background of a person who will participate in the target conversation. Specifically, processor 26 sends a prompt (e.g., the prompt in Figure 8A) to generation AI device 3 requesting an answer from an occupation with knowledge of the topic of the conversation (e.g., the topic of the last utterance), and receives an output result (e.g., the example output result in Figure 8B) output by generation AI device 3 after processing this prompt. Alternatively / in addition, the processor 26 may request the generation AI device 3 to estimate the characteristics of the person who will be adding to the target conversation based on the content and / or topic of the speech. Alternatively or additionally, the processor 26 may request the generation AI device 3 to estimate the role in the conversation (or discussion) that promotes the comment based on the content and / or topic of the comment.

[0058] (Step S50) Next, processor 26 classifies the content of the utterance (here, as an example, the last utterance). For example, processor 26 sends a prompt (e.g., the prompt in Figure 9A) to generation AI device 3 requesting that the content of the utterance be scored for relevance (or applicability) for each predetermined classification, and receives an output result (e.g., the example output result in Figure 9B) output by generation AI device 3 after processing this prompt. In this case, processor 26 determines the classification based on the score in the output result, for example. Specifically, processor 26 determines the classification with the highest score, for example.

[0059] (Step S60) Next, processor 26 analyzes the speaker's emotion or state. In this case, for example, processor 26 sends a prompt to generation AI device 3 requesting that the speaker's emotion or state be analyzed based on the voice or writing style and / or the content of the speech, and receives the output result that generation AI device 3 processes and outputs this prompt. Specifically, for example, processor 26 sends a prompt (e.g., the prompt in FIG. 10A) requesting that the AI ​​device 3 score the relevance for each emotion or state, and receives the output result (e.g., the example output result in FIG. 10B) that generation AI device 3 processes and outputs this prompt.

[0060] (Step S70) After steps S50 and S60 are completed, processor 26 selects a set of candidate behaviors for the next speaker and / or candidate attitudes expected when the next speaker performs the selected behaviors, based on, for example, the content of the speech (e.g., the content of the speech made by the last speaker) and the speaker's emotions or state. At this time, processor 26 may also obtain the priorities of the candidate behaviors. Specifically, for example, processor 26 transmits to generation AI device 3 a prompt (e.g., the prompt in FIG. 11A) requesting candidate behaviors for the next speaker, candidate attitudes expected when the next speaker performs the selected behaviors, and the priorities of the candidate behaviors, based on the content of the speech (e.g., the content of the speech made by the last speaker) and the speaker's emotions or state, and receives an output result (e.g., the example output result in FIG. 11B) output by generation AI device 3 after processing the prompt.

[0061] Next, processor 26 executes a second request procedure (see step S80) to request generation AI device 3 to output an action related to the next utterance based on the analysis result and the utterance content based on the information obtained from generation AI device 3 as a result of the request made in the first request procedure. Here, the analysis result based on the information obtained from generation AI device 3 (e.g., the topic obtained in step S30 and / or the score obtained in step S50) is, for example, "personality information indicating what kind of person would be appropriate to add to the target conversation" obtained from generation AI device 3 in step S40, and / or "candidate behaviors for the next speaker and / or the attitude expected when taking said behaviors, and the priority of said candidate behaviors (however, priority is not required)" obtained from generation AI device 3 in step S70.

[0062] (Step S80) Next, processor 26 uses generation AI device 3 to determine an action for the next utterance. Here, the action for the next utterance may include at least one of the following: the next speaker to be encouraged to speak by a facilitator agent who serves as the facilitator of the conversation or discussion; the content of the facilitator agent's utterance; the necessity of participation of a guest agent who has knowledge of the topic or a role in the conversation or discussion that encourages utterances; the tension or nuance of the facilitator agent's utterance; and the execution of additional commands (e.g., creating or determining avatar motion, generating audio, searching for and pausing supplementary information or reference materials, playing a video, drawing, etc.). With this configuration, at least one of the above can be obtained as an action for the next utterance, allowing for better progress.

[0063] Here, as an example, the processor 26 sends a prompt (e.g., the prompt in Figure 12A) to the generation AI device 3 and receives the output result (e.g., the example output results in Figures 12B and 12C) that the generation AI device 3 processes and outputs this prompt. Here, the prompt includes a sentence requesting the next speaker to be encouraged to speak by a facilitator agent who acts as the facilitator in the conversation or discussion, the content of the facilitator agent's speech (e.g., the behavior and speech that the facilitator should take toward the last speaker), and whether to add the proposed participant to the conversation. In this case, the prompt may also include a request for the tone or nuance of the facilitator agent's speech.

[0064] (Step S90) Processor 26 then determines whether the output result of step S80 includes invoking a guest agent. If the output result does not include invoking a guest agent, the process proceeds to step S120.

[0065] (Step S100) If the output result in step S90 includes calling a guest agent, processor 26, for example, uses generation AI device 3 to create a guest agent (e.g., an AI agent) that has knowledge about the topic or plays a role in the conversation that encourages speech. Specifically, for example, processor 26 generates a prompt requesting an answer that plays a role in encouraging speech. If the role involves expertise, processor 26, for example, uses the Web to collect information about the expertise (e.g., the latest information) from the Internet, thereby creating premise data that collects information about the expertise. Then, processor 26 creates a new AI agent as a guest agent, for example, by setting the created prompt, premise data, and summary data of the conversation so far.

[0066] (Step S110) Next, processor 26 adds the guest agent to the conversation.

[0067] (Step S120) Next, processor 26 generates output information using the action related to the next utterance obtained from generation AI device 3 and transmits this output information to each terminal 1-i. Here, the output information relates to the action related to the next utterance, and may be audio, other text, or an image converted from the action related to the next utterance, or the action related to the next utterance itself. The output information may also specify the type of motion for the avatar based on the action related to the next utterance, or may be video information of the avatar reacting based on this specification.

[0068] Thus, processor 26 may, for example, perform a capture procedure to capture the speech content of one or more people and / or agents.

[0069] Then, the processor 26, for example, sends the content of the statement to the generation AI device 3 and executes a first request procedure to request the generation AI device 3 to output, based on the content of the statement, at least one of "matters related to the content of the statement," a classification of at least one of the content of the statement contained in the content of the statement, or the emotion and / or state of the speaker in at least one of the statements contained in the content of the statement.

[0070] Then, the processor 26 executes a second request procedure, for example, to request the generation AI device 3 to output "analysis results based on information obtained from the generation AI device 3" and "actions related to the next utterance" based on the content of the utterance, as a result of the request made in the first request procedure.

[0071] Here, for example, before making a request to the generation AI device 3 in this second request procedure, the processor 26 requests the generation AI device 3 for the next behavior and / or attitude to be taken based on the classification of at least one utterance obtained from the generation AI device 3 and the speaker's emotions and / or state. The above-mentioned "analysis results based on information obtained from the generation AI device 3" includes the next behavior and / or attitude to be taken obtained from the generation AI device 3 as a result of the above request. With this configuration, it is possible to determine an action regarding the next utterance based on the next behavior and / or attitude to be taken, thereby enabling better progress. Here, for example, when requesting the next behavior and / or attitude to be taken, the processor 26 may further request a priority for each next behavior and / or attitude from the generation AI device 3. In this case, the "analysis result based on information obtained from the generation AI device 3" further includes the priority for each next behavior and / or attitude requested and obtained from the generation AI device 3. With this configuration, an action regarding the next utterance can be determined based on the priority for each next behavior and / or attitude to be taken, thereby enabling better progress.

[0072] Here, the "matters related to the content of the utterance" obtained from the generation AI device 3 may include, for example, the topic of the utterance. In this case, before making a request to the generation AI device 3 in the second request step, the processor 26 requests the generation AI device 3 to output personality information indicating what kind of person would be suitable for the target conversation or discussion based on the topic and / or content of the utterance obtained from the generation AI device 3. In this case, the above-mentioned "analysis result based on information obtained from the generation AI device" includes personality information indicating what kind of person would be suitable to add to the target conversation, obtained by making a request to the generation AI device 3. With this configuration, personality information appropriate for adding to the target conversation can be obtained, and an action regarding the next utterance can be determined based on this personality information, thereby enabling better progress.

[0073] The biographical information that indicates what kind of person is suitable to be included in the target conversation is the person's background, characteristics of the person, or role within the conversation or discussion. Furthermore, for example, when requesting the generation AI device 3 to estimate the person's background, the person's characteristics, or the role, the processor 26 may further request from the generation AI device 3 a score representing the degree of relevance to the topic for each of the person's background, the person's characteristics, or the role. In this case, the above-mentioned "analysis result based on information obtained from the generation AI device" further includes the score obtained from the generation AI device 3 in response to the request. With this configuration, an action regarding the next utterance can be determined based on the score for each person's background, the score for each person's characteristics, or the score for each role in the conversation that promotes the utterance, thereby enabling better progress.

[0074] Here, the action related to the next utterance may include the next speaker and, if there is no suitable candidate for the next speaker, a statement indicating that there is no suitable candidate. In this case, if the processor 26 receives a request from the generation AI device 3 indicating that there is no suitable candidate for the next speaker as a result of the request in the second request step, the processor 26 executes a process of creating a guest agent that fulfills the job type or role and adding the created guest agent to the conversation or discussion. This allows an appropriate guest agent to participate in the conversation or discussion, thereby facilitating the progress of the conversation or discussion.

[0075] Then, the processor 26 executes an output procedure that outputs output information using an action related to the next utterance obtained from the generation AI device 3 as a result of the request made in the second request procedure.

[0076] As a result, this output information suggests an action for the next utterance, thereby promoting the progress of the conversation or discussion.

[0077] <Example 2> Next, Example 2 will be described. In Example 1, there was no agent at the beginning of the conversation, but in Example 2, an example will be described in which an agent is present from the beginning of the conversation. Figure 13 is a schematic diagram of the participants in the conversation in Example 2. As shown in Figure 13, in Example 2, as an example, in addition to a facilitator agent that plays the role of facilitator, two people, User A and User B, and two agents participate in the conversation. Here, the facilitator agent is realized by the information processing device 2 using the generation AI device 3. Agent C is an AI assistant that plays the role of an event planner, and Agent D is a business AI assistant that assists with business operations.

[0078] Fig. 14 is a flowchart showing an example of the conversation flow in Example 2. Fig. 15 is a flowchart continuing from Fig. 14. Fig. 16 is a flowchart continuing from Fig. 15. Below, the processing of the processor 26 of the information processing device 2 will be described along with an example of the conversation flow in Example 2 using Figs. 14 to 16.

[0079] (Step S410) The processor 26 receives user A's message from the terminal 1 of user A, thereby acquiring user A's message.

[0080] (Step S510) Processor 26 analyzes the statement made by user A and determines the content of the next statement made by the facilitator and the next speaker to whom the facilitator will encourage a statement.

[0081] (Step S420) Processor 26 transmits the facilitator's next utterance content determined in step S510 to each of terminals 1 used by user A and user B. This causes the facilitator's next utterance content to be presented.

[0082] (Step S520) Since the next speaker is agent C, processor 26 generates a prompt requesting agent C to generate an answer, and sends the generated prompt to generation AI device 3. Here, the prompt includes the content of the conversation so far and the question.

[0083] (Step S430) When processor 26 receives the output result (including the utterance content) output by generation AI device 3 after processing this prompt, it transmits the utterance content included in this output result to each of terminals 1 used by user A and user B. This presents the utterance content of agent C (event planner).

[0084] (Step S530) Processor 26 analyzes the utterances of agent C (event planner) (specifically, for example, the processes of steps S20 to S80 in the first embodiment).

[0085] (Step S440) During the execution of this analysis, the message of user A of terminal 1 is received and transmitted to information processing device 2.

[0086] (Step S540) If processor 26 acquires a comment from user A during the analysis in step S530, processor 26 cancels the analysis in progress.

[0087] (Step S550) Processor 26 analyzes the utterance of user A acquired in step S540 (specifically, for example, the processing of steps S20 to S80 in Example 1), determines the next speaker, and generates a prompt requesting a response directly from agent C (event planner), who is the next speaker. Processor 26 transmits the generated prompt to generation AI device 3. This prompt includes a question for agent C (event planner).

[0088] (Step S560) Processor 26 determines whether the conversation or question and answer continues from the previous conversation between the same people. Here, as an example, the conversation between user A and agent C continues from the previous conversation. Therefore, processor 26 determines that the conversation or question and answer continues from the previous conversation between the same people, and uses generation AI device 3 to process agent C to generate a response directly without a facilitator's response. Specifically, for example, when processor 26 receives the output result (including the utterance content) output by generation AI device 3 after processing the prompt generated in step S550, processor 26 outputs the utterance content of agent C (event planner) included in this output result. Note that, as an example, processor 26 has configured the agent to speak without a facilitator's utterance if the conversation or question and answer continues from the previous conversation between the same people. However, this is not limited to this. If processor 26 determines that the user is asking the agent a question, the agent may speak without a facilitator's utterance.

[0089] (Step S450) Processor 26 transmits the content of the utterance of agent C (event planner) to each of terminals 1 used by user A and user B. As a result, the content of the utterance of agent C (event planner) is presented.

[0090] (Step S570) Processor 26 analyzes the utterances of agent C (event planner) presented in step S450 (specifically, the processes of steps S20 to S80 in the first embodiment, for example).

[0091] (Step S460) During the execution of this analysis, the terminal 1 of the user A receives the message from the user A and transmits it to the information processing device 2.

[0092] (Step S580) If processor 26 acquires a comment from user A during the analysis in step S570, processor 26 cancels the analysis in progress.

[0093] (Step S590) Processor 26 analyzes the remarks of user A acquired in step S580 (specifically, for example, the processing of steps S20 to S80 in Example 1), determines the next speaker, and determines the content of the facilitator's remarks and the next speaker to whom the facilitator will encourage them to speak.

[0094] (Step S470) Processor 26 transmits this utterance content to each of terminals 1 used by user A and user B. As a result, the facilitator's next utterance content is presented.

[0095] (Step S600) Since the next speaker is agent D, processor 26 generates a prompt requesting that an answer be generated for agent D, and sends the generated prompt to generation AI device 3. Here, the prompt includes the content of the conversation so far and the question.

[0096] (Step S480) When processor 26 receives the output result (including the utterance content) output by generation AI device 3 after processing this prompt, processor 26 transmits the utterance content included in this output result to terminals 1 used by user A and user B. This causes the utterance content of agent D (business assistant) to be presented.

[0097] (Step S610) Processor 26 analyzes the utterance of agent D (task assistant) presented in step S480 (specifically, for example, the processes of steps S20 to S80 in the first embodiment).

[0098] (Step S490) During the execution of this analysis, the terminal 1 of the user B receives the message from the user B and transmits it to the information processing device 2.

[0099] (Step S620) If processor 26 acquires a message from user B during the analysis in step S610, processor 26 cancels the analysis in progress.

[0100] <Modification 1 of Example 1> Variation 1 of Example 1 is an example in which an agent whose speech content and behavior are intentionally biased is added. The processing flow is the same as in Example 1, and is the processing flow of Figure 5, but the conversation content, prompts, and output results are different. Figure 17 is an example of the content of a conversation in Variation 1 of Example 1. The conversation example in Figure 17 assumes a situation in which there are a lot of fillers and the conversation flow is not very good. Here, fillers are words such as "um," "hmm," and "ah," which are inserted between utterances.

[0101] FIG. 18 is an example of an output result in step S30 of FIG. 5 in Modification 1 of Example 1. FIG. 19 is an example of an output result in step S40 of FIG. 5 in Modification 1 of Example 1. FIG. 20 is an example of an output result in step S70 of FIG. 5 in Modification 1 of Example 1. FIG. 21A is an example of a prompt in step S80 of FIG. 5 in Modification 1 of Example 1. FIG. 21B is an example of an output result in step S80 of FIG. 5 in Modification 1 of Example 1.

[0102] <Modification 2 of Example 1> The second modification of the first embodiment is an example in which it is determined that a guest agent should be added in a meeting regarding the development of a fitness app. The processing flow is the same as that of the first embodiment, as shown in FIG. 5, but the content of the conversation, prompts, and output results are different. FIG. 22 is an example of the content of the conversation in the second modification of the first embodiment. In the example of the conversation in FIG. 22, it is assumed that there are some shortcomings in the preliminary investigation (specifically, for example, the opinions of women).

[0103] FIG. 23 is an example of an output result in step S30 of FIG. 5 in Modification 2 of Example 1. FIG. 24 is an example of an output result in step S40 of FIG. 5 in Modification 2 of Example 1. FIG. 25 is an example of an output result in step S70 of FIG. 5 in Modification 2 of Example 1. FIG. 26A is an example of a prompt in step S80 of FIG. 5 in Modification 2 of Example 1. FIG. 26B is an example of an output result in step S80 of FIG. 5 in Modification 2 of Example 1.

[0104] Although each embodiment has been described using an example of a conversation between multiple people, the present invention is not limited to this and can also be applied to a discussion of idea generation between one or multiple people.

[0105] <Example of guest agent creation process> A specific example of the guest agent creation process in step S100 of Fig. 5 will be described below with reference to Fig. 27. Fig. 27 is a sequence diagram showing an example of the guest agent creation process. In the following sequence diagram, the processing of a processor is described, but the processor is omitted from the description.

[0106] (Step S705) The information processing device 2 sets initial information of the agent. Here, the initial information includes, for example, at least one of a summary of the conversation flow, the agent's professional information, the agent's role in the conversation, and the agent's personality and / or characteristics.

[0107] (Step S710) Next, the information processing device 2 generates a summary prompt requesting a summary of the utterance content (e.g., the conversation content). Fig. 28 shows an example of the summary prompt. As shown in Fig. 28, the summary prompt 11 includes natural language instructing that the conversation content so far be summarized in accordance with constraints.

[0108] (Step S715) Next, the information processing device 2 transmits the generated summary prompt P11 to the generation AI device 3, for example.

[0109] (Step S805) When the generation AI device 3 receives the summary prompt P11, it generates a summary in response to this summary prompt P11.

[0110] (Step S810) Next, the generation AI device 3 transmits the generated summary to the information processing device 2.

[0111] (Step S720) Next, the information processing device 2 generates, for example, a character image setting prompt that requests the setting of a character image of an agent to be added to the conversation. Figure 29 is an example of the character image setting prompt. The character image setting prompt P12 in Figure 29 includes natural language that instructs the user to create detailed persona settings in accordance with a template based on an assumed character image.

[0112] (Step S725) Next, the information processing device 2 transmits the generated person image setting prompt P12 to the generation AI device 3, for example.

[0113] (Step S815) When the generation AI device 3 receives the character image setting prompt P12, it generates a character image setting in response to this character image setting prompt P12.

[0114] (Step S820) The generation AI device 3 transmits the generated character image setting to the information processing device 2.

[0115] (Step S725) Next, the information processing device 2 generates, for example, a specialized data collection prompt requesting search keywords for collecting specialized data of the agent to be added to the conversation. Fig. 30 is an example of a specialized data collection prompt. The specialized data collection prompt P13 in Fig. 30 includes natural language that instructs the agent to create multiple patterns of combinations of search keywords according to an output format when there is information that needs to be investigated regarding the deep knowledge and know-how that the agent has based on his or her career history.

[0116] (Step S730) Next, the information processing device 2 transmits the generated specialized data collection prompt P13 to the generation AI device 3, for example.

[0117] (Step S825) When the generation AI device 3 receives the specialized data collection prompt P13, it generates search keywords for specialized data collection in response to the specialized data collection prompt P13.

[0118] (Step S830) Next, the generation AI device 3 transmits the generated search keywords to the information processing device 2.

[0119] (Step S735) When the information processing device 2 receives the search keyword from the generation AI device 3, it uses the Internet search API with the search keyword to collect URLs (Uniform Resource Locators) of web pages.

[0120] (Step S740) Next, the information processing device 2 crawls the web pages of the collected URLs to collect data from these web pages.

[0121] (Step S745) Next, the information processing device 2 records the collected data in a form that allows vector search in the database of the storage device 23. Details of this process will be described later.

[0122] (Step S750) Next, the information processing device 2 generates a decision prompt to determine whether the collected data is sufficient. Fig. 31 is an example of a decision prompt. The decision prompt P14 in Fig. 31 includes natural language that instructs the agent to determine whether the collected information is sufficient as the knowledge and know-how that the agent has in-depth knowledge of based on the career history set for the agent, and to output the decision result as a status. If the decision result means that it is sufficient, a word that indicates that it is sufficient (for example, "sufficient") is set in the status, and if the decision result does not mean that it is sufficient, a word that indicates that more is needed (for example, "need more") is set in the status.

[0123] (Step S755) Next, the information processing device 2 transmits the generated determination prompt P14 to the generation AI device 3, for example.

[0124] (Step S835) When the generation AI device 3 receives the determination prompt P14, it generates a determination result in response to this determination prompt P14.

[0125] (Step S840) Next, the generation AI device 3 transmits the generated judgment result to the information processing device 2.

[0126] (Step S760) When the information processing device 2 receives the determination result, it determines whether the determination result means that the information processing device 2 is sufficient. If the determination result does not mean that the information processing device 2 is sufficient, the process returns to step S735 and the processes from step S735 onwards are repeated.

[0127] (Step S765) If the determination result in step S760 indicates that the information is sufficient, an agent is created by aggregating the various pieces of information and setting it as a system prompt. A detailed example of this process will be described later. In this way, as an example, processor 26 may create an agent to be added to the conversation using the summary, character profile settings, and collected data. However, this is not limiting, and processor 26 may also create an agent to be added to the conversation using at least one of the summary, character profile settings, and collected data.

[0128] In this way, processor 26 may further execute the steps of sending a decision prompt to AI generation device 3 requesting it to decide whether the collected data is sufficient, collecting the URL if the decision result obtained in response to the sending of the decision prompt means sufficient, and creating the agent if the decision result does not mean sufficient. With this configuration, sufficient data can be collected as the knowledge and know-how expected from the history set for the agent.

[0129] In this way, processor 26 may execute the steps of sending a character image setting prompt to generation AI device 3 requesting the setting of a character image of an agent to be added to a conversation, and creating an agent using at least the character image setting obtained from generation AI device 3 in response to the sending of the character image setting prompt. With this configuration, an agent can be created according to the character image setting.

[0130] The processor 26 may also execute the following steps: sending a specialized data collection prompt to the generation AI device 3 requesting search keywords for collecting specialized data for an agent to add to a conversation; collecting URLs by performing an Internet search using the search keywords obtained from the generation AI device 3 in response to the sending of the specialized data collection prompt; collecting data from the web pages of the collected URLs; and creating an agent using at least the collected data. This configuration allows the creation of an agent with appropriate specialized data.

[0131] The processor 26 may also execute the steps of sending a summary prompt to the AI ​​generation device 3 requesting a summary of the utterance, and creating an agent to add to the conversation using at least the summary obtained from the AI ​​generation device in response to the summary prompt. This configuration makes it possible to create an agent that can generate a response based on the utterance content up to that point.

[0132] Fig. 32 is a flowchart showing an example of the details of the process in step S745 in Fig. 27. The process in step S745 in Fig. 27 is a process of recording the collected data in the database of the storage device 23 in a form that allows vector search based on the collected data.

[0133] (Step S905) The processor 26 of the information processing device 2 requests the embeddings API to convert the collected data into a numerical array.

[0134] (Step S910) Next, the processor 26 of the information processing device 2 stores the numerical value array in the vector search database of the storage device 23 in association with the data ID.

[0135] (Step S915) Next, the processor 26 of the information processing device 2 stores the detailed data for information recording with the same data ID.

[0136] Fig. 33 is a flowchart showing an example of the details of the process in step S765 in Fig. 27. The process in step S765 in Fig. 27 is a process for creating an agent in which each piece of information is collected and set as a system prompt.

[0137] (Step S925) The processor 26 of the information processing device 2 saves the collected data in, for example, an appropriate format. Here, an appropriate format may be, for example, data in Markdown format, and the data may be saved by storing it in a file or writing it to a database. FIG. 34 shows an example of the contents of a Markdown format file. File F1 in FIG. 34 contains a combination of the title of a web page, its URL, and a summary of the information on that web page, organized for each web page.

[0138] (Step S930) Next, the processor 26 of the information processing device 2 saves the conversation history up to this point in, for example, an appropriate format. Here, saving in an appropriate format may be saving in a text file or writing in a database.

[0139] (Step S935) Next, the processor 26 of the information processing device 2 creates a data search command and acquires detailed information. An example of specific processing will be described later.

[0140] (Step S940) Next, the processor 26 of the information processing device 2 creates instruction text OT1 for the agent. Fig. 35 is an example of instruction text. The instruction text OT1 in Fig. 35 includes the information acquired by the search as background knowledge, and includes a sentence instructing the agent to respond to the conversation as an AI assistant that conducts dialogues by voice in accordance with persona setting, which is an example of character setting.

[0141] (Step S945) Next, the processor 26 of the information processing device 2 sends information including the instruction text to the agent creation API to create an agent. Specifically, for example, processor 26 may upload a combination of at least one Markdown file of background knowledge information and a file ID identifying the Markdown file, and a conversation history text file. Processor 26 may also send the agent's instruction text, the file ID of the Markdown file of background knowledge information to be used, a data search command to be executed by the agent, and the name of the LLM (Large Scale Language Model) model to be used to the agent creation API. Furthermore, processor 26 may send agent information held by the facilitator (e.g., agent name, role, occupation, etc.) to the agent creation API. Thereby, the processor 26 may obtain, from the agent creation API, an assistant ID that identifies the assistant, which is the specific agent that has been created, and a thread ID that identifies a thread that represents a conversation session.

[0142] (Step S950) Next, processor 26 of information processing device 2 adds the agent to the participating member information on the facilitator system side.

[0143] (Step S955) Next, as an example, processor 26 of information processing device 2 adds the created agent as a member to the remote conference system. However, this is not limiting, and in some cases, the created agent may be added as a member to a chat system.

[0144] FIG. 36 is a flowchart showing an example of the details of the process in step S935 of FIG.

[0145] (Step S965) The processor 26 of the information processing device 2 requests a specific embedding API to convert the search text into a vector (numerical array). Here, the embedding API is an API that converts data such as text, images, or videos into a vector (numerical array). The search text is, for example, information related to a question about planning a cherry blossom viewing event.

[0146] (Step S970) Next, the processor 26 of the information processing device 2 calculates the distance based on the vector returned from the embedding API and acquires the data ID of the vector based on the distance in the vector search database of the storage device 23. Specifically, for example, the processor 26 acquires the data ID of the vector that is closest. Here, the vector that is closest may be the vector that is closest, or may be some or all of the vectors within a predetermined threshold distance.

[0147] (Step S975) Next, the processor 26 of the information processing device 2 acquires detailed information from the information recording database using the acquired data ID as a key.

[0148] At least a part of the information processing device 2 described in the above embodiment may be configured with hardware or software. If configured with software, a program that realizes at least a part of the functions of the information processing device 2 may be stored in a computer-readable recording medium and read and executed by a computer. The recording medium is not limited to removable media such as magnetic disks and optical disks, but may also be fixed recording media such as hard disk drives and memories.

[0149] In addition, a program that realizes at least a part of the functions of the information processing device 2 may be distributed via a communication line (including wireless communication) such as the Internet. Furthermore, the program may be encrypted, modulated, or compressed and distributed via a wired line or wireless line such as the Internet, or stored on a recording medium.

[0150] Furthermore, the information processing device 2 may be functioned by one or more information devices. When multiple information devices are used, one of the devices may be a computer, and the computer may execute a predetermined program to realize the functions of at least one means of the information processing device 2.

[0151] In the method invention, all processes (steps) may be realized by automatic control using a computer. Alternatively, each process may be performed by a computer, with progress control between processes being performed manually. Furthermore, at least some of the processes may be performed manually.

[0152] As described above, the present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0153] 1-1, ..., 1-N terminals 2. Information processing equipment 21 Input Interface 22 Communication Module 23 Storage device 24 memory 25 Output Interface 26 processors 3 Generation AI device S Information Processing System

Claims

1. An information processing device capable of exchanging information with a generation AI device that executes a generation AI that can output information according to input, at least one processor; The at least one processor: an acquisition step of acquiring utterances of one or more people and / or agents; a first request procedure for transmitting the content of the utterance to the generating AI device and requesting the generating AI device to output, based on the content of the utterance, at least one of a matter relating to the content of the utterance, a classification of at least one content of the utterance contained in the content of the utterance, or the emotion and / or state of the speaker in at least one utterance contained in the content of the utterance; a second request procedure for requesting the generating AI device to take an action regarding the next utterance based on the analysis results and utterance content based on the information obtained from the generating AI device as a result of the request in the first request procedure; an output step for outputting output information using an action related to the next utterance obtained from the generating AI device as a result of the request made in the second request step; An information processing device that executes the above.

2. Before making a request to the generating AI device in the second request procedure, the at least one processor requests the generating AI device to determine a next behavior and / or attitude to be taken based on the classification of at least one utterance obtained from the generating AI device and the emotion and / or state of the utterance; The analysis results based on the information obtained from the generating AI device include the requested results, the next behavior and / or attitude obtained from the generating AI device. The information processing device according to claim 1 .

3. The at least one processor, when requesting the next behavior and / or attitude, further requests a priority for each of the next behavior and / or attitude from the generating AI device; The analysis results based on the information obtained from the generating AI device further include a priority for each next behavior and / or attitude requested and obtained from the generating AI device. The information processing device according to claim 2 .

4. The matters relating to the utterance content obtained from the generating AI device include the topic of the utterance, Before making a request to the generating AI device in the second request procedure, the at least one processor requests the generating AI device to output personality information indicating what kind of person is suitable for the target conversation or discussion based on the topic of the utterance and / or the content of the utterance obtained from the generating AI device; The analysis results based on the information obtained from the generating AI device include person profile information indicating what kind of person is suitable to be added to the conversation of the target, obtained by requesting the generating AI device. The information processing device according to claim 1 .

5. The person profile information indicating what kind of person is suitable to be added to the conversation of the target is a person's background, a person's characteristics, or a role in the conversation or discussion; When requesting the generating AI device to estimate the person's background, the person's characteristics, or the role, the at least one processor further requests the generating AI device to provide a score representing the relevance to the topic for each of the person's background, the person's characteristics, or the role; The analysis results based on the information obtained from the generating AI device further include the score obtained from the generating AI device in response to the request. The information processing device according to claim 4 .

6. The action regarding the next utterance includes the next utterance and, if there is no person who corresponds to the next utterance, a statement that there is no person who corresponds to the next utterance; When the at least one processor receives from the generation AI device as a result of the request in the second request procedure that there is no suitable person for the next speaker, the processor creates a guest agent who fulfills the job type or the role and adds the created guest agent to the conversation or consideration. The information processing device according to claim 4 .

7. The action regarding the next utterance includes at least one of the following: the next speaker whom a facilitating agent who plays a role as a facilitator who is a progressor in a conversation or a discussion encourages to speak; the content of the facilitating agent's utterance; whether or not a guest agent who has a job title with knowledge of the topic or a role in the conversation or discussion that encourages utterances is to participate; and the tension or nuance of the facilitating agent's utterance. The information processing device according to claim 1 .

8. The at least one processor sends a character setting prompt to the generating AI device requesting setting of a character of an agent to be added to the conversation; creating an agent using at least the character profile settings obtained from the generating AI device in response to the sending of the character profile setting prompt; The information processing apparatus according to claim 1 , wherein the information processing apparatus executes the following steps:

9. The at least one processor sends an expert data collection prompt to the generating AI device requesting search keywords for collecting expert data about the agent to add to the conversation; collecting URLs by performing an internet search using search keywords obtained from the generating AI device in response to sending the specialized data collection prompt; A procedure for collecting data from the web page of the collected URL; creating an agent using at least the collected data; 9. The information processing apparatus according to claim 1, wherein the information processing apparatus executes the above.

10. the at least one processor sending a decision prompt to the generating AI device requesting it to decide whether the collected data is sufficient; If the determination result obtained in response to the sending of the determination prompt is not sufficient, further performing the step of collecting the URL, and if the determination result is sufficient, creating the agent; The information processing apparatus according to claim 9 , which executes the following:

11. The at least one processor sends a summary prompt to the generating AI device requesting a summary of the utterances; creating an agent to add to the conversation using at least the summary obtained from the generating AI device in response to the summary prompt; 9. The information processing apparatus according to claim 1, wherein the information processing apparatus executes the above.

12. an acquisition step in which an acquisition means acquires speech content of one or more people and / or agents; a first request procedure in which a first request means transmits the utterance content to a generating AI device that executes a generating AI capable of outputting information according to input, and requests the generating AI device to output at least one of a matter relating to the utterance content, a classification of at least one utterance content included in the utterance content, or the emotion and / or state of the speaker in at least one utterance included in the utterance content, based on the utterance content; a second request procedure in which a second request means requests the generating AI device to take an action regarding the next utterance based on the analysis results and utterance content based on the information obtained from the generating AI device as a result of the request made in the first request procedure; an output step in which an output means outputs output information using an action related to the next utterance obtained from the generating AI device as a result of the request made in the second request step; An information processing method comprising:

13. On the computer, an acquisition step of acquiring utterances of one or more people and / or agents; a first request procedure for sending the content of the utterance to a generating AI device that executes a generating AI capable of outputting information according to input, and requesting the generating AI device to output at least one of a matter relating to the content of the utterance, a classification of at least one content of the utterance contained in the content of the utterance, or the emotion and / or state of the speaker in at least one utterance contained in the content of the utterance, based on the content of the utterance; a second request procedure for requesting the generating AI device to take an action regarding the next utterance based on the analysis results and utterance content based on the information obtained from the generating AI device as a result of the request in the first request procedure; an output step for outputting output information using an action related to the next utterance obtained from the generating AI device as a result of the request made in the second request step; A program to execute.

Citation Information

Patent Citations

  • Conference support system

    JP2021163405A