Agent system

The hierarchical agent system with generative AI components addresses the limitations of existing agent-based AI by enabling autonomous task decomposition and execution, enhancing user interaction and long-term memory, thus facilitating business transformation.

JP2026002395APending Publication Date: 2026-01-08NOMURA RESEARCH INSTITUTE

Patent Information

Application Number
JP2024100361
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing agent-based AI systems using generative AI struggle to comprehensively and flexibly interpret user intentions, autonomously break down abstract instructions into specific tasks, and execute them without detailed user instructions, lacking long-term memory and realistic judgment capabilities.

Method used

A hierarchical agent system with a task layer and meta layer, utilizing multiple generative AI components for dialogue, task design, execution, and monitoring, incorporating long-term memory and interactive dialogue to understand user context and execute tasks autonomously.

Benefits of technology

Enables agent-based AI to understand user work background and situation, autonomously decompose abstract instructions into specific tasks, and execute them effectively, reducing user workload and facilitating business transformation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026002395000001_ABST
    Figure 2026002395000001_ABST
Patent Text Reader

Abstract

To understand the job background and job conditions of a user, to autonomously decompose an abstract instruction into concrete tasks and to execute them in an agent type AI.SOLUTION: A dialogue unit 11, a resolving unit 12, an executing unit 13, and a monitoring unit 14, each of which can individually use generated AI, the dialogue unit 11 grasping a work instruction through a dialogue with a user 2 using the generated AI and storing contents of the dialogue in a short-term memory 16 as a dialogue log, the resolving unit 12 decomposing the work instruction into tasks using the generated AI to create a task list, passing an instruction to execute each task to the executing unit 13, and presenting a result of executing each task to the user 2 via the dialogue unit 11, the monitoring unit 14 refers to the dialogue log at any time, grasps the context of the dialogue by the generated AI, predicts the content to be dealt with next, and stores the content as a summary in the short-term memory 16 so that the content can be referred to at any time. In the example of FIG. AI.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technology for utilizing generative AI (Artificial Intelligence), and in particular to a technology that is effective when applied to an agent system that realizes agent-type AI. [Background technology]

[0002] Advances in IT technology have led to a situation in which a vast amount of information is exchanged in various workplaces, exceeding the capacity of human judgment, and multiple business tasks are being performed in parallel. To assist with such cognitively demanding work situations, researchers are exploring the use of agent-based AI, including ChatGPT (registered trademark), and large-scale language models (LLMs) (hereinafter collectively referred to as "generative AI"). Agent-based AI here refers to AI that autonomously sets goals to be achieved according to human instructions and autonomously executes the necessary tasks toward those goals. Research has been conducted into whether such agent-based AI can be realized using generative AI. However, existing agent-based AI faces various challenges when used in actual workplaces, preventing it from achieving fundamental business transformation.

[0003] As an example of a technology related to agent-based AI, Japanese Patent Application Laid-Open No. 2018-81444 (Patent Document 1) describes a method in which multiple dialogue agent means for providing various services are provided for each service, and highly specialized services are provided by specializing each of them in that service, while when each dialogue agent is unable to respond on its own, the dialogue is transferred to another dialogue agent, thereby guiding the user to a dialogue agent that can provide a more appropriate response. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-81444 Summary of the Invention [Problem to be solved by the invention]

[0005] According to the conventional technology described in Patent Document 1, by providing multiple highly specialized AIs for each service and matching them with AIs that can handle the issues, it is possible to handle a variety of business operations and tasks, and also reduce installation and operation costs compared to building a universal agent system.

[0006] However, with conventional agent systems, it has been difficult to realize an agent-type AI that can respond comprehensively and flexibly, for example, by interpreting a user's intentions from natural interactions with the user, autonomously constructing necessary tasks, and executing each task without requiring detailed instructions from the user.

[0007] Therefore, the object of this invention is to provide an agent system that utilizes generative AI to understand the user's work background and work situation, and autonomously breaks down abstract instructions into specific tasks and executes them.

[0008] The above and other objects and novel features of the present invention will become apparent from the description of this specification and the accompanying drawings. [Means for solving the problem]

[0009] Among the inventions disclosed in this application, the outline of representative inventions will be briefly explained as follows.

[0010] The agent system, which is a representative embodiment of the present invention, is an agent system that understands business issues through dialogue with a user and executes tasks related to the business instructions, and has a dialogue unit, a solution unit, an execution unit, and a monitoring unit, each of which can individually use a generation AI.

[0011] The dialogue unit then grasps the work instructions through a dialogue with the user using the generation AI and stores the content of the dialogue with the user as a dialogue log in the short-term memory unit; the resolution unit breaks down the work instructions grasped by the dialogue unit into tasks using the generation AI to create a task list, passes execution instructions for each task to the execution unit, and presents the execution results by the execution unit to the user via the dialogue unit; the execution unit performs work using the corresponding data source using the generation AI corresponding to the task related to the execution instruction passed from the resolution unit, and passes the execution results to the resolution unit; the monitoring unit refers to the dialogue log at any time, extracts information related to predetermined matters, grasps the context of the dialogue using the generation AI, predicts the next response content, stores the summary in the short-term memory unit, and makes it possible to refer to the dialogue unit, the resolution unit, and the execution unit at any time. [Effects of the Invention]

[0012] The effects obtained by the representative inventions disclosed in this application can be briefly explained as follows.

[0013] In other words, according to a representative embodiment of the present invention, it is possible to realize an agent system in which an agent-type AI utilizing a generative AI understands the user's work background and work situation, and autonomously breaks down abstract instructions into specific tasks and executes them. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a diagram showing an overview of an example of the configuration of an agent system according to an embodiment of the present invention; [Figure 2] FIG. 1 is a diagram showing an overview of an example of the architecture of an agent system according to an embodiment of the present invention. [Figure 3] FIG. 1 is a diagram showing an outline of an example of how "Mechanism 1" solves [Problem 1] in one embodiment of the present invention. [Figure 4]FIG. 10 is a diagram showing an outline of an example of how "Mechanism 2" solves [Problem 2] in one embodiment of the present invention. [Figure 5] FIG. 10 is a diagram showing an outline of an example of how "Mechanism 3" solves [Problem 3] in one embodiment of the present invention. [Figure 6] FIG. 10 is a diagram showing an overview of an example of how "Mechanism 4" solves [Problem 4] in one embodiment of the present invention. [Figure 7] FIG. 1 is a diagram showing an outline of an example of a processing flow in an embodiment of the present invention. [Figure 8] FIG. 1 is a diagram showing an outline of an example of a processing flow in an embodiment of the present invention. [Figure 9] FIG. 1 is a diagram showing an outline of an example of a processing flow in an embodiment of the present invention. [Figure 10] FIG. 1 is a diagram showing an outline of an example of a processing flow in an embodiment of the present invention. [Figure 11] FIG. 1 is a diagram showing an outline of an example of a processing flow in an embodiment of the present invention. [Figure 12] FIG. 1 is a diagram showing an outline of an example of a processing flow in an embodiment of the present invention. [Figure 13] FIG. 1 is a diagram showing an outline of an example of a processing flow in an embodiment of the present invention. [Figure 14] FIG. 10 is a diagram outlining an example of a prompt of a dialogue unit that generates a response to a user in one embodiment of the present invention. [Figure 15] FIG. 10 is a diagram outlining an example of a resolver prompt for creating a task list in one embodiment of the present invention. [Figure 16] FIG. 10 is a diagram outlining an example of a monitor prompt summarizing a dialogue log in accordance with one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In all drawings used to explain the embodiments, the same parts are generally designated by the same reference numerals, and repeated explanations will be omitted. However, parts that have been designated and explained in one drawing may be referred to by the same reference numerals in the explanation of other drawings, although they will not be shown again.

[0016] <Summary> Currently available generative AI can, for example, perform various tasks (such as information acquisition and information processing) in response to text instructions and respond in text. Attempts are also being made to use generative AI to realize agent-type AI that can respond autonomously.

[0017] On the other hand, realizing agent-type AI faces the following challenges that cannot be resolved using existing generative AI technology alone.

[0018] [Task 1] Existing agent-type AI using generative AI can autonomously execute tasks according to user instructions, but it cannot control whether or not a task can be executed and the execution status according to the user's interactive dialogue.

[0019] [Task 2] Existing agent-type AI using generative AI can remember the content of the most recent conversation, but cannot accumulate, store, or reference the content of past conversations over the long term. As a result, it is unable to reference the accumulated knowledge gained through the accumulation of individual conversations, such as past work content with the user, current progress, and prerequisite knowledge required for the work. Instead, the user must provide the necessary background knowledge each time a conversation begins. In this situation, it is impossible to understand work instructions based on the work context through dialogue.

[0020] [Task 3] Agent-type AIs using existing generative AIs are designed to execute individual tasks or a predetermined series of tasks, and when the user provides specific and detailed instructions, they execute the individual tasks or a predetermined series of tasks that match the instructions. In this case, agent-type AIs using existing generative AIs determine whether the user's individual instructions match the implementation requirements of the specified task and then decide what to execute. Therefore, they cannot respond to instructions that are not specified in the implementation requirements. For this reason, users must provide specific and detailed instructions that are in line with the implementation requirements and that the agent-type AI can determine, and must execute the instructions while taking into account the task judgment criteria of the agent-type AI.

[0021] [Task 4] Existing agent-based AI using generative AI compares the user's instructions with the implementation requirements to determine whether to execute the task. However, this determination often deviates from realistic judgment. This is because the generative AI makes its determination based on the semantic similarity of the text, and is unable to determine the feasibility of the task, or because it does not recognize whether the instructions are within its own range of solvability (a similar problem has been raised as the AI ​​frame problem).

[0022] In the agent system according to one embodiment of the present invention, mechanisms for solving the above-mentioned problems are introduced as follows (details of each mechanism will be described later). "Mechanism 1": Controlling agents through interactive dialogue (corresponding to [Problem 1]) "Mechanism 2" Long-term memory maintenance (corresponding to [Problem 2]) "Mechanism 3": Autonomous decomposition of abstract and complex instructions (corresponding to [Problem 3]) "Mechanism 4": Addressing the frame problem (addressing [Problem 4])

[0023] Since it is not realistic to use a single agent (generative AI) to solve the above-mentioned problems of existing generative AI or agent-based AI using generative AI, in this embodiment, multiple agents are provided, each with a set of roles, and a hierarchy is created. That is, the entire system is divided into two layers: a "task layer" that designs and executes tasks, and a "meta layer" that monitors and stores the overall situation. In the task layer, the processing steps are divided into "problem extraction," "task design," and "task execution," and agents are assigned to handle each process, thereby realizing an agent-based AI that can properly understand the user's business problems and execute them by breaking them down into individual, specific tasks.

[0024] FIG. 2 is a diagram outlining the architecture of an agent-based AI in one embodiment of the present invention. In the example of FIG. 2, the lower task layer is composed of the generation AIs, i.e., the dialogue AI (11a), the solution AI (12a), and the execution AI (13a), while the upper meta layer is composed of the generation AI, i.e., the monitoring AI (14a). The task layer has the function of breaking down complex tasks into parts for each generation AI, executing them, and integrating the results. Meanwhile, the meta layer has the function of extracting, understanding, and accumulating business issues and the design know-how of specific tasks required to solve them from dialogue with the user, and appropriately referencing them to each generation AI in the task layer.

[0025] The task-layer dialogue AI (11a) has the function of interacting with User 2 and responding in accordance with User 2's implicit situation. The solution-making AI (12a) receives abstract task instructions from User 2, breaks them down into specific tasks based on accumulated business know-how, and integrates the results of the tasks. In other words, the aforementioned "mechanism 3" realizes autonomous task design from highly abstract instructions. The execution AI (13a) is composed of individual generation AIs for each specialized task, each of which receives instructions from the solution-making AI (12a), executes the task of retrieving and processing data from the corresponding data source 3, and returns the execution results to the solution-making AI (12a). The solution-making AIs (12a) interact with each other to recognize the feasible range, thereby achieving the aforementioned "mechanism 4" for dealing with the frame problem.

[0026] The monitoring AI (14a) in the meta layer has the function of monitoring the context from the dialogue content with the user, understanding the context, and making each generation AI in the task layer operate in accordance with the context. At this time, if the dialogue content is insufficient in understanding the context, it interactively queries the user 2 via the dialogue AI (11a) to supplement the information and lead to the modification of the task. In other words, as the above-mentioned "mechanism 1," it realizes interactive task design and modification through dialogue. It also has the function of acquiring business know-how, etc. from the dialogue content, storing it as explicit knowledge and making it available for reference.

[0027] In this embodiment, the history of the content of the dialogue with user 2 is classified and stored into four types consisting of combinations of long-term / short-term / objective / subjective, and the timing and method of use for each type are organized. For example, the current dialogue with user 2 (dialogue log) is stored in memory as short-term memory 16 so that it can be referenced in real time, while past dialogues (session logs) are stored in a database as long-term memory 15 and can be referenced by retrieval when necessary to generate a response (RAG: Retrieval-Augmented Generation).

[0028] Furthermore, the system distinguishes between objective factual memory such as dialogue logs and subjective, variable memory such as dialogue summaries and information related to the next response predicted from the summary content. The former, which is retained in real time during the dialogue, is stored as short-term objective memory 16a, and the latter as short-term subjective memory 16b. After the dialogue, the text data of the objective dialogue log is stored in a database as long-term objective memory 15a, and the text data of the subjective dialogue record summary is stored in a database as long-term subjective memory 15b, so that they can be referenced appropriately during future responses. In other words, these enable the aforementioned "Mechanism 2" to realize the functions of dialogue summarization and long-term memory management.

[0029] This architecture overcomes the challenges of existing agent-based AI that use generative AI, as described above, and makes it possible to realize agent-based AI that can properly understand the execution situation of business operations and autonomously design and execute tasks.

[0030] In other words, the agent-based AI in this embodiment has a mechanism for extending agent-based AI that utilizes existing generative AI. It understands the business background and situation through dialogue with the user, clarifies business issues, and recognizes the goal to be achieved, without requiring specific and detailed instructions from the user. It then autonomously lists the individual tasks required to achieve the goal based on the user's abstract instructions, breaks them down into specific instructions, and, after agreeing with the user on the work content, autonomously executes each task to achieve the goal. This assists the user's work, reduces their workload, contributes to innovation in business flow, and realizes fundamental business transformation.

[0031] <System configuration> 1 is a diagram showing an overview of an example of the configuration of an agent system according to one embodiment of the present invention. The agent system 1 is configured, for example, with server equipment or a virtual server built on a cloud computing service, and realizes various functions of an agent-based AI by using a central processing unit (CPU) (not shown), which uses middleware such as an operating system (OS) and a database management system (DBMS), a web server program, and other software running on the OS and the responsive application programming interfaces (APIs) of various LLM services, or an LLM model built in a local environment.

[0032] The agent system 1 has various parts (AI mechanisms) implemented as software, such as a dialogue unit 11, a solution unit 12, an execution unit 13, and a monitoring unit 14. It also has various data stores, such as a long-term memory 15, a short-term memory 16, and setting information 17, implemented as databases, files, etc.

[0033] The dialogue unit 11 is an AI mechanism that includes (or uses) the dialogue AI (11a) in FIG. 2 described above, and has the function of engaging in dialogue with the user 2 via input text from the user 2. Through the dialogue with the user 2, the dialogue unit 11 has the function of extracting issues and interpreting and clarifying work instructions. When engaging in dialogue with the user 2, the dialogue content can be made unique to the user 2, for example, by referring to various information set in the setting information 17. The content of the dialogue is stored in the short-term memory 16 in real time as a dialogue log.

[0034] The solving unit 12 is an AI mechanism that includes (or uses) the solving AI (12a) in Figure 2 described above, and has the function of breaking down the work instructions clarified by the dialogue unit 11 into specific work work items (tasks) and instructing the execution unit 13, described below, to execute each task. This eliminates the need for specific and detailed instructions from the user 2 for each task, and even complex and abstract work instructions can be converted into appropriate work instructions by the solving unit 12 to obtain the desired work results, thereby resolving the above-mentioned [Problem 3], which is the need for specific and detailed instructions.

[0035] The solving unit 12 then formats the execution instructions for each task as text information including information required by the execution AI (13a) corresponding to the task, lists it, and passes the work instruction text for each target execution AI (13a) to the execution unit 13 (described later), thereby executing each task and obtaining the results. The solving unit 12 summarizes the content of the work performed by the execution unit 13 and presents it to the user 2 through the dialogue unit 11.

[0036] The execution unit 13 is a collection of AI mechanisms that includes (or uses) multiple execution AIs (13a) shown in FIG. 2, and each AI mechanism has the function of executing each of the decomposed business work items (tasks). Each AI mechanism has a corresponding individual data source 3. The data source 3 is, for example, various databases, business management applications, or APIs of external data sources. Each AI mechanism understands the work procedure for the corresponding data source 3 (for example, information acquisition or information processing), executes the work by receiving appropriate instructions from the solution unit 12, and acquires data from the data source 3.

[0037] If the instruction input passed from the solving unit 12 does not include information necessary for the AI ​​mechanism to execute processing, the information recorded in the short-term memory 16 is referenced via the monitoring unit 14, which will be described later, to check whether the necessary information exists. If the information does not exist in the short-term memory 16, the monitoring unit 14 may instruct the dialogue unit 11 to make an inquiry to the user 2, and the dialogue unit 11 may acquire the necessary information by inquiring about the missing information from the user 2.

[0038] Furthermore, if the instructions from the resolving unit 12 cannot be resolved using the data source 3 of the target execution AI (13a), the resolving unit 12 responds to that effect and requests a change to the work content. As a result, the resolving unit 12 refers to the indications from the execution AI (13a) regarding the work instruction and corrects the work instruction. When correcting the work instruction, the resolving unit 12 refers to the dialogue history information (short-term memory 16) created by the monitoring unit 14, which will be described later, and makes corrections in accordance with the work instruction of the user 2. Through dialogue between the resolving unit 12 and the execution unit 13, the unexecutable task in the above-mentioned [Problem 4] is addressed.

[0039] The monitoring unit 14 is an AI mechanism that includes (or uses) the monitoring AI (14a) of Figure 2 described above, and has the function of interpreting text information exchanged between the user 2, the dialogue unit 11, the resolution unit 12, and the execution unit 13, organizing the information, and storing it in the long-term memory 15 and the short-term memory 16.

[0040] That is, the monitoring unit 14 sequentially refers to the text information exchanged between the user 2, the dialogue unit 11, the resolution unit 12, and the execution unit 13, extracts the status of information that is considered important when executing a task, and stores it as a summary. This information is stored as text data in the short-term memory 16 and is referenced in real time by the dialogue unit 11, the resolution unit 12, and the execution unit 13. Important information that the monitoring unit 14 should extract is, for example, defined in advance in the setting information 17, and the generation AI function extracts this information using sentences and keywords from the text information exchanged between each unit, grasps the context of the dialogue session, and predicts and suggests the next response content.

[0041] The monitoring unit 14 then organizes information such as the work instructions and work items given by the user 2 via the dialogue unit 11, and the execution results of the execution unit 13 as business know-how, and stores it in the long-term subjective memory 15b. The long-term subjective memory 15b holds the organized information as text data and embeddings (embedded representations) obtained by the monitoring AI (14a). For example, in subsequent dialogues with the user 2, this information is used to understand the user 2's current situation, work orientation, and work background, and is also used as reference information when the user 2 executes a work instruction similar to a past task. This allows the dialogue unit 11, the solution unit 12, and the execution unit 13 to interpret the user 2's intentions and orientations, overcome the above-mentioned [Problem 1], and improve understanding of the business background and work instructions.

[0042] In the configuration example of Figure 1, the dialogue unit 11, resolution unit 12, execution unit 13, and monitoring unit 14 are each configured as an AI mechanism equipped with a generation AI (or using a generation AI), but these generation AIs may each use a different generation AI system or service, or one or more generation AIs may use the same generation AI system or service. Also, in the configuration example of Figure 1, the above-mentioned units are shown as being all included within the server system of agent system 1, but this is merely a logical configuration, and it goes without saying that one or more units may physically be configured as subsystems by different server systems and function in conjunction with each other.

[0043] <Methods for solving problems> As described above, this embodiment is an agent-type AI that, by being equipped with the above-mentioned "Mechanism 1" to "Mechanism 4," overcomes the above-mentioned [Problem 1] to [Problem 4] that exist in existing generative AI and agent-type AI that uses generative AI, and enables the realization of fundamental business reform.

[0044] ■ Solving [Problem 1] with "Mechanism 1" As an improvement to solve [Problem 1], in which existing agent-type AI using generative AI cannot control whether a task can be executed and the execution status according to an interactive dialogue with user 2, in this embodiment, the dialogue unit 11 receives the task list created by the solution unit 12 and presents it to user 2, and if user 2 agrees to execute the task, it instructs the solution unit 12 to execute the task. At that time, user 2 can request additional tasks to be added to the proposed task list or request modifications to the content of the proposed task.

[0045] Such a dialogue between the dialogue unit 11 and the user 2 is recorded as an dialogue log as needed, and a summary is created by the monitoring unit 14. Then, the resolution unit 12 creates instructions for the work to be performed. The resolution unit 12 also references the matters to be addressed, the dialogue log, and the dialogue summary created by the monitoring unit 14, and if the user 2 requests a revision of the task list, it carries out the revision. Furthermore, when the execution unit 14 actually executes the task list with the consent of the user 2, the dialogue unit 11 can also inquire of the user 2 about information necessary to execute the task, as necessary.

[0046] In this way, in this embodiment, the dialogue unit 11 and the monitoring unit 14 cooperate to store the response items requested by the user 2, i.e., the response items that the agent system 1 should carry out next, in the short-term subjective memory 16b, thereby enabling the requested task to be carried out autonomously while maintaining an interactive dialogue with the user 2.

[0047] Figure 3 is a diagram outlining an example of how "Mechanism 1" solves [Problem 1] in one embodiment of the present invention. In the diagram, the left side shows an example screen of the content of a dialogue between User 2 and a general generation AI, and the right side shows an example of a solution attempted by Agent System 1 of this embodiment in a similar situation (the same applies to the following Figures 4 to 6). In the general dialogue on the left, the agent AI unilaterally interprets an abstract inquiry from User 2 and proceeds with the dialogue, whereas in the dialogue of this embodiment on the right side, the user and Agent System 1 interactively converse to refine and concretize the content of the task.

[0048] Specifically, the monitoring unit 14 (monitoring AI (14a)) aggregates important information about the dialogue from the most recent dialogue log (short-term objective memory 16a) as short-term memory 16 (short-term subjective memory 16b), searches past dialogue history (long-term subjective memory 15b) based on that information, and determines in real time when the topic is changing by referring to the most similar history, allowing user 2 to modify the content of the task as needed during the dialogue and interactively input instructions that match his or her intentions to the dialogue unit 11 (dialogue AI (11a)). Furthermore, since the monitoring unit 14 previously grasps the minimum information required to execute each task by referring to the setting information 17, etc., even if the instructions from user 2 lack information required to execute the task, it is possible to design an executable task by inquiring of user 2 before executing the task.

[0049] ■ Solving [Problem 2] with "Mechanism 2" In order to overcome [Problem 2], in which an agent-type AI using an existing generative AI can remember the contents of the most recent dialogue but cannot accumulate, store, and refer to the contents of past dialogues over the long term, this embodiment enables the agent-type AI to continue to remember the work background and work situation over the long term from the contents of the dialogue with user 2. That is, when the dialogue with user 2 ends, the monitoring unit 14 extracts and summarizes information that is considered important for understanding the current work background and future work instructions from the contents of the dialogue, such as the user 2's request, the task design content by the agent, the task execution results, the user 2's evaluation of the execution results, or the user 2's request orientation to the agent, and stores this information in a database as long-term subjective memory 15b.

[0050] This summary contains extracted keywords that are important for understanding the business background and situation, and is stored in a database so that they can be used when searching dialogue records. This mechanism allows the agent-based AI to appropriately refer to summary information from past dialogues similar to the current dialogue when it next engages in dialogue with User 2, helping to ensure smooth dialogue with User 2.

[0051] Figure 4 is a diagram outlining an example of how "Mechanism 2" solves [Problem 2] in one embodiment of the present invention. In the dialogue with a general generating AI on the left, the agent-type AI does not remember the answer it gave in the dialogue on April 1st and is unable to respond in the dialogue on May 2nd, but in the dialogue according to this embodiment on the right, by searching past dialogues, it is able to respond in the dialogue on May 2nd based on the answer it gave in the dialogue on April 1st.

[0052] Specifically, the dialogue unit 11 (dialogue AI (11a)) acquires and references "user preferences" and "recent dialogue history" from the long-term memory 15 using RAG, thereby realizing a conversation that inherits past memories. In this case, by performing RAG on the summary recorded in the long-term subjective memory 15b rather than the raw dialogue log of the dialogue session recorded in the long-term objective memory 15a, the accuracy of the search can be improved. This is because the raw dialogue log contains a majority of statements made by the dialogue AI (11a), making it difficult to reference the user's statements and intentions. By summarizing the dialogue content at the end of the dialogue session in the monitoring unit 14 (monitoring AI (14a)), necessary information can be consolidated. For example, by summarizing the characteristics of user 2, such as "Mr. A prefers concise responses," it is possible to provide personalized responses according to each user 2's preferences.

[0053] ■ Solving [Problem 3] with "Mechanism 3" In order to solve [Problem 3], which is that agent-type AI using existing generative AI cannot handle abstract or complex instructions that are not specified in the implementation requirements, we will make it possible for User 2 to give abstract or complex work instructions in the form of natural dialogue, without taking into account the agent-type AI's judgment criteria.

[0054] That is, in this embodiment, when the dialogue unit 11 and the agent-type AI reach an agreement on highly abstract instructions and complex task content through dialogue, the dialogue unit 11 issues an execution queue to the resolution unit 12, and the resolution unit 12 breaks down the work instructions of user 2 into specific work items by referring to the dialogue summary created by the monitoring unit 14, the next work to be performed, user 2's instructions, and past instructions and task design results of user 2 that are similar to the instructions of user 2 obtained from the long-term subjective memory 15b.

[0055] Regarding the method of breaking down work, the prompts input to the solving unit 12 provide examples of work content that the executing unit 13 can perform, and examples of work goals and task lists that should be created based on instructions from the user 2, so the solving unit 12 can autonomously create a task list that the user 2 desires based on this information. This task list indicates the name of the execution AI (13a) that should perform the task in the executing unit 13, and the specific work content to be given to the execution AI (13a), and by sequentially executing this, it is possible to execute the abstract and complex work instructions given by the user 2.

[0056] Figure 5 is a diagram outlining an example of how "Mechanism 3" solves [Problem 3] in one embodiment of the present invention. In the general dialogue on the left, a specific task request is received from the user (in the example shown, "inquiry about sales performance"), the agent-type AI determines the appropriate task to be performed ("extraction of sales performance"), and has the target generation AI ("sales performance AI") carry out that task.

[0057] In contrast, the dialogue in this embodiment on the right shows a situation in which, upon receiving an abstract task request from user 2 (in the example in the figure, "I want to create a sales strategy"), the solution unit 12 (solution AI (12a)) of agent system 1 specifies the task to be done ("check sales performance"), determines the execution unit 13 (execution AI (13a)) ("sales performance AI") that will execute the task, and causes it to be executed.

[0058] In this way, general generative AI and agent-type AI can respond appropriately to specific task execution instructions from the user, but cannot respond appropriately to abstract task execution instructions.In contrast, in this embodiment, by dividing the roles between the solving unit 12 (solving AI (12a)), which breaks down abstract tasks into concrete tasks, and the executing unit 13 (executing AI (13a)), which receives and executes the concrete task instructions, it is possible to appropriately interpret and execute abstract task execution instructions.

[0059] ■ Solving [Problem 4] with "Mechanism 4" In order to solve [Problem 4] (frame problem) where an agent-type AI using an existing generation AI often deviates from a realistic judgment when comparing the instructions of User 2 with the implementation requirements to determine whether or not to execute a task, in this embodiment, the agent-type AI modifies the task list it has created through dialogue with other AIs to determine whether the tasks are solvable using its own external data sources and executable functions (tools), and modifies them to within the range of feasibility.

[0060] The solving unit 12 creates a list of tasks to be performed, checks whether or not the user 2 has made any corrections through the dialogue unit 11, and if no corrections are necessary, sequentially reads out the task list and passes the created specific work instructions to the relevant execution AI (13a) of the listed execution unit 13. Each execution AI (13a) of the execution unit 13 has a data source 3 and an executable function (tool) for the data source 3.

[0061] Each execution AI (13a) corresponds to an agent-type AI that uses an existing generation AI, and in accordance with the specific and detailed instructions, it checks the execution criteria of its own executable functions against the input instructions and executes the corresponding function (tool). At this time, the target execution AI (13a) determines whether the passed work instruction is a task that it can execute and whether the instructions contain the parameters required for execution. If the work instruction is not executable, it points out to the resolution unit 12 that it is an inexecutable task and requests that the task be modified. In response to the request from the execution AI (13a), the resolution unit 12 modifies the corresponding task so that it becomes executable.

[0062] Furthermore, if the work instructions from user 2 are appropriate but do not include the parameters necessary for the work, the execution AI (13a) refers to the short-term subjective memory 16b created by the monitoring unit 14 and attempts to extract the parameters necessary for the work from the summary of the dialogue, and if the relevant information can be referenced, it autonomously complements the parameters and executes the given task.

[0063] If information capable of completing the relevant parameters is not provided even when referring to the short-term subjective memory 16b, the execution AI (13a) asks the user 2 about the necessary parameters via the dialogue unit 11 through the solution unit 12. When the user 2 answers the necessary parameters, the execution AI (13a) executes the task using the parameters. This makes it possible for the agent-type AIs to correct any deficiencies in the task instructions through interactive dialogue between other agent-type AIs or between the agent-type AI and the user 2, and to execute each task while avoiding the frame problem.

[0064] Figure 6 is a diagram outlining an example of how "Mechanism 4" solves Problem 4 in one embodiment of the present invention. The general dialogue on the left shows a situation in which a user inquires about sales performance for "males," but the task cannot be executed because the reference database does not contain a "gender" field. However, the agent-type AI is unable to grasp the limits (frame) of the tasks it can execute, and therefore designs an unexecutable task. In contrast, the dialogue on the right, in this embodiment, shows that the agent system 1 understands that the reference database does not contain a "gender" field, which is its own limit (frame), and therefore allows the agent-type AIs to interact with each other to design and execute appropriate tasks.

[0065] As described above, a typical agent-based AI may design a task that exceeds its own capabilities (frame problem). However, in this embodiment, in response to a user request, the AI ​​can refer to "business knowledge (past findings)" using RAG to acquire similar tasks and the details of tasks previously performed, understand its own limitations (the range of feasibility), and design an appropriate task. Furthermore, at the end of the session, the created task design is stored in the long-term objective memory 15a by the monitoring unit 14, linking the instructions from user 2, the created task design, and user 2's evaluation of the execution results. This information is then used for future task design. This allows for task design based on user 2's preferences when similar instructions are executed next time.

[0066] <Processing flow> 7 to 13 are diagrams outlining an example of a processing flow in an embodiment of the present invention. Here, as an example of the flow of task execution through dialogue between user 2 and agent system 1, a situation is assumed in which user 2, who is a person in charge of a franchise store in a franchise store management organization, uses agent system 1 to collect related information in advance when discussing with the store he is in charge of what kind of sales promotion activities should be carried out for a specific event.

[0067] Specifically, the system engages in a dialogue with User 2 while referencing the history of past dialogues with User 2, designs a task in line with User 2's requests while understanding User 2's work background and work situation, and modifies the task design while engaging in an interactive dialogue. The designed task is then executed by each execution AI (13a), and if there are any deficiencies or insufficiencies in the execution instruction input, User 2 is queried. The task is then executed, the results are collected, and the collected data is summarized and presented to User 2. When the dialogue ends, the monitoring unit 14 stores the content of the dialogue with User 2 for use in future responses, and also stores the work results and User 2's evaluation of them, thereby extracting User 2's preferences and contributing to long-term storage and reference of the dialogue with User 2.

[0068] In Fig. 7, first, user 2 uses an information processing terminal such as a PC or smartphone (not shown) to start an application that accesses agent system 1 (or accesses the services of agent system 1 using a web browser). At the time of start-up, dialogue unit 11 acquires various information such as information about user 2, date, store information, and event information from setting information 17. In addition, a summary log of user 2's previous session is acquired from long-term memory 15.

[0069] The dialogue unit 11 then creates a prompt to generate a dialogue for the user 2, and the dialogue AI (11a) generates a greeting at the start of the session and presents it to the user 2. In the example shown in the figure, the work status of the most recent session from the current date is confirmed, and information about upcoming events and the previous year is presented, along with a question about which event the user would like to discuss.

[0070] The greeting presented by the dialogue unit 11 and the text information of the response from the user 2 to this greeting ("For Christmas at Store A...") are stored as a dialogue log in the short-term memory 16 (short-term objective memory 16a). After that, the monitoring unit 14 uses the monitoring AI (14a) to create a summary from the contents of the dialogue log recorded in the short-term memory 16, and stores this as a summary log in the short-term memory 16 (short-term subjective memory 16b). At the end of the summary, the content determined as the next action to be taken based on the dialogue content ("Propose a plan for Christmas at Store A") is written.

[0071] 8, the dialogue unit 11 acquires and refers to the input contents including the response from the user 2, the summary created by the monitoring unit 14 (monitoring AI (14a)), and the content of the next action to be taken described in the summary from the short-term memory 16, creates a prompt to determine whether the next action will be taken by the resolution unit 12 or whether the dialogue unit 11 itself will continue to take the action, and inputs this to the dialogue AI (11a) to determine the AI ​​mechanism that will take the next action. In the example of FIG. 8, it is assumed that it has been determined that the dialogue unit 11 will continue to take the action next time.

[0072] As the next response, the dialogue unit 11 searches past dialogue logs stored in the long-term memory 15 based on the input content from user 2, and acquires summaries of similar past dialogues (sessions) based on the search results. Then, based on the input content from user 2, the dialogue log and summary of the current session, and summaries of past sessions, it creates a prompt to generate a response content for user 2, and inputs this to the dialogue AI (11a), which then creates and presents the response content to user 2. In the example shown in the figure, it shows the response content from last year and asks whether it will be considered again this year.

[0073] FIG. 14 is a diagram outlining an example of a prompt template for the dialogue unit 11 that generates a response to a user in one embodiment of the present invention. In the diagram, "{summary_text}" is read from the short-term subjective memory 16b and replaced. Also, "{user_profile}" is replaced by reading the corresponding value from the setting information 17. Furthermore, a summary of a similar past dialogue is acquired from the long-term subjective memory 15b and embedded in the variable "{reference_text}." By substituting variables in this way, a prompt can be dynamically created and input to the dialogue AI (11a), allowing a personalized response to user 2 to be generated, as shown in the example of FIG. 8.

[0074] Returning to Figure 8, the response content presented by the dialogue unit 11 and the text information of the response from user 2 to that response are each stored as a dialogue log in short-term memory 16 (short-term objective memory 16a). Thereafter, in the monitoring unit 14, a monitoring AI (14a) creates a summary from the content of the dialogue log recorded in short-term memory 16, and stores this as a summary log in short-term memory 16 (short-term subjective memory 16b). At the end of the summary, the content determined as the next action to be taken based on the dialogue content ("Propose the content of the task to consider the number of Christmas products to be purchased at Store A") is written.

[0075] Next, moving to Figure 9, the dialogue unit 11 acquires and refers to the input contents including the response from the user 2, the summary created by the monitoring unit 14 (monitoring AI (14a)), and the content of the next action to be taken described in the summary from the short-term memory 16, creates a prompt to determine whether the solution unit 12 will take the next action or whether the dialogue unit 11 itself will continue to take the next action, and inputs this to the dialogue AI (11a), thereby determining the AI ​​mechanism that will take the next action. In the example of Figure 9, it is assumed that the solution unit 12 has determined that task design will be performed next, and the dialogue unit 11 notifies the user 2 of this ("Task design will be performed").

[0076] FIG. 15 is a diagram outlining an example of a prompt of the solving unit 12 that creates a task list in one embodiment of the present invention. The diagram instructs the execution unit 13 to break down the task while referring to the "user instructions" and create a task list that can be executed. For tasks that can be executed by the execution unit 13, the "available data" contains the name of each execution AI (13a) and a description of the functions (tools) that can be executed. In addition, the "past cases" contains task lists created in response to past user instructions similar to the input content from user 2.

[0077] The task list is specified in a format such as JSON (JavaScript Object Notation), and is output as a list that can be decomposed so that the application can execute the tasks sequentially. In this example, the user's instructions are decomposed into individual tasks, and specific instructions for functions (tools) that can be executed by each execution AI (13a) are output as subtasks required for each task. This task decomposition procedure makes it possible to break down abstract and complex instructions from User 2 into concrete ones.

[0078] Returning to FIG. 9, when the resolution unit 12 receives an instruction from the dialogue unit 11 to take the next action, it refers to past dialogue (session) logs and summary logs stored in the long-term memory 15 based on the input content from the user 2, searches for similar input from the user 2 in the past, and reads out and acquires from the long-term memory 15 examples of task design that correspond to the obtained past input (dialogue log).

[0079] The acquired task design examples are then referenced, a prompt for executing the task design corresponding to the input of user 2 is created, and the prompt is input to the solution AI (12a), thereby performing task design. The results of the task design are stored in the short-term memory 16 as a task list in JSON format, which describes what instructions should be given to which execution AI (13a) for each work item. The JSON format task list is then formatted into text and presented to user 2 via the dialogue unit 11.

[0080] In the example of FIG. 9, in response to the task list presented by the dialogue unit 11, the user 2 responds by saying, "Please limit the target products to cakes," and this text information is stored as a dialogue log in the short-term memory 16 (short-term objective memory 16a). Thereafter, the monitoring unit 14 uses the monitoring AI (14a) to create a summary from the contents of the dialogue log recorded in the short-term memory 16, and stores this as a summary log in the short-term memory 16 (short-term subjective memory 16b). At the end of the summary, the content determined as the next action to be taken based on the dialogue content ("As a next action, the agent needs to modify the task") is described.

[0081] 10, the solution unit 12 acquires and refers to the input contents including the response from the short-term memory 16, the summary created by the monitoring unit 14 (monitoring AI (14a)), and the content of the next action to be taken described in the summary, and determines whether the task needs to be modified. In the example of FIG. 10, it is assumed that it has been determined that the task needs to be modified, and a message to that effect ("Task modification will be implemented") is presented to the user 2 via the dialogue unit 11. The task is then modified by referring to the dialogue log stored in the short-term memory 16 and the designed task list, and the content of the modified task is presented to the user 2 via the dialogue unit 11.

[0082] The content of the task presented by the dialogue unit 11 and the text information of the response from the user 2 to that task are stored as a dialogue log in the short-term memory 16 (short-term objective memory 16a). After that, the monitoring unit 14 uses the monitoring AI (14a) to create a summary from the content of the dialogue log recorded in the short-term memory 16, and stores this as a summary log in the short-term memory 16 (short-term subjective memory 16b). At the end of the summary, the content determined as the next action to be taken based on the dialogue content ("perform the task list") is written.

[0083] 11, the solution unit 12 retrieves and refers to the input contents, including the response from the user 2, the summary created by the monitoring unit 14 (monitoring AI (14a)), and the next action to be taken described in the summary from the short-term memory 16, and determines whether the task needs to be revised. In the example of FIG. 11, it is assumed that the solution unit 12 determines that the task does not need to be revised, and the solution unit 12 notifies the user 2 via the dialogue unit 11 that the task should be executed. The solution unit 12 then passes the tasks in the task list to the execution unit 13 in order (or in parallel, if possible) to instruct the execution unit 13 to execute them. Specifically, the solution unit 12 passes text containing instructions for executing the target task to the execution unit 13, and also presents the task contents (in the example of FIG. 11, "Task 1-1: Extract product information about cakes for last year's Christmas...") to the user 2 via the dialogue unit 11.

[0084] In the example of FIG. 11, the execution of "task 1-1" is instructed, and the execution unit 13 creates a prompt for executing the task by referring to the task execution instruction passed from the solution unit 12 and the dialogue log and summary log stored in the short-term memory 16, and inputs this to the corresponding execution AI (13a), thereby executing the task. The execution AI (13a) performs work using the corresponding data source 3 (e.g., information acquisition or information processing), formats the execution result into text, and presents it to the user 2 via the dialogue unit 11. The content of the presented dialogue is stored by the dialogue unit 11 in the short-term memory 16 as a dialogue log.

[0085] 12, the solving unit 12 passes the next task to the executing unit 13 and instructs it to be executed. Specifically, the solving unit 12 passes a text containing instructions for executing the target task to the executing unit 13, and also presents the content of the task ("Task 2: Last year's sales performance of each extracted product...") to the user 2 via the dialogue unit 11.

[0086] In the example of Figure 12, the execution of "Task 2" is instructed, and the execution unit 13 creates a prompt for executing the task by referring to the task execution instruction passed from the solution unit 12 and the dialogue log and summary log stored in the short-term memory 16, and inputs it into the corresponding execution AI (13a), thereby executing the task. However, in the example of Figure 12, it is assumed that the task execution instruction, dialogue log, summary log, etc. do not contain information necessary for task execution and are therefore insufficient. In this case, for example, in accordance with the indications and instructions contained in the response from the execution AI (13a), the execution unit 13 presents a prompt to the user 2 via the dialogue unit 11 to instruct them to fill in the missing information (in the example shown in the figure, "aggregation period").

[0087] The content of the replenishment instruction presented by the dialogue unit 11 and the text information of the response from the user 2 are stored as dialogue logs in the short-term memory 16 (short-term objective memory 16a). Then, the execution AI (13a) executes the task based on the content replenished by the user 2. Thereafter, the subsequent tasks in the task list are executed sequentially in the same manner.

[0088] 13, when the execution of all tasks in the task list has been completed, the solving unit 12 compiles the input contents including the response from the user 2 from the short-term memory 16, the summary created by the monitoring unit 14 (monitoring AI (14a)), and the external data information acquired by all the relevant execution AIs (13a) of the execution unit 13, to create a prompt for generating an answer, and inputs this into the solving AI (12a) to obtain an answer. The acquired answer is presented to the user 2 via the dialogue unit 11. The content of the answer presented to the user 2 by the dialogue unit 11 and the text information of the response from the user 2 to this are each stored as a dialogue log in the short-term memory 16 (short-term objective memory 16a).

[0089] Thereafter, when User 2 instructs to end the interactive session, the monitoring unit 14 creates a summary of the entire session from the entire interaction log of the session stored in the short-term memory 16. For example, a prompt for generating a summary including information such as what User 2 wanted, what tasks were performed, and what User 2's evaluation was is created, and input into the monitoring AI (14a) to obtain a summary of the session.

[0090] FIG. 16 is a diagram outlining an example of a prompt of the monitoring unit 14 that summarizes the dialogue log in one embodiment of the present invention. The diagram instructs the system to retrieve knowledge about User 2's profile, work situation, and work background from the long-term subjective memory 16b while referring to the dialogue log, and to make corrections where necessary. It also instructs the system to extract keywords for important information. This extracted information is stored in the long-term subjective memory 16b for future reference. Separately, the dialogue log is extracted and summarized as text data to be used in future sessions based on User 2's instructions, the created task list, a concise summary of the entire session, and the like, and stored in the long-term subjective memory 16b.

[0091] 13, the acquired information such as the session summary and dialogue log is stored in the long-term memory 15 for future dialogue sessions as described above, and if any changes are detected in the information about user 2 or the store information registered in the setting information 17, these are modified and saved. Through the above series of processes, the agent system 1 designs, modifies, and executes a task while having an interactive dialogue with user 2.

[0092] As explained above, according to the agent system 1 which is one embodiment of the present invention, the entire system is divided into a task layer (roles are shared among the dialogue unit 11, solution unit 12, and execution unit 13) which executes tasks, and a meta layer (monitoring unit 14) which monitors and stores the business status. By combining these, it is possible to realize an agent-type AI which can properly understand the execution status of tasks and autonomously design and execute tasks.

[0093] The invention made by the inventor has been specifically described above based on the embodiments, but it goes without saying that the present invention is not limited to the above embodiments and can be modified in various ways without departing from the spirit of the invention. Furthermore, the above embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the described configurations. Furthermore, it is possible to add, delete, or replace part of the configuration of the above embodiments with other configurations.

[0094] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a storage device such as a memory, hard disk, or SSD, or in a storage medium such as an IC card, SD card, or DVD.

[0095] In addition, in the above figures, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily show all the control lines and information lines that are actually implemented. In reality, it can be assumed that almost all components are interconnected. [Industrial Applicability]

[0096] The present invention can be used in an agent system that realizes agent-type AI. [Explanation of symbols]

[0097] 1. Agent system, 2. User, 3. External data source, 11...Dialogue unit, 11a...Dialogue AI, 12...Resolution unit, 12a...Resolution AI, 13...Execution unit, 13a...Execution AI, 14...Monitoring unit, 14a...Monitoring AI, 15...Long-term memory, 15a...Long-term objective memory, 15b...Long-term subjective memory, 16...Short-term memory, 16a...Short-term objective memory, 16b...Short-term subjective memory, 17...Setting information

Claims

1. An agent system that grasps a work instruction from a dialogue with a user and executes a task related to the work instruction, Each of the units has a dialogue unit, a solution unit, an execution unit, and a monitoring unit that can independently use the generation AI; The dialogue unit grasps the task instruction through a dialogue with the user by the generation AI, and stores the content of the dialogue with the user as a dialogue log in a short-term memory unit; The solving unit breaks down the work instructions grasped by the dialogue unit into tasks using a generation AI to create a task list, passes execution instructions for each task to the execution unit, and presents execution results by the execution unit to the user via the dialogue unit; The execution unit executes a task using a corresponding data source by a generation AI corresponding to the task related to the execution instruction passed from the resolution unit, and passes the execution result to the resolution unit; The monitoring unit refers to the dialogue log at any time, extracts information related to predetermined matters, grasps the context of the dialogue using a generation AI, predicts the next response to be made, and stores the summary in a short-term memory unit, making it possible to refer to the dialogue unit, the resolution unit, and the execution unit at any time. This is an agent system.

2. 2. The agent system according to claim 1, An agent system in which the dialogue unit, through dialogue with the user, confirms with the user whether or not the task list created by the resolution unit needs to be modified, and if modification is required, passes the modification content instructed by the user to the resolution unit to instruct it to modify the task list, and if modification is not required, instructs the resolution unit to execute the task list.

3. 2. The agent system according to claim 1, When the dialogue unit has finished dialogue with the user, the monitoring unit uses a generation AI to extract predetermined information relating to the user's work background and / or work situation based on the dialogue log, stores the information in a long-term memory unit, and makes it possible for the dialogue unit and the resolution unit to refer to the information at any time.

4. 4. The agent system according to claim 3, An agent system in which, when the solving unit breaks down the work instructions grasped by the dialogue unit into tasks using a generation AI to create the task list, the solving unit uses as input information to the generation AI a summary of the dialogue with the user stored in the short-term memory unit and information obtained from the long-term memory unit that includes past work instructions similar to the work instructions and the task designs at that time.

5. 2. The agent system according to claim 1, An agent system in which the execution unit determines whether each task related to the execution instruction passed from the resolution unit can be executed by the corresponding generation AI, and if it cannot be executed, requests the resolution unit to modify the target task.

6. 6. The agent system according to claim 5, An agent system in which, if the target task is executable, the execution unit determines whether the information necessary to perform the work is missing using the corresponding data source, and if so, references the information in the short-term memory unit to supplement the information.

7. 7. The agent system according to claim 6, An agent system wherein, when information required to perform a task using a data source corresponding to the target task is lacking, the execution unit queries the user via the dialogue unit to fill the gap.

Citation Information

Patent Citations

  • User support system, user support program, and user support method

    JP2018081444A

Cited By

  • Information processing device, information processing method, program, and storage medium

    JP7872085B1