Method for improving reasoning capability of reasoning system comprising LLM-based multiple agents, by using conversation with external system, reasoning system performing same, and computer-readable recording medium implementing same

The method enhances the inference ability of LLM-based multi-agent systems through inter-system dialogue and dynamic learning, ensuring objective evaluation and accurate feedback without fine-tuning, addressing the lack of objectivity in existing assessment methods.

WO2026095204A1PCT designated stage Publication Date: 2026-05-07SELECT STAR INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SELECT STAR INC
Filing Date
2024-12-13
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing evaluation methods for large language models (LLMs) lack objectivity when systems surpass human capabilities, necessitating a method to objectively assess and improve the inference ability of LLM-based multi-agent systems.

Method used

A method involving inter-system dialogue and dynamic learning with an external system, where the inference system generates questions and answers, adjusts evaluation scores, and learns using pre-set objective functions to enhance its own capabilities.

Benefits of technology

Enables objective evaluation and dynamic evolution of inference system capabilities, leveraging collective intelligence without requiring fine-tuning of individual LLMs, and providing granular feedback for accurate credit assignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096927_07052026_PF_FP_ABST
    Figure KR2024096927_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for improving reasoning capability of a reasoning system comprising large language model (LLM)-based multiple agents by using conversation with an external system, a reasoning system performing same, and a computer-readable recording medium implementing same, and, more specifically, to a method for improving reasoning capability of a reasoning system comprising LLM-based multiple agents, by using conversation with an external system, a reasoning system performing same, and a computer-readable recording medium implementing same, the method comprising, when an LLM- or AI-based system is to be evaluated through a reasoning system that includes one or more language model-based agents, performing conversation with the system, which is to be evaluated, so as to dynamically evolve reasoning capability of the reasoning system, thereby ensuring evaluation objectivity.
Need to check novelty before this filing date? Find Prior Art

Description

A method for improving the inference capability of an inference system including an LLM-based multi-agent system using dialogue with an external system, an inference system performing the same, and a computer-readable recording medium implementing the same

[0001] The present invention relates to a method for improving the inference ability of an inference system including an LLM-based multi-agent system using dialogue with an external system, an inference system for performing the same, and a computer-readable recording medium for implementing the same. More specifically, when evaluating a system based on a large language model (LLM) or AI through an inference system including one or more language model-based agents, the inference ability of the inference system can be dynamically evolved by performing dialogue with the system being evaluated, thereby ensuring objectivity in the evaluation. The invention relates to a method for improving the inference ability of an LLM-based multi-agent system using dialogue with an external system, an inference system for performing the same, and a computer-readable recording medium for implementing the same.

[0002]

[0003] Large Language Models (LLMs) are a type of artificial intelligence system that utilizes large-scale neural network structures with billions to hundreds of billions of parameters. By pre-training on vast amounts of text data, they can identify language patterns, grammar, and semantics, thereby enabling them to perform various natural language processing tasks such as text generation, translation, summarization, and question answering. Representative examples of LLMs include Google's BERT and OpenAI's GPT.

[0004] Meanwhile, as related technologies have advanced, various large language models and artificial intelligence technologies have become increasingly numerous. Each has highlighted its distinct strengths and sought its own application; consequently, numerous studies have been conducted on methods to evaluate or verify whether these large language models or AI systems actually possess those strengths (capabilities). One of the representative indicators for evaluating generative language models, such as large language models, is ROUGE (Recall-Oriented Understudy for Gisting Evaluation), which is used to assess the quality of document summaries. ROUGE evaluates language models by measuring the degree of agreement between automatically generated summaries and standard summaries written by humans.

[0005] As such, while Human Evaluation, in which humans directly participate to assess large language models or artificial intelligence systems, is adopted as the most credible method for evaluation, a problem arises in that objective assessment becomes impossible using this approach when the performance of the systems or models being evaluated exceeds human capabilities in specific areas. Therefore, there is a need for technology to evaluate large language models or artificial intelligence systems that surpass the capabilities of humans and existing evaluation systems in specific areas or fields.

[0006]

[0007] The present invention aims to provide a method for improving the inference ability of an inference system including an LLM-based multi-agent system using dialogue with an external system, an inference system for performing the same, and a computer-readable recording medium for implementing the same. More specifically, when evaluating a system based on a large language model (LLM) or AI through an inference system including one or more language model-based agents, the inference ability of the inference system can be dynamically evolved by performing dialogue with the system being evaluated, thereby ensuring objectivity in the evaluation. The invention aims to provide a method for improving the inference ability of an LLM-based multi-agent system using dialogue with an external system, an inference system for performing the same, and a computer-readable recording medium for implementing the same.

[0008]

[0009] In order to solve the above problems, one embodiment of the present invention provides a method for improving inference ability using a conversation with an external system, which is performed in an inference system including one or more agents based on a language model, comprising: an inter-system conversation step in which the inference system performs a conversation with an external system to be evaluated, generates an answer to a question received from the system to be evaluated, derives an evaluation score for the answer received from the system to be evaluated, and generates a new question by varying the question received from the system to be evaluated; and a dynamic learning step in which the inference agent included in the one or more agents learns in a direction that satisfies a pre-set first objective function and a second objective function, wherein the first objective function is intended to make the evaluation score currently derived by the system to be evaluated higher than the evaluation score previously derived by the system to be evaluated, and the second objective function is intended to make the evaluation score currently derived by the inference system lower than the evaluation score previously derived by the inference system.

[0010] In one embodiment of the present invention, the inter-system conversation step may include a question variation step of generating a new question to be delivered to the evaluation target system based on at least two of: a question previously generated by the inference system or the evaluation target system; an answer to the question; and an evaluation score for the answer.

[0011] In one embodiment of the present invention, the inter-system conversation step includes a verification step performed by generating an initial verification question by the system to be verified, and the verification step comprises: a first verification step of receiving a first verification question corresponding to the initial verification question, generating a first verification answer, and generating a second verification question that is a variation of the first verification question; a second verification step of receiving a second verification answer for the second verification question, generating a second verification answer evaluation score, and generating a third verification question that is a variation of the second verification question; and a third verification step of receiving a third verification answer for the third verification question, generating a third verification answer evaluation score, and receiving a fourth verification question that is a variation of the third verification question, generating a fourth verification answer. and may include a fourth verification step of receiving a fifth verification question that is a variation of the fourth verification question, generating a fifth verification answer, and generating a sixth verification question that is a variation of the fifth verification question.

[0012] In one embodiment of the present invention, the inference system performs a conversation with the system to be evaluated by sequentially repeating the same process as the verification step 2, the verification step 3, and the verification step 4 until the conversation ends when the evaluation score derived by the system to be evaluated is less than or equal to the evaluation score derived by the inference system, and may terminate the inter-system conversation step when the evaluation score derived by the system to be evaluated is greater than the evaluation score derived by the inference system.

[0013] In one embodiment of the present invention, the inter-system conversation step includes an evaluation step performed by generating an initial evaluation question by the inference system, and the evaluation step may include: an evaluation first step of generating a first evaluation question corresponding to the initial evaluation question and transmitting it to the verification target system; an evaluation second step of receiving a first evaluation answer to the first evaluation question to generate a first evaluation answer score and receiving a second evaluation question that is a variation of the first evaluation question to generate a second evaluation answer; an evaluation third step of receiving a third evaluation question that is a variation of the second evaluation question to generate a third evaluation answer and generating a fourth evaluation question that is a variation of the third evaluation question; and an evaluation fourth step of receiving a fourth evaluation answer to the fourth evaluation question to generate fourth evaluation answer evaluation information and generating a fifth evaluation question that is a variation of the fourth evaluation question.

[0014] In one embodiment of the present invention, the inference system performs a conversation with the system to be evaluated by sequentially repeating the same process as the second evaluation step, the third evaluation step, and the fourth evaluation step until the conversation ends when the evaluation score derived by the system to be evaluated is less than or equal to the evaluation score derived by the inference system, and may terminate the inter-system conversation step when the evaluation score derived by the system to be evaluated is greater than the evaluation score derived by the inference system.

[0015] In one embodiment of the present invention, the inference agent may include: an actor module that generates and outputs first action information including one or more of a question, an answer, and an evaluation score; an evaluator module that generates and outputs internal feedback information based on the reward information when reward information corresponding to the first action information is generated by the evaluation target system; and a self-reflection module that receives external feedback information including the reward information and the internal feedback information, and outputs self-reflection information that enables the actor module to extract meaningful information on its own.

[0016] In one embodiment of the present invention, the inference agent may further include: a short-term memory storage unit that receives and stores the compensation information from the evaluation target system; and a long-term memory storage unit that receives and stores the self-reflection information from the self-reflection module.

[0017] In one embodiment of the present invention, the one or more agents further include a management agent corresponding to a sub-agent of the inference agent; and a working agent corresponding to a sub-agent of the management agent; and the agent of the upper layer can perform information exchange with the agent of the lower layer.

[0018] In one embodiment of the present invention, the inference agent may include: an action module that outputs second action information to the management agent, wherein the second action information includes information related to sub-tasks that divide the input task; an evaluator module that generates and outputs internal feedback information regarding the observation information based on observation information corresponding to the second action information from the evaluation target system, wherein the observation information includes memory information of the management agent; and a self-reflection module that receives external feedback information including the observation information and the internal feedback information, and outputs self-reflection information that enables the action module to extract meaningful information on its own.

[0019] In one embodiment of the present invention, a short-term memory storage unit that stores observation information including information about the current state of the management agent and memory information; and a long-term memory storage unit that extracts and stores meaningful information among the information stored in the short-term memory storage unit may be further included.

[0020] In one embodiment of the present invention, the method for improving reasoning ability may further include a system evaluation step in which, when the evaluation score derived by the system to be evaluated exceeds the evaluation score derived by the reasoning system, the reasoning system evaluates the system to be evaluated and derives the corresponding evaluation result.

[0021] In order to solve the above problems, in one embodiment of the present invention, an inference system that performs a method for improving inference ability using a conversation with an external system is provided, wherein the inference system includes one or more agents based on a language model and performs a conversation with an external system to be evaluated, and generates an answer to a question received from the system to be evaluated, derives an evaluation score for the answer received from the system to be evaluated, and generates a new question by varying the question received from the system to be evaluated; and a dynamic learning unit in which an inference agent included in the one or more agents learns in a direction that satisfies a pre-set first objective function and a second objective function; wherein the first objective function is intended to make the evaluation score currently derived by the system to be evaluated higher than the evaluation score previously derived by the system to be evaluated, and the second objective function is intended to make the evaluation score currently derived by the inference system lower than the evaluation score previously derived by the inference system.

[0022] To solve the above-mentioned problem, in one embodiment of the present invention, a computer-readable recording medium for implementing a method for improving reasoning ability using dialogue with an external system, which is performed in a computing system comprising one or more processors and one or more memories, wherein the computing system comprises a reasoning system comprising one or more agents based on a language model, and the computer-readable recording medium stores instructions for causing the computing system to perform the following steps, wherein the following steps are: a method for improving reasoning ability using dialogue with an external system, which is performed in a reasoning system comprising one or more agents based on a language model, wherein the reasoning system performs a dialogue with an external system to be evaluated, generates an answer to a question received from the system to be evaluated, derives an evaluation score for the answer received from the system to be evaluated, and generates a new question by varying the question received from the system to be evaluated; A computer-readable medium is provided, comprising: a dynamic learning step in which an inference agent included in one or more agents learns in a direction that satisfies a preset first objective function and a second objective function by means of an inference system; wherein the first objective function is intended to make the evaluation score currently derived by the evaluation target system higher than the evaluation score previously derived by the evaluation target system, and the second objective function is intended to make the evaluation score currently derived by the inference system lower than the evaluation score previously derived by the inference system.

[0023]

[0024] According to one embodiment of the present invention, when the capability of a system to be evaluated is superior to the capability of a reasoning system intended to evaluate the system to be evaluated, the capability of the reasoning system can be dynamically evolved through dialogue with the system to be evaluated, thereby enabling the objective evaluation of the system to be evaluated.

[0025] According to one embodiment of the present invention, unlike the conventional Knowledge Distillation technique using a Teacher-Student Model in which a Student Model learns the knowledge of a higher-level model with fixed model parameters, the language model of an inference agent calls various LLM instances through planning and assigns a role to each, thereby establishing an organizational hierarchy and enabling the effect of surpassing the knowledge of the system under evaluation by utilizing the collective intelligence of the organization.

[0026] According to one embodiment of the present invention, fine tuning is not required for each of the large language models used, thereby enabling lighter execution than conventional LLM-based technologies.

[0027] According to one embodiment of the present invention, a more granular form of feedback is possible compared to scalar or vector compensation, and such compensation can have the effect of enabling more accurate credit assignment.

[0028]

[0029] FIG. 1 schematically illustrates the configuration of an evaluation target system and an inference system according to one embodiment of the present invention.

[0030] FIG. 2 schematically illustrates the steps of performing a conversation between an evaluation target system and an inference system according to one embodiment of the present invention.

[0031] FIG. 3 schematically illustrates the process in which the capability of a reasoning system dynamically evolves through a conversation with a system under evaluation according to one embodiment of the present invention.

[0032] FIG. 4 schematically illustrates the process of the question variation step according to one embodiment of the present invention.

[0033] FIG. 5 schematically illustrates the steps of performing a verification step according to one embodiment of the present invention.

[0034] FIG. 6 schematically illustrates the execution steps of the process in which the inter-system conversation step is terminated according to one embodiment of the present invention.

[0035] FIG. 7 schematically illustrates the execution steps of an evaluation step according to an embodiment of the present invention.

[0036] FIG. 8 schematically illustrates the execution steps of a system evaluation step according to one embodiment of the present invention.

[0037] FIG. 9 schematically illustrates the process of a conversation being performed between an evaluation target system and an inference system according to one embodiment of the present invention.

[0038] FIG. 10 schematically illustrates the internal configuration of an inference agent according to one embodiment of the present invention.

[0039] FIG. 11 schematically illustrates the process of generating self-reflection information according to one embodiment of the present invention.

[0040] FIG. 12 schematically illustrates one or more agents included in an inference system according to one embodiment of the present invention.

[0041] FIG. 13 schematically illustrates the process of an inference agent and a management agent exchanging information according to one embodiment of the present invention.

[0042] FIG. 14 schematically illustrates the internal configuration of an inference agent according to one embodiment of the present invention.

[0043] FIG. 15 schematically illustrates the internal configuration of a computing device according to one embodiment of the present invention.

[0044]

[0045] According to one embodiment of the present invention, unlike the conventional Knowledge Distillation technique using a Teacher-Student Model in which a Student Model learns the knowledge of a higher-level model with fixed model parameters, the language model of the inference agent calls various LLM instances through planning and assigns a role to each, thereby establishing an organizational hierarchy and utilizing the collective intelligence of the organization to achieve the effect of surpassing the knowledge of the system under evaluation.

[0046] According to one embodiment of the present invention, fine tuning is not required for each of the large language models used, thereby enabling lighter execution than conventional LLM-based technologies.

[0047] According to one embodiment of the present invention, a more granular form of feedback is possible compared to scalar or vector compensation, and such compensation can have the effect of enabling more accurate credit assignment.

[0048]

[0049] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0050] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

1. A method for improving reasoning ability using dialogue with an external system, performed in a reasoning system comprising one or more agents based on a language model, wherein A system-to-system dialogue step, wherein a dialogue is performed with an external system to be evaluated by means of an inference system, wherein an answer to a question received from the system to be evaluated is generated, an evaluation score for the answer received from the system to be evaluated is derived, and a new question is generated by varying the question received from the system to be evaluated; and A dynamic learning step in which an inference agent included in the one or more agents learns in a direction that satisfies a preset first objective function and a second objective function by means of an inference system; The above first objective function is, The purpose is to ensure that the evaluation score currently derived by the aforementioned evaluation target system is higher than the evaluation score previously derived by the aforementioned evaluation target system, and The above second objective function is, A method for improving reasoning ability, which aims to make the evaluation score currently derived by the above-mentioned reasoning system lower than the evaluation score previously derived by the above-mentioned reasoning system.

2. In Claim 1, The above inter-system dialogue step is, A method for improving reasoning ability, comprising a question variation step of generating a new question to be delivered to the evaluation target system based on at least two of: a question previously generated by the inference system or the evaluation target system; an answer to the question; and an evaluation score for the answer.

3. In Claim 1, The above inter-system dialogue step includes a verification step performed by generating an initial verification question by the system to be verified, and The above verification step is, A first verification step of receiving a first verification question corresponding to the above initial verification question, generating a first verification answer, and generating a second verification question that is a variation of the above first verification question; Verification step 2, which involves receiving a second verification answer to the second verification question, generating a second verification answer evaluation score, and generating a third verification question that is a variation of the second verification question; A third verification step of receiving a third verification answer to the above third verification question to generate a third verification answer evaluation score, and receiving a fourth verification question that is a variation of the above third verification question to generate a fourth verification answer; and A method for improving reasoning ability, comprising: a verification step 4 of receiving a fifth verification question that is a variation of the fourth verification question, generating a fifth verification answer, and generating a sixth verification question that is a variation of the fifth verification question.

4. In Claim 3, The above inference system is, If the evaluation score derived by the evaluation target system is less than or equal to the evaluation score derived by the inference system, a conversation with the evaluation target system is conducted by sequentially repeating the same process as the verification step 2, verification step 3, and verification step 4 until the conversation ends. A method for improving reasoning ability, wherein if the evaluation score derived by the evaluation target system exceeds the evaluation score derived by the inference system, the inter-system dialogue step is terminated.

5. In Claim 1, The above inter-system dialogue step includes an evaluation step performed by generating an initial evaluation question by the inference system, and The above evaluation step is, Evaluation Step 1: generating a first evaluation question corresponding to the above initial evaluation question and transmitting it to the above verification target system; Evaluation step 2, which involves receiving a first evaluation answer to the first evaluation question to generate a first evaluation answer score, and receiving a second evaluation question that is a variation of the first evaluation question to generate a second evaluation answer; Evaluation step 3, which involves receiving a third evaluation question that is a variation of the second evaluation question and generating a third evaluation answer, and generating a fourth evaluation question that is a variation of the third evaluation question; and A method for improving reasoning ability, comprising: a fourth evaluation step of receiving a fourth evaluation answer to the above fourth evaluation question to generate fourth evaluation answer evaluation information, and generating a fifth evaluation question that is a variation of the above fourth evaluation question.

6. In Claim 5, The above inference system is, If the evaluation score derived by the evaluation target system is less than or equal to the evaluation score derived by the inference system, a conversation with the evaluation target system is conducted by sequentially repeating the same process as the evaluation step 2, evaluation step 3, and evaluation step 4 until the conversation ends. A method for improving reasoning ability, wherein if the evaluation score derived by the evaluation target system exceeds the evaluation score derived by the inference system, the inter-system dialogue step is terminated.

7. In Claim 1, The above inference agent is, An actor module that generates and outputs first action information including one or more of a question, an answer, and an evaluation score; An evaluator module that generates and outputs internal feedback information based on the compensation information when compensation information corresponding to the first action information is generated by the evaluation target system; and A method for improving reasoning ability, comprising: a self-reflection module that receives external feedback information including the above-mentioned reward information and the above-mentioned internal feedback information, and outputs self-reflection information that enables the above-mentioned behavior module to extract meaningful information on its own.

8. In Claim 7, The above inference agent is, A short-term memory storage unit that receives and stores the compensation information from the above-mentioned evaluation target system; and A method for improving reasoning ability, further comprising a long-term memory storage unit that receives and stores self-reflection information from the self-reflection module.

9. In Claim 1, The above one or more agents, A management agent corresponding to the sub-agent of the above-mentioned inference agent; and It further includes a working agent corresponding to a sub-agent of the above-mentioned management agent; and A method for improving reasoning ability in which an agent in a higher layer performs information exchange with an agent in a lower layer.

10. In Claim 9, The above inference agent is, An action module that outputs second action information to the above-mentioned management agent, wherein the second action information includes information related to sub-tasks divided from the input task; An evaluator module that generates and outputs internal feedback information regarding the observation information based on the observation information corresponding to the second action information from the evaluation target system, wherein the observation information includes memory information of the management agent; and A method for improving reasoning ability, comprising: a self-reflection module that receives external feedback information including the observation information and internal feedback information, and outputs self-reflection information that enables the behavior module to extract meaningful information on its own.

11. In Claim 10, The above inference agent is, A short-term memory storage unit that stores the observation information including information on the current state of the management agent and memory information; and A method for improving reasoning ability, further comprising a long-term memory storage unit that extracts and stores meaningful information among the information stored in the short-term memory storage unit.

12. In Claim 1, The above method for improving reasoning ability is, A method for improving reasoning ability, further comprising: a system evaluation step in which, if the evaluation score derived by the evaluation target system exceeds the evaluation score derived by the inference system, the inference system evaluates the evaluation target system and derives the corresponding evaluation result.

13. As a reasoning system that performs a method for improving reasoning ability using dialogue with an external system, The above inference system includes one or more agents based on a language model, and An inter-system dialogue unit that performs a dialogue with an external system subject to evaluation, generates an answer to a question received from the system subject to evaluation, derives an evaluation score for the answer received from the system subject to evaluation, and generates a new question by varying the question received from the system subject to evaluation; and A dynamic learning unit comprising an inference agent included in the above one or more agents that learns in a direction satisfying a preset first objective function and a second objective function; and The above first objective function is, The purpose is to ensure that the evaluation score currently derived by the aforementioned evaluation target system is higher than the evaluation score previously derived by the aforementioned evaluation target system, and The above second objective function is, An inference system that aims to make the evaluation score currently derived by the above inference system lower than the evaluation score previously derived by the above inference system. A computer-readable recording medium for implementing a method for improving reasoning ability using dialogue with an external system, which is performed on a computing system comprising 14.1 or higher processors and 1 or more memories, The above computing system includes an inference system comprising one or more agents based on a language model, and The computer-readable recording medium stores instructions that cause the computing system to perform the following steps, and The steps below are: A method for improving reasoning ability using dialogue with an external system, performed in a reasoning system comprising one or more agents based on a language model, wherein A system-to-system dialogue step, wherein a dialogue is performed with an external system to be evaluated by means of an inference system, wherein an answer to a question received from the system to be evaluated is generated, an evaluation score for the answer received from the system to be evaluated is derived, and a new question is generated by varying the question received from the system to be evaluated; and A dynamic learning step in which an inference agent included in the one or more agents learns in a direction that satisfies a preset first objective function and a second objective function by means of an inference system; The above first objective function is, The purpose is to ensure that the evaluation score currently derived by the aforementioned evaluation target system is higher than the evaluation score previously derived by the aforementioned evaluation target system, and The above second objective function is, A computer-readable recording medium intended to ensure that the evaluation score currently derived by the above inference system is lower than the evaluation score previously derived by the above inference system.

Citation Information

Patent Citations

  • Bed

    KR1020240175211A

  • Method for learning using tree-based search technology, and computer program recorded on record-medium for executing method therefor

    KR102694634B1

  • Methods and systems for responding to a natural language query

    US20230196033A1

  • Guided dialogue using language generation neural networks and search

    US20240104336A1

  • KR20240020519A