LLM-based dialectic decision-making method between multiple agents and method for synthesizing, by using same, datasets for evaluating artificial intelligence system

The LLM-based multi-agent dialectical decision-making method addresses the inefficiencies of user-labeling by iteratively refining answers through thesis, antithesis, and synthesis, enhancing the quality and reducing the cost of AI evaluation datasets.

WO2026095205A1PCT designated stage Publication Date: 2026-05-07SELECT STAR INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SELECT STAR INC
Filing Date
2024-12-13
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional methods for generating datasets to evaluate artificial intelligence models are time-consuming and costly due to the need for user labeling processes.

Method used

An LLM-based multi-agent dialectical decision-making method that utilizes thesis, antithesis, and synthesis to generate and refine answers, allowing for iterative refinement of information until validity is achieved.

Benefits of technology

This approach reduces the time and cost associated with dataset generation by enabling efficient, iterative refinement of answers, ensuring high-quality evaluation data for AI systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096928_07052026_PF_FP_ABST
    Figure KR2024096928_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to an LLM-based dialectic decision-making method between multiple agents and a method for synthesizing, by using the same, datasets for evaluating an artificial intelligence system, the method involving: implementing, by using an LLM, a dialectic decision-making through first information corresponding to a thesis as an answer to a query, second information corresponding to an antithesis as a refutation to the corresponding answer, and third information corresponding to a synthesis as a final conclusion considering the first information and the second information; and applying the dialectic decision-making in a process for synthesizing datasets for evaluating an artificial intelligence system.
Need to check novelty before this filing date? Find Prior Art

Description

LLM-based multi-agent dialectical decision-making method and method for synthesizing a dataset for evaluating an artificial intelligence system using the same

[0001] The present invention relates to an LLM-based multi-agent dialectical decision-making method and a method for synthesizing a dataset for evaluating an artificial intelligence system using the same, wherein the method implements dialectical decision-making by utilizing LLM through first information corresponding to the thesis as an answer to a query, second information corresponding to the antithesis as a rebuttal to the answer, and third information corresponding to the synthesis as a final conclusion considering the first and second information, and applies the dialectical decision-making during the process of synthesizing a dataset for evaluating an artificial intelligence system.

[0002] A Large Language Model (LLM) refers to a generative AI that recognizes prompts containing text and outputs responses. Meanwhile, attempts to utilize LLMs have recently been made in various fields such as law, education, and medicine. Examples include implementing chatbots via LLMs to handle customer consultations, performing searches or queries in fields requiring specialized knowledge such as law or medicine, and evaluating applicants during the recruitment process.

[0003] Meanwhile, various methods for evaluating the performance of artificial intelligence models are also being discussed in this regard. For example, before launching services based on AI systems, companies may want to comprehensively evaluate factors such as the accuracy of responses, context, appropriate responses to various topics, and adaptability to new problems.

[0004] For such an evaluation, a dataset including questions and answers is required; specifically, it may require specific evaluations of evaluation items described by users, redefinitions of those evaluation items, and various sets of questions and answers tailored to the redefined evaluation items.

[0005] However, conventionally, generating datasets for evaluating the performance of artificial intelligence models required a user labeling process, which posed a significant problem in terms of time and cost for creating large-scale datasets.

[0006] The present invention aims to provide an LLM-based multi-agent dialectical decision-making method and a method for synthesizing a dataset for evaluating an artificial intelligence system using the same, wherein the invention implements dialectical decision-making by utilizing LLM through first information corresponding to the thesis as an answer to a query, second information corresponding to the antithesis as a rebuttal to the answer, and third information corresponding to the synthesis as a final conclusion considering the first and second information, and applies the dialectical decision-making during the process of synthesizing a dataset for evaluating an artificial intelligence system.

[0007] To solve the above problem, a method for dialectically making decisions using multiple agents based on a Large Language Model (LM) and performed in a computing system comprising one or more processors and one or more memories, comprising: an initial answer prompt input step of inputting an initial answer prompt into an LLM including an input query and a phrase requesting the generation of an answer to said query, and an initial first information derivation step of deriving initial first information regarding said query; an initial second information derivation step of inputting an initial rebuttal prompt into an LLM including said query, said initial first information, and a phrase requesting a rebuttal to said initial first information, and an initial second information derivation step of deriving initial second information corresponding to a rebuttal to said initial first information; and said query; said initial first information; said initial second information; The present invention provides a method for dialectically making decisions using an LLM-based multi-agent, comprising: an initial decision step including an initial sum prompt input step for inputting an initial sum prompt into an LLM that includes a phrase requesting the generation of initial third information corresponding to the information combining the above query, the above initial first information, and the above second information; and an initial third information derivation step for deriving initial third information combining the above initial first information and the above initial second information.

[0008] In one embodiment of the present invention, the initial third information derivation step further includes a third information validity review step for deriving the initial third information validity by reviewing the validity of the initial third information by the LLM model; and the method includes a repeat decision step for deriving repeat first information, repeat second information, and repeat third information by inputting one or more pieces of information derived in the initial decision step into the LLM when the initial third information validity derived in the initial decision step is below a preset standard; and the repeat third information may be information obtained by combining the initial first information and the initial second information by the LLM.

[0009] In one embodiment of the present invention, the iterative decision-making step may include: a iterative first information derivation step, which inputs the query, the initial first information, the initial second information, and the initial third information into an LLM to derive iterative first information by synthesizing the query, the initial first information, the initial second information, and the initial third information; an iterative second information derivation step, which inputs information including the query and the iterative first information into an LLM to derive iterative second information corresponding to a rebuttal to the iterative first information; and an iterative third information derivation step, which inputs information including the query, the iterative first information, and the iterative second information into an LLM to derive iterative third information by synthesizing the iterative first information and the iterative second information.

[0010] In one embodiment of the present invention, the initial sum prompt may further include a phrase requesting a validity review of the initial third information generated based on the initial sum prompt input to the LLM.

[0011] In one embodiment of the present invention, the initial first information derivation step may include: a first information validity review step comprising the step of inputting into an LLM a prompt including the generated initial first information and a phrase requesting a validity review of the initial first information; a first alternative information derivation step comprising the step of inputting into an LLM a prompt including the initial first information, a first error derived during the validity review process of the initial first information, and a phrase requesting to derive first alternative information when the validity of the initial first information is below a preset standard; and an initial first information replacement step comprising inputting into an LLM a prompt including the first alternative information and a phrase requesting a validity review of the first alternative information, and replacing the initial first information based on the first alternative information when the validity of the first alternative information exceeds a preset standard.

[0012] In one embodiment of the present invention, the first alternative information may correspond to an answer output to the LLM from a different perspective than when the LLM outputs the initial first information corresponding to the answer to the query in the initial first information derivation step.

[0013] In one embodiment of the present invention, the initial first information replacement step further includes a second replacement information derivation step performed when it is determined that the validity of one or more of the first replacements is below a preset standard; and the second replacement information derivation step may include a second replacement information derivation step comprising a step of inputting a prompt into an LLM that includes a first replacement information whose validity is below a preset standard, a second error derived during the validity review process for the first replacement information, and a phrase requesting the derivation of the second replacement information; and a first replacement information replacement step that replaces the first replacement information based on the second replacement information.

[0014] To solve the above-mentioned problem, a method for generating questions and answers for evaluating an artificial intelligence system including an LLM comprises: a request receiving step for receiving a request regarding one or more of evaluation items and evaluation criteria for an artificial intelligence system to be evaluated by a user; a request analysis step for analyzing the request and deriving a request analysis result by means of a request analysis model including or connected to an LLM; and an evaluation rubric derivation step for deriving an evaluation rubric including score-based evaluation criteria based on the request analysis result by means of a rubric generation model including or connected to an LLM. A question-answer generation step that generates questions and answers according to the evaluation rubric by means of a question-answer generation model that includes or is connected to an LLM; wherein one or more of the request analysis step, the evaluation rubric derivation step, and the question-answer generation step include an initial correct prompt input step that inputs an initial correct prompt into the LLM, the initial correct prompt including an input query and a phrase requesting the generation of an answer to the query, and an initial first information derivation step that derives initial first information regarding the query; an initial second information derivation step that inputs an initial rebuttal prompt into the LLM, the initial first information including the query, the initial first information, and a phrase requesting a rebuttal to the initial first information, and an initial second information derivation step that derives initial second information corresponding to a rebuttal to the initial first information; and the query; the initial first information; the initial second information; The present invention provides a method for generating questions and answers for evaluating an artificial intelligence system, comprising: an initial decision step; a step of inputting an initial sum prompt into an LLM, the initial sum prompt including a phrase requesting the generation of initial third information corresponding to the information combining the above query, the above initial first information, and the above second information; and an initial third information derivation step of deriving initial third information combining the above initial first information and the above initial second information.

[0015] According to one embodiment of the present invention, for a query including a user's request, a final conclusion can be derived by using first information corresponding to an answer generated by an LLM for said query, second information corresponding to a rebuttal to said first information, and third information that comprehensively considers said first information and second information.

[0016] According to one embodiment of the present invention, if it is determined that the third information derived in the initial decision-making stage by LLM is invalid, the first information, the second information, and the third information derived in the initial decision-making stage are combined to derive a new first information, a second information corresponding to a rebuttal to the re-derived first information is derived, and the third information is derived by comprehensively considering the re-derived first information and the re-derived second information to derive a final conclusion.

[0017] According to one embodiment of the present invention, an iterative decision-making step of rederiving the first information, the second information, and the third information is performed until the third information is determined to be valid by the LLM, and a final conclusion can be derived based on the third information determined to be valid by the LLM.

[0018] According to one embodiment of the present invention, when the first information is determined to be invalid by the LLM, a first alternative information is derived based on a first error derived when the first information is determined to be invalid, and when the first alternative information is determined to be valid by the LLM, the first information can be replaced with the first alternative information.

[0019] According to one embodiment of the present invention, based on a first error derived when the LLM determines that the first information is invalid, the LLM can derive first alternative information that is output for a query from a different perspective than when the LLM outputs the first information corresponding to the answer to the query.

[0020] According to one embodiment of the present invention, when each of one or more first alternative information is determined to be valid by LLM, the first information can be replaced based on information that combines the one or more first alternative information determined to be valid.

[0021] According to one embodiment of the present invention, based on a second error derived when the LLM determines that the first alternative information is invalid, the LLM can derive second alternative information that is output for a query from a different perspective than when the LLM outputs the first information.

[0022] According to one embodiment of the present invention, when the second information is determined to be invalid by the LLM, a third alternative information is derived based on the error derived when the second information is determined to be invalid, and when the third alternative information is determined to be valid by the LLM, the second information can be replaced with the third alternative information.

[0023] According to one embodiment of the present invention, based on an error derived when the second information is determined to be invalid by the LLM, the third alternative information that the LLM outputs for a query can be derived from a different perspective than when the LLM outputs the second information corresponding to the answer to the query.

[0024] According to one embodiment of the present invention, a dialectical decision-making process using multiple agents can be applied to a process of analyzing a request received from a user, generating an evaluation rubric for the request analysis results regarding the request, and generating questions and answers based on the generated evaluation rubric.

[0025] FIG. 1 illustrates the components of a decision-making system according to one embodiment of the present invention.

[0026] FIG. 2 schematically illustrates a decision-making method according to one embodiment of the present invention.

[0027] FIG. 3 illustrates the steps of an initial decision-making step and an iterative decision-making step according to one embodiment of the present invention.

[0028] FIG. 4 illustrates an initial decision-making step according to one embodiment of the present invention.

[0029] FIG. 5 illustrates an iterative decision-making step according to one embodiment of the present invention.

[0030] FIG. 6 illustrates a third information validity review step according to an embodiment of the present invention.

[0031] FIG. 7 illustrates the steps of generating first replacement information and second replacement information and replacing the first information according to an embodiment of the present invention.

[0032] FIG. 8 illustrates details regarding the first alternative information according to one embodiment of the present invention.

[0033] FIG. 9 illustrates details regarding second alternative information according to one embodiment of the present invention.

[0034] FIG. 10 illustrates data according to one embodiment of the present invention.

[0035] FIG. 11 illustrates the steps of a method for generating questions and answers for evaluating an artificial intelligence system including an LLM according to an embodiment of the present invention.

[0036] FIG. 12 illustrates an evaluation rubric according to one embodiment of the present invention.

[0037] FIG. 13 schematically illustrates the internal configuration of a computing device according to one embodiment of the present invention.

[0038] Hereinafter, various embodiments and / or aspects are disclosed with reference to the drawings. For illustrative purposes, numerous specific details are disclosed in the following description to aid in a general understanding of one or more aspects. However, it will also be recognized by those skilled in the art that these aspects may be practiced without such specific details. The following description and the accompanying drawings describe specific exemplary aspects of one or more aspects in detail. However, these aspects are exemplary, and some of the various methods in the principles of the various aspects may be used, and the description is intended to include all such aspects and their equivalents.

[0039] In addition, various aspects and features will be presented by a system that may include multiple devices, components and / or modules, etc. It should also be understood and recognized that various systems may include additional devices, components and / or modules, etc., and / or may not include all of the devices, components, modules, etc. discussed in relation to the drawings.

[0040] Terms such as “embodiment,” “example,” “aspect,” “example,” etc., as used herein, may not be interpreted as implying that any aspect or design described is superior or advantageous to other aspects or designs. Terms used below, such as “part,” “component,” “module,” “system,” “interface,” etc., generally refer to computer-related entities and may, for example, refer to hardware, a combination of hardware and software, or software.

[0041] Additionally, the terms “comprising” and / or “comprising” should be understood to mean that the relevant feature and / or component is present, but not to exclude the presence or addition of one or more other features, components and / or groups thereof.

[0042] Additionally, terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0043] Furthermore, in the embodiments of the present invention, all terms used herein, including technical or scientific terms, unless otherwise defined, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in the embodiments of the present invention.

[0044] FIG. 1 illustrates the components of a decision-making system according to one embodiment of the present invention, and FIG. 2 schematically illustrates a decision-making method according to one embodiment of the present invention.

[0045] As illustrated in FIGS. 1 and 2, a method for dialectically making decisions using multiple agents based on a Large Language Model (LM) and performed in a computing system comprising one or more processors and one or more memories, comprising: an initial answer prompt input step of inputting an initial answer prompt into an LLM including an input query and a phrase requesting the generation of an answer to said query, and an initial first information derivation step of deriving initial first information regarding said query; an initial second information derivation step of inputting an initial rebuttal prompt into an LLM including said query, said initial first information, and a phrase requesting a rebuttal to said initial first information, and deriving initial second information corresponding to a rebuttal to said initial first information; and said query; said initial first information; said initial second information; The method may include an initial decision step, comprising: an initial sum prompt input step for inputting an initial sum prompt into an LLM that includes a phrase requesting the generation of initial third information corresponding to the information combining the above query, the above initial first information, and the above second information; and an initial third information derivation step for deriving initial third information combining the above initial first information and the above initial second information.

[0046] The decision-making described in the present invention can be understood as outputting an answer to a query entered by a user using an LLM. Preferably, it can be understood as causing the LLM to output an answer to a query entered by a user through a dialectical decision-making process.

[0047] Specifically, dialectics is a logical flow of thought that derives a 'thesis' (these) regarding the first thought or perspective presented in response to a claim, inquiry, or request, a 'antithesis' (antithese) corresponding to an opinion, claim, or refutation opposed or in opposition to the said 'thesis', and a 'synthesis' derived by comprehensively considering the mutually contradictory claims of the said 'thesis' and the said 'antithesis'. In this invention, the LLM can be enabled to sequentially derive the 'thesis', 'antithesis', and 'synthesis' in response to a query to make a final decision.

[0048] For example, in response to the query "Explain photosynthesis" entered by a user, LLM can derive a final conclusion regarding the query by outputting information corresponding to "Thesis" such as "Photosynthesis refers to the process in which plants use sunlight to produce nutrients themselves," information corresponding to "Antithesis" such as "Photosynthesis is not merely a simple energy conversion process but is also deeply involved in environmental issues such as climate change," and information corresponding to "Synthesis" such as "Photosynthesis refers to an important ecological process that regulates the Earth's environment and climate, in addition to being an energy conversion process for plant survival."

[0049] In one embodiment of the present invention, the final conclusion may be determined based on information corresponding to 'thesis', 'antithesis', and 'synthesis'. Or, in one embodiment of the present invention, the final conclusion may be determined based on information corresponding to 'synthesis'.

[0050] The decision-making system of the present invention, which performs such a process, may include the components illustrated in FIG. 1. Specifically, the decision-making system may include an initial decision-making unit (1) that performs an initial decision-making step for a query, and a repeating decision-making unit (2) that performs an initial decision-making step when the 'initial third information' output for the initial decision-making step is not valid.

[0051] In addition, the repetitive decision-making unit (2) can repeatedly perform the repetitive decision-making steps described below.

[0052]

[0053] The initial decision unit (1) may include a data receiving unit that receives data including a query from a user, an initial first information generating unit (11) that derives 'initial first information' for the query, an initial second information generating unit (12) that derives 'initial second information', and an initial third information generating unit (13) that derives 'initial third information'.

[0054] Additionally, the initial first information generation unit (11) may include a first replacement information generation unit (110) that generates first replacement information to replace the 'initial first information' depending on whether the 'initial first information' is valid, and a second replacement information generation unit (111) that generates second replacement information to replace the 'first replacement information' depending on whether the 'first replacement information' is valid. Preferably, it may include n nth (where n is a natural number greater than or equal to 1) replacement information generation units for generating replacement information at a given step based on information derived from a preceding step or configuration in a recursive manner. A detailed explanation thereof will be provided later.

[0055] Meanwhile, the initial second information generation unit (12) may include a third replacement information generation unit (120) that generates a "third replacement information" to replace the "initial second information" depending on the validity of the "initial second information," and a fourth replacement information generation unit (121) that generates a "fourth replacement information" to replace the "third replacement information" depending on the validity of the "third replacement information." Preferably, the operation method of the third replacement information generation unit (120) may be substantially the same as the operation method of the first replacement information generation unit (110) in the initial first information generation unit (11). Likewise, the operation method of the fourth replacement information generation unit (121) may be substantially the same as the operation method of the second replacement information generation unit (111) in the initial first information generation unit (11).

[0056] Meanwhile, the decision-making system may include a repeating decision-making unit (2) that generates 'repeat first information', 'repeat second information', and 'repeat third information' based on information derived from the initial decision-making unit (1).

[0057] Specifically, the iteration decision unit (2) can generate ‘iteration first information’, ‘iteration second information’, and ‘iteration third information’ when the LLM determines that the ‘initial third information’ is not valid.

[0058] In one embodiment of the present invention, the decision-making system may include one or more iterative decision-making units (2), and a sequential decision-making process may be performed by one or more iterative decision-making units (2). Specifically, when the first iterative decision-making unit (2) determines that the 'initial third information' derived from the initial decision-making unit (1) is not valid, it may generate 'iterative first information', 'iterative second information', and 'iterative third information', and when the second iterative decision-making unit (2) determines that the 'iterative third information' derived from the first iterative decision-making unit (2) is not valid, it may regenerate 'iterative first information', 'iterative second information', and 'iterative third information', and in this manner, one or more iterative decision-making units (2) may sequentially generate 'iterative first information', 'iterative second information', and 'iterative third information'.

[0059] Consequently, in one embodiment of the present invention, until it is determined that the ‘initial third information’ or ‘repetition third information’ corresponding to the ‘synthesis’ in dialectics is valid, the (initial decision unit (1)) - (first repetition decision unit) - (second repetition decision unit) - (n-th repetition decision unit) can sequentially generate the ‘initial third information’ or ‘repetition third information’.

[0060] In other words, if the 'initial third information' derived from the initial decision unit (1) is determined to be valid, the repeated decision unit may not perform a separate decision-making process.

[0061] Meanwhile, one or more repetition decision units (2) may each include a 'repetition first information generation unit (20)', a 'repetition second information generation unit (21)', and a 'repetition third information generation unit (22)'.

[0062] Additionally, the ‘repetition first information generation unit (20)’ may include a first replacement information generation unit (200) that generates ‘first replacement information’ to replace ‘repetition first information’ based on the validity of ‘repetition first information’, and a second replacement information generation unit (201) that generates ‘second replacement information’ to replace ‘first replacement information’ based on the validity of ‘first replacement information’.

[0063] Preferably, the method by which the 1.1 replacement information generation unit (200) generates replacement information for the 'repeated first information' may be substantially the same as the method by which the 1 replacement information generation unit (110) in the initial first information generation unit (11) generates replacement information for the 'initial first information'.

[0064] In addition, the method by which the 2.1 replacement information generation unit (201) generates replacement information for the '1.1 replacement information' may be substantially the same as the method by which the 2 replacement information generation unit (111) in the initial 1 information generation unit (11) generates replacement information for the '1 replacement information'.

[0065] Similarly, the repeating second information generation unit (21) may further include a 3.1 replacement information generation unit (210) that generates replacement information for the 'repeating second information' in the same manner as the 3 replacement information generation unit (120) in the initial second information generation unit (12) generates replacement information for the 'initial second information', and a 4.1 replacement information generation unit (211) that generates replacement information for the '3.1 replacement information' in the same manner as the 4 replacement information generation unit (121) in the initial second information generation unit (12) generates replacement information for the '3 replacement information'.

[0066] Meanwhile, the decision-making system of the present invention may include an LLM internally or be connected (communicate) with an externally established LLM. Specifically, an LLM may be utilized during the execution process of each component of the aforementioned decision-making system.

[0067] For example, 'initial first information', 'initial second information', 'initial third information', 'first alternative information', and 'second alternative information' can be generated by LLM.

[0068] The decision-making illustrated in FIG. 2 can be achieved through such components.

[0069] Specifically, the decision-making system may generate 'initial first information' for the received query (S1), generate 'initial second information' (S2), and generate 'initial third information' (S3). In this way, steps S1 to S3 performed in the initial decision-making stage may be a decision-making process performed sequentially.

[0070] Meanwhile, if it is determined that the 'initial third information' is not valid, the decision-making system may generate 'repetition first information' (S4), generate 'repetition second information' (S5), and generate 'repetition third information' (S6). In this way, steps S4 through S6 performed in the repetitive decision-making stage may be a decision-making process performed sequentially.

[0071] In the same way, if it is determined that the 'repetition third information' is invalid, the decision system can regenerate the 'repetition first information' (S7), generate the 'repetition second information' (S8), and generate the 'repetition third information' (S9).

[0072] To summarize, steps S1 to S12 in Fig. 2 can be understood as steps in which the decision-making system sequentially generates information corresponding to 'thesis', 'antithesis', and 'synthesis'.

[0073] Meanwhile, in the initial decision-making stage or the iterative decision-making stage, a decision-making process can be carried out in a recursive manner when deriving information corresponding to the 'thesis' or 'opposition'.

[0074] To explain based on the initial decision-making stage, the decision-making system can generate one or more first alternative information (first alternative information a and first alternative information b) when it is determined that the 'initial first information' corresponding to 'correct' is not valid (S1.1).

[0075] Again, when the decision-making system determines that any one of the 'first alternative information', for example, first alternative information b, is not valid, it may generate one or more second alternative information (second alternative information a and second alternative information b) to replace the first alternative information b (S1.2).

[0076] Subsequently, if the decision-making system determines that one or more second alternative information (second alternative information a and second alternative information b) are all valid, it may replace the first alternative information b based on the one or more second alternative information (second alternative information a and second alternative information b) (S1.3). That is, the first alternative information b, which was previously determined to be invalid, may be replaced based on the second alternative information (second alternative information a and second alternative information b) which was determined to be valid.

[0077] Subsequently, the decision-making system can replace the 'initial first information' (S1.4) based on one or more first alternative informations (first alternative information a and first alternative information b replaced through step S1.3) as first alternative information a is determined to be valid and first alternative information b is replaced validly.

[0078]

[0079] To summarize, the recursive decision-making process in which 'initial first information' is replaced with valid information when deemed invalid can be summarized as follows.

[0080] ('Initial 1st Information' is invalid) -> (Generate 1st Alternative Information a and 1st Alternative Information b for 'Initial 1st Information') -> (1st Alternative Information b is invalid) -> (Generate 2nd Alternative Information a and 2nd Alternative Information b for 1st Alternative Information b) -> (2nd Alternative Information a and 2nd Alternative Information b are valid) -> (Replace 1st Alternative Information b based on 2nd Alternative Information a and 2nd Alternative Information b) -> (Replace 'Initial 1st Information' based on 1st Alternative Information a and the replaced 1st Alternative Information b) -> ('Initial 1st Information' is determined to be valid)

[0081] Consequently, until the 'initial first information' is determined to be valid, the first replacement information, the second replacement information (and any additional replacement information generated thereafter) are generated, and the 'initial first information' can be replaced.

[0082] In a similar manner, the decision-making system may generate one or more third alternative information (third alternative information a and third alternative information b) when it is determined that the 'initial second information' corresponding to 'half' is invalid (S2.1).

[0083] Subsequently, if the decision-making system determines that one or more third alternative information (third alternative information a and third alternative information b) are all valid, it may replace the 'initial second information' based on the one or more third alternative information (third alternative information a and third alternative information b) (S2.2). That is, the 'initial second information' previously determined to be invalid may be replaced based on the one or more third alternative information (third alternative information a and third alternative information b) determined to be valid.

[0084] Consequently, until the 'initial second information' is determined to be valid, third replacement information, fourth replacement information (and additional replacement information generated thereafter) are generated, and the 'initial second information' can be replaced.

[0085] To summarize, steps S1.1 to S1.4 and S2.1 to S2.2 in Fig. 2 can be understood as a recursive decision-making process in which the decision-making system replaces previously generated information with newly generated information until the information corresponding to 'correct' or 'incorrect' is valid.

[0086] Meanwhile, although not illustrated in FIG. 2, as described above, a recursive decision-making process such as steps S1.1 to S1.4 and S2.1 to S2.2 can also be performed for 'repetition first information' and 'repetition second information'.

[0087] As such, according to one embodiment of the present invention, a sequential decision-making process (steps S1 to S12) and a recursively performed decision-making process (steps S1.1 to S1.4 and steps S2.1 to S2.2) can be performed simultaneously.

[0088] More specifically, the decision-making process that proceeds sequentially may be determined whether to perform based on third information corresponding to 'synthesis', and the decision-making process that proceeds recursively may be determined whether to perform based on first information corresponding to 'correction' and second information corresponding to 'rejection'.

[0089] As such, according to one embodiment of the present invention, when making a decision (drawing a conclusion) on a query using an LLM, the LLM can be guided to perform a lot of analysis on the input prompt, thereby inducing the LLM to output the most appropriate answer to the prompt.

[0090] In the following Figures 3 through 6, a decision-making process that proceeds sequentially will be explained, and in Figures 7 through 9, a decision-making process that is performed recursively will be explained.

[0091] FIG. 3 illustrates the steps of an initial decision-making step and an iterative decision-making step according to one embodiment of the present invention.

[0092] As illustrated in FIG. 3, the iterative decision-making step is repeated one or more times when the initial third information validity derived from the initial decision-making step is below a preset standard, and the iterative first information in the iterative decision-making step can be derived based on the information derived from the initial decision-making step or the information derived from the iterative decision-making step performed in the previous iterative step.

[0093] Specifically, steps S1, S2, S3, and Q1 may be steps included in the initial decision-making stage performed by the initial decision-making unit (1).

[0094] Specifically, the initial decision-making unit (1) sequentially derives the 'initial first information', 'initial second information', and 'initial third information' in steps S1, S2, and S3, and can derive a conclusion in step Q1 if the 'initial third information' is valid.

[0095] Meanwhile, if the above 'initial third information' is not valid in step Q1, a repetition decision step may be performed by the repetition decision unit (2). Specifically, steps S4, S5, S6, and Q2 may be steps included in the repetition decision step performed first by the repetition decision unit (2).

[0096] Specifically, the repetition decision unit (2) sequentially derives the 'repetition first information', 'repetition second information', and 'repetition third information' in steps S4, S5, and S6, and can derive a conclusion in step Q2 if the 'repetition third information' is valid.

[0097] Meanwhile, if the above 'repetition third information' is not valid even in the Q2 stage, a repetition decision stage may be additionally performed by the repetition decision unit (2). Specifically, stages S7, S8, S9, and Q3 may be stages included in the repetition decision stage performed a second time by the repetition decision unit (2).

[0098] In the same way, the repetition decision unit (2) sequentially re-derives the 'repetition first information', 'repetition second information', and 'repetition third information' in steps S7, S8, and S9, and can derive a conclusion in step Q3 if the 'repetition third information' is valid.

[0099] As described above, according to one embodiment of the present invention, when the 'initial third information' or 'repeated third information' is determined to be valid, an additional iterative decision-making step may not be performed. In other words, the iterative decision-making step may be performed repeatedly until the 'initial third information' or 'repeated third information' is determined to be valid.

[0100] FIG. 4 illustrates an initial decision-making step according to one embodiment of the present invention.

[0101] As illustrated in FIG. 4(a), the decision system can derive 'initial first information' for a query by inputting an initial decision prompt containing an input query and a phrase requesting the generation of an answer to the query into the LLM.

[0102] Specifically, a query may be information that a user queries or requests. For example, a query may be information requesting a user's question or analysis regarding a certain topic, and its form may include one or more of text, voice, images, and video, but is not limited to any one of them.

[0103] For example, the initial prompt may include the user's request, 'Explain photosynthesis,' and a phrase requesting the generation of an answer to that question ('Generate an answer to the entered question').

[0104] LLM can receive an initial prompt and generate 'initial first information' corresponding to the answer to the query, and as previously mentioned, 'initial first information' can be understood as information corresponding to 'thesis' dialectically.

[0105] Next, as illustrated in FIG. 4(b), the decision system can derive 'initial second information' by inputting an initial counter-prompt containing a query, 'initial first information', and a phrase requesting a rebuttal to 'initial first information' into the LLM.

[0106] Specifically, the decision-making system can input 'initial first information' into the LLM and derive 'initial second information' while asserting an opinion refuting the 'initial first information'.

[0107] In other words, 'initial second information' can be a claim that opposes, contradicts, or conflicts with 'initial first information,' and can be understood dialectically as information corresponding to 'antithesis.'

[0108] In one embodiment of the present invention, the phrase requesting a rebuttal to the 'initial first information' may be information in the form of text, such as “Rebuttal to the 'initial first information'”.

[0109] Next, as illustrated in Fig. 4(c), the decision-making system can derive the 'initial third information' by inputting an initial sum prompt into the LLM that includes a query, 'initial first information', 'initial second information', and a phrase requesting the generation of initial third information corresponding to the information combining the query, initial first information, and initial second information.

[0110] Specifically, the decision-making system can input 'initial first information' and 'initial second information' into the LLM, and then synthesize the 'initial first information' and 'initial second information' to derive 'initial third information'.

[0111] In other words, 'initial third information' can be a new dimension of argument derived by comprehensively considering 'initial first information' and 'initial second information,' and can be understood as information corresponding to a 'synthesis' dialectically.

[0112] In one embodiment of the present invention, the phrase requesting the generation of 'initial third information' may be information in the form of text, such as “Synthesize 'initial first information' and 'initial second information' to derive a final conclusion.”

[0113] FIG. 5 illustrates an iterative decision-making step according to one embodiment of the present invention.

[0114] As illustrated in FIG. 5, the initial third information derivation step further includes a third information validity review step for deriving initial third information validity by reviewing the validity of the initial third information by the LLM model; and the method further includes a repeating decision step for deriving repeating first information, repeating second information, and repeating third information by inputting one or more pieces of information derived in the initial decision step into the LLM when the initial third information validity derived in the initial decision step is below a preset standard; and the repeating third information may be information obtained by combining the initial first information and the initial second information by the LLM.

[0115] Additionally, the iterative decision-making step may include: a iterative first information derivation step, which inputs the query, the initial first information, the initial second information, and the initial third information into an LLM to derive iterative first information by synthesizing the query, the initial first information, the initial second information, and the initial third information; an iterative second information derivation step, which inputs information including the query and the iterative first information into an LLM to derive iterative second information corresponding to a rebuttal to the iterative first information; and an iterative third information derivation step, which inputs information including the query, the iterative first information, and the iterative second information into an LLM to derive iterative third information by synthesizing the iterative first information and the iterative second information.

[0116] Additionally, the above initial synopsis prompt may further include a phrase requesting a validity review of the initial third information generated based on the above initial synopsis prompt input into the LLM.

[0117] As shown in FIG. 5(a) following FIG. 4(c), the decision-making system can review the validity of the 'initial third information'.

[0118] In one embodiment of the present invention, the decision-making system may perform a step S3 for deriving 'initial third information' and a step Q1 for reviewing the validity of 'initial third information' together. Specifically, the initial prompt may further include a phrase requesting a review of the validity of 'initial third information'.

[0119] For example, an initialization prompt may include phrases such as, “Synthesize ‘initial first information’ and ‘initial second information’ to derive a final conclusion, and review whether the derived conclusion is valid.”

[0120]

[0121] An LLM, having been requested by a decision-making system to review the validity of 'initial third information,' can independently examine whether the 'initial third information' derived by itself is valid. Specifically, the LLM can determine whether the 'initial third information' is valid or not. In other words, at the Q1 stage, a reflection process can be performed in which the LLM provides feedback and analyzes whether the 'initial third information' is logically correct and free of errors.

[0122] As mentioned above, whether to perform the iterative decision-making step may be determined based on the LLM's judgment regarding the validity of the 'initial third information'. Specifically, if the 'initial third information' is valid, the initial decision-making system may derive a final conclusion based on the 'initial third information', and said conclusion may be an answer to the query.

[0123] In one embodiment of the present invention, the LLM quantifies the validity of information for which a validity review has been requested, and determines that the information is valid if the quantified validity exceeds a preset standard, and determines that the information is not valid if it is below the preset standard.

[0124] Meanwhile, if the 'initial third information' is not valid, the initial decision-making system can perform an iterative decision-making step based on the 'initial third information'.

[0125] Specifically, as illustrated in FIG. 5(b), the decision-making system can derive ‘repetition first information’ by inputting a repetition prompt containing a query, ‘initial first information’, ‘initial second information’, ‘initial third information’, and a request phrase for generating repetition first information into the LLM.

[0126] That is, in stage S4, the information derived from the initial decision stage ('Initial 1st Information', 'Initial 2nd Information', and 'Initial 3rd Information') can be combined to derive 'Repetitive 1st Information' corresponding to a new 'decision'.

[0127] In this way, while the 'initial first information' corresponding to the 'correct' output in the initial decision-making stage is generated based on a query, the 'repetitive first information' corresponding to the 'correct' output in the iterative decision-making stage can be generated by additionally considering information derived in the initial decision-making stage in addition to the query.

[0128] Meanwhile, as described above, in one embodiment of the present invention, the iterative decision-making step may be performed one or more times, and the decision-making system may derive 'iteration first information' in the iterative decision-making step performed in that step based on the information derived from the previously performed iterative decision-making step.

[0129] For example, the decision-making system can derive 'Iteration 1 Information' in the second iteration decision-making step by inputting 'Iteration 1 Information', 'Iteration 2 Information', and 'Iteration 3 Information' derived in the first iteration decision-making step into the LLM.

[0130] Next, as illustrated in Fig. 5(c), the decision-making system can derive ‘repetition second information’ by inputting a repetition prompt containing a query, ‘repetition first information’, and a request phrase for generating ‘repetition second information’ into the LLM.

[0131] That is, in step S5, based on the 'repetition first information' newly derived in the repetition decision step, 'repetition first information' corresponding to a new 'repetition' can be derived.

[0132] Next, as illustrated in FIG. 5 (d), the decision system can derive ‘repetition 3 information’ by inputting a repetition prompt containing a query, ‘repetition 1 information’, ‘repetition 2 information’, and ‘repetition 3 information’ to the LLM.

[0133] That is, in step S6, 'repetition information 3' corresponding to a new 'sum' can be derived based on the 'repetition information 1' and 'repetition information 2' newly derived in the repetition decision-making step.

[0134] Subsequently, the decision-making system examines the validity of the 'repeated third information' and, when it is determined that the 'repeated third information' is valid, draws a conclusion based on the 'repeated third information', and when it is not valid, it can re-derive the 'repeated first information', 'repeated second information', and 'repeated third information'.

[0135] FIG. 6 illustrates a third information validity review step according to an embodiment of the present invention.

[0136] As illustrated in FIG. 6, the third information validity review step may include a step of deriving the initial third information validity by inputting a prompt containing the initial third information and a phrase requesting a validity review of the initial third information into the LLM.

[0137] Specifically, when the decision-making system examines the validity of 'initial third information', it may sequentially perform the S3 step of deriving 'initial third information' and the Q1 step of examining the validity of 'initial third information'.

[0138] Specifically, the decision-making system derives 'initial third information' through the LLM, and can review the validity of the 'initial third information' by inputting the derived 'initial third information' and a phrase requesting a validity review of the 'initial third information' into the LLM.

[0139] FIG. 7 illustrates the steps of generating first replacement information and second replacement information and replacing the first information according to an embodiment of the present invention.

[0140] As illustrated in FIG. 7, the decision-making system determines whether the 'initial first information' is valid (Q10), generates one or more first replacement information to replace the 'initial first information' determined to be invalid (S10), and can replace the 'initial first information' based on the first replacement information (S11).

[0141] Meanwhile, the decision-making system may determine whether each of the one or more first alternative information is valid (Q11), generate one or more second alternative information to replace the first alternative information determined to be invalid (S100), and when the one or more second alternative information is determined to be valid (Q12), replace the first alternative information based on the second alternative information (S102).

[0142] FIG. 8 illustrates details regarding the first alternative information according to one embodiment of the present invention.

[0143] As illustrated in FIG. 8, the initial first information derivation step may include: a first information validity review step comprising the step of inputting a prompt into an LLM that includes the generated initial first information and a phrase requesting a validity review of the initial first information; a first alternative information derivation step comprising the step of inputting a prompt into an LLM that includes the initial first information, a first error derived during the validity review process of the initial first information, and a phrase requesting to derive first alternative information when the validity of the initial first information is below a preset standard; and an initial first information replacement step comprising inputting a prompt into an LLM that includes the first alternative information and a phrase requesting a validity review of the first alternative information, and replacing the initial first information based on the first alternative information when the validity of the first alternative information exceeds a preset standard.

[0144] In addition, the initial first information replacement step may replace the initial first information based on the one or more first replacement information when the validity of each of the one or more first replacement information generated for the initial first information exceeds a preset standard.

[0145] As illustrated in FIG. 8(a), the decision-making system can review the validity of the 'initial first information' by inputting a prompt into the LLM that includes the generated 'initial first information' and a phrase requesting a validity review of the 'initial first information'.

[0146] Meanwhile, a first error may be generated during the process in which the LLM examines the validity of the 'initial first information'. Specifically, a first error may arise from the finding that the 'initial first information' is invalid, and it may correspond to factual errors, logical errors, interpretive errors, linguistic errors, biased opinions, etc., regarding the 'initial first information'.

[0147] In other words, the first error may be the basis for the judgment that the 'initial first information' is invalid when it is judged to be invalid.

[0148] As illustrated in FIG. 8(b), when the decision system determines that the 'initial first information' is invalid, it inputs a prompt to the LLM containing phrases requesting the generation of the 'initial first information', the first error, and the first alternative information, thereby generating one or more first alternative information for the 'initial first information'.

[0149] Specifically, the phrase requesting the generation of the first alternative information may be a phrase requesting the LLM to derive a new answer to the query based on the first error.

[0150] In one embodiment of the present invention, the LLM can generate first alternative information corresponding to a new answer to the query by analyzing the query from a different perspective than when outputting the 'initial first information' based on the first error. Specifically, the LLM can generate one or more first alternative information by analyzing the query from one or more perspectives.

[0151] For example, in response to the query “Tell me about photosynthesis,” LLM analyzes from a scientific perspective to generate ‘initial first information,’ and when generating first alternative information, it can generate first alternative information from an ecological perspective and first alternative information from a philosophical perspective, respectively.

[0152] Alternatively, the first alternative information may correspond to an answer that corrects the errors of the 'initial first information' by taking the first error into account. Specifically, the LLM may correct factual errors, logical errors, and interpretive errors of the 'initial first information' respectively, and generate first alternative information corresponding to each.

[0153] As illustrated in FIG. 8(c), the decision-making system can review the validity of each of one or more first alternative information. Specifically, the decision-making system can review the validity of each of one or more first alternative information by inputting a prompt containing the first alternative information and a phrase requesting a review of the validity of the first alternative information into the LLM.

[0154] Meanwhile, if the decision-making system determines that each of the one or more first alternative informations is valid, it may replace the 'initial first information' based on the one or more first alternative informations.

[0155] In one embodiment of the present invention, the decision-making system may replace the 'initial first information' with information generated by synthesizing one or more first alternative information using LLM. Or, in another embodiment of the present invention, the decision-making system may replace the 'initial first information' with any one of one or more first alternative information.

[0156] According to one embodiment of the present invention, 'initial first information' that was determined to be invalid by LLM is replaced based on one or more first alternative information that is determined to be valid by LLM, thereby enabling more accurate decision-making.

[0157] FIG. 9 illustrates details regarding second alternative information according to one embodiment of the present invention.

[0158] As illustrated in FIG. 9, the initial first information replacement step further includes a second replacement information derivation step performed when it is determined that the validity of one or more of the first replacements is below a preset standard; and the second replacement information derivation step may include a second replacement information derivation step comprising a step of inputting a prompt into an LLM that includes the first replacement information whose validity is below a preset standard, a second error derived during the validity review process for the first replacement information, and a phrase requesting the derivation of the second replacement information; and a first replacement information replacement step that replaces the first replacement information based on the second replacement information.

[0159] As illustrated in FIG. 9(a), the decision-making system can review the validity of each of one or more first alternative informations, and a second error may be generated for the first alternative information that is determined to be invalid. Specifically, the first alternative information shaded in gray in FIG. 9(a) is information determined to be invalid by the LLM, and a second error may be generated for the said first alternative information.

[0160] Preferably, the second error is generated during the process in which the LLM examines the validity of the first alternative information, and may correspond to reasons why the first alternative information is invalid, factual errors, logical errors, interpretive errors, linguistic errors, biased opinions, etc.

[0161] As illustrated in FIG. 9(b), the decision-making system can generate second alternative information for first alternative information determined to be invalid by inputting a prompt to the LLM that includes first alternative information determined to be invalid, a second error generated for said first alternative information, and a phrase requesting to derive second alternative information.

[0162] Specifically, the decision-making system may generate one or more second alternative information for each of the one or more first alternative information that is determined to be invalid among the generated one or more first alternative information.

[0163] Similar to how the first replacement information is information that corrects the first error of the 'initial first information', the second replacement information may be information that corrects the second error of the first replacement information. Additionally, the second replacement information may be generated in one or more quantities for the first replacement information.

[0164] Additionally, a phrase requesting the generation of second alternative information may be a phrase requesting the LLM to derive an answer different from the first alternative information based on the second error.

[0165] As illustrated in FIG. 9(c), the decision system can review the validity of each of one or more second alternative information. Specifically, the decision system can review the validity of each of one or more second alternative information by inputting a prompt containing the second alternative information and a phrase requesting a review of the validity of the second alternative information into the LLM.

[0166] The decision-making system may replace the first alternative information based on the one or more second alternative information when it is determined that each of the one or more second alternative informations is valid.

[0167] In one embodiment of the present invention, the decision-making system may replace the first alternative information with information generated by synthesizing one or more second alternative information using LLM. Or, in another embodiment of the present invention, the decision-making system may replace the first alternative information with any one of one or more second alternative information.

[0168] According to one embodiment of the present invention, a first alternative information that was determined to be invalid by LLM is replaced based on one or more second alternative information that was determined to be valid by LLM, and then the 'initial first information' that was determined to be invalid is replaced based on one or more first alternative information that was determined to be valid, thereby making a more accurate decision.

[0169] Meanwhile, if there is a second alternative information among one or more second alternative information that is judged to be invalid by the LLM, the decision system inputs the 'second alternative information judged to be invalid' into the LLM again and generates alternative information to replace the second alternative information, and when the generated alternative information is valid, it can replace the second alternative information based on the said alternative information.

[0170] The decision-making system can repeat such recursive decision-making steps until the 'initial first information' is finally replaced with valid information.

[0171] FIG. 10 illustrates data according to one embodiment of the present invention.

[0172] As illustrated in FIG. 10, the first alternative information may correspond to an answer output to the LLM from a different perspective than when the LLM outputs the initial first information corresponding to the answer to the query in the initial first information derivation step.

[0173]

[0174] As described above, the decision-making system of the present invention can sequentially generate information corresponding to the thesis, antithesis, and synthesis when making a decision on a query through a dialectical process.

[0175] As illustrated in FIG. 10 (a), the information corresponding to the 'correct' output by the LLM that received the query as an answer to the query may differ from the ground truth for the query. Specifically, the information corresponding to the 'correct' output by the LLM may be understood as having a difference (error) from the ground truth.

[0176] In other words, conceptually, 'Jeong' may be information that does not correspond to ground truth, or information that does not contain information that corresponds to ground truth.

[0177] Meanwhile, an LLM that receives information corresponding to 'thesis' and a phrase requesting a rebuttal to the 'thesis' can refute the content, opinions, etc. in the 'thesis' and output information corresponding to 'the rebuttal'.

[0178] In other words, conceptually, 'antithesis' can include information omitted from 'thesis'—that is, information that corresponds to ground truth but is not included in 'thesis'. Also, 'antithesis' can include incorrect information from 'thesis'—that is, information that does not correspond to ground truth but is included in 'thesis'.

[0179] Consequently, LLM can derive a 'synthesis' by combining the 'thesis' and the 'antithesis'; conceptually, it can derive a final conclusion ('synthesis') regarding a query by excluding incorrect information from the initial response ('thesis') output and additionally considering missing information ('antithesis').

[0180] As shown in Fig. 10(b), the query can be analyzed from various perspectives by LLM.

[0181] For example, regarding a query asking about photosynthesis, LLM can initially analyze it from a scientific perspective and derive 'initial first information,' which is an answer from a scientific perspective.

[0182] Meanwhile, if it is determined that the 'initial first information' is not valid, one or more first replacement information may be generated to replace the 'initial first information' with a different point from when the 'initial first information' was derived.

[0183] For example, the first alternative information may correspond to an answer from an ecological or philosophical perspective, rather than a scientific perspective, to a query asking about photosynthesis in the LLM.

[0184] As described above, if the 'first alternative information' is determined to be invalid, one or more second alternative information may be generated to replace the first alternative information with a different point from when the 'first alternative information' was derived.

[0185] For example, in one embodiment of the present invention, the LLM may derive 'initial first information' from a scientific perspective, derive first alternative information from an ecological perspective, and derive second alternative information from a philosophical perspective.

[0186] In this manner, according to one embodiment of the present invention, regarding input information, the LLM can be configured to analyze it from multiple perspectives rather than a single perspective to output an answer.

[0187] The following describes a method for generating questions and answers to evaluate an artificial intelligence system including LLM. Specifically, the “method for generating questions and answers to evaluate an artificial intelligence system including LLM” can be implemented by applying the aforementioned “method for dialectically making decisions using multiple agents based on LLM (Large Language Model).”

[0188] That is, “a method for generating questions and answers for evaluating an artificial intelligence system including LLM and a system for performing the same” may include “a method for making decisions dialectically using multiple agents based on LLM (Large Language Model) and a decision-making system for performing the same.”

[0189] FIG. 11 illustrates the steps of a method for generating questions and answers for evaluating an artificial intelligence system including an LLM according to an embodiment of the present invention.

[0190] As illustrated in FIG. 11, a method for generating questions and answers for evaluating an artificial intelligence system including an LLM comprises: a request receiving step for receiving a request regarding one or more of evaluation items and evaluation criteria for an artificial intelligence system to be evaluated by a user; a request analysis step for analyzing the request and deriving a request analysis result by means of a request analysis model including or connected to an LLM; and an evaluation rubric derivation step for deriving an evaluation rubric including score-based evaluation criteria based on the request analysis result by means of a rubric generation model including or connected to an LLM. A question-answer generation step that generates questions and answers according to the evaluation rubric by means of a question-answer generation model that includes or is connected to an LLM; wherein one or more of the request analysis step, the evaluation rubric derivation step, and the question-answer generation step include an initial correct prompt input step that inputs an initial correct prompt into the LLM, the initial correct prompt including an input query and a phrase requesting the generation of an answer to the query, and an initial first information derivation step that derives initial first information regarding the query; an initial second information derivation step that inputs an initial rebuttal prompt into the LLM, the initial first information including the query, the initial first information, and a phrase requesting a rebuttal to the initial first information, and an initial second information derivation step that derives initial second information corresponding to a rebuttal to the initial first information; and the query; the initial first information; the initial second information; The method may include an initial decision step, comprising: an initial sum prompt input step for inputting an initial sum prompt into an LLM that includes a phrase requesting the generation of initial third information corresponding to the information combining the above query, the above initial first information, and the above second information; and an initial third information derivation step for deriving initial third information combining the above initial first information and the above initial second information.

[0191] “A method and system for generating questions and answers for evaluating an artificial intelligence system including LLM (hereinafter referred to as the question-answer generation system)” is intended to generate a dataset containing questions and answers to determine whether an artificial intelligence system including LLM meets design objectives and required performance requirements. It can generate a dataset containing various multiple questions and answers that match the redefined evaluation items by performing a specific analysis of evaluation items described by the user.

[0192] Specifically, at step S1000, the question-answer generation system may receive a request from a user. More specifically, the request may be information related to a dataset that the user intends to generate or an artificial intelligence system that the user intends to evaluate. Preferably, the request may include information regarding one or more evaluation items and evaluation criteria for evaluating the artificial intelligence system that the user intends to evaluate.

[0193] For example, regarding a chatbot based on an artificial intelligence system that provides information related to semiconductor process collaboration as illustrated in FIG. 11, it may include one or more of the following: an evaluation item that the user intends to evaluate (whether a potential solution to a problem faced in semiconductor process collaboration is appropriately provided), an evaluation criterion (a criterion for evaluating whether a potential solution to a problem faced in semiconductor process collaboration is appropriately provided), or examples related to the evaluation (cases and grounds where the evaluation was appropriately performed, cases and grounds where the evaluation was not appropriately performed), and other additional information related to the evaluation.

[0194] In one embodiment of the present invention, the request may be a prompt containing information on one or more of evaluation items and evaluation criteria for evaluating an artificial intelligence system that the user intends to evaluate.

[0195] In step S2000, the question-answer generation system can analyze a request received from a user. Specifically, the question-answer generation system may include an LLM or a request analysis model connected to an LLM, and the request received from the user may be analyzed by the request analysis model to derive a request analysis result.

[0196] In one embodiment of the present invention, the request analysis model inputs a prompt containing a request received from a user and a phrase for analyzing the request into an LLM, and can derive a request analysis result for the request.

[0197] For example, in Fig. 11, the request analysis model inputs a prompt containing the request into the LLM, and as a result of the request analysis, can derive that the user's request is related to "reasoning ability to present a potential solution" and "whether it causes hallucination."

[0198] In step S3000, the question-answer generation system can derive an evaluation rubric including score-based evaluation criteria based on the request analysis results. Specifically, the evaluation rubric is a score-based evaluation criterion for evaluating an artificial intelligence system, and preferably, it may include information regarding each score-based evaluation criterion.

[0199] Preferably, the question-answer generation system may include an LLM or a rubric generation model connected to an LLM, and an evaluation rubric including score-based evaluation criteria may be generated based on the request analysis results by the rubric generation analysis model.

[0200] For example, in FIG. 11, the rubric generation model inputs a prompt containing the request analysis results into the LLM to derive evaluation criteria that can evaluate “reasoning ability to present potential solutions” and “whether it causes hallucination” on a scale of 1 to 5 as an evaluation rubric.

[0201] At step S4000, the question-answer generation system can generate questions and answers based on an evaluation rubric. Specifically, the question-answer generation system may include an LLM or a question-answer generation model connected to an LLM, and questions and answers based on the evaluation rubric can be generated by the question-answer generation analysis model.

[0202] Preferably, the question-answer generation model can generate questions and answers based on the highest score or grade in the evaluation rubric. In other words, the question-answer generation model can generate model questions and answers that receive the highest evaluation (score or grade) according to the evaluation rubric.

[0203] For example, in Fig. 11, the question-answer generation model can generate questions and answers that can measure “reasoning ability to suggest potential solutions” and “whether to cause hallucination.”

[0204] In step S5000, the question-answer generation system can generate multiple questions and answers corresponding to lower scores or grades based on the questions and answers generated for the highest score or grade in the evaluation rubric. Specifically, the question-answer generation system may include an LLM or an answer diversification model connected to an LLM, and questions and answers for each of the remaining scores or grades, excluding the highest score or grade in the evaluation rubric, can be generated by the answer diversification model.

[0205] For example, in FIG. 11, the question-answer generation model can generate questions and answers that measure “reasoning ability to present potential solutions” and “whether to cause hallucination” corresponding to each of the highest scores or grades in the evaluation rubric. Specifically, the question-answer generation model can generate one or more questions and answers that are evaluated lower than the highest evaluation according to the evaluation rubric by referring to the best questions and answers that receive the highest evaluation (score or grade) according to the evaluation rubric.

[0206] In step S6000, the question-answer generation system can synthesize one or more generated questions and answers to generate a final dataset containing questions and answers. Specifically, based on multiple questions and answers generated for each evaluation criterion distinguished by score in the evaluation rubric, it can generate a final question and answer to evaluate an artificial intelligence system according to a user's request.

[0207] In this way, steps S2000 to S6000 can derive conclusions at the corresponding steps using LLM, and as described above, the content described in the “method for dialectically making decisions using multiple agents based on LLM (Large Language Model) and a decision-making system for performing the same” can be applied to each step using LLM.

[0208] For example, in the request analysis stage, when the request analysis model uses LLM to analyze a request entered by a user and derives a request analysis result, a sequential decision-making process that derives information corresponding to 'thesis', 'antithesis', and 'synthesis' for the request, and a recursive decision-making process that validly replaces information corresponding to 'thesis' and 'antithesis' may be performed.

[0209] That is, the request input into the LLM during the request analysis stage may correspond to the query input into the decision system in Fig. 2, and the final conclusion derived from the query (information based on 'sum') may correspond to the request analysis result.

[0210] Alternatively, the request analysis result input into the LLM in the evaluation rubric derivation stage may correspond to the query input into the decision-making system in Fig. 2, and the final conclusion derived from the query (information based on 'sum') may correspond to the evaluation rubric.

[0211] FIG. 12 illustrates an evaluation rubric according to one embodiment of the present invention.

[0212] As illustrated in FIG. 12, when a user wants to evaluate an 'artificial intelligence system that analyzes case law and writes related reports' or to generate a dataset for evaluation, the question-answer generation system can generate an evaluation rubric that includes evaluation criteria by score for each evaluation item, including case law analysis, case law criticism, and report writing.

[0213] In addition, based on such an evaluation rubric, the question and answer generation system can generate questions and answers for the highest grade (5 points) (question and answer generation stage), and then generate questions and answers for each of the remaining grades (1, 2, 3, 4 points) by referring to the questions and answers generated for the highest grade (5 points) (answer diversification stage).

[0214] FIG. 13 schematically illustrates the internal configuration of a computing device according to one embodiment of the present invention.

[0215] The decision-making system illustrated in FIG. 1 described above may include the components of the computing device (11000) illustrated in FIG. 13 above.

[0216] As illustrated in FIG. 13, the computing device (11000) may include at least one processor (11100), memory (11200), peripheral interface (11300), input / output subsystem (I / O subsystem) (11400), power circuit (11500), and communication circuit (11600). In this case, the computing device (11000) may correspond to the decision-making system or question-answer generation system illustrated in FIG. 1.

[0217] The memory (11200) may include, for example, high-speed random access memory, magnetic disk, SRAM, DRAM, ROM, flash memory, or non-volatile memory. The memory (11200) may include software modules, instruction sets, or various other data required for the operation of the computing device (11000).

[0218] At this time, access to memory (11200) from other components such as the processor (11100) or peripheral device interface (11300) can be controlled by the processor (11100).

[0219] The peripheral device interface (11300) can connect input and / or output peripheral devices of the computing device (11000) to the processor (11100) and memory (11200). The processor (11100) can perform various functions for the computing device (11000) and process data by executing software modules or instruction sets stored in memory (11200).

[0220] The input / output subsystem can connect various input / output peripherals to the peripheral interface (11300). For example, the input / output subsystem may include a controller for connecting peripherals such as a monitor, keyboard, mouse, printer, or, if necessary, a touchscreen or sensor to the peripheral interface (11300). According to another aspect, input / output peripherals may be connected to the peripheral interface (11300) without passing through the input / output subsystem.

[0221] The power circuit (11500) can supply power to all or part of the components of the terminal. For example, the power circuit (11500) may include one or more power sources such as a power management system, a battery or alternating current (AC), a charging system, a power failure detection circuit, a power converter or inverter, a power status indicator, or any other components for power generation, management, and distribution.

[0222] The communication circuit (11600) can enable communication with another computing device using at least one external port.

[0223] Alternatively, as described above, the communication circuit (11600) may enable communication with other computing devices by including an RF circuit and transmitting and receiving an RF signal, also known as an electromagnetic signal.

[0224] The embodiment of FIG. 13 is merely an example of a computing device (11000), and the computing device (11000) may have some components shown in FIG. 13 omitted, additional components not shown in FIG. 13 added, or a configuration or arrangement that combines two or more components. For example, a computing device for a communication terminal in a mobile environment may include a touchscreen or sensors, etc., in addition to the components shown in FIG. 13, and the communication circuit (11600) may include a circuit for RF communication of various communication methods (WiFi, 3G, LTE, Bluetooth, NFC, Zigbee, etc.). The components that can be included in the computing device (11000) may be implemented as hardware, software, or a combination of both hardware and software, including one or more integrated circuits specialized for signal processing or applications.

[0225] The methods according to the embodiments of the present invention may be implemented in the form of program instructions that can be executed through various computing devices and recorded on a computer-readable medium. In particular, the program according to the present embodiment may be configured as a PC-based program or an application dedicated to a mobile terminal. An application to which the present invention is applied may be installed on a computing device (11000) through a file provided by a file distribution system. For example, the file distribution system may include a file transmission unit (not shown) that transmits the file in response to a request from the computing device (11000).

[0226] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0227] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave in order to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computing devices and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0228] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0229] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents. Therefore, other implementations, other embodiments, and equivalents to the claims below are also within the scope of the claims.

[0230]

Claims

A method for making decisions dialectically using multiple agents based on a Large Language Model (LM), performed on a computing system comprising 1.1 or more processors and 1 or more memories, wherein A step of inputting an initial prompt into an LLM, the initial prompt including an input query and a phrase requesting the generation of an answer to the query, and a step of deriving initial first information for the query; A step of inputting an initial prompt into an LLM, comprising the above query, the above initial first information, and a phrase requesting a rebuttal to the above initial first information, and a step of deriving initial second information corresponding to a rebuttal to the above initial first information; and A method for dialectically making decisions using an LLM-based multi-agent, comprising: an initial sum prompt input step for inputting an initial sum prompt into an LLM, the initial sum prompt including the above query; the above initial first information; the above initial second information; and a phrase requesting the generation of initial third information corresponding to information combining the above query, the above initial first information, and the above initial second information; and an initial third information derivation step for deriving initial third information combining the above initial first information and the above initial second information; and an initial decision step.

2. In Claim 1, The above initial third information derivation step is, It further includes a third information validity review step for deriving the validity of the initial third information by reviewing the validity of the initial third information through the above LLM; and The above method is, If the validity of the initial third information derived in the initial decision-making step is below a preset standard, the method includes a repeating decision-making step in which one or more pieces of information derived in the initial decision-making step are input into an LLM to derive repeating first information, repeating second information, and repeating third information; The above repetition third information is, A method of dialectically deciding a decision using an LLM-based multi-agent, which is information obtained by synthesizing the initial first information and the initial second information by the above LLM.

3. In Claim 2, The above iterative decision-making step is, A step for deriving iterative first information, wherein the above query, the above initial first information, the above initial second information, and the above initial third information are input into an LLM to derive iterative first information by synthesizing the above query, the above initial first information, the above initial second information, and the above initial third information; A step for deriving second repetition information, wherein information including the above query and the above repetition first information is input into an LLM to derive second repetition information corresponding to a rebuttal to the above repetition first information; and A method for dialectically making decisions using an LLM-based multi-agent, comprising: a step of deriving third iteration information by inputting information including the above query, the above iteration first information, and the above iteration second information into an LLM to derive third iteration information by synthesizing the above iteration first information and the above iteration second information.

4. In Claim 1, The above initial sum prompt is, A method for dialectically deciding using an LLM-based multi-agent, further comprising a phrase requesting a validity review of the initial third information generated based on the initial syntax prompt input into the LLM.

5. In Claim 1, The above initial first information derivation step is, A first information validity review step comprising the step of inputting into an LLM a prompt including the generated initial first information and a phrase requesting a validity review of the initial first information; A first alternative information derivation step comprising the step of inputting into an LLM a prompt including the initial first information, a first error derived during the validity review process of the initial first information, and a phrase requesting the derivation of first alternative information when the validity of the initial first information is below a preset standard; and A method for dialectically deciding a decision using an LLM-based multi-agent, comprising: inputting a prompt into an LLM that includes the first alternative information and a phrase requesting a validity review of the first alternative information, and an initial first information replacement step that replaces the initial first information based on the first alternative information when the validity of the first alternative information exceeds a preset standard.

6. In Claim 5, The above first alternative information is, A method of dialectically deciding a decision using an LLM-based multi-agent, corresponding to an answer output to the LLM from a different perspective from when the LLM outputs initial first information corresponding to an answer to the query in the initial first information derivation step.

7. In Claim 5, The above initial first information replacement step is, It further includes a second alternative information derivation step performed when the validity of one or more of the first alternatives is determined to be below a pre-established standard; and The above second alternative information derivation step is, A second alternative information derivation step comprising the step of inputting into an LLM a prompt including a first alternative information whose validity is below a preset standard, a second error derived during the validity review process of the first alternative information, and a phrase requesting the derivation of the second alternative information; and A method for dialectically deciding a decision using an LLM-based multi-agent, comprising: a first alternative information replacement step of replacing the first alternative information based on the second alternative information.

8. A method for generating questions and answers for evaluating an artificial intelligence system including LLM, A request receiving step for receiving a request for one or more of the evaluation items and evaluation criteria for an artificial intelligence system that the user intends to evaluate; A request analysis step that analyzes a request and derives a request analysis result by means of a request analysis model that includes or is connected to an LLM; An evaluation rubric derivation step for deriving an evaluation rubric including score-based evaluation criteria based on the request analysis results by means of a rubric generation model that includes or is connected to an LLM; and A question-answer generation step that generates questions and answers according to the evaluation rubric by means of a question-answer generation model that includes or is connected to an LLM; At least one of the above request analysis step, the above evaluation rubric derivation step, and the above question-and-answer generation step is, A step of inputting an initial prompt into an LLM, the initial prompt including an input query and a phrase requesting the generation of an answer to the query, and a step of deriving initial first information for the query; A step of inputting an initial prompt into an LLM, comprising the above query, the above initial first information, and a phrase requesting a rebuttal to the above initial first information, and a step of deriving initial second information corresponding to a rebuttal to the above initial first information; and A method for generating questions and answers for evaluating an artificial intelligence system, comprising: an initial decision step; a step of inputting an initial sum prompt into an LLM, the initial sum prompt including the above query; the above initial first information; the above initial second information; and a phrase requesting the generation of initial third information corresponding to the information combining the above query, the above initial first information, and the above second information; and a step of deriving initial third information by combining the above initial first information and the above initial second information.

Citation Information

Patent Citations

  • Reverse discussion statement retrieval method and equipment based on BERT model

    CN116361439A

  • Method and system for evaluating effect of critical thinking learning mode, and storage medium

    CN118014440A

  • Standard detection method, system and terminal based on remote supervision

    CN118626841A

  • Traditional Chinese medicine intelligent dialectical information processing system based on big data

    CN118824569A

  • Dataset generation using large language models

    US20240185001A1