Method and system for automatically generating plurality of sub-tasks for answering single query by using (s)LLM, automatically selecting plurality of generative artificial intelligence models according to detailed task objectives, and performing parallel processing
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LINKBRICKS HORIZON-AI INC
- Filing Date
- 2025-12-08
- Publication Date
- 2026-07-30
Smart Images

Figure KR2025021002_30072026_PF_FP_ABST
Abstract
Description
A method and system for automatically generating multiple subtasks for answering a single question using (S)LLM and automatically selecting and parallelizing multiple generative artificial intelligence models according to detailed task objectives.
[0001] The present invention relates to a method and system for automatically generating multiple subtasks for answering a single question using (s)LLM and automatically selecting and parallelizing multiple generative artificial intelligence models according to the detailed task objectives.
[0002]
[0003] Large Language Models (LLMs) are artificial intelligence models capable of understanding and generating language by learning from a vast amount of text data. As a core technology that enables AI chatbot technology, they perform various Natural Language Processing (NPL) tasks to understand the meaning and context of a user's question and generate appropriate sentences as answers. Representative LLMs include OpenAI's GPT and Google's BERT.
[0004] However, as LLM stores vast amounts of data and learns more diverse patterns, it requires more resources and storage space for detailed language understanding and content generation, and takes a long time to learn. Consequently, this can lead to the problem of token limits, which increase costs and response times, as well as hallucinations that generate inappropriate answers that differ from the facts.
[0005] In contrast, sLLM (Small Large Language Model) basically performs the same functions as LLM, but the model size is relatively smaller compared to LLM, and it is characterized by reducing the number of model parameters and improving accuracy through fine-tuning. Representative sLLMs include Meta's LLaMA and Stanford University researchers' Alpaca.
[0006] In other words, the difference between LLM and sLLM lies in the size of the model and the amount of training data. LLM learns from a large amount of data and can understand and generate various contexts, but it requires significant resources for training and has the disadvantage of being weak for specific purposes and tasks. On the other hand, sLLM is a lightweight model trained on relatively small amounts of data that is suitable for performing tasks more specialized for specific purposes and has the advantages of fast processing speed and high reliability.
[0007]
[0008] The problem that the present invention aims to solve is to provide a method and system for automatically generating multiple subtasks for answering a single question using (s)LLM, and automatically selecting and parallel processing multiple generative artificial intelligence models according to the specific task objectives.
[0009] The problems that the present invention aims to solve are not limited to those mentioned above, and other unmentioned problems will be clearly understood by a person skilled in the art from the description below.
[0010]
[0011] The method of the present invention for solving the above-mentioned problem may include the steps of: an input module receiving a single question from a user; a subtask generation module generating a plurality of subtasks for the single question; an execution module selecting and executing a generative language model that matches the plurality of subtasks to generate one or more sub-answers; and a summary module summarizing one or more sub-answers to generate a final answer.
[0012] The above subtask generation module is a language model trained with a single question as a feature value and multiple subtasks as label values, and when any single question is input, it generates multiple subtasks, and the feature value and label value may be characterized by being generated from a search term of an external search engine and an associated search list for that search term.
[0013] The above execution module may be characterized by selecting and executing generative language models that match multiple subtasks according to predefined criteria to generate one or more sub-answers.
[0014] The above execution module is a language model that determines the domains of multiple subtasks, and may be characterized by selecting and executing a generative language model trained on a corpus of the corresponding domain according to the domains of the multiple subtasks to generate one or more sub-answers.
[0015] The above execution module may be characterized by executing a selected language model for each of a plurality of subtasks in a distributed computing environment.
[0016] The above execution module may be characterized by executing selected language models in parallel for each of a plurality of subtasks.
[0017] The above summary module is a language model that resolves conflicts or duplicates of one or more sub-answers and may be characterized by generating a final answer that assigns weights according to the reliability and importance of one or more sub-answers.
[0018] The orchestration module may further include a step of determining whether the final answer to a single question is a single answer or a multiple subtask type, and generating the final answer immediately if it is a single answer, and generating multiple subtasks if it is a multiple subtask type.
[0019] The above orchestration module may further include a step of determining whether to re-execute the step of generating multiple subtasks by determining the normal output of the final answer.
[0020] The system of the present invention for solving the above-described problem may include: an input module that receives a single question from a user; a subtask generation module that generates a plurality of subtasks for the single question; an execution module that selects and executes a generative language model that matches the plurality of subtasks to generate one or more sub-answers; a summary module that summarizes one or more sub-answers to generate a final answer; and an orchestration module that determines whether the single question is a single-answer type or a multiple-subtask type to determine whether to execute the subtask generation module, and determines whether to re-execute the subtask generation module by determining the normal output of the final answer.
[0021]
[0022] The present invention has the effect of resolving the disadvantages of hallucination and token limitations when using generative (s)LLM for answers to single questions and generating high-quality, sufficient answers by dividing a single question into subtasks using a RAG (Retrieval Augmented Generation) orchestration technique and generating a final answer through the selection, parallel processing, and summarization of generative (s)LLM suitable for the domain of the subtasks.
[0023] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below.
[0024]
[0025] FIG. 1 is a configuration diagram showing a system for performing the method of the present invention.
[0026] FIG. 2 is a flowchart showing the process in which the method of the present invention is performed.
[0027] FIG. 3 is a conceptual diagram showing the operation of the method of the present invention.
[0028] FIG. 4 is a conceptual diagram showing the operation of the sub-task generation module of the present invention.
[0029] FIG. 5 is a conceptual diagram showing the operation of the execution module of the present invention.
[0030] Figure 6 is a conceptual diagram showing the operation of the orchestration model of the present invention.
[0031]
[0032] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the present invention, and the present invention is defined only by the scope of the claims.
[0033] The terms used in this specification are for describing embodiments and are not intended to limit the invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. The terms "comprises" and / or "comprising" used in this specification do not exclude the presence or addition of one or more other components in addition to the components mentioned. Throughout the specification, the same reference numerals refer to the same components, and "and / or" includes each of the mentioned components and all combinations of one or more. Although terms such as "first," "second," etc., are used to describe various components, these components are not limited by these terms. These terms are used merely to distinguish one component from another. Therefore, the first component mentioned below may be the second component within the technical scope of the invention.
[0034] Unless otherwise defined, all terms used herein (including technical and scientific terms) may be used in a meaning commonly understood by those skilled in the art to which the present invention pertains. Additionally, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.
[0035]
[0036] Hereinafter, the system (10) and method (20) of the present invention will be described with reference to the drawings. FIG. 1 is a configuration diagram showing a system for performing the method of the present invention, FIG. 2 is a flowchart showing a process for performing the method of the present invention, FIG. 3 is a conceptual diagram showing the operation of the method of the present invention, FIG. 4 is a conceptual diagram showing the operation of the sub-task generation module of the present invention, FIG. 5 is a conceptual diagram showing the operation of the execution module of the present invention, and FIG. 6 is a conceptual diagram showing the operation of the orchestration model of the present invention.
[0037] In describing the system (10) and method (20) of the present invention below, the language model can be interpreted as an artificial intelligence model capable of learning text data to understand and generate language, and, for example, may be an LLM, sLLM, SLM, etc., but is not limited thereto and should be broadly interpreted as a comprehensive concept that includes all engines that perform various natural language processing (NPL) tasks.
[0038]
[0039] The system (10) of the present invention may include an input module (11) for receiving a single question from a user, a subtask generation module (12) for generating a plurality of subtasks for the single question, an execution module (13) for generating one or more sub-answers by selecting and executing a generative language model that matches the plurality of subtasks, a summary module (14) for generating a final answer by summarizing one or more sub-answers, and an orchestration module (15) for determining whether the single question is a single-answer type or a plurality of subtask type to determine whether to execute the subtask generation module (12), and determining whether the execution module (13) generates one or more sub-answers within a preset time to determine whether to re-execute the subtask generation module (12).
[0040] Accordingly, the method (20) of the present invention may include the steps of: an input module (11) receiving a single question from a user (21); an orchestration module (15) determining whether the final answer to the single question is a single answer or a plurality of subtasks, and proceeding with the step (22) of generating a final answer immediately if it is a single answer and generating a plurality of subtasks if it is a plurality of subtasks; a subtask generation module (12) generating a plurality of subtasks for the single question (23); an execution module (13) selecting and executing a generative sLLM that matches the plurality of subtasks to generate one or more sub-answers (24); a summary module (14) summarizing one or more sub-answers to generate a final answer (25); and an orchestration module (15) determining whether to proceed with the step of generating a plurality of subtasks again by determining the normal output of the final answer (26).
[0041] The step (21) in which the input module (11) receives a single question from the user may be a step in which the user inputs a question they wish to answer, and may be expressed in the form of a single sentence such as, for example, "What is the capital of South Korea?" or "What is the competitiveness of South Korea?", but is not limited thereto.
[0042] The step (22) in which the orchestration module (15) determines whether the final answer to a single question is a single answer or a multiple subtask type, and in the case of a single answer, immediately generates the final answer, and in the case of a multiple subtask type, generates multiple subtasks, may be a step in which the orchestration module (15), which is a language model capable of understanding the meaning and intent of a single question and answering, analyzes the single question and, if the answer is a single answer, the orchestration module (15) directly answers to finalize the final answer, or, if it is a multiple subtask type, proceeds to the next step to generate the final answer through the selection of a suitable generative language model, parallel work, and summarization.
[0043] In this case, the orchestration module (15) can determine the single question as a single answer type if the content of the answer is not classified into multiple domains by interpreting the meaning and / or intent of the single question, and can determine the single question as a multiple sub-task type if the content of the answer is classified into multiple domains by interpreting the meaning and / or intent of the single question.
[0044] To this end, the orchestration module (15) may be a language model that has been trained with a single question as a feature value and a single answer and multiple sub-task types as label values, and can determine whether a single question is a single answer or multiple sub-task types when any single question is input, but is not limited thereto.
[0045] For example, the orchestration module (15) can determine the question "What is the capital of South Korea?" as a single answer and generate a final answer "Seoul" to it to complete the procedure, and can determine the question "What is the competitiveness of South Korea?" as a multiple sub-task type and proceed to the next procedure.
[0046] The step (23) in which the subtask generation module (12) generates multiple subtasks for a single question may be a step in which the orchestration module (15) automatically generates multiple subtasks (prompts) according to the meaning and / or intent of the single question when the single question is determined to be of the type of multiple subtasks.
[0047] The subtask generation module (12) is a language model that has been trained with a single question as a feature value and multiple subtasks as label values, and can be a language model capable of generating multiple correlated subtasks when any single question is input, but is not limited thereto.
[0048] In this case, the feature values and label values for training the subtask generation module (12) can be generated from external data using the RAG (Retrieval Augmented Generation) technique. For example, the feature values and label values can be generated from search terms of an external search engine and related search lists for those search terms, and can be generated from social postings, blogs, and original articles of external websites.
[0049] For example, the subtask creation module (12) can receive a search term and a related search list for the search term in real time from an external search engine, process the search term to generate a feature value, and process the related search list (People also ask) to generate a label value.
[0050] Additionally, the sub-task generation module (12) can extract a main topic (e.g., the title of an article) and a sub-topic (e.g., the content of an article) from the original text of social postings, blogs, and articles on an external website, process the main topic to generate a feature value, and process the sub-topic to generate a label value. For example, when processing, a single question can be generated using a keyword extracted from the main topic, and multiple sub-tasks can be generated using a keyword extracted from the sub-topic and a keyword extracted from the main topic or a keyword highly related to the keyword extracted from the main topic, but are not limited thereto.
[0051] As described above, highly relevant data from the outside can be used as feature values and label values to train the subtask generation module (12). Meanwhile, depending on the design conditions, the meaning (content) of multiple subtasks may be set to contain the meaning (content) of a single question and have a subdivided meaning (content), but is not limited thereto.
[0052] For example, the subtask creation module (12) can extract keywords (search terms) "South Korea" and "competitiveness" from a single question "What is South Korea's competitiveness?" and create multiple subtasks as related search lists, such as "What is South Korea's manufacturing share?", "What is South Korea's K-POP?", and "What is South Korea's number of college graduates?".
[0053] The step (24) in which the execution module (13) selects and executes a generative language model that matches a plurality of subtasks to generate one or more sub-answers may be a step of selecting a generative language model specialized for the purpose of a plurality of subtasks and processing it in parallel in a distributed computing environment to generate one or more independent sub-answers for each of the plurality of subtasks.
[0054] In this case, the execution module (13) may generate one or more sub-answers by selecting and executing a generative language model that matches a plurality of sub-tasks according to predefined criteria, but is not limited thereto.
[0055] Additionally, the execution module (13) is a language model that determines the domain of multiple subtasks, and can generate one or more sub-answers by selecting and executing a generative language model trained on a corpus of the corresponding domain according to the domain of multiple subtasks, but is not limited thereto.
[0056] To this end, the execution module (13) may be a language model that has been trained with subtasks as feature values and domains as label values, and can determine the domain when any subtask is input, but is not limited thereto.
[0057] Meanwhile, the database of the execution module (13) has a plurality of specialized generative language models mapped to each domain, and a generative language model can be selected and matched according to the domain of each of the plurality of subtasks.
[0058] For example, the execution module (13) can determine the domain of the subtask "What is the proportion of manufacturing in South Korea?" as the economic field and select and execute a dedicated generative language model trained on an economic field corpus, determine the domain of the subtask "What is K-POP in South Korea?" as the cultural field and select and execute a dedicated generative language model trained on an cultural field corpus, and determine the domain of the subtask "What is the number of college graduates in South Korea?" as the education field and select and execute a dedicated generative language model trained on an education field corpus.
[0059] In this case, the execution module (13) executes selected language models for each of the multiple subtasks in parallel in a distributed computing environment, and as described above, the method (20) of the present invention can increase processing speed and reliability.
[0060] Meanwhile, the orchestration module (15) can supervise whether the execution module (13) answers at a given time according to a preset time, and if one or more sub-answers are generated at the preset time, it can transmit them to the summary module (14) to proceed to the next step, and if one or more sub-answers are not generated at the preset time, it can loop the step (23) of generating multiple sub-tasks from a single question again.
[0061] When the subtask generation module (12) performs the step (23) of generating multiple subtasks for a single question again, it may generate multiple subtasks that are at least partially different from the preceding step. For example, a technique for adjusting the weights or parameters of highly relevant data provided from the outside as feature values and label values to train the subtask generation module (12) may be used, but is not limited thereto.
[0062] The step (25) in which the summary module (14) summarizes one or more sub-answers to generate a final answer may be a step of summarizing to resolve conflicts or duplicates of one or more sub-answers and generating a final answer with increased reliability.
[0063] Accordingly, the summary module (14) is a language model that resolves conflicts or duplicates of one or more sub-answers and can generate a final answer with weights assigned according to the reliability and importance of one or more sub-answers.
[0064] For example, the summary module (14) can remove duplicate content from one or more sub-answers and generate a final answer by removing content with low reliability and importance from one or more sub-answers by providing feedback on evaluation results from past users.
[0065] The step (26) in which the orchestration module (15) determines whether to re-run the step of generating multiple subtasks by determining the normal output of the final answer may be a step in which the orchestration module (15) determines the content and contextual completeness of the final answer, and if the content of the final answer does not correspond to the answer to a single question or the contextual completeness is below a preset value, the step of generating subtasks may be re-run (Loop).
[0066] When the subtask generation module (12) performs the step (23) of generating multiple subtasks for a single question again, it may generate multiple subtasks that are at least partially different from the preceding step. For example, a technique for adjusting the weights or parameters of highly relevant data provided from the outside as feature values and label values to train the subtask generation module (12) may be used, but is not limited thereto.
[0067]
[0068] The method (20) of the present invention described above can be implemented as a program (or application) to be executed in combination with a server, which is hardware, and stored on a medium.
[0069] The aforementioned program may include code encoded in a computer language such as C, C++, JAVA, or machine language, which can be read by the computer's processor (CPU) through the computer's device interface, in order for the computer to read the program and execute the methods implemented in the program. Such code may include functional code related to functions that define the necessary functions for executing the methods, and may include control code related to execution procedures necessary for the computer's processor to execute the functions according to a predetermined procedure. Additionally, such code may further include memory reference code regarding where (address) additional information or media necessary for the computer's processor to execute the functions should be referenced in the computer's internal or external memory. In addition, if the processor of the computer needs to communicate with any other computer or server located remotely in order to execute the above functions, the code may further include communication-related code regarding how to communicate with any other computer or server located remotely using the communication module of the computer, and what information or media to transmit or receive during communication.
[0070] The above-mentioned storage medium refers to a medium that stores data semi-permanently and is readable by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, examples of the above-mentioned storage medium include, but are not limited to, ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. That is, the above-mentioned program may be stored on various recording media on various servers that the computer can access, or on various recording media on the user's computer. Additionally, the above-mentioned medium may be distributed across networked computer systems, and computer-readable code may be stored in a distributed manner.
[0071] The steps of the method or algorithm described in connection with embodiments of the present invention may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), Flash Memory, a hard disk, a removable disk, a CD-ROM, or any form of computer-readable recording medium well known in the art to which the present invention belongs.
[0072] Although embodiments of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention may be implemented in other specific forms without altering its technical concept or essential features. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive.
Claims
1. A step in which the input module receives a single question from the user; A step in which a subtask generation module generates multiple subtasks for a single question; A step in which an execution module selects and executes a generative language model that matches a plurality of subtasks to generate one or more sub-answers; and A method comprising a step in which a summary module summarizes one or more sub-answers to generate a final answer.
2. In Paragraph 1, The above subtask generation module is a language model trained with a single question as a feature value and multiple subtasks as label values, and generates multiple subtasks when an arbitrary single question is input, and A method characterized in that feature values and label values are generated from a search term of an external search engine and an associated search list for that search term.
3. In Paragraph 1, A method characterized by the above execution module selecting and executing a generative language model that matches a plurality of subtasks according to predefined criteria to generate one or more sub-answers.
4. In Paragraph 1, The above execution module is a language model that determines the domains of a plurality of subtasks, and is characterized by a method that selects and executes a generative language model trained on a corpus of the corresponding domain according to the domains of the plurality of subtasks to generate one or more sub-answers.
5. In Paragraph 3 or 4, A method characterized by the above execution module executing a selected language model for each of a plurality of subtasks in a distributed computing environment.
6. In Paragraph 3 or 4, A method characterized by the above execution module executing selected language models in parallel for each of a plurality of subtasks.
7. In Paragraph 1, The above summary module is a language model that resolves conflicts or duplicates of one or more sub-answers, and is characterized by generating a final answer that assigns weights according to the reliability and importance of one or more sub-answers.
8. In Paragraph 1, A method characterized by further including a step in which an orchestration module determines whether the final answer to a single question is a single answer or a multiple subtask type, and if it is a single answer, immediately generates the final answer, and if it is a multiple subtask type, generates multiple subtasks.
9. In Paragraph 8, A method characterized by further including a step in which the orchestration module determines the normal output of the final answer and determines whether to re-proceed with the step of generating a plurality of subtasks.
10. Input module for receiving a single question from the user; Subtask generation module that generates multiple subtasks for a single question; An execution module that selects and executes a generative language model matching multiple subtasks to generate one or more sub-answers; A summary module that generates a final answer by summarizing one or more sub-answers; and A system comprising an orchestration module that determines whether a single question is a short-answer type or a multiple-subtask type to determine whether to execute the subtask generation module, and determines whether to re-execute the subtask generation module by determining the normal output of the final answer.