A task execution method and system based on the context of purpose-oriented dialogue.
Patent Information
- Application Number
- JP2024212170
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-11-29
- Filing Date
- 2024-12-05
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2044-12-05
AI Technical Summary
【0037】 本開示の多様な実施形態によって、対話グラフと対話モデルを通じた目的指向対話のデータを分析して該当対話の文脈を決定し、決定された対話の文脈に基づいてユーザが要求するタスクを決定し、複数のタスク実行モデルうち決定されたタスクに最適化したモデルを利用することによりユーザ要求タスクを正確に実行できる対話モデルを通じた目的指向対話の文脈に基づくタスク実行方法を提供することができる。
Smart Images

Figure 0007927048000005 
Figure 0007927048000006 
Figure 0007927048000007
Abstract
Description
Technical Field
[0001] The present disclosure relates to a task execution method and system based on the context of goal-oriented dialogue, and more particularly, to a task execution method and system for performing a user-requested task determined based on the context of goal-oriented dialogue through a dialogue graph and a dialogue model.
Background Art
[0002] Recently, with the development of artificial intelligence (AI) technology, various services utilizing artificial intelligence have been commercialized in various industrial fields. In such artificial intelligence technology, an artificial neural network model can learn a huge amount of data, output and provide information desired by a user.
[0003] On the other hand, interactive artificial intelligence models that can interact with users are known, which go beyond simply allowing an artificial intelligence model to learn data patterns and provide output values for input values. Artificial intelligence secretaries utilizing such interactive artificial intelligence models can be applied to smart devices to serve as personal secretaries, can be applied to chatbots that answer questions from enterprise customers, and can be applied to smart home systems that can control the operation of home appliances connected to the Internet of Things according to user requirements. Thereby, the interactive artificial intelligence model can provide users with more convenient and user-friendly experiences in various technical fields.
[0004] In addition, active research is being conducted on artificial intelligence technology that can execute tasks requested by users, and for this purpose, many attempts have been made to apply analysis methods based on goal-oriented (Task-oriented) dialogue graphs to interactive artificial intelligence models.
[0005] However, most research related to task-oriented dialogue graphs is limited to methods that attempt to generate dialogue models using dialogue graphs drawn directly by humans or utilize rule-based systems that automatically infer dialogue graphs based on already learned dialogue policies. This has limitations, as it prevents the provision of dialogue graph-based dialogue models that are automatically constructed based solely on dialogue datasets without human intervention, and it cannot model diverse dialogue flows.
[0006] On the other hand, artificial intelligence is generally realized through a large number of AI models and deep learning based on them.
[0007] Such artificial intelligence is being developed to provide diverse services by taking into account the user's context (e.g., context, environment, and / or intent).
[0008] However, when attempting to process a specific task based on large amounts of data, there is a limitation in that the computational costs and time required are considerable.
[0009] As a result, there are certain limitations on the use of AI models in on-device environments, which have recently been attracting attention.
[0010] To solve this problem, traditionally, model architectures such as MoE (Mixture of Experts) have been utilized.
[0011] Here, MoE refers to a machine learning model architecture that solves complex problems by combining numerous expert models.
[0012] Such a MoE may include expert models, which are multiple small networks designed to learn different parts and / or different characteristics of a given set of data and perform data processing operations accordingly, and a gating network that evaluates the performance of each expert model and, based on this, determines which expert model is best suited to a particular task given the data.
[0013] Therefore, according to the MoE architecture, a gating network that acquires predetermined input data determines a probabilistic or deterministic work assignment for each expert model, and the selected expert models perform their respective tasks and return the results, thereby performing data processing for a specific task.
[0014] By utilizing such MoE (Moment of Engineering), AI models can improve overall efficiency and performance by activating only specific parts and concentrating computing resources when dealing with complex tasks or large datasets.
[0015] However, with conventional MoE systems, not only is a high level of VRAM required, but there are also a considerable number of issues that must be resolved during the fine-tuning process.
[0016] In addition, the traditional MoE method is designed to efficiently manage large-sized models, but it has limitations in supporting the efficiency of remaining resources that were not activated in response to the given task.
[0017] Furthermore, while conventional AI technologies in this field utilize mostly general-purpose AI models to provide services, they face the problem of difficulty in quickly and easily ensuring AI analysis performance that is best suited to a given context. [Overview of the Initiative] [Problems that the invention aims to solve]
[0018] Through various embodiments of this disclosure, we aim to provide a task execution method based on the context of a purpose-oriented dialogue via a dialogue model that can accurately execute user-requested tasks by determining the context of a purpose-oriented dialogue through a dialogue graph and a dialogue model, and determining the tasks requested by the user based on the determined dialogue context.
[0019] However, the technical challenges that the various embodiments of this disclosure aim to address are not limited to those described above, and other technical challenges may exist. [Means for solving the problem]
[0020] One embodiment provides a task execution method based on the context of a purpose-oriented dialogue through an interaction model by a computing system including memory and a processor, the method comprising the steps of: receiving user dialogue input; determining and providing a response dialogue action for the user dialogue input based on a dialogue graph; determining the context of the purpose-oriented dialogue by analyzing data of a series of purpose-oriented dialogues including the user dialogue input and the response dialogue action; determining the type of task requested by the user based on the context of the purpose-oriented dialogue; and performing the task of which the type has been determined.
[0021] In other aspects, in the step of determining the context of the purpose-oriented dialogue, a plurality of keywords may be extracted from the data of the purpose-oriented dialogue, and the context may be determined by analyzing at least one of the correlations between the plurality of dialogue acts included in the purpose-oriented dialogue, the intention of the purpose-oriented dialogue, or the purpose of the purpose-oriented dialogue based on the plurality of keywords.
[0022] In other respects, the task execution method may include a step of determining at least one task execution model from among a plurality of task execution models that is optimized for the task, whose type is determined based on the context of the goal-oriented dialogue, and in the step of performing the task, the task may be performed using the determined at least one task execution model.
[0023] In another aspect, in the step of determining the type of said task and the step of executing said task, said computing system may determine the type of said task and support predetermined calculations required for executing said task.
[0024] In another aspect, in the step of executing said task, analysis may be performed on programming code associated with said task, programming code for executing said task may be generated based on a result of said analysis, and said task may be executed by executing said programming code.
[0025] In another aspect, said task execution method further comprises the step of, by receiving user interaction input, capturing a screen of an electronic device used by a user to obtain a user screen screenshot, and in the step of determining the type of said task, the type of said task may be determined based on information about the determined context of said goal-oriented interaction and analysis of said user screen screenshot.
[0026] In another aspect, the step of determining and providing said responsive interaction act may comprise: generating an interaction graph modeling at least one conditional relationship for an interaction dataset; sampling a plurality of interaction act groups for responding to said user interaction input using a pre-trained interaction model; adjusting said plurality of interaction act groups based on said interaction graph; and selecting any one interaction act group that satisfies a predetermined condition from among said plurality of interaction act groups.
[0027] In another aspect, said at least one conditional relationship may comprise at least any one of: a first conditional relationship relating to what kind of utterance should be made for one utterance in an interaction flow, a second conditional relationship relating to what kind of utterance can be made for one utterance, and a third conditional relationship relating to what kind of utterance must not be made for one utterance.
[0028] In another aspect, in the step of selecting any one of the dialogue act groups, any one dialogue act group that most satisfies said at least one conditional relationship among said plurality of dialogue act groups may be selected.
[0029] In another aspect, each of said plurality of dialogue act groups may include at least one dialogue act for said user dialogue input.
[0030] In another aspect, in the step of adjusting said plurality of dialogue act groups, for each of said plurality of dialogue act groups, a dialogue act that satisfies said first conditional relationship among said at least one dialogue act may be added, a dialogue act that does not satisfy the second conditional relationship may be removed, and a dialogue act that does not satisfy the third conditional relationship may be removed.
[0031] In another aspect, in the step of selecting any one of the dialogue act groups, a dialogue act group that includes the largest number of dialogue acts satisfying said at least one conditional relationship among said plurality of dialogue act groups may be selected.
[0032] In another aspect, said goal-oriented dialogue method may further comprise a response dialogue act determining step of providing any one dialogue act included in any one of the selected dialogue act groups as a response to said user dialogue input.
[0033] In another aspect, in said response dialogue act determining step, a dialogue act having the highest relevance to said user dialogue input among at least one dialogue act included in any one of the already selected dialogue act groups may be provided as a response to said user dialogue input.
[0034] One embodiment includes at least one memory and at least one processor that reads at least one instruction word stored in the memory and performs a task execution method based on the context of a goal-oriented dialogue, wherein the at least one processor may receive a user dialogue input, determine and provide a response dialogue action to the user dialogue input based on a dialogue graph, analyze data of a series of goal-oriented dialogues including the user dialogue input and the response dialogue action to determine the context of the goal-oriented dialogue, determine the type of task requested by the user based on the context of the goal-oriented dialogue, and perform the determined task.
[0035] One embodiment provides a task execution system based on the context of a purpose-oriented dialogue, comprising an electronic device for receiving user dialogue input, and a computing device including at least one memory, and at least one processor for reading at least one instruction word stored in the at least one memory and performing a task execution method based on the context of a purpose-oriented dialogue, wherein the at least one processor receives user dialogue input, determines and provides a response dialogue action to the user dialogue input based on a dialogue graph, analyzes data of a series of purpose-oriented dialogues including the user dialogue input and the response dialogue action to determine the context of the purpose-oriented dialogue, determines the type of task requested by the user based on the context of the purpose-oriented dialogue, and performs the determined task.
[0036] In other aspects, the electronic device may receive the user interaction input in the form of at least one of text, voice, gesture, and touch. [Effects of the Invention]
[0037] Through various embodiments of this disclosure, it is possible to provide a task execution method based on the context of a purpose-oriented dialogue via a dialogue model, which analyzes data from a purpose-oriented dialogue through a dialogue graph and a dialogue model to determine the context of the dialogue, determines the task requested by the user based on the determined dialogue context, and accurately executes the user-requested task by utilizing the model optimized for the determined task from among multiple task execution models.
[0038] However, the effects that can be obtained by the various embodiments of this disclosure are not limited to those mentioned above, and other effects not mentioned can be clearly understood from the following description. [Brief explanation of the drawing]
[0039] [Figure 1] This shows an example block diagram of a computing system that implements a purpose-oriented dialogue service and a task execution service according to one embodiment. [Figure 2] This is a conceptual diagram illustrating how a computing system according to one embodiment performs user-requested tasks based on a purpose-oriented dialogue service. [Figure 3] This figure shows a simplified structure of a neuromorphic circuit that may be included in a processor according to one embodiment. [Figure 4] This is a block diagram of a computing device that implements a purpose-oriented dialogue service and a task execution service according to one embodiment. [Figure 5] This is a block diagram of a computing device that implements purpose-oriented dialogue services and task execution services according to another embodiment. [Figure 6] This is a block diagram of a computing device that realizes a purpose-oriented dialogue service according to another embodiment. [Figure 7] This is a conceptual diagram illustrating the process by which a purpose-oriented dialogue service is performed according to one embodiment. [Figure 8] This shows an internal block diagram of an AI agent model according to one embodiment. [Figure 9] This is a flowchart of a purpose-oriented dialogue method according to one embodiment. [Figure 10] This is a conceptual diagram illustrating each step of a purpose-oriented dialogue method according to one embodiment. [Figure 11] This table illustrates the performance evaluation results of a purpose-oriented dialogue method according to one embodiment. [Figure 12] This is a flowchart of a task execution method based on the context of goal-oriented dialogue through a dialogue model according to one embodiment. [Figure 13] This is a flowchart illustrating a MoE-based model identification method according to one embodiment. [Figure 14] This is a conceptual diagram illustrating a MoE-based model identification method according to one embodiment. [Figure 15] This figure illustrates an example of specialized model characteristic information according to one embodiment. [Figure 16] This is a flowchart illustrating a method for providing an AI agent based on MoE-applied LLM according to one embodiment. [Figure 17] This is a conceptual diagram illustrating a method for providing an AI agent based on MoE-applied LLM according to one embodiment. [Modes for carrying out the invention]
[0040] The present invention can be modified in various ways and has many different embodiments; therefore, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. The effects and features of the present invention, and how to achieve them, will become clear when you refer to the embodiments described in detail below, along with the drawings. However, the present invention is not limited to the embodiments disclosed below and can be realized in many different forms. In the following embodiments, terms such as "first," "second," etc., are not restrictive and are used to distinguish one component from another. Also, singular expressions include plural expressions unless they have a clearly different meaning in context. Also, terms such as "includes" or "has" mean that the features or components described in the specification exist, and do not preclude the possibility that one or more other features or components may be added. Also, in the drawings, the sizes of components may be exaggerated or reduced for the sake of illustration. For example, the size and thickness of each component in the drawings are shown arbitrarily for the sake of illustration, so the present invention is not necessarily limited to what is shown.
[0041] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components will be given the same reference numerals, and redundant explanations will be omitted.
[0042] - System 1000 that provides purpose-oriented dialogue services
[0043] In one embodiment, system 1000 generates a task-oriented dialogue graph based on an analysis of a dialogue dataset, and selects a final response dialogue action by verifying and adjusting various response dialogue actions sampled by a pre-trained dialogue model based on the generated dialogue graph, and provides this as a response to the user dialogue input.
[0044] In this case, the system 1000 can provide the user with more reliable response dialogue behavior by modeling predetermined conditional relationships for the dialogue dataset to generate a dialogue graph, and by verifying and adjusting the sampled response dialogue behavior based on the predetermined conditional relationships included in the dialogue graph.
[0045] Furthermore, the system 1000 can provide the user with a more convenient task processing experience by determining the type of task requested by the user based on the user dialogue input and the ultimately generated response dialogue action, and by efficiently executing the task based on the determined task type.
[0046] Figure 1 shows an example block diagram of a computing system 1000 that implements a purpose-oriented dialogue service and a task execution service according to one embodiment.
[0047] As shown in Figure 1, a computing system 1000 that realizes a purpose-oriented dialogue service and a task execution service according to one embodiment includes a user computing device 110, a server computing system 130, and a training computing system 150, and the devices are able to communicate via a network 170.
[0048] A purpose-oriented dialogue method according to one embodiment may be implemented and provided locally by a user computing device 110, 2) implemented and provided in the form of a web service by a server computing system 130 that communicates with the user computing device 110, or 3) implemented and provided by the user computing device 110 and the server computing system 130 working together.
[0049] In this embodiment, the user computing device 110 and / or the server computing system 130 can train the machine learning models 120 and / or 140 through interaction with a training computing system 150 which is communicatively connected via a network 170. The training computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130.
[0050] Furthermore, the artificial intelligence model can be trained in three ways: 1) by the user computing device 110 directly locally; 2) by the server computing system 130 and the user computing device 110 interacting with each other via the network 170; and 3) by a separate training computing system 150 using various training and learning techniques. The training computing system 150 can also transmit and provide / update the trained artificial intelligence model to the user computing device 110 and / or the server computing system 130 via the network 170.
[0051] In some embodiments, the training computing system 150 may be part of the server computing system 130 or part of the user computing device 110.
[0052] Figure 2 is a conceptual diagram illustrating how a computing system 1000 according to one embodiment performs tasks requested by a user based on a purpose-oriented dialogue service.
[0053] As shown in Figure 2, the computing system 1000 can receive various forms of user interaction input, provide appropriate response interaction actions to the received user interaction input, and further perform tasks requested by the user based on the user interaction input and response interaction actions. Here, tasks may include various types of tasks determined based on goal-oriented dialogue, including user interaction input and response interaction actions.
[0054] The computing system 1000 grasps the context of a series of goal-oriented interactions consisting of user interaction inputs and corresponding response interactions.
[0055] The computing system 1000 can analyze patterns of purpose-oriented dialogue, which consist of various types of user dialogue inputs and corresponding dialogue responses, and based on this, can determine the context related to the intent, purpose, etc., of the dialogue in question.
[0056] The computing system 1000 can extract multiple keywords contained in the data of a purpose-oriented dialogue, and based on the extracted keywords, it analyzes the correlation between multiple dialogue actions contained in the dialogue, the intention and purpose of the dialogue, and ultimately determines the context of the purpose-oriented dialogue.
[0057] Furthermore, the computing system 1000 can determine the type of task requested by the user based on information regarding the context of the goal-oriented dialogue.
[0058] For example, multiple keywords (e.g., hotel, reservation, date, number of people, 5 stars, etc.) can be extracted from data in a series of dialogues that include user dialogue input requesting a hotel reservation and response dialogue actions requesting information about the hotel reservation (e.g., reservation date, number of guests, hotel grade, etc.).
[0059] Furthermore, based on the extracted keywords, the context of the dialogue can be determined by analyzing at least one of the following: the correlation between the user dialogue actions included in the dialogue and the response dialogue actions provided by the computing system 1000, the intent of the dialogue, and its purpose.
[0060] For example, following a keyword-based context determination process, the context of a conversation that includes multiple dialogue actions such as "Please book a hotel," "When are your check-in dates?", "Check-in on December 31, 2024, check-out on January 3, 2025," "How many people are staying?", "3 people" will be used to determine that the user is requesting a hotel reservation and therefore requires an automated hotel reservation service.
[0061] Furthermore, the type of task requested by the user can be determined based on the context of the determined dialogue.
[0062] For example, based on the context of the conversation in which it was determined that an automated hotel reservation service is needed in response to the user's hotel reservation request, the task requested by the user is determined to be "hotel reservation".
[0063] On the other hand, the number of task types determined based on the context of the goal-oriented dialogue may be multiple. The goal-oriented dialogue may include various types of user dialogue inputs and various response dialogue actions thereto, and the context of such a goal-oriented dialogue may be associated with various tasks. Thus, multiple task types can be determined based on the context of the goal-oriented dialogue according to one embodiment.
[0064] System 1000 executes a task whose type is determined based on a goal-oriented dialogue that includes user dialogue input and response dialogue actions, and provides the user with the result of that execution.
[0065] For example, if the task type is determined to be "hotel reservation," the system 1000 can complete a hotel reservation task that meets the user's request based on data related to a series of goal-oriented interactions.
[0066] In this case, system 1000 determines the type of task and supports predetermined calculations necessary to perform the task of the determined type.
[0067] For example, the operating system of system 1000 can control processors 111 and 131 to directly determine the type of task and to perform predetermined operations necessary to carry out the task of the determined type.
[0068] Furthermore, for example, the operating system of system 1000 can determine the type of task and provide a predetermined API (Application progRAMMing interface) and / or SDK (Software development kit) to support predetermined operations necessary to perform the determined task.
[0069] System 1000 generates the necessary programming code to perform a task whose type is determined based on the context of a goal-oriented dialogue, and then executes this code to perform the task. Here, the programming code may be code written using a programming language (e.g., Java, Python, JavaScript, etc.) to perform a specific task.
[0070] On the other hand, the system 1000 can capture the screen of an electronic device used by the user and obtain a user screen screenshot by receiving user interaction input.
[0071] Thereafter, when determining the type of task, system 1000 can determine the type of task based on information regarding the context of the determined goal-oriented dialogue and an analysis of the user screen screenshot.
[0072] For example, if the context of the dialogue is determined to be "hotel reservation" and the user screen screenshot includes a screen of the homepage that provides the hotel reservation service, the system 1000 can determine the task type to be "hotel reservation via the relevant hotel reservation service homepage".
[0073] In this case, the system 1000 can automatically perform a series of actions necessary to carry out the task determined on the user screen (e.g., cursor movement, clicking, text input, etc.).
[0074] However, the task may be determined based on purpose-oriented dialogue, depending on the content of user dialogue input and response dialogue actions, into countless types. For example, the task may be determined into various types such as purchasing a product, sending an email, retrieving information, or creating a document.
[0075] Furthermore, if the type of task is determined based on the context of a goal-oriented dialogue according to one embodiment, the system 1000 determines at least one task execution model from among a plurality of task execution models that is optimized for the task whose type was determined based on the context of the goal-oriented dialogue, and performs the task using the determined at least one task execution model.
[0076] In this case, the method by which system 1000 determines at least one task execution model optimized for the task and uses it to perform the task is substantially the same as the "MoE-based model identification method" described later, and therefore, an explanation of this will be omitted.
[0077] Furthermore, for example, system 1000 may include a first type first user computing device 180, a second type second user computing device 181, a third type third user computing device 182, and a fourth type fourth user computing device 183, all capable of receiving various forms of user interaction input from the user.
[0078] Here, user interaction input may take the form of at least one of text, voice, gesture, and touch. However, it is not limited to these forms, and user interaction input can take a variety of forms other than those exemplified above.
[0079] The first user computing device 180 of the first type may be a virtual reality electronic device, the second user computing device 181 of the second type may be a mobile electronic device, the third user computing device 182 of the third type may be an augmented reality electronic device, and the fourth user computing device 183 of the fourth type may be a desktop computer.
[0080] However, the system 1000 may include a variety of user computing devices other than those exemplified above that can receive user interaction input, without being limited thereto.
[0081] - User Computing Device (110: User Computing Device)
[0082] The user computing device 110 may include all other types of computing devices, such as smartphones, mobile phones, digital broadcasting devices, PDAs (personal digital assistants), PMPs (portable multimedia players), desktops, wearable devices, embedded computing devices, tablet PCs, augmented reality (VR) devices, and / or virtual reality (AR) devices.
[0083] Such a user computing device 110 may include at least one processor 111 and memory 112. Here, the processor 111 may consist of at least one or more electrically connected processors from among a central processing unit (CPU), graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (progRAMMable logic devices), FPGAs (field progRAMMable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.
[0084] For example, ASICs may have an array-like neuromorphic circuit structure that includes multiple neuronal circuits.
[0085] As shown in Figure 3, for example, the neuromorphic circuit 300 may include a plurality of presynaptic neuron circuits 310, a plurality of presynaptic lines 311 extending laterally from the plurality of presynaptic neuron circuits 310, a plurality of postsynaptic neuron circuits 320, a plurality of postsynaptic lines 321 extending vertically from the plurality of postsynaptic neuron circuits 320, and synaptic circuits 330 provided at the intersections of the plurality of presynaptic lines 311 and the plurality of postsynaptic lines 321.
[0086] Multiple presynaptic neuron circuits 310 can transmit signals input from an external source to multiple synaptic circuits 330 in the form of electrical signals via multiple presynaptic lines 311.
[0087] Furthermore, multiple post-synaptic neuron circuits 320 can receive electrical signals from multiple synaptic circuits 330 via multiple post-synaptic lines 321.
[0088] Furthermore, multiple post-synaptic neuron circuits 320 can also transmit electrical signals to multiple synaptic circuits 330 via multiple post-synaptic lines 321.
[0089] Multiple synaptic circuits 330 store weight values included in the layers that constitute the neural network system realized by the neuromorphic circuit 300, and can perform predetermined calculations based on the weight values and input data.
[0090] For example, each of the multiple synaptic circuits 330 may include a resistive memory cell having a variable resistance. In this case, the resistance of the multiple synaptic circuits 330 changes with the voltage applied through the multiple presynaptic neuron circuits 310 or the multiple postsynaptic neuron circuits 320, and weighted value data resulting from such resistance changes can be stored.
[0091] The neuromorphic circuit 300 is formed by mimicking the neuron and synapse structure, which are essential elements of the human brain. When a deep neural network (DNN) is implemented using the neuromorphic circuit 300, data processing speed can be improved and power consumption reduced compared to using existing von Neumann structures.
[0092] The memory 112 may include one or more non-temporary / temporary computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof, and may also include web storage of a server that performs memory storage functions over the internet. Such memory 112 can store data 113 and instruction words 114 necessary for at least one of the processors 111 to train an artificial intelligence model or to perform functional operations such as performing vision checks through the artificial intelligence model.
[0093] In one embodiment, the user computing device 110 can store at least one or more machine learning models 120.
[0094] For example, the machine learning model 120 may be a variety of machine learning models, such as multiple neural networks (e.g., deep neural networks) for performing goal-oriented dialogue and task execution methods, or other types of machine learning models including nonlinear and / or linear models, or a combination thereof.
[0095] For example, machine learning models can store linear regression, decision trees, random forests, gradient boosting, pre-trained language models, and / or deep learning models. And neural networks may include at least one of the following: feed-forward neural networks, recurrent neural networks (e.g., long-short-term memory recurrent neural networks), convolutional neural networks, and / or other forms of neural networks.
[0096] Furthermore, in various embodiments, the user computing device 110 can also store models and prompt templates that form the basis of input to the models, which are used in each process to execute at least a portion of the processes performed for purpose-oriented interaction methods and task execution methods through a large-scale language model (LLM).
[0097] In one embodiment, the user computing device 110 receives at least one or more machine learning models 120 from the server computing system 130 via the network 170, stores them in the memory 112, and then executes the stored machine learning models 120 using the processor 111 to perform conversational dataset analysis and the like.
[0098] In another embodiment, the server computing system 130 includes at least one machine learning model 140 and performs operations through the machine learning model 140, and can work in conjunction with the user computing device 110 to provide purpose-oriented dialogue services and task execution services to the user by communicating with the user computing device 110 and related data.
[0099] For example, the user computing device 110 can perform a purpose-oriented dialogue service in which the server computing system 130 provides output in response to user input via the web using a machine learning model 140.
[0100] Furthermore, the artificial intelligence model can also be implemented in a manner in which at least a portion of the machine learning models 120 and / or 140 are executed on the user computing device 110, and the remainder is executed on the server computing system 130.
[0101] Furthermore, the user computing device 110 may include at least one input component 121 that senses user input. For example, the user input component 121 may include a touch sensor (e.g., a touchscreen and / or touchpad) that senses touch from the user's input medium (e.g., a finger or stylus), an image sensor that senses user motion input, a microphone that senses user voice input, buttons, a mouse and / or keyboard, etc. Also, if the user input component 121 receives input to an external controller (e.g., a mouse and / or keyboard) via an interface, it may include an interface and an external controller.
[0102] - Server Computing System (130: Server Computing System)
[0103] The server computing system 130 performs a series of processes to provide purpose-oriented conversational services.
[0104] Furthermore, the server computing system 130 may perform a series of processes to provide task execution services based on the context of goal-oriented dialogue.
[0105] In more detail, in this embodiment, the server computing system 130 can provide purpose-oriented dialogue services by exchanging data with an external device, such as a user computing device 110, that is necessary to drive purpose-oriented dialogue services and task execution service processes on the external device.
[0106] More specifically, in this embodiment, the server computing system 130 can provide an environment on the user computing device 110 in which applications for providing purpose-oriented dialogue services and task execution services can operate.
[0107] For this purpose, the server computing system 130 may include application programs, data, and / or instructions for the application to run, and can send and receive various data based thereon with the external device.
[0108] The server computing system 130 may include at least one processor 131 and memory 132. Here, the processor 131 may consist of at least one or more electrically connected processors from among a central processing unit (CPU), graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (progRAMMable logic devices), FPGAs (field progRAMMable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.
[0109] For example, ASICs may have an array-like neuromorphic circuit structure containing multiple neuronal circuits (see Figure 3).
[0110] The memory 132 may also include one or more non-temporary / temporary computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, and magnetic disks, or combinations thereof. Such a memory 132 can store data 133 and instruction words 134 necessary for the processor 131 to train an artificial intelligence model or to perform functional operations such as performing goal-oriented dialogue methods and task execution methods through the artificial intelligence model.
[0111] In one embodiment, the server computing system 130 may be implemented including at least one computing device. For example, the server computing system 130 may be implemented so that multiple computing devices operate in a sequential computing architecture, a parallel computing architecture, or a combination thereof. The server computing system 130 may also include multiple computing devices connected by a network 170.
[0112] Furthermore, the server computing system 130 can store at least one or more machine learning models 140. For example, the server computing system 130 may include neural networks and / or other multilayer nonlinear models as machine learning models 140. Exemplaryly, the neural networks may include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.
[0113] In one embodiment, the server computing system 130 may further include a data store computing system (hereinafter referred to as "data store") which is storage for continuously storing and managing the raw data that forms the basis of the purpose-oriented dialogue service.
[0114] Such data stores may include various forms of data storage, ranging from file systems to cloud storage. For example, a data store may include relational databases that use a structured query language (SQL) to define and manipulate data, NoSQL databases designed for flexibility and scalability to process unstructured and semi-structured data, data warehouses used for reporting and data analysis that centralize large volumes of data from multiple sources and optimize them for querying and analysis, data warehouses that store large amounts of raw data in basic forms such as structured data, semi-structured data, and unstructured data, and at least one database among local storage devices and NAS (Network Attached Storage) that store data in files in a format that can generally be accessed by computer operating systems.
[0115] - Training Computing System (150)
[0116] The training computing system 150 may include at least one processor 151 and memory 152. Here, the processor 151 may consist of at least one or more electrically connected processors from among central processing units (CPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), digital signal processors (DSSPs), digital signal processing devices (DSPDs), progRAMMable logic devices (PLDs), field progRAMMable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions.
[0117] For example, ASICs may have an array-like neuromorphic circuit structure containing multiple neuronal circuits (see Figure 3).
[0118] The memory 152 may also include one or more non-temporary / temporary computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, and magnetic disks, or combinations thereof. Such a memory 152 can store data 153 and instruction words 154 necessary for the processor 151 to perform tasks such as training an artificial intelligence model.
[0119] For example, the training computing system 150 may include a model trainer 160 that trains machine learning models 120 and / or 140 stored in the user computing device 110 and / or server computing system 130 using a variety of training or learning techniques, such as backward propagation of errors.
[0120] For example, such a model trainer 160 can perform backpropagation-based updates to one or more parameters of machine learning models 120 and / or 140 for goal-oriented dialogue services based on a defined loss function.
[0121] In some implementations, error back propagation may include truncated back propagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight reduction, dropout, and / or knowledge distillation) to improve the generalization capabilities of the machine learning models 120 and / or 140 being trained.
[0122] For example, a model trainer 160 can train machine learning models 120 and / or 140 based on a set of training data 161, where the training data 161 may include data of different forms, such as images, audio samples, and / or text.
[0123] Furthermore, the training data 161 may include, for example, various types of dialogue data. In this case, the dialogue data may be related to task-oriented dialogues in which a specific task is requested and a response is provided for the requested task.
[0124] Examples of usable image types may include video frames, LiDAR point clouds, X-ray images, computed tomography scans, superspectroscopic images, and / or other diverse forms of images.
[0125] Such training data 161 can be provided by the user computing device 110 and / or the server computing system 130. When the training computing device trains the machine learning models 120 and / or 140 on specific data from the user computing device 110, the machine learning models 120 and / or 140 can be characterized into personalized models.
[0126] Furthermore, the Model Trainer 160 includes computer logic that is used to provide the desired functionality.
[0127] Furthermore, the model trainer 160 can be implemented using hardware, firmware, and / or software that control a general-purpose processor. In one implementation, the model trainer 160 includes a program file stored in a storage device, which is loaded into memory 152 and can be executed by one or more processors 151. In another implementation, the model trainer 160 includes one or more sets of computer-executable data 153 and instruction words 154 stored in a computer-readable storage medium of the type of RAM hard disk or optical or magnetic media.
[0128] Network 170 includes, but is not limited to, 3GPP® (3rd Generation Partnership Project) networks, LTE (Long Term Evolution) networks, WiMAX (World Interoperability for Microwave Access) networks, the Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth® (Bluetooth) networks, satellite broadcasting networks, analog broadcasting networks, and / or DMB (Digital Multimedia Broadcasting) networks.
[0129] Generally, communication over network 170 may be conducted using any type of wired and / or wireless connection, via various communication protocols (e.g., TCP / IP, HTTP, SMTP, and / or FTP), encoding or formatting (e.g., HTML and / or XML), and / or protection schemes (e.g., VPN, secure HTTP, and / or SSL).
[0130] Figure 4 is a block diagram of a computing device 100 that implements a purpose-oriented dialogue service and a task execution service according to one embodiment.
[0131] As shown in Figure 4, the computing device 100 included in the user computing device 110, the server computing system 130, and the training computing system 150 contains a number of applications (e.g., application 1 to application N). Each application may include a machine learning library and one or more machine learning models. For example, applications may include image processing applications (e.g., detection, classification, and / or segmentation), text messaging applications, email applications, writing applications, virtual keyboard applications, browser applications, and / or chatbot applications.
[0132] In one embodiment, the computing device 100 may include a model trainer 160 for training an artificial intelligence model, and by storing and operating the trained artificial intelligence model, it can provide output data corresponding to predetermined input data (such as a dialogue dataset in one embodiment).
[0133] Each application of the computing device 100 can communicate with a number of other components of the computing device 100, such as at least one sensor, a context manager, a device state component, and / or additional components. In one embodiment, each application can communicate with each device component using an API (e.g., a public API). In one embodiment, the API used by each application may be specific to that application.
[0134] Figure 5 is a block diagram of a computing device 200 that implements purpose-oriented dialogue services and task execution services according to other embodiments.
[0135] As shown in Figure 5, the computing device 200 includes a number of applications (e.g., Application 1 to Application N). Each application can communicate with the central intelligence layer. For example, applications may include an image processing application, a text messaging application, an email application, a writing application, a virtual keyboard application, and / or a browser application. In one embodiment, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0136] The central intelligence layer may contain multiple machine learning models. For example, as shown in Figure 5, at least a portion of each machine learning model may be provided to each application and managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model to all applications. In some implementations, the central intelligence layer may be contained within the operating system of the computing device 200, or implemented separately.
[0137] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data storage for the computing device 200. As shown in Figure 5, the central device data layer can communicate with many other components of the computing device 200, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0138] The technologies described herein can refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information transmitted to or from such systems. The inherent flexibility of computer-based systems will be recognized as allowing for a wide range of possible configurations, combinations, and divisions of work and functionality between and from components. For example, the processes described herein can be implemented using a single device or component, or multiple devices or components operating in combination. Databases and applications may be implemented in a single system or in a distributed system across multiple systems. Distributed components can operate sequentially or in parallel.
[0139] Figure 6 is a block diagram illustrating the functions of a computing device 400 that realizes a purpose-oriented dialogue service according to one embodiment.
[0140] As shown in Figure 6, the computing device 400 may include a dialogue graph generation module 10, a dialogue action group sampling module 20, a dialogue action group adjustment module 30, a dialogue action group selection module 40, and a dialogue action selection module 50.
[0141] The dialogue graph generation module 10 can generate a dialogue graph based on a dialogue dataset received from an external source.
[0142] Dialogue datasets may include data related to conversations between diverse speakers in diverse environments. For example, a dialogue dataset may include data related to a series of diverse conversations, such as a conversation between a customer and a reservation agent when booking a hotel, a conversation between a buyer and a seller when purchasing goods, or a conversation between a tutor and a tuty discussing academic assignments.
[0143] A dialogue graph is a graphical representation of the relationships between various types of utterances included in a dialogue dataset. It may include multiple nodes corresponding to multiple dialogue acts that represent the function of a particular utterance, and multiple edges that show various information about the relationships between multiple dialogue acts.
[0144] Here, for example, the act of dialogue may include various types of speech functions such as questioning, informing, requesting, and confirming.
[0145] Furthermore, the relationships between multiple dialogue acts may include predetermined conditional relationships between two utterances. For example, a predetermined conditional relationship may include at least one of the following: a first conditional relationship (Should relationship) concerning what utterance should be made in response to an utterance; a second conditional relationship (Can relationship) concerning what utterance can be made in response to an utterance in the flow of dialogue; and a third conditional relationship (Should-not relationship) concerning what utterance should not be made in response to an utterance.
[0146] For example, as shown in Figure 7, the dialogue graph generation module 10 generates a task-oriented dialogue flow (TOD-Flow) graph that includes multiple nodes, including A, B, C, and D, and multiple edges corresponding to the relationships between those nodes (S10).
[0147] The dialogue graph generation module 10 can vectorize the dialogue dataset by embedding words and sentences included in the dialogue dataset. The dialogue graph generation module 10 can analyze the vectorized dialogue dataset to learn relevant information about the dialogue dataset, such as the intent of the dialogue utterances, dialogue actions, slots, values, contextual information, information about the relationships between utterances, and patterns of sequential dialogue flow.
[0148] For example, the dialogue graph generation module 10 can learn relevant information from a dialogue dataset using an artificial intelligence model based on at least one of the following: a transformer-based model, a recurrent neural network (RNN), or a longshort-term memory (LSTM).
[0149] Furthermore, the dialogue graph generation module 10 can generate a dialogue graph by modeling at least one conditional relationship for the dialogue dataset based on the relevant information of the learned dialogue dataset.
[0150] In this case, the dialogue graph generation module 10 models at least one relationship to the dialogue dataset by maximizing a loss function optimized for at least one conditional relationship.
[0151] For example, the dialogue graph generation module 10 generates a dialogue graph by utilizing a dialogue graph generation model that has been trained to maximize the expected value when the nth dialogue action satisfies the first condition relation (Should relation) in the dialogue context.
[0152] In more detail, the dialogue graph generation module 10 can generate a dialogue graph by utilizing a dialogue graph generation model that has been trained to maximize the expected value when the nth dialogue action satisfies the first condition relation (Should relation) in a situation where the first condition relation (Should relation) of the dialogue context must be satisfied.
[0153] For example, the dialogue graph generation module 10 can model the first conditional relationship (Should relationship) regarding what utterances should be made in response to a single utterance in the flow of a dialogue by maximizing the loss function defined in [Equation 1] below.
[0154]
number
[0155] JPEG0007927048000002.jpg48151
[0156] Furthermore, for example, the dialogue graph generation module 10 can generate a dialogue graph by utilizing a dialogue graph generation model that has been trained to maximize the expected value when the nth dialogue action in the dialogue context satisfies the second condition relation (Can relation) and at the same time does not satisfy the third condition relation (Should-not relation).
[0157] In detail, the dialogue graph generation module 10 generates a dialogue graph by utilizing a dialogue graph generation model that has been trained to maximize the sum of the first expected value when the nth dialogue action in the dialogue context satisfies the second condition relation (Can relation) and does not satisfy the third condition relation (Should-not relation) in a situation where the nth dialogue action does not occur and does not satisfy the second condition relation (Can relation) and does satisfy the third condition relation (Should-not relation). For example, the dialogue graph generation module 10 can model the second condition relation (Can relation) regarding what utterances can be made in response to an utterance in the flow of dialogue, and the third condition relation (Should-not) regarding what utterances should not be made in response to an utterance, by maximizing the loss function defined in [Equation 2] below.
[0158]
number
[0159] JPEG0007927048000004.jpg58151
[0160] The dialogue action group sampling module 20 can sample multiple dialogue action groups to respond to user dialogue input received from an external source.
[0161] For example, the dialogue action group sampling module 20 can sample multiple appropriate dialogue action groups as responses to user dialogue input received via an electronic device. As shown in Figure 7, it can predict the next dialogue action for a user dialogue input. For example, for a user dialogue input related to "hotel reservation," the dialogue action group sampling module 20 samples dialogue action groups that include various types of dialogue actions such as "Confirm Information" and "Book Hotel" as dialogue action groups that are responses following the user dialogue input (S20).
[0162] The dialogue behavior group sampling module 20 may include an artificial neural network structure that has already learned to extract features from various types of dialogue datasets and provide appropriate output data for input data. For example, the dialogue behavior group sampling module 20 may include a Transformer-based neural network architecture (e.g., GPT-3, GPT-4, BERT-based models, etc.).
[0163] The dialogue action group adjustment module 30 can adjust multiple sampled dialogue action groups based on the generated dialogue graph.
[0164] Each of the multiple dialogue action groups generated by the dialogue action group sampling module 20 may include at least one dialogue action related to user dialogue input.
[0165] The dialogue action group adjustment module 30 can adjust multiple dialogue action groups based on whether each of the multiple dialogue action groups generated by the dialogue action group sampling module 20 in response to user dialogue input satisfies at least one conditional relationship included in the dialogue graph.
[0166] For example, the dialogue action group adjustment module 30 can determine whether each of the multiple dialogue action groups satisfies a first conditional relationship (Should relationship) with respect to the user dialogue input, and can add a dialogue action that satisfies the first conditional relationship (Should relationship) with respect to the user dialogue input to a dialogue action group that does not satisfy the conditional relationship.
[0167] Furthermore, the dialogue action group adjustment module 30 determines whether each of the multiple dialogue action groups satisfies a second conditional relationship (Can relationship) and a third conditional relationship (not should not) with respect to the user dialogue input, and can remove at least one dialogue action from a dialogue action group that does not satisfy the second conditional relationship (Can relationship) and the third conditional relationship (not should not).
[0168] In this way, the multiple dialogue action groups sampled by the dialogue action group adjustment module 30 are adjusted to conform to the dialogue graph, thereby improving the reliability of the purpose-oriented dialogue service provided by the system 1000 and its control over the dialogue model (S20).
[0169] The dialogue action group selection module 40 can select one dialogue action group from among several coordinated dialogue action groups that meets predetermined conditions.
[0170] For example, the dialogue action group selection module 40 ranks multiple dialogue action groups in descending order of the number of dialogue action groups that satisfy at least one conditional relationship of the dialogue graph (S30).
[0171] For example, the first dialogue action group, the second dialogue action group, and the third dialogue action group can be coordinated by the dialogue action group coordination module 30, and after the coordination work is completed, the first dialogue action group will include three dialogue actions, the second dialogue action group will include two dialogue actions, and the third dialogue action group will include four dialogue actions.
[0172] In this case, the dialogue action group selection module 40 can select the third dialogue action group as the first priority, which contains the most dialogue actions after being adjusted based on at least one conditional relationship of the dialogue graph among the first to third dialogue action groups.
[0173] The dialogue action selection module 50 can select one of at least one dialogue action included in one dialogue action group selected from multiple dialogue action groups and provide a response output to the user dialogue input.
[0174] For example, the dialogue action selection module 50 can calculate the relevance of at least one dialogue action included in a selected group of dialogue actions to the user dialogue input as a probability. Thereafter, the dialogue action selection module 50 can select one dialogue action from the selected group of dialogue actions that has the highest probability of corresponding to the relevance to the user dialogue input and provide it to the user as a response.
[0175] - AI Agent Specialization Model (AIAM: Artificial Intelligence Agent Specialization Model)
[0176] In other respects, the computing system 1000 described above may include an AI agent specialization model (AIAM) according to one embodiment.
[0177] Here, the AI agent specialist model (AIAM) according to one embodiment is an AI agent model that applies the MoE (Mixture of Experts) architecture realized by one embodiment, and may be an artificial intelligence model that includes a data processing algorithm that can act autonomously in a specific environment, solve tasks, and achieve goals. Here, MoE refers to an architecture of a machine learning model that solves complex problems by combining a large number of expert models.
[0178] Such AI agent specialization models (AIAMs) may include data processing algorithms to realize cognitive abilities to collect and interpret data from a given environment, decision mechanisms to determine the optimal action based on the collected data, execution abilities to carry out the decided action, and learning abilities to improve the action through experience.
[0179] As an embodiment, an AI agent specialization model (AIAM) can acquire predetermined input data (e.g., text, audio, images, videos, and / or specific sensor-based sensing data) and provide output data (e.g., response data to specific questions and / or control signals using specific command words) by performing predetermined tasks based on the acquired input data.
[0180] Figure 8 shows an internal block diagram of an AI agent model according to one embodiment.
[0181] More specifically, as shown in Figure 8, an AI agent model according to one embodiment may include at least one router (RT: Router, Gating Network), an orchestrator (OCT: Orchestrator), a small-large language model (sLLM), a normal MoE model (NM: Normal MoE Model), an external model (EM: External Model), and / or a specialized model (SM: Specialized Model).
[0182] In Figure 8, we explain that the AI agent model includes the aforementioned components in order to prevent the characteristics of various embodiments from becoming unclear.
[0183] However, it is obvious to a person of ordinary skill in the art that, depending on the embodiment, other general-purpose components may be included in addition to those shown in Figure 8, or some of the components shown in Figure 8 may be omitted.
[0184] More specifically, in one embodiment, the router (RT: Router, Gating Network) may be an artificial intelligence module that performs tasks such as assigning tasks and / or traffic coordination to multiple models within the MoE architecture.
[0185] Specifically, a router (RT) can analyze the given input data and / or requested tasks to determine which model is best suited for processing that data.
[0186] In this embodiment, the router (RT) can determine the model optimized for the given data processing based on the performance, expertise, and / or prior experience of each model.
[0187] Furthermore, the router (RT) can support efficient data processing by distributing a given task to at least one or more models, taking system load into consideration.
[0188] Furthermore, the router (RT) can flexibly respond to changes in the real-time system and adjust the tasks assigned to a particular model.
[0189] In one embodiment, such a router (RT) may be an artificial intelligence module that selectively determines a model (hereinafter referred to as a domain-specific model) that performs data processing operations optimized for a given domain.
[0190] In other words, in this embodiment, the router (RT) may be an artificial intelligence module that selects a model (i.e., a domain-specific model) from among multiple models included in the AI agent specialization model (AIAM) that is determined to perform the most suitable data processing (such as deep learning in this embodiment) for a given domain.
[0191] For reference, the domain in this embodiment means the data, rules, terminology, problem definitions, and / or processes that a given AI system uses to perform a specific task.
[0192] In one embodiment, the router (RT) can perform data analysis based on predetermined input data characteristics (e.g., user input and / or specific sensing data) and / or requested tasks, and based on this, it can grasp the data processing characteristics optimized for the task, detect a predetermined model that realizes this, and determine a domain-specific model.
[0193] In other words, the router (RT) in this embodiment may be an artificial intelligence module that detects the model that can most effectively perform data processing for a given domain, assigns / distributes the corresponding task processing work, and manages it.
[0194] Here, the router (RT) according to the embodiment may include a router (RT) already learned by a predetermined algorithm disclosed, a router (RT) further learned by one embodiment, and / or a router (RT) newly learned by a new method.
[0195] On the other hand, an orchestrator (OCT) according to one embodiment may be an artificial intelligence module that comprehensively controls and manages the overall configuration of an AI agent specialization model (AIAM).
[0196] In a detailed embodiment, the orchestrator (OCT) assigns various tasks arising from the overall system to appropriate resources (such as a router (RT) and / or a predetermined model, in an embodiment).
[0197] Furthermore, the orchestrator (OCT) manages resources such as available models and hardware resources (e.g., CPU and / or GPU) to ensure efficient use.
[0198] Furthermore, the orchestrator (OCT) monitors the performance of the overall system and adjusts specific parameters as needed, or optimizes the network configuration, etc.
[0199] Furthermore, the orchestrator (OCT) manages the coordination between multiple routers (RTs) and / or models, and controls the data flow and processing processes.
[0200] In other words, in this embodiment, the orchestrator (OCT) controls and manages the entire AI agent specialization model (AIAM) system and acts as the main router (RT), controlling at least one router (RT).
[0201] In this embodiment, the orchestrator (OCT) and router (RT) can work closely together to support the efficient operation of the MoE system.
[0202] Specifically, Orchestration (OCT) acts as the administrator of the entire system, monitoring the performance of the routers (RTs) and adjusting the router strategies as needed.
[0203] On the other hand, the router (RT) can effectively achieve efficient system control by assigning data processing tasks in accordance with the instructions of the orchestrator (OCT) and / or its own algorithms.
[0204] In some embodiments, the aforementioned orchestrator (OCT) and / or router (RT) may also be a master model (P) capable of controlling and managing the rest of the overall system and / or AI agent model (i.e., sLLM, general MoE model (NM), external model (EM), and / or specialized model (SM), etc.).
[0205] On the other hand, sLLM (small Large Language Model) according to one embodiment is an artificial intelligence module implemented as a lightweight version of LLM (Large Language Model).
[0206] In other words, sLLM is an artificial intelligence module built to achieve performance similar to large models like LLM with fewer resources.
[0207] In some embodiments, such an sLLM may include a multi-specialty model (SM) and router (RT) coupling-based MoE model (MoELM in some embodiments) according to one embodiment. Furthermore, the sLLM may include a domain-specific specialty model-based MoE model (DMoE model in some embodiments) according to one embodiment.
[0208] Furthermore, a Normal MoE Model (NM) according to an embodiment of the present invention may mean a predetermined MoE model realized by the disclosed universal formula.
[0209] For example, a general MoE model (NM) may include Switch Transformers, Conditional Computation in Neural Networks, Sparse Mixture of Experts, and / or Megatron-LM.
[0210] Furthermore, an external model (EM) according to one embodiment may refer to a predetermined artificial intelligence model realized by the various algorithms disclosed.
[0211] For example, the external model (EM) may include ChatGPT, Gemini, and / or Llama.
[0212] In embodiments, such external models (EMs) can be selectively used as needed to support the processing of a given task.
[0213] Furthermore, a Specialized Model (SM) according to one embodiment may refer to an artificial intelligence model that has undergone optimized learning for a specific purpose, and which has been trained using training data and methods specialized for that purpose.
[0214] In other words, a specialized model (SM) can be an artificial intelligence model that has been trained using specialized training data and methods to achieve a predetermined objective.
[0215] In embodiments, such a specialized model (SM) may include a predetermined learned sLLM (including MoELM and / or DMoE models), a general MoE model (NM), and / or an external model (EM). Furthermore, the specialized model (SM) may include a specialized module model according to one embodiment disclosed in the "MoE-based model identification method" described later. A detailed explanation of this is provided in the "MoE-based model identification method."
[0216] In embodiments, the aforementioned sLLM, general MoE model (NM), external model (EM), and / or specialized model (SM) may be secondary models (S) capable of performing specific tasks through the control and management of the master model (P) of the AI agent model (i.e., an orchestrator (OCT) and / or router (RT), etc.).
[0217] -Goal-oriented dialogue method (S100)
[0218] The following describes in detail a purpose-oriented dialogue method (S100) that can provide more accurate responses to user dialogue input by extracting and learning the features of various types of dialogue datasets, generating a dialogue graph that models predetermined conditional relationships to the dialogue datasets based on this, and then selecting the optimal dialogue action from among multiple dialogue actions sampled by a pre-trained dialogue model based on the dialogue graph and providing it as a response to user dialogue input.
[0219] Dialogue datasets may include data related to dialogues between diverse speakers taking place in diverse environments. Therefore, dialogue datasets may contain diverse types of data depending on the nature of the dialogues between speakers.
[0220] For example, a first dialogue dataset and a second dialogue dataset related to different types of tasks may contain different types of data, and the data structure of the first dialogue graph generated based on the first dialogue dataset and the data structure of the second dialogue graph generated based on the second dialogue dataset may be different from each other.
[0221] A dialogue graph is a graphical representation of the relationships between various types of utterances included in a dialogue dataset. It can be a structured dataset containing data on multiple dialogue actions corresponding to the functions and intentions of multiple utterances, as well as data on the conditional relationships between those dialogue actions.
[0222] The objective-oriented dialogue method (S100) can select and provide the optimal dialogue action group as a response to user dialogue input from among multiple dialogue action groups sampled by the dialogue model based on the dialogue graph.
[0223] Furthermore, based on the user interaction input and the selected interaction action groups, a task requested by the user is determined, and the computing system 1000 according to one embodiment can perform the determined task and provide the user with a task execution service.
[0224] In the following, we will describe in detail a purpose-oriented dialogue method (S100) in which a computing system 1000 according to one embodiment provides an appropriate response to user dialogue input based on a dialogue graph that models at least one conditional relationship to a dialogue dataset, and the dialogue model performs a task requested by the user.
[0225] Figure 9 is a flowchart of a purpose-oriented dialogue method (S100) according to one embodiment.
[0226] As shown in Figure 9, a purpose-oriented dialogue method (S100) according to one embodiment may include the steps of: generating a dialogue graph that models at least one conditional relationship for a dialogue dataset (S101); receiving user dialogue input (S103); sampling a plurality of dialogue action groups to respond to the user dialogue input using an already learned dialogue model (S105); adjusting the plurality of dialogue action groups based on the dialogue graph (S107); and selecting any one of the plurality of dialogue action groups that satisfies predetermined conditions (S109).
[0227] In step (S101), the processors 111 and 131 of the system 1000 generate a dialogue graph based on various types of dialogue datasets.
[0228] For example, processors 111 and 131 generate a first dialogue graph that models at least one conditional relationship for a first type of dialogue dataset related to the first task. Processors 111 and 131 also generate a second dialogue graph that models at least one conditional relationship for a second type of dialogue dataset related to the first task and other second tasks.
[0229] Here, at least one conditional relation for a dialogue dataset may include at least one of the following: a first conditional relation (Should relation) concerning what utterances should be made for an utterance included in the dialogue dataset; a second conditional relation (Can) concerning what utterances can be made for an utterance; and a third conditional relation concerning what utterances should not be made for an utterance.
[0230] In step (S103), the processors 111 and 131 of the system 1000 receive user interaction input.
[0231] For example, processors 111 and 131 can receive user interaction input data received via user input component 121 by a user computing device 110, which can be implemented with various types of electronic devices.
[0232] In step (S105), the processors 111 and 131 of the system 1000 sample a group of dialogue actions to be provided as responses to user dialogue input, utilizing a dialogue model that has already been learned.
[0233] For example, as shown in Figure 10, processors 111 and 131 can sample multiple dialogue action groups a1, a2, ... using a dialogue model (π) already learned by the model trainer 160 of the training computing system 150. In this case, the number of dialogue action groups a1, a2, ... sampled by the dialogue model (π) may be several tens of times, but is not limited to this.
[0234] For example, the sampled first dialogue action group a1 may include four dialogue actions A, B, C, and F, and the second dialogue action group a2 may include three dialogue actions A, C, and G.
[0235] In step (S107), the processors 111 and 131 of the system 1000 coordinate a plurality of dialogue action groups sampled based on the dialogue graph.
[0236] Processors 111 and 131 can coordinate multiple dialogue action groups based on whether each of the multiple dialogue action groups satisfies at least one conditional relationship included in the dialogue graph in response to user dialogue input.
[0237] First, processors 111 and 131 determine one of several dialogue graphs generated for various dialogue datasets that corresponds to the type of user dialogue input.
[0238] Subsequently, processors 111 and 131 coordinate multiple dialogue action groups based on the determined dialogue graph.
[0239] As shown in Figure 10, the determined dialogue graph includes G as a dialogue action that satisfies the first conditional relationship (Should relationship) with respect to the user dialogue input, A, C, F, and G as dialogue actions that satisfy the second conditional relationship (Can relationship), and A as a dialogue action that satisfies the third conditional relationship (Should-not relationship).
[0240] Processors 111 and 131 remove B from the first dialogue action group a1 based on the determined dialogue graph so that the first dialogue action group a1 satisfies the second conditional relationship (Can relationship).
[0241] Furthermore, processors 111 and 131 add G to the first dialogue action group a1 based on the determined dialogue graph so that the first dialogue action group a1 satisfies the first condition relation (Should relation).
[0242] Furthermore, processors 111 and 131 remove A from the first dialogue action group a1 based on the determined dialogue graph so that the first dialogue action group a1 satisfies the third condition relation (should-not relation).
[0243] Similarly, processors 111 and 131 remove A from the second dialogue action group a2 based on the determined dialogue graph so that the second dialogue action group a2 satisfies the first to third conditional relationships.
[0244] In this case, after the coordination work for multiple dialogue action groups is completed, the first dialogue action group a1 includes three dialogue actions C, F, and G, and the second dialogue action group a2 includes two dialogue actions C and G.
[0245] By adjusting the sampled dialogue action groups to correspond to the dialogue graph, the reliability of the goal-oriented dialogue service provided by system 1000 and its control over the dialogue model can be improved.
[0246] In step (S109), the processors 111 and 131 of the system 1000 select one of the coordinated dialogue action groups that satisfies predetermined conditions and provide it as a response to the user dialogue input.
[0247] Processors 111 and 131 select the dialogue action group from among the coordinated dialogue action groups that contains the most dialogue actions that satisfy at least one conditional relationship of the dialogue graph.
[0248] For example, processors 111 and 131, after an adjustment process, can select from a first dialogue action group (a1) containing three dialogue actions C, F, and G, and a second dialogue action group (a2) containing two dialogue actions C and G, the first dialogue action group (a1) containing the most dialogue actions that satisfy at least one conditional relationship of the dialogue graph, and provide it as a response to user dialogue input.
[0249] Figure 11 shows the response prediction performance index (F-1 score) for user dialogue input of various dialogue models (FLAN-T5, GPT-turbo) to which a purpose-oriented dialogue method (S100) based on a dialogue graph according to one embodiment of this disclosure is applied to various dialogue datasets (SGD, MultiWOZ).
[0250] In this case, the response prediction performance index of the dialogue model changes by changing the method of selecting one of several dialogue action groups as a response, which are generated from various dialogue models (FLAN-T5, GPT-turbo) and adjusted based on the dialogue graph.
[0251] Furthermore, when selecting one of several dialogue action groups, the response prediction performance index of the dialogue model changes depending on whether or not adjustment work is performed based on at least one conditional relationship in the dialogue graph.
[0252] For example, processors 111 and 131 select the dialogue action group that the dialogue model has determined to have the highest response probability from among multiple dialogue action groups and provide it as a response to the user dialogue input (Greedy).
[0253] Furthermore, processors 111 and 131 either select the dialogue action group that contains the most dialogue actions satisfying at least one conditional relationship in the dialogue graph (Compliance), or select the dialogue action group that is least adjusted based on at least one conditional relationship and provide it as a response to the user dialogue input (Violation).
[0254] Furthermore, processors 111 and 131 select the dialogue behavior group that has been sampled the most from among multiple dialogue behavior groups sampled by the dialogue model and provide it as a response to the user dialogue input (Majority).
[0255] As shown in Figure 11, the dialogue model achieves the highest response prediction performance index when multiple dialogue action groups are adjusted based on all of the first conditional relationships (Should relationship), second conditional relationships (Can relationship), and third conditional relationships (Should-not relationship) of the dialogue graph, and the dialogue action group that contains the most dialogue actions satisfying at least one conditional relationship of the dialogue graph is selected (Compliance).
[0256] In this way, by appropriately adjusting multiple dialogue action groups based on the dialogue graph, and selecting one of the adjusted dialogue action groups that best satisfies the conditional relationships of the dialogue graph, it is possible to provide a more reliable response to the user dialogue input.
[0257] On the other hand, method (S100) may further include a response dialogue action determination step of providing one of the selected dialogue actions included in any one dialogue action group as a response to the user dialogue input.
[0258] For example, if the user dialogue input is text data converted from the spoken phrase "Please book a hotel," then the first dialogue action group a1 selected from among several coordinated dialogue action groups includes three dialogue actions: "How many people will be staying? (C)", "What date would you like to book? (F)", and "What grade of hotel would you prefer? (G)".
[0259] In the response dialogue action determination step, processors 111 and 131 select one of several dialogue actions (C, F, G) included in the first dialogue action group a1 and provide it as a response to the user dialogue input. In this case, processors 111 and 131 can calculate the degree of relevance of each of the several dialogue actions (C, F, G) included in the first dialogue action group a1 to the user dialogue input, and can select and provide the one dialogue action with the highest degree of relevance as a response to the user dialogue input.
[0260] Furthermore, processors 111 and 131 can determine the type of task requested by the user based on the user interaction input and the group of interaction actions ultimately selected as the response, and can perform the determined task.
[0261] For example, if the task type is determined to be "hotel reservation," the processors 111 and 131 can complete a hotel reservation that matches the user's request based on various information related to the hotel reservation determined during a series of conversations, including user dialogue input and selected dialogue action groups.
[0262] In this way, the system 1000 can provide the user with a more reliable response by adjusting the dialogue actions sampled by the dialogue model based on a dialogue graph in which predetermined conditional relationships are structured and modeled based on a dialogue dataset, and by providing the adjusted dialogue actions as a response to the user dialogue input.
[0263] Furthermore, the type of task requested by the user is accurately determined based on a series of dialogues, including dialogue actions that are adjusted by user dialogue input and a dialogue graph and determined as the final response under predetermined conditions. By having the system 1000 perform the task thus determined, a task execution service that satisfies the user can be provided.
[0264] Figure 12 is a flowchart of a task execution method (S200) based on the context of a goal-oriented dialogue through a dialogue model according to one embodiment.
[0265] As shown in Figure 12, a task execution method (S200) according to one embodiment may include the steps of: receiving user dialogue input (S201); determining and providing a response dialogue action to the user dialogue input based on a dialogue graph (S203); analyzing data from a series of purpose-oriented dialogues including the user dialogue input and the response dialogue action to determine the context of the purpose-oriented dialogue (S205); determining the type of task requested by the user based on the context of the purpose-oriented dialogue (S207); and performing the task of which the type has been determined (S209).
[0266] In step (S201), the processors 111 and 131 of the system 1000 receive user interaction input.
[0267] For example, processors 111 and 131 can receive user interaction input data received via user input component 121 by a user computing device 110, which can be implemented with various types of electronic devices.
[0268] In step (S203), the processors 111 and 131 of the system 1000 determine and provide a response dialogue action to the user dialogue input based on the dialogue graph related to the user dialogue input.
[0269] Step (S203) is substantially identical to the purpose-oriented dialogue method (S100) described with reference to Figures 9 and 10, and therefore no further explanation is provided.
[0270] In step (S205), the processors 111 and 131 of the system 1000 analyze data from a series of goal-oriented dialogues, including user dialogue inputs and response dialogue actions, to determine the context of the goal-oriented dialogue.
[0271] The processors 111 and 131 of system 1000 can grasp the context of a series of purpose-oriented dialogues consisting of user dialogue inputs and corresponding response dialogue actions.
[0272] For example, the processors 111 and 131 of system 1000 extract multiple keywords contained in the data of a purpose-oriented dialogue, and based on the extracted keywords, analyze the correlation between multiple dialogue actions contained in the dialogue, the intention and purpose of the dialogue, etc., to ultimately determine the context of the purpose-oriented dialogue.
[0273] In step (S207), the processors 111 and 131 of the system 1000 determine the type of task requested by the user based on the context of the goal-oriented dialogue.
[0274] The processors 111 and 131 of system 1000 determine the type of task corresponding to the context of the purpose-oriented dialogue, which is determined based on data related to the task corresponding to the context of the purpose-oriented dialogue.
[0275] In this case, data related to the task corresponding to the context of the goal-oriented dialogue may already be stored in the memory 112, 132 of the system 1000, or may have already been learned by the machine learning models 120, 140.
[0276] For example, if the processors 111 and 131 of system 1000 determine that an automated hotel reservation service is needed in response to the user's request for a hotel reservation, they can determine that the task requested by the user is "hotel reservation".
[0277] On the other hand, the processors 111 and 131 of system 1000 can determine multiple task types based on the context of a goal-oriented dialogue. A goal-oriented dialogue includes various types of user dialogue inputs and various response dialogue actions thereto, and the context of such a goal-oriented dialogue is associated with various tasks. Thus, multiple task types can be determined based on the context of a goal-oriented dialogue according to one embodiment.
[0278] In step (S209), processors 111 and 131 of system 1000 perform a task of a determined type.
[0279] Processors 111 and 131 of the system 1000 execute a task whose type is determined based on a goal-oriented dialogue including user dialogue input and response dialogue act, and provide the execution result to the user.
[0280] In this case, processors 111 and 131 of the system 1000 generate programming code necessary to perform a task whose type is determined based on the context of the goal-oriented dialogue, and execute the programming code to perform the task.
[0281] Also, according to an embodiment, processors 111 and 131 of the system 1000 can capture a screen of an electronic device used by a user and obtain a user screen screenshot by receiving a user dialogue input.
[0282] Thereafter, when determining the type of the task, processors 111 and 131 of the system 1000 determine the type of the task based on information about the determined context of the goal-oriented dialogue and an analysis of the user screen screenshot.
[0283] In this case, when performing the task, processors 111 and 131 of the system 1000 can automatically execute a series of actions (e.g., cursor movement, clicking, text input, etc.) necessary to perform the determined task on the user screen.
[0284] Furthermore, when the type of a task is determined based on the context of a goal-oriented dialogue according to an embodiment, processors 111 and 131 of the system 1000 determine, among a plurality of task execution models, at least one task execution model optimized for the task whose type is determined based on the context of the goal-oriented dialogue, and perform the task using the determined at least one task execution model.
[0285] In this case, the method by which the processors 111 and 131 of system 1000 determine at least one task execution model optimized for the task and use it to perform the task is substantially the same as the "MoE-based model identification method" described later, and therefore, an explanation of this is omitted here.
[0286] If multiple task types are determined, the processors 111 and 131 of system 1000 can determine multiple task execution models optimized for each of the multiple tasks and use these to perform multiple tasks.
[0287] - MoE-based model identification method
[0288] The following describes in detail, with reference to the attached drawings, how a computing system 1000 according to one embodiment realizes a MoE architecture-based model provision service that enables modularization for predetermined expert models (SMs) within a MoE (Mixture of Experts) model.
[0289] Figure 13 is a flowchart illustrating a MoE-based model identification method according to one embodiment, and Figure 14 is a conceptual diagram illustrating a MoE-based model identification method according to one embodiment.
[0290] As shown in Figures 13 and 14, a method for realizing a MoE architecture-based model provision service in which a computing system 1000 according to one embodiment modularizes a specialized model (SM) including a MoE model may include the steps of: performing MoELM-based MoE learning (S301); acquiring specialized model (SM) characteristic information by MoE learning (S303); generating a specialized module model based on the acquired specialized model (SM) characteristic information (S305); acquiring predetermined domain information (S307); determining a domain-specific specialized model based on the acquired domain information (S309); constructing a MoE model based on the determined domain-specific specialized model (S311); and providing output data based on the constructed MoE model (S313).
[0291] Specifically, in many cases, it is difficult to distinguish or understand which domain a general, pre-trained specialization model (SM) specializes in.
[0292] This may impose certain constraints on selecting and utilizing specialized models (SMs) optimized for specific tasks (Tasks).
[0293] To address this, in one embodiment, the computing system 1000 can perform the following process to identify the role and / or function of each specialty model (SM) and modularize it, thereby effectively selecting and utilizing customized specialty models (SMs) optimized for a specific domain.
[0294] In detail, the computing system 1000 according to one embodiment performs MoELM-based MoE learning (S301).
[0295] In other words, in this embodiment, the computing system 1000 can perform MoE learning based on the aforementioned multiple expert model (SM) and router (RT) coupling-based MoELM.
[0296] At this time, through the learning process, the computing system 1000 can perform learning for each of the multiple specialized models (SMs) included in MoELM.
[0297] In other words, through the learning process described above, multiple specialized models (SMs) within MoELM can be trained individually.
[0298] In other words, the specialized model (SM) in this embodiment is an artificial intelligence model that has undergone optimized learning for a specific purpose, and can mean an artificial intelligence model that has been trained using training data and methods specific to that purpose.
[0299] In embodiments, such a specialized model (SM) may include a learned predetermined sLLM (including MoELM and / or DMoE models), a general MoE model (NM), an external model (EM), and / or a specialized module model (MM) according to embodiments disclosed below.
[0300] Furthermore, the computing system 1000 according to one embodiment acquires specialized model characteristic information (SMFI) through MoE learning (S303).
[0301] Here, the Specialized Model Feature Information (SMFI) according to the embodiment may mean information that identifies the role and / or function of a given specialized model (SM).
[0302] Referring further to Figure 8, in one embodiment, the computing system 1000 may further include a Model Specialization Module (MSM) according to one embodiment.
[0303] Further, the computing system 1000 can acquire the aforementioned specialized model characteristic information (SMFI) through linkage with a model specialization module (MSM).
[0304] Here, the model specialization module (MSM: Model Specialization Module) according to one embodiment may be an artificial intelligence module that generates and outputs specialized model characteristic information (SMFI) corresponding to a predetermined specialized model (SM) based on Mixture of Experts learning.
[0305] Specifically, in the embodiment, when the aforementioned Mixture of Experts learning is performed, the model specialization module (MSM) can monitor and track the task assignment status of the router (RT) for each specialized model (SM).
[0306] That is, in the embodiment, as a Mixture of Experts learning-based model is trained and operated, the model specialization module (MSM) can grasp what kind of tasks the router (RT) allocates and assigns to which specialized models (SM).
[0307] According to an embodiment, the model specialization module (MSM) can also generate a tag for identifying each tracked task assignment status and perform matching management on the tags.
[0308] Accordingly, in the embodiment, the model specialization module (MSM) can determine the specialization of each of the plurality of specialized models (SM).
[0309] In addition, in the embodiment, the model specialization module (MSM) generates specialized model characteristic information (SMFI) corresponding to each specialized model (SM) based on the determined specialization of each specialized model (SM).
[0310] FIG. 15 is a diagram showing an example of specialized model characteristic information (SMFI) according to an embodiment.
[0311] Here, as shown in Figure 15, in this embodiment, the Model Identification Module (MSM) can generate the aforementioned Specialized Model Feature Information (SMFI) in at least one of the following forms.
[0312] [Form 1] Specialized Model Feature Information (SMFI) in the form of selecting one of the pre-configured specialized model (SM) role and / or function-specific categories (e.g., Q&A or equipment control) in response to user input.
[0313] [Second form] Specialized model characteristic information (SMFI) in a natural language form that identifies the role and / or function of the specialized model (SM).
[0314] [Third form] Specialized model characteristic information (SMFI) in a form that identifies the role and / or function of a specialized model (SM) in at least one of the first and second forms, and further defines the input and output data of said specialized model (SM).
[0315] Next, in this embodiment, the Model Identification Module (MSM) can provide the Specialized Model Specific Information (SMFI) generated as described above to the computing system 1000 as output data.
[0316] Therefore, in this embodiment, the computing system 1000 can acquire characteristic information for each specialized model (SM) by linking with a model identification module (MSM).
[0317] Furthermore, the computing system 1000 according to one embodiment generates a Specialized Module Model (MM) based on the acquired Specialized Model Characteristic Information (SMFI) (S305).
[0318] Here, a specialized module model (MM) according to one embodiment may mean a specialized model (SM) that is matched with predetermined specialized model characteristic information (SMFI) and is independently separated.
[0319] In a more detailed embodiment, the computing system 1000 matches the acquired specialized model characteristic information (SMFI) as described above to the corresponding specialized model (SM).
[0320] In addition, in this embodiment, the computing system 1000 independently separates and databases each professional model (SM) that has been matched with professional model characteristic information (SMFI).
[0321] In other words, in this embodiment, the computing system 1000 performs modularization by matching each specialized model (SM) with its corresponding specialized model characteristic information (SMFI), and storing and managing them separately.
[0322] Therefore, the computing system 1000 can generate specialized module models (MMs), which are independently separated specialized models (SMs), at the same time that specialized model characteristic information (SMFI) is matched.
[0323] Thus, in this embodiment, the computing system 1000 grasps the characteristics of each specialty model (SM) within a given MoE model (MoELM in this embodiment) and modifies each specialty model (SM) to a small size that is reusable and shareable, reflecting these characteristics.
[0324] This allows the computing system 1000 to quickly and efficiently select specialized models (SMs) that realize data processing processes optimized for specific domains with greater accuracy, and to easily support the flexible expansion or contraction of the MoE model based on these.
[0325] Furthermore, the computing system 1000 according to one embodiment can acquire predetermined domain information (S307).
[0326] Here, the domain information according to the embodiment may be information that defines a domain that identifies data, rules, terminology, problem definitions and / or processes, etc., used by a given AI system to perform a given task.
[0327] In a more detailed embodiment, the computing system 1000 acquires predetermined input data (e.g., text, audio, images, video, and / or specific sensor-based sensing data).
[0328] In addition, in this embodiment, the computing system 1000 determines the domain corresponding to the acquired input data.
[0329] In this embodiment, the method by which the computing system 1000 determines the domain of the input data may be based on various disclosed algorithms that are executable, and the embodiments of the present invention do not limit or restrict the algorithms themselves.
[0330] Therefore, the implementation system 1000 can obtain domain information for the task it intends to process.
[0331] Furthermore, the computing system 1000 according to one embodiment determines a domain-specific professional model based on the acquired domain information (S309).
[0332] Here, a domain-specific specialized model according to the embodiment may mean a specialized model (SM) that performs data processing operations (such as deep learning, as an embodiment) optimized for a given domain.
[0333] Referring further to Figure 14, in an embodiment, the computing system 1000 determines at least one domain-specific professional model based on the domain information and professional model characteristic information (SMFI) obtained as described above.
[0334] More specifically, in an embodiment, the computing system 1000 can detect at least one Specialized Model Feature Information (SMFI) having characteristics corresponding to the acquired domain information.
[0335] For example, when computing system 1000 confirms the "characteristics of a task that outputs response data to predetermined question data" based on first domain information, it can detect at least one specialized model characteristic information (SMFI) from among multiple database-stored specialized model characteristic information (SMFI) that is identified as a "role and / or function specialized in question and answer."
[0336] In this embodiment, the computing system 1000 can detect at least one expert model characteristic information (SMFI) corresponding to domain information based on multiple tags generated by the model identification module (MSM) for each task assignment state of the router (RT) to multiple expert models (SM) during the aforementioned MoE architecture-based learning.
[0337] In other words, according to the embodiment, the computing system 1000 can detect at least one expert model characteristic information (SMFI) corresponding to the domain information by comparing the multiple tags and domain information generated as described above.
[0338] In this embodiment, the computing system 1000 filters the tags to be compared according to the generation time of each tag.
[0339] Specifically, the computing system 1000 sets at least one tag generated at the time of a particular task assignment as the comparison tag, according to user input and / or a pre-configured proprietary process.
[0340] For example, the computing system 1000, noting that the accuracy of task assignment improves as the learning rate increases, can set at least one tag generated for task assignment states performed after a predetermined time during the total learning time as the comparison tag.
[0341] Therefore, the computing system 1000 can detect at least one Specialized Model Feature Information (SMFI) corresponding to the relevant domain information by comparing it with at least one filtered tag and domain information, which ensures higher accuracy.
[0342] In addition, in the embodiment, the computing system 1000 extracts a specialized model (SM) (i.e., a specialized module model (MM)) that matches each of the detected specialized model characteristic information (SMFI).
[0343] In one embodiment, the computing system 1000 determines that at least one extracted specialized module model (MM) is a domain-specific specialized model.
[0344] Furthermore, a computing system 1000 according to one embodiment of the present invention constructs a determined domain-specific professional model-based MoE model (S311).
[0345] Referring further to Figure 14, in this embodiment, the computing system 1000 can construct a model (hereinafter referred to as the DMoE model) that operates like an MoE architecture based on at least one domain-specific professional model determined as described above.
[0346] In other words, computing system 1000 can construct a MoE model (i.e., a DMoE model) that enables data processing optimized for a specific domain by utilizing at least some of the multiple modularized, small-sized specialty models (SMs) (i.e., domain-specific specialty models).
[0347] In more detail, in the embodiment, the computing system 1000 can construct the aforementioned DMoE model by combining at least one domain-specific professional model and a predetermined router (RT).
[0348] Therefore, in this embodiment, the computing system 1000 can construct a domain-specific specialized model and a DMoE model including a router (RT).
[0349] Here, depending on the embodiment, the DMoE model may be included in the sLLM according to the embodiment of the present invention.
[0350] In other words, an sLLM according to one embodiment may include a DMoE model constructed according to one embodiment.
[0351] Furthermore, the computing system 1000 according to one embodiment of the present invention provides output data based on the constructed MoE model (S313).
[0352] In other words, in this embodiment, the computing system 1000 uses the DMoE model constructed as described above to provide output data (e.g., response data to a specific question and / or control signals using specific command words) for predetermined input data (e.g., text, audio, images, video, and / or sensing data based on specific sensors).
[0353] As described above, in the embodiment, the computing system 1000 can identify the role and / or function of each specialized model (SM), and at the same time separate and modularize them to a reusable and shareable level, thereby enabling the rapid and flexible construction of customized MoE models (i.e., DMoE models) optimized for specific domains, and providing predetermined output data through efficient task processing using the constructed models.
[0354] In other words, in this embodiment, the computing system 1000 can realize and provide a MoE model with further improved data processing (and / or computation) speed and inference performance, thereby supporting a variety of services, and thus effectively improving its performance and quality.
[0355] - Method for providing MoE-applied LLM-based AI agents
[0356] The following describes in detail, with reference to the attached drawings, how a computing system 1000 according to one embodiment of the present invention realizes a MoE architecture-based model provision service that determines an application model optimized for the domain based on an LLM (Large Language Model) that applies MoE (Mixture of Experts), and provides an on-device specialized AI agent (Artificial Intelligence Agent) that executes output based on the determined application model.
[0357] Figure 16 is a flowchart illustrating a method for providing an AI agent based on MoE-applied LLM according to one embodiment, and Figure 17 is a conceptual diagram illustrating a method for providing an AI agent based on MoE-applied LLM according to one embodiment.
[0358] As shown in Figures 16 and 17, a method for realizing a MoE architecture-based model provision service in which a computing system 1000 according to one embodiment determines an application model optimized for a domain based on an external environment based on an LLM that applies MoE, and provides an on-device specialized AI agent specialized model (AIAM) that outputs based on the determined application model, may include the steps of executing an on-device AI agent service (S401), acquiring predetermined input data (S403), determining a domain based on the acquired input data (S405), determining an application model based on the determined domain (S407), and providing output data based on the determined application model (S409).
[0359] Specifically, a computing system 1000 according to one embodiment of the present invention performs an on-device AI agent service (S401).
[0360] For reference, on-device AI can refer to a technology that directly performs artificial intelligence-based data processing within the user's device, rather than on the cloud and / or external servers. This can offer advantages such as protection of personal information, real-time processing, and reduced reliance on internet connectivity, as all processing is completed within the device without sending data externally.
[0361] Therefore, in this context, on-device AI agent services can refer to various services that are realized by utilizing on-device AI.
[0362] Exemplary examples of on-device AI agent services may include smartphone voice assistant services (e.g., Google Assistant, Apple Siri, Samsung Bixby, etc.), smart camera services (e.g., Google Pixel HDR+, Apple Deep Fusion, etc.), fitness tracker and smartwatch services (e.g., Apple Watch, Fitbit, etc.), autonomous driving services for automobiles (e.g., Tesla Autopilot, etc.), and / or home security services (e.g., Nest Secure, Ring, etc.).
[0363] In one embodiment, the computing system 1000 can execute a predetermined on-device AI agent service based on its interaction with an AI agent specialization model (AIAM) and / or a predetermined application, etc.
[0364] Furthermore, the computing system 1000 according to one embodiment of the present invention acquires predetermined input data (S403).
[0365] In more detail, in the embodiment, the computing system 1000 can acquire at least one input data (e.g., predetermined text, voice, image, video, and / or sensing data) based on user input and / or interaction with an external device (e.g., a predetermined sensor) based on the on-device AI agent service performed as described above.
[0366] In this embodiment, the input data acquired as described above may include predetermined data that can identify the target task of data processing.
[0367] Furthermore, the computing system 1000 according to one embodiment determines the domain based on the acquired input data (S405).
[0368] In other words, the domain in this embodiment may mean data, rules, terminology, problem definitions, and / or processes that a given AI system uses to perform a specific task.
[0369] In a more detailed embodiment, the computing system 1000 determines the domain corresponding to the acquired input data.
[0370] In this embodiment, the method by which the computing system 1000 determines the domain of the input data may be based on various disclosed algorithms that are executable, and in one embodiment, the algorithm itself is not limited or restricted.
[0371] Therefore, in this embodiment, the computing system 1000 can obtain domain information corresponding to the task it intends to process.
[0372] Furthermore, the computing system 1000 according to one embodiment determines the application model based on the determined domain (S407).
[0373] Here, the application model according to the embodiment may mean a model that performs a predetermined task processing based on given input data.
[0374] In an embodiment, such an application model may be at least one of the secondary models (S) described above.
[0375] In other words, a secondary model (S) in an embodiment may mean a model that can perform a specific task through the control and management of a master model (P) (i.e., an orchestrator (OCT) and / or router (RT), etc.) that is responsible for controlling and managing a given AI system operation.
[0376] In embodiments, such a secondary model (S) may include at least one model of an sLLM (including a MoELM and / or DMoE model), a general MoE model (NM), an external model (EM) and / or a specialized model (SM) (including a specialized module model (MM)).
[0377] In a more detailed embodiment, the computing system 1000 determines at least one application model based on the domain information obtained as described above.
[0378] More specifically, in one embodiment, the computing system 1000 works in conjunction with a master model (P) according to one embodiment (i.e., an orchestrator (OCT) and / or a router (RT), etc.) to detect at least one of the secondary models (S) described above that performs data processing operations (in one embodiment, such as deep learning) optimized for given domain information (i.e., domain-specific models).
[0379] Here, in one embodiment, the specific method by which the computing system 1000 detects a domain-specific model in conjunction with the master model (P) is omitted by applying the explanation of the router (RT) and orchestrator (OCT) disclosed in the aforementioned "AI Agent Specialization Model (AIAM)".
[0380] In addition, in the embodiment, the computing system 1000 determines that at least one detected domain-specific model is the applicable model.
[0381] Furthermore, the computing system 1000 according to one embodiment provides output data based on the determined application model (S409).
[0382] In other words, in this embodiment, the computing system 1000 can generate and provide output data (e.g., response data to a specific question and / or control signals using specific command words) for predetermined input data (e.g., text, voice, images, videos and / or specific sensor-based sensing data) based on at least one application model determined through the AI agent specialization model (AIAM) as described above.
[0383] In other words, the computing system 1000 performs a predetermined requested task based on given input data using the application model determined as described above, and provides output data resulting from the performed data processing.
[0384] In this embodiment, the computing system 1000 can provide the output data based on the on-device AI agent service described above.
[0385] As described above, in the embodiment, the computing system 1000 can effectively determine a model optimized for data processing according to a given domain, even in an on-device environment, based on an AI agent specialization model (AIAM) which includes models realized by applying the MoE architecture in various embodiments (such as MoELM, DMoE model and / or Specialized Module Model (MM) in the embodiment), and can provide output through efficient data processing via the determined model.
[0386] In other words, the computing system 1000 can realize and provide an artificial intelligence model (i.e., an AI agent specialist model (AIAM)) that can better listen to, understand, execute, and respond to given tasks in any environment.
[0387] Therefore, in this embodiment, the computing system 1000 can directly and significantly improve the quality and performance of a variety of AI agent-based services (e.g., smartphone voice assistant services, smart camera services, fitness tracker and smartwatch services, autonomous driving services for automobiles and / or home security services).
[0388] The embodiments of the present invention described above can be implemented in the form of program instructions that can be executed by a variety of computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., individually or in combination. The program instructions recorded on the computer-readable recording medium may be specifically designed or configured for the present invention, or may be publicly known and usable by those skilled in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memory. Examples of program instructions include not only machine code that can be produced by a compiler, but also advanced language code that can be executed by a computer using an interpreter or the like. Hardware devices can be modified into one or more software modules to perform the processing according to the present invention, and vice versa.
[0389] The specific executions described herein are merely embodiments and do not in any way limit the scope of the invention. For the sake of brevity, descriptions of conventional electronic configurations, control systems, software, and other functional aspects of said systems may be omitted. Furthermore, the connections of lines or connecting members between components shown in the drawings are illustrative examples of functional and / or physical or circuit connections and may be substituted or shown as a variety of additional functional, physical, or circuit connections in actual devices. Also, components that are not necessarily required for the application of the invention may not be necessary unless specifically mentioned, such as "essential" or "important."
[0390] Furthermore, while the detailed description of the present invention has been provided with reference to preferred embodiments, a person skilled in the art or with ordinary knowledge in the art will understand that the present invention can be modified and altered in various ways, within the scope of the concept and technical domain of the invention as described in the claims below. Therefore, the technical scope of the present invention is not limited to what is described in the detailed description of the specification, but must be defined by the claims. [Explanation of Symbols]
[0391] 10. Interactive Graph Generation Module 20 Dialogue Action Group Sampling Module 30 Dialogue Action Group Coordination Module 40 Dialogue Action Group Selection Module 50 Dialogue Action Selection Module 110, 180, 181, 182, 183 User Computing Devices 111, 131, 151 processors 112, 132, 152 memory 113, 133, 153 Data 114, 134, 154 imperative words 120, 140 Machine Learning Models 121 User Input Components 160 Model Trainer 161 training data 100, 200, 400 computing devices 300 Neuromorphic Circuits 310 Presynaptic Neuron Circuits 311 Presynaptic Line 320 Postsynaptic Neuron Circuits 330 Synaptic Circuits
Claims
1. A task execution method based on the context of purpose-oriented dialogue through an interaction model by a computing system including memory and a processor, Steps to receive user interaction input, A step of generating a response dialogue action to the user dialogue input using the dialogue model that has already been trained based on a dialogue graph that models at least one conditional relationship for the dialogue dataset, A step of determining the context of the purpose-oriented dialogue by analyzing data from a series of purpose-oriented dialogues, including the user dialogue input and the response dialogue actions. The steps include determining the type of task requested by the user based on the context of the aforementioned purpose-oriented dialogue, and A task execution method based on the context of a purpose-oriented dialogue, including the step of performing the task whose type has been determined.
2. In the step of determining the context of the aforementioned purpose-oriented dialogue, A method for executing a task based on the context of a purpose-oriented dialogue according to claim 1, comprising extracting a plurality of keywords from the data of the purpose-oriented dialogue, and determining the context by analyzing at least one of the correlations between a plurality of dialogue actions included in the purpose-oriented dialogue, the intention of the purpose-oriented dialogue, or the purpose of the purpose-oriented dialogue based on the plurality of keywords.
3. The process further includes the step of determining at least one task execution model from among several task execution models that is optimized for the task, whose type is determined based on the context of the goal-oriented dialogue, A task execution method based on the context of a purpose-oriented dialogue according to claim 1, wherein in the step of performing the task, the task is performed using the determined at least one task execution model.
4. The steps include determining the type of task and performing the task, A task execution method based on the context of a purpose-oriented dialogue according to claim 1, wherein the computing system determines the type of task and supports predetermined calculations necessary to perform the task.
5. In the step of performing the aforementioned task, We will analyze the programming code related to the aforementioned task, Based on the analysis results, generate programming code to perform the above task. A task execution method based on the context of a purpose-oriented dialogue according to claim 1, comprising executing the aforementioned programming code to perform the aforementioned task.
6. The method further includes the step of receiving the user interaction input and capturing the screen of an electronic device used by the user to obtain a user screen screenshot, A task execution method based on the context of a purpose-oriented dialogue according to claim 1, wherein in the step of determining the type of task, the type of task is determined based on information regarding the context of the purpose-oriented dialogue determined and an analysis of the user screen screenshot.
7. The step of generating the aforementioned response dialogue action is: A step of generating a dialogue graph that models at least one conditional relationship for the dialogue dataset, A step of sampling multiple dialogue action groups to respond to the user dialogue input using a dialogue model that has already been trained, A step of adjusting the plurality of dialogue action groups based on the dialogue graph, The steps include selecting one of the aforementioned group of dialogue actions that satisfies predetermined conditions, and A task execution method based on the context of a purpose-oriented dialogue according to claim 1, comprising the step of determining one dialogue action included in any one of the selected dialogue action groups as a response to the user dialogue input.
8. A task execution method based on the context of a purpose-oriented dialogue according to claim 7, wherein the at least one conditional relationship includes at least one of the following: a first conditional relationship concerning what utterance should be made in response to an utterance in the flow of dialogue; a second conditional relationship concerning what utterance may be made in response to an utterance; and a third conditional relationship concerning what utterance may not be made in response to an utterance.
9. In the step of selecting one of the aforementioned dialogue action groups, A task execution method based on the context of a purpose-oriented dialogue according to claim 7, comprising selecting one of the dialogue action groups that best satisfies the at least one conditional relationship.
10. A task execution method based on the context of a purpose-oriented dialogue according to claim 8, wherein each of the plurality of dialogue action groups includes at least one dialogue action in response to the user dialogue input.
11. In the step of coordinating the aforementioned multiple dialogue groups, A task execution method based on the context of a purpose-oriented dialogue according to claim 8, comprising adding at least one dialogue action that satisfies the first condition relationship to each of the plurality of dialogue action groups, removing dialogue actions that do not satisfy the second condition relationship, and removing dialogue actions that do not satisfy the third condition relationship.
12. In the step of selecting one of the aforementioned dialogue action groups, A task execution method based on the context of a purpose-oriented dialogue according to claim 11, comprising selecting from among the plurality of dialogue action groups the dialogue action group that contains the most dialogue actions that satisfy at least one conditional relationship.
13. In the aforementioned response dialogue action decision step, A task execution method based on the context of a purpose-oriented dialogue according to claim 7, wherein the dialogue action that is most relevant to the user dialogue input among at least one dialogue action included in any one of the selected dialogue action groups is provided as a response to the user dialogue input.
14. At least one memory, and Includes at least one processor that reads at least one instruction word stored in the memory and performs a task execution method based on the context of a goal-oriented dialogue, The aforementioned at least one processor is Receive user interaction input, Using the dialogue model already trained based on a dialogue graph that models at least one conditional relationship for the dialogue dataset, a response dialogue action to the user dialogue input is generated. The data of a series of purpose-oriented dialogues, including the user dialogue input and the response dialogue actions, is analyzed to determine the context of the purpose-oriented dialogue. Based on the context of the aforementioned goal-oriented dialogue, determine the type of task requested by the user. A task execution system based on the context of goal-oriented dialogue, which performs the determined task.
15. An electronic device for receiving user interaction input, and A computing device including at least one memory, at least one processor that reads at least one instruction word stored in the at least one memory and performs a task execution method based on the context of a purpose-oriented dialogue, The aforementioned at least one processor is Receive user interaction input, Using the dialogue model already trained based on a dialogue graph that models at least one conditional relationship for the dialogue dataset, a response dialogue action to the user dialogue input is generated. The data of a series of purpose-oriented dialogues, including the user dialogue input and the response dialogue actions, is analyzed to determine the context of the purpose-oriented dialogue. Based on the context of the aforementioned goal-oriented dialogue, determine the type of task requested by the user. A task execution system based on the context of goal-oriented dialogue, which performs the determined task.
16. The task execution system based on the context of a purpose-directed dialogue according to claim 15, wherein the electronic device receives the user dialogue input in at least one form of text, voice, gesture, and touch.
Citation Information
Patent Citations
Real scene information classification method and system for task-oriented dialogue
CN113901213A
Systems and methods for assisting agents through artificial intelligence
JP2022525362A
Entity-Level Data Augmentation in Chatbots for Robust Named Entity Recognition
JP2023530423A
System and method for generating dialogue graphs
US20200004878A1
Utilizing rule specificity in conversational ai
US20200227029A1