Method and system for performing task based on context of task-oriented dialogue
The task execution method addresses the challenges of high computing costs and resource inefficiencies in AI models by using a context-oriented dialogue graph to determine and execute tasks efficiently, enhancing performance and adaptability in on-device environments.
Patent Information
- Application Number
- JP2024212170
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-29
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-08
AI Technical Summary
Existing AI models face challenges in efficiently processing large datasets and adapting to specific contexts, leading to high computing costs and limitations in on-device environments, while conventional MoE architectures require significant VRAM and have issues with fine-tuning and resource management.
A task execution method based on context-oriented dialogue through a dialogue graph and model, which includes extracting keywords, analyzing intent, and determining the required task using optimized task execution models, enabling efficient task performance with reduced resource usage.
This approach allows for accurate and efficient execution of user-requested tasks by optimizing resource usage and adapting to specific contexts, improving computing efficiency and performance in on-device environments.
Smart Images

Figure 2025102692000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a task execution method and system based on the context of goal-oriented dialogue, and more particularly, to a task execution method and system for performing a user-requested task determined based on the context of goal-oriented dialogue through a dialogue graph and a dialogue model.
Background Art
[0002] Recently, as artificial intelligence (AI) technology has developed, various services utilizing artificial intelligence in diverse industrial fields have been commercialized. Such artificial intelligence technology enables an artificial neural network model to learn a vast amount of data and output and provide the information desired by the user.
[0003] On the other hand, there is a known dialogue-type artificial intelligence model that can interact with a user, rather than simply having an artificial intelligence model learn data patterns and provide an output value for an input value. An artificial intelligence secretary utilizing such a dialogue-type artificial intelligence model can be applied to smart devices to serve as a personal secretary, can be applied to a chatbot that answers questions from corporate customers, and can be applied to a smart home system that can control the operation of home appliances connected to the Internet of Things according to user requests. As a result, the dialogue-type artificial intelligence model can provide a more convenient and user-friendly experience to users in various technical fields.
[0004] In addition, research on artificial intelligence technology capable of executing the operations requested by a user has been actively conducted, and for this purpose, applications to dialogue-type artificial intelligence models of an analysis method based on a task-oriented dialogue graph have been attempted many times.
[0005] However, most of the research related to task-oriented dialogue graphs has remained at the level of trying to generate dialogue models using dialogue graphs directly drawn by humans or leveraging rule-based systems that automatically infer dialogue graphs based on already learned dialogue policies. As a result, it has been unable to provide a dialogue graph-based dialogue model automatically constructed based solely on dialogue datasets without human intervention, or has limitations in modeling diverse dialogue flows.
[0006] On the other hand, generally, artificial intelligence is realized through a large number of AI models and deep learning based on them.
[0007] Such artificial intelligence has been developed to provide diverse services considering the user's context (such as context, environment, and / or intention, etc.).
[0008] However, when trying to process a specific task based on large-capacity data, there are limitations in that the required computing costs and time are considerable.
[0009] As a result, there are also certain constraints in the use of AI models in the recently spotlighted on-device environment.
[0010] To solve this, conventionally, model architectures such as MoE (Mixture of Experts) have been utilized.
[0011] Here, MoE refers to the architecture of a machine learning model that combines a number of expert models to solve complex problems.
[0012] Such a MoE may include an expert model which is a plurality of small networks designed to learn different parts and / or different characteristics of predetermined data and perform corresponding data processing operations, and a gating network that evaluates the performance of each expert model and determines which expert model is most suitable for assigning a specific task according to the predetermined data based on this evaluation.
[0013] Therefore, according to the MoE architecture, the gating network that obtains the predetermined input data determines a probabilistic or deterministic task assignment for each expert model, and the selected expert models perform their respective tasks and return the results to perform data processing for a specific task.
[0014] By utilizing such a MoE, the AI model can enhance the overall efficiency and performance by activating only specific parts and concentrating computing resources in cases such as handling complex tasks or large datasets.
[0015] However, in the case of a conventional MoE, not only is a high level of VRAM required, but there are also quite a few issues that need to be solved in the fine-tuning process.
[0016] In addition to this, the conventional MoE method is for efficiently managing large-sized models, and has limitations in supporting the efficiency of the remaining resources that are not activated according to a given task.
[0017] Also, in the conventional art in this technical field, services are mostly provided using generally realized AI models, but there is a problem that it is difficult to quickly and easily ensure the AI analysis performance most suitable for a given context.
Summary of the Invention
Problems to be Solved by the Invention
[0018] According to various embodiments of the present disclosure, by determining the context of goal-oriented dialogue through a dialogue graph and a dialogue model, and determining the task required by the user based on the determined context of the dialogue, a task execution method based on the context of goal-oriented dialogue through a dialogue model that can accurately execute the user request task is to be provided.
[0019] However, the technical problems to be achieved by various embodiments of the present disclosure are not limited to the above technical problems, and other technical problems may exist.
Means for Solving the Problems
[0020] One embodiment provides a task execution method based on the context of goal-oriented dialogue through a dialogue model by a computing system including a memory and a processor, the method including the steps of receiving a user dialogue input, determining and providing a response dialogue act for the user dialogue input based on a dialogue graph, analyzing a series of goal-oriented dialogue data including the user dialogue input and the response dialogue act to determine the context of the goal-oriented dialogue, determining the type of task required by the user based on the context of the goal-oriented dialogue, and performing the task whose type has been determined.
[0021] In another aspect, in the step of determining the context of the goal-oriented dialogue, a plurality of keywords may be extracted from the data of the goal-oriented dialogue, and at least one of the correlation relationship between a plurality of dialogue acts included in the goal-oriented dialogue, the intention of the goal-oriented dialogue, and the goal may be analyzed based on the plurality of keywords to determine the context.
[0022] In another aspect, the task execution method includes the step of determining at least one task execution model optimized for the task whose type is determined based on the context of the goal-oriented dialogue among a plurality of task execution models, and in the step of performing the task, the determined at least one task execution model may be used to perform the task.
[0023] On the other hand, in the steps of determining the type of the task and performing the task, the computing system may determine the type of the task and support predetermined operations necessary for performing the task.
[0024] On the other hand, in the step of performing the task, an analysis of the programming code related to the task may be performed, programming code for performing the task may be generated based on the analysis result, and the task may be performed by executing the programming code.
[0025] On the other hand, the task execution method may further include the step of capturing a screen of an electronic device used by a user by receiving user interaction input to obtain a user screen shot, and determining the type of the task based on information regarding the determined context of the goal-oriented dialogue and the analysis of the user screen shot in the step of determining the type of the task.
[0026] On the other hand, the step of determining and providing the response dialogue act may include the step of generating a dialogue graph modeling at least one conditional relationship for a dialogue data set, the step of sampling a plurality of dialogue act groups for responding to the user dialogue input using a previously learned dialogue model, the step of adjusting the plurality of dialogue act groups based on the dialogue graph, and the step of selecting any one dialogue act group satisfying a predetermined condition among the plurality of dialogue act groups.
[0027] On the other hand, the at least one conditional relationship may include at least any one of a first conditional relationship regarding what utterance should be made for a single utterance, a second conditional relationship regarding what utterance can be made for a single utterance, and a third conditional relationship regarding what utterance should not be made for a single utterance in the flow of the dialogue.
[0028] On the other hand, in the step of selecting any one of the dialogue act groups, any one of the dialogue act groups that most satisfies the at least one conditional relationship among the plurality of dialogue act groups may be selected.
[0029] On the other hand, each of the plurality of dialogue act groups may include at least one dialogue act for the user dialogue input.
[0030] On the other hand, in the step of adjusting the plurality of dialogue act groups, for each of the plurality of dialogue act groups, a dialogue act that satisfies the first conditional relationship among at least one dialogue act may be added, a dialogue act that does not satisfy the second conditional relationship may be removed, and a dialogue act that does not satisfy the third conditional relationship may be removed.
[0031] On the other hand, in the step of selecting any one of the dialogue act groups, any one of the dialogue act groups that includes the most dialogue acts that satisfy the at least one conditional relationship among the plurality of dialogue act groups may be selected.
[0032] On the other hand, the goal-oriented dialogue method may further include a response dialogue act determination step of providing, as a response to the user dialogue input, any one of the dialogue acts included in any one of the selected dialogue act groups.
[0033] On the other hand, in the response dialogue act determination step, any one of the dialogue acts with the highest relevance to the user dialogue input among at least one dialogue act included in any one of the already selected dialogue act groups may be provided as a response to the user dialogue input.
[0034] One embodiment includes at least one memory and at least one processor that reads at least one instruction word stored in the memory and performs a task execution method based on the context of goal-directed dialogue. The at least one processor receives user dialogue input, determines and provides a response dialogue act for the user dialogue input based on a dialogue graph, analyzes a series of goal-directed dialogue data including the user dialogue input and the response dialogue act to determine the context of the goal-directed dialogue, determines the type of task required by the user based on the context of the goal-directed dialogue, and may perform the determined task.
[0035] One embodiment includes an electronic device that receives user dialogue input, and a computing device including at least one memory and at least one processor that reads at least one instruction word stored in the at least one memory and performs a task execution method based on the context of goal-directed dialogue. The at least one processor receives user dialogue input, determines and provides a response dialogue act for the user dialogue input based on a dialogue graph, analyzes a series of goal-directed dialogue data including the user dialogue input and the response dialogue act to determine the context of the goal-directed dialogue, determines the type of task required by the user based on the context of the goal-directed dialogue, and performs the determined task, providing a task execution system based on the context of goal-directed dialogue.
[0036] In another aspect, the electronic device may receive the user dialogue input in at least one of the forms of text, voice, gesture, and touch.
Advantages of the Invention
[0037] According to various embodiments of the present disclosure, data of goal-oriented dialogue is analyzed through a dialogue graph and a dialogue model to determine the context of the corresponding dialogue, a task required by the user is determined based on the determined dialogue context, and a dialogue model capable of accurately executing the user-requested task is obtained by using a model optimized for the determined task among a plurality of task execution models, thereby providing a task execution method based on the context of goal-oriented dialogue.
[0038] However, the effects obtained by various embodiments of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood from the following description.
Brief Description of Drawings
[0039]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying Out the Invention
[0040] Since the present invention can be applied with various conversions and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. The effects and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described in detail below together with the drawings. However, the present invention is not limited to the embodiments disclosed below and can be realized in various forms. In the following embodiments, terms such as first and second are not used in a limiting sense but are used to distinguish one component from another. Also, the singular expression includes plural expressions unless the context clearly has a different meaning. Also, terms such as "comprising" or "having" mean that the features or components described in the specification exist, and do not preclude in advance the possibility of adding one or more other features or components. Also, in the drawings, the sizes of the components may be exaggerated or reduced for convenience of explanation. For example, the sizes and thicknesses of each configuration in the drawings are arbitrarily shown for convenience of explanation, so the present invention is not necessarily limited to the illustrated content.
[0041] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. When describing with reference to the drawings, the same or corresponding components are given the same reference numerals, and duplicate descriptions thereof are omitted.
[0042] - System 1000 for providing a goal-oriented dialogue service
[0043] According to an embodiment, system 1000 generates a task-oriented dialogue graph based on an analysis of a dialogue dataset, and based on the generated dialogue graph, an already learned dialogue model verifies and adjusts various response dialogue acts sampled for a user dialogue input, and finally selects a response dialogue act and provides this as a response to the user dialogue input.
[0044] In this case, the system 1000 can provide the user with a more reliable response dialogue act by modeling a predetermined conditional relationship for the dialogue dataset to generate a dialogue graph and verifying and adjusting the response dialogue acts sampled based on the predetermined conditional relationship included in the dialogue graph.
[0045] In addition, the system 1000 can determine the type of task required by the user based on the user dialogue input and the finally generated response dialogue act, and provide the user with a more convenient task processing experience by efficiently executing the task based on the determined task type.
[0046] FIG. 1 shows an example of a block diagram of a computing system 1000 that implements an object-oriented dialogue service and a task execution service according to an embodiment.
[0047] As shown in FIG. 1, a computing system 1000 that implements an object-oriented dialogue service and a task execution service according to an embodiment includes a user computing device 110, a server computing system 130, and a training computing system 150, and the devices can communicate via a network 170.
[0048] The object-oriented dialogue method according to an embodiment may be realized and provided locally by the user computing device 110, may be realized and provided in the form of a web service by the server computing system 130 that communicates with the user computing device 110, or may be realized and provided in cooperation between the user computing device 110 and the server computing system 130.
[0049] Here, in the embodiment, the user computing device 110 and / or the server computing system 130 can learn the machine learning models 120 and / or 140 through interaction with a training computing system 150 communicatively connected via a network 170. The training computing system 150 may be separate from the server computing system 130 or may be a part of the server computing system 130.
[0050] And here, the artificial intelligence model can be learned in the following ways: 1) directly learned locally by the user computing device 110, 2) learned through interaction between the server computing system 130 and the user computing device 110 via the network 170, and 3) learned by a separate training computing system 150 using various training techniques and learning techniques. And it can also be realized in a way that the artificial intelligence model learned by the training computing system 150 is transmitted to the user computing device 110 and / or the server computing system 130 via the network 170 for providing / updating.
[0051] In some embodiments, the training computing system 150 may be a part of the server computing system 130 or a part of the user computing device 110.
[0052] FIG. 2 is a conceptual diagram for explaining a method by which a computing system 1000 according to an embodiment performs a task requested by a user based on an object-oriented dialogue service.
[0053] As shown in FIG. 2, the computing system 1000 can receive various forms of user interaction input, provide appropriate response interaction behaviors for the received user interaction input, and further perform tasks required by the user based on the user interaction input and the response interaction behaviors. Here, the tasks may include various types of tasks determined based on goal-oriented dialogues including user interaction input and response interaction behaviors.
[0054] The computing system 1000 grasps the context of a series of goal-oriented dialogues consisting of user interaction input and response interaction behaviors thereto.
[0055] The computing system 1000 can analyze the patterns of goal-oriented dialogues composed of various types of user interaction input and response interaction behaviors thereto, and based on this, determine the context related to the intention, purpose, etc. of the corresponding dialogue.
[0056] The computing system 1000 can extract a plurality of keywords included in the data of the goal-oriented dialogue, and analyze the correlation between the plurality of dialogue behaviors included in the dialogue, the intention and purpose of the dialogue, etc. based on the extracted plurality of keywords, and finally determine the context of the goal-oriented dialogue.
[0057] In addition, the computing system 1000 can determine the type of task required by the user based on the information related to the context of the goal-oriented dialogue.
[0058] For example, a plurality of keywords (such as hotel, reservation, date, number of people, five-star, etc.) are extracted from a series of dialogue data including user interaction input requesting hotel reservation and response interaction behaviors requesting information related to hotel reservation (such as reservation date, number of overnight guests, hotel grade, etc.).
[0059] Further, based on the plurality of extracted keywords, at least one of the correlation between the user dialogue behavior included in the corresponding dialogue and the response dialogue behavior provided by the computing system 1000, the intention of the corresponding dialogue, and the purpose can be analyzed to determine the context of the corresponding dialogue.
[0060] For example, according to the keyword-based context determination process, the context of a dialogue including a plurality of dialogue behaviors such as "Please make a hotel reservation", "What is the check-in date?", "Check-in on December 31, 2024, and check-out on January 3, 2025", "How many people will stay?", and "3 people" is determined that the user requests a hotel reservation and thus an automatic hotel reservation service is required.
[0061] Furthermore, based on the determined context of the dialogue, the type of task requested by the user can be determined.
[0062] For example, based on the context of the dialogue determined that an automatic hotel reservation service is required in response to the user's hotel reservation request, the task requested by the user is determined to be "hotel reservation".
[0063] On the other hand, the number of types of tasks determined based on the context of goal-oriented dialogue may be plural. The goal-oriented dialogue may include various types of user dialogue inputs and various response dialogue behaviors corresponding thereto, and the context of such goal-oriented dialogue may be related to various tasks. Thereby, based on the context of the goal-oriented dialogue according to an embodiment, the types of a plurality of tasks can be determined.
[0064] The system 1000 executes the task whose type is determined based on the goal-oriented dialogue including the user dialogue input and the response dialogue behavior, and provides the execution result to the user.
[0065] For example, when the type of the task is determined to be "hotel reservation", the system 1000 can complete the hotel reservation task in response to the user's request based on the data related to a series of goal-oriented dialogues.
[0066] In this case, the system 1000 determines the type of the task and supports the predetermined operations required to perform the task with the determined type.
[0067] For example, the operating system of the system 1000 can control the processors 111 and 131 to directly determine the type of the task and execute the predetermined operations required to perform the task with the determined type.
[0068] Also, for example, the operating system of the system 1000 can provide a predetermined API (Application Programming Interface) and / or SDK (Software Development Kit) for determining the type of the task and supporting the predetermined operations required to perform the task with the determined type.
[0069] The system 1000 generates the programming code required to perform the task with the type determined based on the context of the goal-oriented dialogue, and executes this to perform the task. Here, the programming code may be code created using a programming language (for example, Java, Python, JavaScript, etc.) to perform a specific task.
[0070] On the other hand, the system 1000 can capture the screen of the electronic device used by the user by receiving the user dialogue input to obtain a user screen screenshot.
[0071] Thereafter, when determining the type of the task, the system 1000 can determine the type of the task based on the information regarding the determined context of the goal-oriented dialogue and the analysis of the user screen screenshot.
[0072] For example, when the context of the conversation is determined to be "hotel reservation" and the user screen screenshot includes the screen of the homepage that provides the hotel reservation service, the system 1000 can determine the type of task as "hotel reservation via the corresponding hotel reservation service-providing homepage".
[0073] In this case, the system 1000 can automatically perform a series of actions (such as cursor movement, clicking, text input, etc.) necessary to perform the task determined on the user screen.
[0074] However, not limited to this, the tasks determined based on the goal-oriented conversation may be determined to be of countless types according to the content of the user interaction input and the response interaction actions. For example, the tasks may be determined to be of various types such as product purchase, email sending, information search, document creation, etc.
[0075] Furthermore, when the type of task is determined based on the context of the goal-oriented conversation according to an embodiment, the system 1000 determines at least one task execution model optimized for the task whose type is determined based on the context of the goal-oriented conversation among a plurality of task execution models, and performs the task using the determined at least one task execution model.
[0076] In this case, the method by which the system 1000 determines at least one task execution model optimized for the task and uses this to perform the task is substantially the same as the "MoE-based model identification method" described later, and the description thereof is omitted.
[0077] Also, for example, the system 1000 may include a first type of first user computing device 180, a second type of second user computing device 181, a third type of third user computing device 182, and a fourth type of fourth user computing device 183 that can receive various forms of user interaction input from the user.
[0078] Here, the user interaction input may have at least one of the forms of text, voice, gesture, and touch. However, without being limited thereto, the user interaction input can have various forms other than the above examples.
[0079] The first user computing device 180 of the first type may be a virtual reality electronic device, the second user computing device 181 of the second type may be a mobile electronic device, the third user computing device 182 of the third type may be an augmented reality electronic device, and the fourth user computing device 183 of the fourth type may be a desktop.
[0080] However, without being limited thereto, the system 1000 may include user computing devices in various forms other than the above examples that can receive user interaction input.
[0081] - User Computing Device (110:User Computing Device)
[0082] The user computing device 110 may include all other types of computing devices such as smart phones, mobile phones, digital broadcast devices, PDAs (personal digital assistants), PMPs (portable multimedia players), desktops, wearable devices, embedded computing devices, tablet PCs, augmented reality (VR) devices, and / or virtual reality (AR) devices.
[0083] Such a user computing device 110 may include at least one or more processors 111 and a memory 112. Here, the processor 111 may be composed of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, micro - controllers, microprocessors and / or at least one of other electrical units for performing functions or a plurality of electrically connected processors.
[0084] For example, the ASICs may have a structure of an array - shaped neuromorphic circuit including a plurality of neuron circuits.
[0085] As shown in FIG. 3, for example, the neuromorphic circuit 300 may include a plurality of presynaptic neuron circuits 310, a plurality of presynaptic lines 311 extending laterally from the plurality of presynaptic neuron circuits 310, a plurality of postsynaptic neuron circuits 320, a plurality of postsynaptic lines 321 extending vertically from the plurality of postsynaptic neuron circuits 320, and synapse circuits 330 provided at intersections of the plurality of presynaptic lines 311 and the plurality of postsynaptic lines 321.
[0086] The plurality of presynaptic neuron circuits 310 can transmit externally input signals in the form of electrical signals to the plurality of synapse circuits 330 via the plurality of presynaptic lines 311.
[0087] Moreover, the plurality of postsynaptic neuron circuits 320 can receive electrical signals from the plurality of synaptic circuits 330 via the plurality of postsynaptic lines 321.
[0088] Furthermore, the plurality of postsynaptic neuron circuits 320 can also transmit electrical signals to the plurality of synaptic circuits 330 via the plurality of postsynaptic lines 321.
[0089] The plurality of synaptic circuits 330 store the weight values included in the layer constituting the neural network system realized by the neuromorphic circuit 300, and can perform a predetermined operation based on the weight values and the input data.
[0090] For example, each of the plurality of synaptic circuits 330 may include a resistive memory cell having a variable resistor. In this case, the resistance value of the plurality of synaptic circuits 330 changes due to the voltage applied via the plurality of presynaptic neuron circuits 310 or the plurality of postsynaptic neuron circuits 320, and weight value data due to such resistance change can be stored.
[0091] The neuromorphic circuit 300 is formed by mimicking the neuron and synapse structures that are essential elements of the human brain. When realizing a deep neural network (DNN) using the neuromorphic circuit 300, the data processing speed can be improved and the power consumption can be reduced compared to the case of utilizing the existing von Neumann architecture.
[0092] The memory 112 may include one or more non-transitory / temporary computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof, and may also include web storage of a server that performs the storage function of the memory on the internet. Such a memory 112 can store the data 113 and instruction words 114 necessary for the at least one or more processors 111 to perform functional operations such as learning an artificial intelligence model or performing vision inspection through the artificial intelligence model.
[0093] In one embodiment, the user computing device 110 can store at least one or more machine learning models 120.
[0094] For example, the machine learning model 120 may be various machine learning models such as multiple neural networks (e.g., Deep neural network) for performing goal-oriented dialogue methods and task execution methods, or non-linear models and / or linear models, or other types of machine learning models including combinations thereof.
[0095] For example, linear regression, decision trees, random forests, gradient boosting, pre-trained language models, or / and deep learning models, etc. can be stored in the machine learning model. And the neural network may include at least one or more of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or / and other forms of neural networks.
[0096] In addition, in various embodiments, the user computing device 110 can also store, for each process, a model and a prompt template that serves as the basis for the input to the model for performing at least a part of the processes for the goal-oriented dialogue method and the task execution method through a large language model (LLM).
[0097] In one embodiment, the user computing device 110 receives at least one or more machine learning models 120 from the server computing system 130 via the network 170, stores them in the memory 112, and then executes the stored machine learning models 120 by the processor 111 to perform dialogue dataset analysis and the like.
[0098] In other embodiments, the server computing system 130 includes at least one or more machine learning models 140, executes operations through the machine learning models 140, and can provide goal-oriented dialogue services and task execution services to the user in a manner that communicates with the user computing device 110 and the data related thereto.
[0099] For example, the user computing device 110 can perform goal-oriented dialogue services in such a way that the server computing system 130 uses the machine learning model 140 to provide an output for the user's input via the web.
[0100] Also, at least a part of the machine learning models 120 and / or 140 can be executed in the user computing device 110, and the remaining part can be executed in the server computing system 130, so that the artificial intelligence model can be realized.
[0101] In addition, the user computing device 110 may include at least one or more input components 121 that sense user input. For example, the user input component 121 may include a touch sensor (such as a touch screen and / or a touch pad, etc.) that senses the touch of a user input medium (such as a finger or a stylus), an image sensor that senses user motion input, a microphone that senses user voice input, buttons, a mouse and / or a keyboard, etc. Further, when the user input component 121 receives input to an external controller (such as a mouse and / or a keyboard, etc.) via an interface, it may include the interface and the external controller.
[0102] - Server Computing System (130:Server Computing System)
[0103] The server computing system 130 performs a series of processes to provide a goal-oriented dialogue service.
[0104] In addition, the server computing system 130 may further perform a series of processes to provide a task execution service based on the context of goal-oriented dialogue.
[0105] Specifically, in an embodiment, the server computing system 130 can provide a goal-oriented dialogue service by exchanging data necessary for driving the goal-oriented dialogue service and task execution service processes in an external device such as the user computing device 110 with the external device.
[0106] More specifically, in an embodiment, the server computing system 130 can provide an environment in which an application for providing a goal-oriented dialogue service and a task execution service can operate in the user computing device 110.
[0107] For this purpose, the server computing system 130 may include application programs, data, and / or instruction words for the application to operate, and can transmit and receive various data based on this to the external device.
[0108] The server computing system 130 may include at least one or more processors 131 and a memory 132. Here, the processor 131 may be a central processing unit (CPU), a graphics processing unit (GPU), application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, and / or at least one of other electrical units for executing functions or a plurality of electrically connected processors.
[0109] For example, the ASICs may have a structure of an array-like neuromorphic circuit including a plurality of neuron circuits (see Figure 3).
[0110] And the memory 132 may include one or more non-temporary / temporary computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. Such a memory 132 can store the data 133 and instruction words 134 necessary for the processor 131 to perform functional operations such as learning an artificial intelligence model or performing an object-oriented dialogue method and a task execution method through the artificial intelligence model.
[0111] In one embodiment, the server computing system 130 may be implemented to include at least one or more computing devices. For example, the server computing system 130 may be implemented to operate multiple computing devices in a sequential computing architecture, a parallel computing architecture, or a combination thereof. Also, the server computing system 130 may include a plurality of computing devices connected by a network 170.
[0112] Further, the server computing system 130 can store at least one or more machine learning models 140. For example, the server computing system 130 may include a neural network and / or other multi-layer non-linear models as the machine learning model 140. Exemplarily, the neural network may include a feedforward neural network, a deep neural network, a recurrent neural network, and a convolutional neural network.
[0113] In an embodiment, the server computing system 130 may further include a data store computing system (hereinafter referred to as a data store), which is a storage for continuously storing and managing the raw data underlying the goal-oriented dialogue service.
[0114] Such a data store may include various forms of data storage, ranging from file systems to cloud storage. For example, the data store may include a relational database that uses Structured Query Language (SQL) to define and manipulate data, a NoSQL database designed for flexibility and scalability to handle unstructured and semi-structured data, a data warehouse used for reporting and data analysis that centralizes large volumes of data from multiple sources and optimizes it for querying and analysis, a data warehouse that stores large amounts of raw data in structured, semi-structured, and unstructured data in its basic form, and at least one database including a local storage device or a NAS (Network Attached Storage) that stores data in a form accessible in a general computer operating system in a file.
[0115] - Training Computing System (150:Training Computing System)
[0116] The training computing system 150 may include at least one or more processors 151 and a memory 152. Here, the processor 151 may be composed of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, micro-controllers, microprocessors, and / or at least one or a plurality of electrically connected electrical units for performing other functions.
[0117] For example, ASICs may have the structure of an array-like neuromorphic circuit including a plurality of neuron circuits (see FIG. 3).
[0118] And the memory 152 may include one or more non-transitory / temporary computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. Such a memory 152 can store the data 153 and instruction words 154 necessary for the processor 151 to execute learning of an artificial intelligence model, etc.
[0119] For example, the training computing system 150 may include a model trainer 160 that uses various training or learning techniques such as error backpropagation to train the machine learning models 120 and / or 140 stored in the user computing device 110 and / or the server computing system 130.
[0120] Exemplarily, such a model trainer 160 can perform updates to one or more parameters of the machine learning models 120 and / or 140 for the purpose-oriented dialogue service in a backpropagation manner based on a defined loss function.
[0121] In some embodiments, performing error backpropagation may include performing truncated back propagation through time. The model trainer 160 can execute a number of generalization techniques (e.g., weight decay, dropout, and / or knowledge distillation, etc.) to improve the generalization ability of the machine learning models 120 and / or 140 to be trained.
[0122] For example, the model trainer 160 can train the machine learning model 120 and / or 140 based on a series of training data 161. Here, the training data 161 may include data in different formats such as, for example, images, audio samples, and / or text.
[0123] Also, the training data 161 may include, for example, various types of dialogue data. In this case, the dialogue data may be data related to task-oriented dialogue that requests a specific task and provides a response to the requested task.
[0124] Examples of image types that can be used may include video frames, LiDAR point clouds, X-ray images, computed tomography scans, hyperspectral images, and / or other diverse forms of images.
[0125] Such training data 161 can be provided by the user computing device 110 and / or the server computing system 130. When the training computing device causes the machine learning model 120 and / or 140 to learn based on specific data of the user computing device 110, the machine learning model 120 and / or 140 can be characterized as a personalized model.
[0126] And the model trainer 160 includes computer logic that is utilized to provide the desired functionality.
[0127] In addition, the model trainer 160 can be implemented by hardware, firmware, and / or software that controls a general-purpose processor. In one implementation example, the model trainer 160 includes program files stored in a storage device, is loaded into the memory 152, and can be executed by one or more processors 151. In other implementation examples, the model trainer 160 includes one or more sets of computer-executable data 153 and instruction words 154 stored in a computer-readable storage medium of a type such as a RAM hard disk or an optical or magnetic medium.
[0128] The network 170 includes, but is not limited to, a 3GPP (Registered Trademark) (3rd Generation Partnership Project) network, an LTE (Long Term Evolution) network, a WIMAX (World Interoperability for Microwave Access) network, the Internet, a LAN (Local Area Network), a Wireless LAN (Wireless Local Area Network), a WAN (Wide Area Network), a PAN (Personal Area Network), a Bluetooth (Registered Trademark) network, a satellite broadcast network, an analog broadcast network, and / or a DMB (Digital Multimedia Broadcasting) network.
[0129] Generally, communication via the network 170 may be performed using any type of wired and / or wireless connection via various communication protocols (e.g., TCP / IP, HTTP, SMTP, and / or FTP, etc.), encoding or formats (e.g., HTML and / or XML, etc.), and / or protection schemes (e.g., VPN, secure HTTP, and / or SSL, etc.).
[0130] FIG. 4 is a block diagram of a computing device 100 that implements an object-oriented dialogue service and a task execution service according to an embodiment.
[0131] As shown in FIG. 4, the computing device 100 included in the user computing device 110, the server computing system 130, and the training computing system 150 includes a number of applications (for example, Application 1 to Application N). Each application may include a machine learning library and one or more machine learning models. For example, the application may include an image processing (for example, Detection, Classification, and / or Segmentation, etc.) application, a text messaging application, an email application, a writing application, a virtual keyboard application, a browser application, and / or a chat-bot application, etc.
[0132] In an embodiment, the computing device 100 may include a model trainer 160 for training an artificial intelligence model, and by storing and operating the trained artificial intelligence model, output data corresponding to predetermined input data (as an embodiment, a dialogue dataset, etc.) can be provided.
[0133] Each application of the computing device 100 can communicate with a number of other components of the computing device 100, such as, for example, at least one or more sensors, a context manager, a device state component, and / or additional components. In one embodiment, each application can communicate with each device component using an API (for example, a public API). In one embodiment, the API used by each application may be specific to that application.
[0134] FIG. 5 is a block diagram of a computing device 200 that implements an intent-based dialogue service and a task execution service according to another embodiment.
[0135] As shown in FIG. 5, the computing device 200 includes a number of applications (e.g., Application 1 to Application N). Each application can communicate with the central intelligence layer. For example, the applications may include an image processing application, a text messaging application, an email application, a writing application, a virtual keyboard application, and / or a browser application, etc. In one embodiment, each application can use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0136] The central intelligence layer may include a number of machine learning models. For example, as shown in FIG. 5, at least a part of each machine learning model can be provided for each application and can be managed by the central intelligence layer. In other implementation examples, two or more applications can share a single machine learning model. For example, in some implementation examples, the central intelligence layer can provide a single model for all applications. In some implementation examples, the central intelligence layer may be included within the operating system of the computing device 200 or may be implemented differently.
[0137] The Central Intelligence Layer can communicate with the Central Device Data Layer. The Central Device Data Layer can be a central centralized data storage for the computing device 200. As shown in FIG. 5, the Central Device Data Layer can communicate with a number of other components of the computing device 200, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the Central Device Data Layer can communicate with each device component using an API (e.g., a private API).
[0138] The techniques described herein can refer to actions taken and information sent to or from the systems, not only servers, databases, software applications, and other computer-based systems. One will recognize the inherent flexibility of computer-based systems to allow for a wide variety of possible configurations, combinations, and divisions of work and functionality between and among components. For example, the processes described herein can be implemented using a single device or component or multiple devices or components operating in combination. The databases and applications may be implemented in a distributed system across a single system or multiple systems. The distributed components can operate sequentially or in parallel.
[0139] FIG. 6 is a block diagram for explaining the functions of a computing device 400 that implements a goal-oriented dialogue service according to an embodiment.
[0140] As shown in FIG. 6, the computing device 400 may include a dialogue graph generation module 10, a dialogue act group sampling module 20, a dialogue act group adjustment module 30, a dialogue act group selection module 40, and a dialogue act selection module 50.
[0141] The dialogue graph generation module 10 can generate a dialogue graph based on a dialogue dataset received from the outside.
[0142] The dialogue dataset may include data related to conversations between various speakers in various environments. For example, the dialogue dataset may include data related to a series of various types of conversation flows, such as a conversation between a customer and a reservation agent in a situation of booking a hotel, a conversation between a buyer and a seller in a situation of purchasing an item, and a conversation between a tutor and a tutee regarding academic tasks.
[0143] A dialogue graph is a graphically represented model of the relationships between various types of utterances included in a dialogue dataset, and may include a plurality of nodes corresponding to a plurality of dialogue acts indicating the functions of a specific utterance, and a plurality of edges indicating various information regarding the relationships between the plurality of dialogue acts.
[0144] Here, for example, the dialogue act may include the functions of various types of utterances such as a question, an inform, a request, and a confirm.
[0145] In addition, the relationship between a plurality of dialogue acts may include a predetermined conditional relationship between two utterances. For example, the predetermined conditional relationship may include at least any one of a first conditional relationship (Should relationship) regarding what utterance should be made in response to one utterance, a second conditional relationship (Can relationship) regarding what utterance can be made in response to one utterance in the flow of the conversation, and a third conditional relationship (Should-not relationship) regarding what utterance should not be made in response to one utterance.
[0146] For example, as shown in FIG. 7, the dialogue graph generation module 10 generates a task-oriented dialogue flow (TOD-Flow) graph including a plurality of nodes including A, B, C, D and a plurality of edges corresponding to the relationships between the plurality of nodes (S10).
[0147] The dialogue graph generation module 10 can vectorize the dialogue dataset by performing embedding on the words and sentences included in the dialogue dataset. The dialogue graph generation module 10 analyzes the vectorized dialogue dataset to learn relevant information of the dialogue dataset, such as the intent of the dialogue utterance, dialogue acts, slots, values, context information, information on the relationship between utterances, and the pattern of the sequential dialogue flow.
[0148] For example, the dialogue graph generation module 10 can learn the relevant information of the dialogue dataset by using an artificial intelligence model based on at least one of a Transformer-based model, a recurrent neural network (RNN), and a long short-term memory (LSTM).
[0149] In addition, the dialogue graph generation module 10 can generate a dialogue graph by modeling at least one conditional relationship for the dialogue dataset based on the learned relevant information of the dialogue dataset.
[0150] In this case, the dialogue graph generation module 10 models at least one relationship for the dialogue dataset by maximizing a loss function optimized for each at least one conditional relationship.
[0151] For example, when the nth dialogue act satisfies the first conditional relationship (Should relationship) in the dialogue context, the dialogue graph generation module 10 generates a dialogue graph by utilizing a dialogue graph generation model learned to maximize the expected value.
[0152] Specifically, in a situation where the first conditional relationship (Should relationship) of the dialogue context must be satisfied, the dialogue graph generation module 10 can generate a dialogue graph by utilizing a dialogue graph generation model learned to maximize the expected value when the nth dialogue act satisfies the first conditional relationship (Should relationship).
[0153] For example, the dialogue graph generation module 10 can model the first conditional relationship (Should relationship) regarding what utterance should be made for a single utterance in the dialogue flow by maximizing the loss function defined in the following [Equation 1].
[0154]
Equation
[0155] JPEG2025102692000003.jpg48151
[0156] Also, for example, when the nth dialogue act of the dialogue context satisfies the second conditional relationship (Can relationship) and at the same time does not satisfy the third conditional relationship (Should - not relationship), the dialogue graph generation module 10 can generate a dialogue graph by utilizing a dialogue graph generation model learned to maximize the expected value.
[0157] Specifically, in the situation where the nth dialogue act occurs, when the nth dialogue act in the dialogue context satisfies the second conditional relationship (Can relationship) and does not satisfy the third conditional relationship (Should-not relationship), the first expected value, and in the situation where the nth dialogue act does not occur, when the nth dialogue act in the dialogue context does not satisfy the second conditional relationship (Can relationship) and satisfies the third conditional relationship (Should-not relationship), the dialogue graph generation module 10 generates a dialogue graph by utilizing a dialogue graph generation model learned to maximize the sum of the second expected values. For example, the dialogue graph generation module 10 models the second conditional relationship (Can relationship) regarding what utterances can be made for a single utterance in the flow of the dialogue and the third conditional relationship (Should-not) regarding what utterances should not be made for a single utterance by maximizing the loss function defined in the following [Equation 2].
[0158]
Equation
[0159] JPEG2025102692000005.jpg58151
[0160] The dialogue act group sampling module 20 can sample a plurality of dialogue act groups for responding to user dialogue inputs received from the outside.
[0161] For example, the dialogue act group sampling module 20 can sample a plurality of dialogue act groups appropriate as responses to user dialogue inputs received via an electronic device. As shown in FIG. 7, the next dialogue act prediction can be performed for the user dialogue input. For example, for a user dialogue input related to "hotel reservation", the dialogue act group sampling module 20 samples a dialogue act group including various types of dialogue acts such as "Confirm Information" and "Book Hotel" as the dialogue act group that is a response following the user dialogue input (S20).
[0162] The dialogue act group sampling module 20 may include an artificial neural network structure that can extract the characteristics of various types of dialogue datasets and has already been trained to provide appropriate output data for the input data. For example, the dialogue act group sampling module 20 may include a neural network architecture based on Transformer (such as GPT-3, GPT-4, BERT-based models, etc.).
[0163] The dialogue act group adjustment module 30 can adjust a plurality of dialogue act groups sampled based on the generated dialogue graph.
[0164] Each of the plurality of dialogue act groups generated by the dialogue act group sampling module 20 may include at least one dialogue act related to the user dialogue input.
[0165] The dialogue act group adjustment module 30 can adjust a plurality of dialogue act groups based on whether each of the plurality of dialogue act groups generated by the dialogue act group sampling module 20 satisfies at least one conditional relationship included in the dialogue graph for the user dialogue input.
[0166] For example, the dialogue act group adjustment module 30 determines whether each of the plurality of dialogue act groups satisfies a first conditional relationship (Should relationship) for the user dialogue input, and can add a pair of dialogue acts that satisfy the first conditional relationship (Should relationship) for the user dialogue input to a pair of dialogue act groups that do not satisfy it.
[0167] In addition, the dialogue act group adjustment module 30 determines whether each of a plurality of dialogue act groups satisfies a second conditional relationship (Can relationship) and a contra-third conditional relationship (not should not) with respect to a user dialogue input, and removes a dialogue act that does not satisfy the second conditional relationship (Can relationship) and the contra-third conditional relationship (not should not) among at least one dialogue act included in a dialogue act group that does not satisfy the condition.
[0168] In this way, by adjusting the plurality of dialogue act groups sampled by the dialogue act group adjustment module 30 so as to conform to the dialogue graph (S20), the reliability of the goal-oriented dialogue service provided by the system 1000 and the control power over the dialogue model are improved.
[0169] The dialogue act group selection module 40 can select any one dialogue act group that satisfies a predetermined condition from among the plurality of adjusted dialogue act groups.
[0170] For example, the dialogue act group selection module 40 determines the ranking of a plurality of dialogue act groups in descending order of the number of dialogue acts that satisfy at least one conditional relationship of the dialogue graph among the plurality of adjusted dialogue act groups (S30).
[0171] For example, the first dialogue act group, the second dialogue act group, and the third dialogue act group can be adjusted by the dialogue act group adjustment module 30. After the adjustment operation is completed, the first dialogue act group includes three dialogue acts, the second dialogue act group includes two dialogue acts, and the third dialogue act group includes four dialogue acts.
[0172] In this case, the dialogue act group selection module 40 can select the third dialogue act group that includes the most dialogue acts as the first rank among the first to third dialogue act groups based on at least one conditional relationship of the dialogue graph after adjustment.
[0173] The dialogue act selection module 50 can select any one of at least one dialogue act included in one dialogue act group selected from a plurality of dialogue act groups and provide a response output to the user dialogue input.
[0174] For example, the dialogue act selection module 50 can calculate the probability of the relevance of at least one dialogue act included in the selected dialogue act group to the user dialogue input. Thereafter, the dialogue act selection module 50 can select any one dialogue act with the highest calculated probability corresponding to the relevance of the user dialogue input from the selected dialogue act group and provide it to the user as a response.
[0175] -AI Agent Specialization Model (AIAM)
[0176] In other aspects, the computing system 1000 as described above may include an AI Agent Specialization Model (AIAM) according to one embodiment.
[0177] Here, the AI Agent Specialization Model (AIAM) according to an embodiment is an AI agent model that applies a MoE (Mixture of Experts) architecture implemented by one embodiment, and can be an artificial intelligence model including a data processing algorithm that can autonomously act in a specific environment, solve tasks, and achieve goals. Here, MoE means an architecture of a machine learning model that combines a number of expert models to solve complex problems.
[0178] Such an AI Agent Specialization Model (AIAM) may include data processing algorithms for realizing cognitive abilities such as collecting and interpreting data from a given environment, decision mechanisms for determining optimal actions based on the collected data, execution abilities for executing the determined actions, and learning abilities for improving actions through experience.
[0179] As an embodiment, the AI Agent Specialization Model (AIAM) can obtain predetermined input data (such as text, voice, image, video, and / or specific sensor-based sensing data, etc.), and provide output data (such as response data to specific queries and / or control signals by specific command words, etc.) by performing predetermined tasks based on the obtained input data.
[0180] FIG. 8 shows an internal block diagram of an AI agent model according to an embodiment.
[0181] Specifically, as shown in FIG. 8, an AI agent model according to an embodiment may include at least one router (RT: Router, Gating Network), an orchestrator (OCT: Orchestrator), a sLLM (small Large Language Model), a general MoE model (NM: Normal MoE Model), an external model (EM: External Model), and / or a specialized model (SM: Specialized Model).
[0182] Here, in FIG. 8, in order to prevent the features according to various embodiments from being unclear, it is described that the AI agent model includes the above-described components.
[0183] However, it is obvious that a person of ordinary skill in the art can understand that, in addition to the components shown in FIG. 8 according to the embodiment, other general-purpose components may be further included, or some of the components shown in FIG. 8 may be omitted.
[0184] More specifically, a router (RT: Router, Gating Network) according to an embodiment may be an artificial intelligence module that performs task assignment and / or traffic adjustment for a plurality of models in the MoE architecture.
[0185] Specifically, the router (RT) can analyze the given input data and / or requested tasks, etc., to determine which model is most suitable for the data processing.
[0186] At this time, as an embodiment, the router (RT) can determine the model optimized for the given data processing based on the performance, expertise, and / or previous experience of each model, etc.
[0187] In addition, the router (RT) can consider the system load and distribute the given task to at least one or more models to support efficient data processing.
[0188] Furthermore, the router (RT) can flexibly respond to changes in the real-time system and adjust the tasks assigned to a specific model.
[0189] In an embodiment, such a router (RT) may be an artificial intelligence module that selectively determines a model (hereinafter, domain-specific model) that executes a data processing operation optimized for a predetermined domain.
[0190] That is, in an embodiment, the router (RT) may be an artificial intelligence module that selects a model (i.e., a domain specialization model) that is determined to perform data processing (such as deep learning in an embodiment) most suitable for a given domain among a plurality of models included in the AI agent specialization model (AIAM).
[0191] For reference, the domain according to an embodiment means data, rules, terms, problem definitions, and / or processes, etc. used by a given AI system to perform a specific task.
[0192] As an embodiment, the router (RT) performs data analysis based on characteristics of given input data (e.g., user input and / or specific sensing data, etc.) and / or a required task, etc., grasps data processing characteristics optimized for the corresponding task based on this, and can detect a given model that realizes this to determine the domain specialization model.
[0193] In other words, the router (RT) according to an embodiment may be an artificial intelligence module that detects a model that can most effectively execute data processing for a given domain, assigns / distributes the corresponding task processing work, and manages this.
[0194] Here, the router (RT) according to an embodiment may include a router (RT) already learned by a disclosed given algorithm, a router (RT) additionally learned according to an embodiment, and / or a router (RT) newly learned in a new manner, etc.
[0195] On the other hand, the orchestrator (OCT: Orchestrator) according to an embodiment may be an artificial intelligence module that overall controls and manages the overall configuration of the AI agent specialization model (AIAM).
[0196] In a detailed embodiment, the orchestrator (OCT) assigns various tasks generated from the entire system to appropriate resources (such as a router (RT) and / or a predetermined model as an embodiment).
[0197] Also, the orchestrator (OCT) manages to efficiently use resources such as available models and hardware resources (e.g., CPU and / or GPU, etc.).
[0198] Also, the orchestrator (OCT) monitors the performance of the entire system and adjusts specific parameters as needed, or optimizes the network configuration, etc.
[0199] Also, the orchestrator (OCT) manages the cooperation between multiple routers (RTs) and / or models, and controls the data flow and processing process.
[0200] That is, in an embodiment, the orchestrator (OCT) controls and manages the entire system of the AI agent specialization model (AIAM), and performs the role of the main router (RT) that controls at least one router (RT).
[0201] Here, the orchestrator (OCT) and the router (RT) according to the embodiment can cooperate closely with each other to support the efficient operation of the MoE system.
[0202] Specifically, the orchestration (OCT) monitors the performance of the router (RT) as the administrator of the entire system, and adjusts the strategy of the router (RT) as needed.
[0203] On the other hand, the router (RT) can substantially perform the assignment of data processing tasks according to the instructions of the orchestrator (OCT) and / or its own algorithms to achieve efficient system control.
[0204] In an embodiment, the aforementioned orchestrator (OCT) and / or router (RT) may be a master model (P:Master Model) that can control and manage the remaining components of the overall system and / or the AI agent model (i.e., sLLM, general MoE model (NM), external model (EM), and / or specialized model (SM), etc.).
[0205] On the other hand, an sLLM (small Large Language Model) according to one embodiment is an artificial intelligence module realized as a lightweight version of an LLM (Large Language Model).
[0206] That is, the sLLM is an artificial intelligence module constructed to achieve performance similar to that of a large model such as an LLM with fewer resources.
[0207] In an embodiment, such an sLLM may include a plurality of specialized models (SM) and a router (RT)-coupled based MoE model (in the embodiment, MoELM) according to one embodiment. Further, the sLLM may include a domain-specialized specialized model-based MoE model (in the embodiment, DMoE model) according to one embodiment.
[0208] Also, the general MoE model (NM:Normal MoE Model) according to an embodiment of the present invention may mean a predetermined MoE model realized by a disclosed universal method.
[0209] Exemplarily, the general MoE model (NM) may include Switch Transformer, Conditional Computation in Neural Networks, Sparse Mixture of Experts, and / or Megatron-LM, etc.
[0210] In addition, an external model (EM) according to an embodiment may refer to a predetermined artificial intelligence model implemented by various disclosed algorithms.
[0211] For example, the external model (EM) may include ChatGPT, Gemini, and / or Llama, etc.
[0212] In an embodiment, such an external model (EM) can selectively be used as needed to support the given task processing.
[0213] In addition, a specialized model (SM) according to an embodiment is an artificial intelligence model in which optimal learning for a specific purpose has been performed, and may refer to an artificial intelligence model learned by training data and methods specialized for the corresponding purpose.
[0214] That is, the specialized model (SM) can be an artificial intelligence model learned by training data and methods specialized to achieve a predetermined purpose.
[0215] In an embodiment, such a specialized model (SM) may include a learned predetermined sLLM (including MoELM and / or DMoE models), a general MoE model (NM), and / or an external model (EM), etc. Further, the specialized model (SM) may include a specialized module model according to an embodiment disclosed in the "MoE-based model identification method" described later. A detailed description thereof will be given in the "MoE-based model identification method".
[0216] In an embodiment, the sLLM, general MoE model (NM), external model (EM), and / or specialized model (SM) as described above can be secondary models (S: Secondary Model) that can execute specific tasks under the control and management of the master model (P) of the AI agent model (i.e., the orchestrator (OCT) and / or router (RT), etc.).
[0217] - Goal - Oriented Dialogue Method (S100)
[0218] The following will explain in detail the goal - oriented dialogue method (S100) that can provide a more accurate response to a user dialogue input by extracting and learning the characteristics of various types of dialogue datasets, generating a dialogue graph that models a predetermined conditional relationship for the dialogue dataset based on this, and selecting the optimal dialogue act from among a plurality of dialogue acts sampled by a dialogue model already learned for the user dialogue input and providing it as a response.
[0219] The dialogue dataset may include data related to dialogues between various speakers conducted in various environments. Therefore, the dialogue dataset includes various types of data according to the nature of the dialogues conducted between speakers.
[0220] For example, the first dialogue dataset and the second dialogue dataset related to different types of tasks may include different types of data, and the data structures of the first dialogue graph generated based on the first dialogue dataset and the second dialogue graph generated based on the second dialogue dataset may be different from each other.
[0221] The dialogue graph models the relationships between various types of utterances included in the dialogue dataset and represents them in graph form. It can be data having a structured form of data related to a plurality of dialogue acts corresponding to the functions, intentions, etc. of a plurality of utterances and data related to the conditional relationships between a plurality of dialogue acts.
[0222] The goal - oriented dialogue method (S100) can select and provide the optimal dialogue act group as a response to the user dialogue input from among a plurality of dialogue act groups sampled by the dialogue model based on the dialogue graph.
[0223] Also, based on the user dialogue input and the dialogue action groups selected as responses, the task required by the user is determined, and the computing system 1000 according to one embodiment can perform the determined task to provide the user with a task execution service.
[0224] In the following, a goal-oriented dialogue method (S100) will be described in detail, in which the dialogue model according to one embodiment provides an appropriate response to the user dialogue input based on the dialogue graph that models at least one conditional relationship for the dialogue dataset, and can execute the task required by the user.
[0225] FIG. 9 is a flowchart of a goal-oriented dialogue method (S100) according to one embodiment.
[0226] As shown in FIG. 9, the goal-oriented dialogue method (S100) according to one embodiment may include a step (S101) of generating a dialogue graph that models at least one conditional relationship for the dialogue dataset, a step (S103) of receiving a user dialogue input, a step (S105) of sampling a plurality of dialogue action groups for responding to the user dialogue input using a previously learned dialogue model, a step (S107) of adjusting the plurality of dialogue action groups based on the dialogue graph, and a step (S109) of selecting any one dialogue group that satisfies a predetermined condition among the plurality of dialogue action groups.
[0227] In step (S101), the processors 111 and 131 of the system 1000 generate a dialogue graph based on various types of dialogue datasets.
[0228] For example, the processors 111 and 131 generate a first dialogue graph that models at least one conditional relationship for a first type of dialogue dataset related to the first task. Also, the processors 111 and 131 generate a second dialogue graph that models at least one conditional relationship for a second type of dialogue dataset related to the first task and another second task.
[0229] Here, at least one conditional relationship for the dialogue dataset may include at least any one of a first conditional relationship (Should relationship) regarding what utterance should be made for a single utterance included in the dialogue dataset, a second conditional relationship (Can) regarding what utterance can be made for a single utterance, and a third conditional relationship regarding what utterance should not be made for a single utterance.
[0230] In step (S103), the processors 111 and 131 of the system 1000 receive user dialogue input.
[0231] For example, the processors 111 and 131 can be transmitted with data of user dialogue input received by the user computing device 110, which can be realized by various types of electronic devices, via the user input component 121.
[0232] In step (S105), the processors 111 and 131 of the system 1000 sample a plurality of dialogue act groups for providing as responses to the user dialogue input by using the already learned dialogue model.
[0233] For example, as shown in FIG. 10, the processors 111 and 131 can sample a plurality of dialogue act groups a1, a2,... by using the dialogue model (π) already learned by the model trainer 160 of the training computing system 150. In this case, the number of the plurality of dialogue act groups a1, a2,... sampled by the dialogue model (π) may be several tens, but is not limited thereto.
[0234] For example, the sampled first dialogue act group a1 may include four dialogue acts A, B, C, F, and the second dialogue act group a2 may include three dialogue acts A, C, G.
[0235] In step (S107), processors 111 and 131 of system 1000 adjust a plurality of dialogue act groups sampled based on the dialogue graph.
[0236] Processors 111 and 131 can adjust the plurality of dialogue act groups based on whether each of the plurality of dialogue act groups satisfies at least one conditional relationship included in the dialogue graph with respect to the user dialogue input.
[0237] First, processors 111 and 131 determine any one dialogue graph corresponding to the type of user dialogue input among the plurality of dialogue graphs generated for various dialogue datasets.
[0238] Thereafter, processors 111 and 131 adjust the plurality of dialogue act groups based on the determined dialogue graph.
[0239] As shown in FIG. 10, the determined dialogue graph includes G as a dialogue act that satisfies the first conditional relationship (Should relationship) with respect to the user dialogue input, includes A, C, F, and G as dialogue acts that satisfy the second conditional relationship (Can relationship), and includes A as a dialogue act that satisfies the third conditional relationship (Should-not relationship).
[0240] Processors 111 and 131 remove B from the first dialogue act group a1 so that the first dialogue act group a1 satisfies the second conditional relationship (Can relationship) based on the determined dialogue graph.
[0241] Also, processors 111 and 131 add G to the first dialogue act group a1 so that the first dialogue act group a1 satisfies the first conditional relationship (Should relationship) based on the determined dialogue graph.
[0242] Furthermore, processors 111 and 131 remove A from the first dialogue act group a1 so that the first dialogue act group a1 satisfies the third conditional relationship (Should-not relationship) based on the determined dialogue graph.
[0243] Similarly, the processors 111 and 131 remove A from the second dialogue act group a2 so that the second dialogue act group a2 satisfies the first to third conditional relationships based on the determined dialogue graph.
[0244] In this case, after the adjustment operations for the plurality of dialogue act groups are completed, the first dialogue act group a1 includes three dialogue acts of C, F, and G, and the second dialogue act group a2 includes two dialogue acts of C and G.
[0245] By adjusting the plurality of sampled dialogue act groups to conform to the dialogue graph in this way, the reliability of the goal-oriented dialogue service provided by the system 1000 and the control power over the dialogue model can be improved.
[0246] In step (S109), the processors 111 and 131 of the system 1000 select any one dialogue act group that satisfies a predetermined condition among the adjusted plurality of dialogue act groups and provide it as a response to the user dialogue input.
[0247] The processors 111 and 131 select the dialogue act group that includes the most dialogue acts satisfying at least one conditional relationship of the dialogue graph among the adjusted plurality of dialogue act groups.
[0248] For example, after the adjustment process, the processors 111 and 131 can select the first dialogue act group a1 that includes three dialogue acts of C, F, and G and the second dialogue act group (a2) that includes two dialogue acts of C and G, and select the first dialogue act group a1 that includes the most dialogue acts satisfying at least one conditional relationship of the dialogue graph and provide it as a response to the user dialogue input.
[0249] FIG. 11 shows the response prediction performance index (F-1 score) of various dialogue models (FLAN-T5, GPT-turbo) for user dialogue inputs with respect to the goal-oriented dialogue method (S100) based on the dialogue graph according to an embodiment of the present disclosure for various dialogue datasets (SGD, MultiWOZ).
[0250] In this case, by changing the method of selecting, as a response, any one of a plurality of dialogue act groups generated from various dialogue models (FLAN-T5, GPT-turbo) and adjusted based on the dialogue graph, the response prediction performance index of the dialogue model changes.
[0251] Also, when selecting any one of the plurality of dialogue act groups, the response prediction performance index of the dialogue model changes depending on whether or not an adjustment operation based on at least one conditional relationship of the dialogue graph is performed.
[0252] For example, the processors 111, 131 select the dialogue act group determined by the dialogue model to have the highest response probability among the plurality of dialogue act groups and provide it as a response to the user dialogue input (Greedy).
[0253] Also, the processors 111, 131 select the dialogue act group that contains the most dialogue acts satisfying at least one conditional relationship of the dialogue graph (Compliance), or select the dialogue act group that is adjusted the least based on at least one conditional relationship and provide it as a response to the user dialogue input (Violation).
[0254] Furthermore, the processors 111, 131 select the most sampled dialogue act group among the plurality of dialogue act groups sampled by the dialogue model and provide it as a response to the user dialogue input (Majority).
[0255] As shown in FIG. 11, when following the method (Compliance) of adjusting a plurality of dialogue act groups based on all of the first conditional relationship (Should relationship), the second conditional relationship (Can relationship), and the third conditional relationship (Should-not relationship) of the dialogue graph, and selecting the dialogue act group that contains the most dialogue acts satisfying at least one conditional relationship of the dialogue graph among the plurality of dialogue act groups, the response prediction performance index of the dialogue model is the highest.
[0256] In this way, by appropriately adjusting a plurality of dialogue act groups based on the dialogue graph, and selecting any one dialogue act group that most satisfies the conditional relationship of the dialogue graph among the adjusted plurality of dialogue act groups and providing it as a response to the user dialogue input, a more reliable response can be provided.
[0257] On the other hand, the method (S100) may further include a response dialogue act determination step of providing, as a response to the user dialogue input, any one dialogue act included in any one selected dialogue act group.
[0258] For example, when the user dialogue input is text data obtained by converting the voice "Please make a hotel reservation" into text, the selected first dialogue act group a1 among the adjusted plurality of dialogue act groups includes three dialogue acts: "How many people will be staying? (C)", "What date do you wish to make a reservation? (F)", and "What grade of hotel do you prefer? (G)".
[0259] In the response dialogue act determination step, the processors 111, 131 select any one of the plurality of dialogue acts (C, F, G) included in the first dialogue act group a1 and provide it as a response to the user dialogue input. In this case, the processors 111, 131 can calculate the relevance of each of the plurality of dialogue acts (C, F, G) included in the first dialogue act group a1 to the user dialogue input, and select and provide as a response to the user dialogue input any one dialogue act with the highest relevance.
[0260] Furthermore, the processors 111 and 131 can determine the type of task requested by the user based on the user interaction input and the dialogue act group finally selected as the response, and perform the determined task.
[0261] For example, when the type of task is determined to be "hotel reservation", based on various information related to the hotel reservation determined in the process of a series of dialogues including the user interaction input and the selected dialogue act group, the processors 111 and 131 can complete a hotel reservation that meets the user's requirements.
[0262] In this way, the system 1000 can adjust the dialogue act sampled by the dialogue model based on the dialogue graph in which the predetermined conditional relationship is structured and modeled based on the dialogue dataset, and provide the adjusted dialogue act as a response to the user interaction input, thereby providing a more reliable response to the user.
[0263] Furthermore, based on a series of dialogues including the user interaction input and the dialogue act adjusted by the dialogue graph and determined as the final response according to the predetermined conditions, the type of task requested by the user is accurately determined, and by the system 1000 performing the task determined in this way, a task execution service that satisfies the user can be provided.
[0264] FIG. 12 is a flowchart of a task execution method (S200) based on the context of goal-oriented dialogue through a dialogue model according to an embodiment.
[0265] As shown in FIG. 12, a task execution method (S200) according to an embodiment may include a step (S201) of receiving a user interaction input, a step (S203) of determining and providing a response interaction act for the user interaction input based on an interaction graph, a step (S205) of analyzing data of a series of goal-oriented dialogs including the user interaction input and the response interaction act to determine the context of the goal-oriented dialog, a step (S207) of determining the type of task requested by the user based on the context of the goal-oriented dialog, and a step (S209) of performing the task whose type has been determined.
[0266] In step (S201), processors 111 and 131 of system 1000 receive a user interaction input.
[0267] For example, processors 111 and 131 can be transmitted data of a user interaction input received by user computing device 110, which can be realized by various types of electronic devices, via user input component 121.
[0268] In step (S203), processors 111 and 131 of system 1000 determine and provide a response interaction act for the user interaction input based on an interaction graph related to the user interaction input.
[0269] Step (S203) is substantially the same as the goal-oriented dialog method (S100) described with reference to FIGS. 9 and 10, and the description thereof is omitted.
[0270] In step (S205), processors 111 and 131 of system 1000 analyze data of a series of goal-oriented dialogs including the user interaction input and the response interaction act to determine the context of the goal-oriented dialog.
[0271] Processors 111 and 131 of system 1000 can grasp the context of a series of goal-oriented dialogs consisting of the user interaction input and the response interaction act thereto.
[0272] For example, the processors 111 and 131 of the system 1000 extract a plurality of keywords included in the data of the goal-oriented dialogue, and analyze the correlation relationship between a plurality of dialogue acts included in the dialogue, the intention and purpose of the dialogue, etc. based on the extracted plurality of keywords, and finally determine the context of the goal-oriented dialogue.
[0273] In step (S207), the processors 111 and 131 of the system 1000 determine the type of task required by the user based on the context of the goal-oriented dialogue.
[0274] The processors 111 and 131 of the system 1000 determine the type of task corresponding to the context of the goal-oriented dialogue determined based on the data related to the task corresponding to the context of the goal-oriented dialogue.
[0275] In this case, the data related to the task corresponding to the context of the goal-oriented dialogue may already be stored in the memories 112 and 132 of the system 1000, or may already be learned by the machine learning models 120 and 140.
[0276] For example, when the processors 111 and 131 of the system 1000 determine that the context of the goal-oriented dialogue requires an automatic hotel reservation service in response to the user's hotel reservation request, the processors 111 and 131 can determine the task required by the user as "hotel reservation".
[0277] On the other hand, the processors 111 and 131 of the system 1000 can determine multiple types of tasks based on the context of the goal-oriented dialogue. The goal-oriented dialogue includes various types of user dialogue inputs and various response dialogue acts in response thereto, and such a context of the goal-oriented dialogue is related to various tasks. Accordingly, based on the context of the goal-oriented dialogue according to an embodiment, multiple types of tasks can be determined.
[0278] In step (S209), the processors 111 and 131 of the system 1000 perform the task whose type has been determined.
[0279] The processors 111 and 131 of the system 1000 execute a task whose type is determined based on an goal-oriented dialogue including user interaction inputs and response interaction acts, and provide the execution result to the user.
[0280] In this case, the processors 111 and 131 of the system 1000 generate programming code necessary to perform a task whose type is determined based on the context of the goal-oriented dialogue, and execute this to perform the task.
[0281] Also, according to one embodiment, the processors 111 and 131 of the system 1000 can capture the screen of the electronic device used by the user by receiving a user interaction input, and obtain a user screen screenshot.
[0282] After that, when determining the type of the task, the processors 111 and 131 of the system 1000 determine the type of the task based on information regarding the determined context of the goal-oriented dialogue and the analysis of the user screen screenshot.
[0283] In this case, when performing the task, the processors 111 and 131 of the system 1000 can automatically execute a series of acts (such as cursor movement, click, text input, etc.) necessary to perform the determined task on the user screen.
[0284] Furthermore, when the type of the task is determined based on the context of the goal-oriented dialogue according to one embodiment, the processors 111 and 131 of the system 1000 determine at least one task execution model optimized for the task whose type is determined based on the context of the goal-oriented dialogue among a plurality of task execution models, and perform the task using the determined at least one task execution model.
[0285] In this case, the method by which the processors 111 and 131 of the system 1000 determine at least one task execution model optimized for a task and use it to perform the task is substantially the same as the "MoE-based model identification method" described later, and the description thereof is omitted here.
[0286] When the types of tasks are determined to be plural, the processors 111 and 131 of the system 1000 can determine a plurality of task execution models optimized for each of the plurality of tasks and use them to perform the plurality of tasks.
[0287] -MoE-based model identification method
[0288] Hereinafter, a method for implementing a model providing service based on a MoE architecture by which a computing system 1000 according to an embodiment realizes modularization for a predetermined specialized model (SM) in a MoE (Mixture of Experts) model will be described in detail with reference to the accompanying drawings.
[0289] FIG. 13 is a flowchart for explaining a MoE-based model identification method according to an embodiment, and FIG. 14 is a conceptual diagram for explaining a MoE-based model identification method according to an embodiment.
[0290] As shown in FIGS. 13 and 14, a method for realizing a MoE architecture-based model providing service for modularizing a specialized model (SM) including a MoE model in a computing system 1000 according to an embodiment may include steps of performing MoE learning based on MoELM (S301), obtaining specialized model (SM) characteristic information by the MoE learning (S303), generating a specialized module model based on the obtained specialized model (SM) characteristic information (S305), obtaining predetermined domain information (S307), determining a domain-specific specialized model based on the obtained domain information (S309), constructing a MoE model based on the determined domain-specific specialized model (S311), and providing output data based on the constructed MoE model (S313).
[0291] Specifically, in many cases, it is difficult to distinguish or grasp for a general already learned specialized model (SM) what domain the model specializes in.
[0292] As a result, there may be certain constraints in selecting and utilizing a specialized model (SM) optimized for a specific task.
[0293] To solve this problem, in one embodiment, the computing system 1000 can execute the following process of identifying and modularizing the role and / or function of each specialized model (SM) and effectively selecting and utilizing a customized specialized model (SM) optimized for a specific domain based on this.
[0294] Specifically, the computing system 1000 according to an embodiment performs MoE learning based on MoELM (S301).
[0295] That is, in the embodiment, the computing system 1000 can perform MoE learning based on the plurality of aforementioned specialized models (SMs) and the router (RT)-coupled MoELM.
[0296] At this time, by performing learning, the computing system 1000 can realize learning for each of the plurality of specialized models (SMs) included in the MoELM.
[0297] In other words, by performing the learning as described above, the plurality of specialized models (SMs) in the MoELM can be learned respectively.
[0298] Here, in other words, the specialized model (SM) according to the embodiment is an artificial intelligence model for which optimal learning for a specific purpose has been performed, and can mean an artificial intelligence model learned by training data and methods specialized for the corresponding purpose.
[0299] In the embodiment, such a specialized model (SM) may include a learned predetermined sLLM (including MoELM and / or DMoE model), a general MoE model (NM), an external model (EM), and / or a specialized module model (MM) according to the embodiments disclosed below.
[0300] Also, the computing system 1000 according to an embodiment acquires specialized model characteristic information (SMFI) by MoE learning (S303).
[0301] Here, the specialized model characteristic information (SMFI) according to the embodiment may mean information for specifying the role and / or function of a predetermined specialized model (SM).
[0302] Specifically, referring further to FIG. 8, in the embodiment, the computing system 1000 may further include a model specialization module (MSM) according to an embodiment.
[0303] And the computing system 1000 can obtain the aforementioned specialized model feature information (SMFI) in conjunction with the model specialization module (MSM).
[0304] Here, the model specialization module (MSM) according to an embodiment may be an artificial intelligence module that generates and outputs specialized model feature information (SMFI) corresponding to a predetermined specialized model (SM) based on MoE learning.
[0305] Specifically, in an embodiment, the model specialization module (MSM) can monitor and track the task allocation status of the router (RT) for each specialized model (SM) when the aforementioned MoE learning is performed.
[0306] That is, in an embodiment, the model specialization module (MSM) can grasp how the router (RT) allocates and assigns tasks to which specialized models (SM) when the MoELM is learned and operated.
[0307] According to an embodiment, the model specialization module (MSM) can also generate a tag for identifying each tracked task allocation status and perform matching management.
[0308] Thereby, in an embodiment, the model specialization module (MSM) can determine the specialization for each of the plurality of specialized models (SM).
[0309] In addition, in an embodiment, the model specialization module (MSM) generates specialized model feature information (SMFI) corresponding to each specialized model (SM) based on the determined specialization for each specialized model (SM).
[0310] FIG. 15 is a diagram showing an example of specialized model feature information (SMFI) according to an embodiment.
[0311] Here, as shown in FIG. 15, as an embodiment, the model specification module (MSM) can generate the aforementioned specialized model characteristic information (SMFI) in at least one of the following forms.
[0312] [First form] Specialized model characteristic information (SMFI) in the form of selecting any one category of the specialized model (SM) role and / or function specific category (for example, question and answer or device control, etc.) that has already been set according to user input
[0313] [Second form] Specialized model characteristic information (SMFI) in the form of specifying the role and / or function of the specialized model (SM) in natural language form
[0314] [Third form] Specialized model characteristic information (SMFI) in the form of specifying the role and / or function of the specialized model (SM) in at least one of the first form and the second form, and further defining the input data and output data of the specialized model (SM)
[0315] Subsequently, in the embodiment, the model specification module (MSM) can provide the specialized model characteristic information (SMFI) generated as described above to the computing system 1000 as output data.
[0316] Therefore, in the embodiment, the computing system 1000 can obtain characteristic information for each specialized model (SM) by interacting with the model specification module (MSM).
[0317] In addition, the computing system 1000 according to an embodiment generates a specialized module model (MM: Specialized Module Module) based on the obtained specialized model characteristic information (SMFI) (S305).
[0318] Here, a specialized module model (MM) according to an embodiment may mean a specialized model (SM) in which predetermined specialized model characteristic information (SMFI) is matched and independently separated.
[0319] Specifically, in an embodiment, the computing system 1000 matches the specialized model characteristic information (SMFI) obtained as described above to a corresponding specialized model (SM).
[0320] Also, in an embodiment, the computing system 1000 independently separates and databases the specialized models (SM) to which the specialized model characteristic information (SMFI) is matched.
[0321] That is, in an embodiment, the computing system 1000 matches the corresponding specialized model characteristic information (SMFI) to each specialized model (SM), and performs modularization to separately classify, store, and manage them.
[0322] Therefore, the computing system 1000 can generate a specialized module model (MM) that is a specialized model (SM) that is independently separated while the specialized model characteristic information (SMFI) is matched.
[0323] In this way, in an embodiment, the computing system 1000 grasps the different characteristics of the specialized models (SM) within a given MoE model (in the embodiment, MoELM), and reflects this to modularize each specialized model (SM) into a small size that can be reused and shared.
[0324] Thereby, the computing system 1000 can quickly and efficiently select and choose specialized models (SM) that realize a data processing process optimized for a specific domain with higher accuracy, and can easily support flexible expansion or contraction of the MoE model based on this.
[0325] Also, a computing system 1000 according to an embodiment can acquire predetermined domain information (S307).
[0326] Here, the domain information according to the embodiment can be information that defines a domain that specifies data, rules, terms, problem definitions, and / or processes, etc. used by a predetermined AI system to perform a predetermined task.
[0327] Specifically, in the embodiment, the computing system 1000 acquires predetermined input data (for example, text, voice, image, video, and / or specific sensor-based sensing data, etc.).
[0328] Also, in the embodiment, the computing system 1000 determines the domain corresponding to the acquired input data.
[0329] Here, in the embodiment, the method by which the computing system 1000 determines the domain for the input data may be performed based on various disclosed algorithms capable of executing this, and in the embodiments of the present invention, the corresponding algorithms themselves are not limited or restricted.
[0330] Therefore, the implementation system 1000 can acquire domain information for the task to be processed.
[0331] Also, a computing system 1000 according to an embodiment determines a domain-specialized model based on the acquired domain information (S309).
[0332] Here, the domain-specialized model according to the embodiment can mean a specialized model (SM) that executes data processing (as an embodiment, deep learning, etc.) operations optimized for a predetermined domain.
[0333] More specifically, referring further to FIG. 14, in an embodiment, the computing system 1000 determines at least one domain-specific specialized model based on the domain information and specialized model characteristic information (SMFI) obtained as described above.
[0334] More specifically, in an embodiment, the computing system 1000 can detect at least one specialized model characteristic information (SMFI) having characteristics corresponding to the obtained domain information.
[0335] For example, when the computing system 1000 confirms the "characteristics of the task of outputting response data for predetermined interrogation data" based on the first domain information, it can detect at least one specialized model characteristic information (SMFI) specified as the "role and / or function specialized for interrogation response" from among the plurality of specialized model characteristic information (SMFI) stored in a database.
[0336] Here, according to an embodiment, the computing system 1000 can detect at least one specialized model characteristic information (SMFI) corresponding to the domain information based on a plurality of tags generated by a model specification module (MSM) for each task assignment state of a router (RT) for a plurality of specialized models (SM) during learning based on the aforementioned MoE architecture.
[0337] That is, according to an embodiment, the computing system 1000 can detect at least one specialized model characteristic information (SMFI) corresponding to the relevant domain information by mutually comparing the plurality of tags generated as described above and the domain information.
[0338] Here, according to an embodiment, the computing system 1000 filters the tags to be compared according to the generation time of each tag.
[0339] Specifically, the computing system 1000 sets at least one tag generated at a specific task assignment time as a comparison target tag according to user input and / or a pre-set unique process.
[0340] Exemplarily, the computing system 1000 can set at least one tag generated for a task assignment state performed after a time set during the overall learning time as a comparison target tag, focusing on the fact that the accuracy of task assignment improves as the learning rate increases.
[0341] Therefore, the computing system 1000 can detect at least one specialized model feature information (SMFI) corresponding to the relevant domain information by comparing at least one filtered tag that guarantees higher accuracy with the domain information.
[0342] Also, in an embodiment, the computing system 1000 extracts a specialized model (SM) (i.e., a specialized module model (MM)) that matches each of the detected at least one specialized model feature information (SMFI).
[0343] And, in an embodiment, the computing system 1000 determines the extracted at least one specialized module model (MM) as a domain-specific specialized model.
[0344] Also, the computing system 1000 according to an embodiment of the present invention constructs a MoE model based on the determined domain-specific specialized model (S311).
[0345] Referring further to FIG. 14, that is, in an embodiment, the computing system 1000 can construct a model (hereinafter, DMoE model) that operates like a MoE architecture based on at least one domain-specific specialized model determined as described above.
[0346] In other words, the computing system 1000 can construct a MoE model (i.e., a DMoE model) that realizes data processing optimized for a specific domain by using at least some of a plurality of specialized models (SMs) that are modularized and of a small size (i.e., domain-specific specialized models).
[0347] Specifically, in an embodiment, the computing system 1000 can construct the aforementioned DMoE model by combining at least one domain-specific specialized model and a predetermined router (RT).
[0348] Therefore, in an embodiment, the computing system 1000 can construct a DMoE model including a domain-specific specialized model and a router (RT).
[0349] Here, according to an embodiment, the DMoE model may be included in the sLLM according to an embodiment of the present invention.
[0350] In other words, the sLLM according to an embodiment may include a DMoE model constructed according to an embodiment.
[0351] Also, the computing system 1000 according to an embodiment of the present invention provides output data based on the constructed MoE model (S313).
[0352] That is, in an embodiment, the computing system 1000 uses the DMoE model constructed as described above to provide output data (e.g., response data to a specific query and / or a control signal by a specific command word, etc.) for predetermined input data (e.g., text, voice, image, video, and / or specific sensor-based sensing data, etc.).
[0353] As described above, in the embodiment, the computing system 1000 can identify the roles and / or functions of each specialized model (SM), and at the same time, separate and modularize them to a reusable and sharable level, and utilize this to quickly and flexibly construct a customized MoE model optimized for a specific domain (i.e., DMoE model), and can provide predetermined output data by efficient task processing using the constructed model.
[0354] In other words, in the embodiment, the computing system 1000 can realize and provide an MoE model with further improved data processing (and / or computing) speed and inference performance, and can support various services thereby, so that the performance and quality improvement can be effectively achieved.
[0355] -MoE Application LLM-Based AI Agent Provision Method
[0356] Hereinafter, a method for realizing a model providing service based on an MoE architecture that provides an on-device specialized AI agent (Artificial Intelligence Agent) that determines an application model optimized for a domain by an external environment based on an LLM (Large Language Model) applying MoE (Mixture of Experts) in an embodiment of the present invention and executes an output based on the determined application model will be described in detail with reference to the accompanying drawings.
[0357] FIG. 16 is a flowchart for explaining an MoE application LLM-based AI agent provision method according to an embodiment, and FIG. 17 is a conceptual diagram for explaining an MoE application LLM-based AI agent provision method according to an embodiment.
[0358] As shown in FIGS. 16 and 17, a method for realizing a model providing service based on a MoE architecture that provides a specialized model (AIAM) for on-device specialized AI agents that determines an application model optimized for a domain by an external environment based on an LLM applying MoE and performs an output based on the determined application model may include steps of executing an on-device AI agent service (S401), obtaining predetermined input data (S403), determining a domain based on the obtained input data (S405), determining an application model based on the determined domain (S407), and providing output data based on the determined application model (S409).
[0359] Specifically, a computing system 1000 according to an embodiment of the present invention executes an on-device AI agent service (S401).
[0360] Here, for reference, on-device AI may mean a technology that directly executes artificial intelligence-based data processing inside a user's device, rather than in the cloud and / or an external server. This can provide advantages such as personal information protection, real-time processing, and reduced dependence on an Internet connection, since all processing is completed inside the device without sending data externally.
[0361] Therefore, in this context, the on-device AI agent service may mean various services realized by utilizing on-device AI.
[0362] Exemplarily, the on-device AI agent service may include a voice assistant service for a smartphone (e.g., Google Assistant, Apple Siri, Samsung Bixby, etc.), a smart camera service (e.g., HDR+ of Google Pixel, Deep Fusion of Apple, etc.), a fitness tracker and a smartwatch service (e.g., Apple Watch, Fitbit, etc.), an automobile autonomous driving service (e.g., Autopilot of Tesla, etc.) and / or a home security service (e.g., Nest Secure, Ring, etc.).
[0363] In an embodiment, the computing system 1000 can execute a predetermined on-device AI agent service based on the linkage with an AI agent specialization model (AIAM) according to an embodiment and / or a predetermined application, etc.
[0364] Also, the computing system 1000 according to an embodiment of the present invention acquires predetermined input data (S403).
[0365] Specifically, in an embodiment, the computing system 1000 can acquire at least one input data (e.g., predetermined text, voice, image, video and / or sensing data, etc.) based on user input based on the on-device AI agent service executed as described above and / or the linkage with an external device (e.g., a predetermined sensor, etc.).
[0366] In an embodiment, the input data acquired as described above may include predetermined data that can identify the target task of data processing.
[0367] Also, the computing system 1000 according to an embodiment determines the domain by the acquired input data (S405).
[0368] Here, in other words, the domain according to the embodiment may mean data, rules, terms, problem definitions, and / or processes used by a given AI system to perform a specific task, and the like.
[0369] Specifically, in the embodiment, the computing system 1000 determines the domain corresponding to the acquired input data.
[0370] Here, in the embodiment, the method by which the computing system 1000 determines the domain for the input data may be performed based on various disclosed algorithms capable of doing so. In one embodiment, the corresponding algorithm itself is not limited or restricted.
[0371] Therefore, in the embodiment, the computing system 1000 can obtain the domain information corresponding to the task to be processed.
[0372] Also, the computing system 1000 according to one embodiment determines an application model according to the determined domain (S407).
[0373] Here, the application model according to the embodiment may mean a model that performs a predetermined task process with the given input data.
[0374] In the embodiment, such an application model may be at least one model of the secondary model (S) as described above.
[0375] Here, in other words, the secondary model (S) according to the embodiment may mean a model that can execute a specific task by the control and management of a master model (P) (that is, an orchestrator (OCT) and / or a router (RT), etc.) responsible for the control and management of a given AI system operation.
[0376] In an embodiment, such a secondary model (S) may include at least one model of a small language model (sLLM) (including a MoELM and / or a DMoE model), a general MoE model (NM), an external model (EM), and / or a specialized model (SM) (including a specialized module model (MM)).
[0377] Specifically, in an embodiment, the computing system 1000 determines at least one application model based on the domain information obtained as described above.
[0378] More specifically, as an embodiment, the computing system 1000, in conjunction with a master model (P) according to an embodiment (i.e., an orchestrator (OCT) and / or a router (RT), etc.), detects at least one or more models that execute data processing (such as deep learning, etc.) operations optimized for the given domain information among the aforementioned secondary models (S) (i.e., domain-specific models).
[0379] Here, the specific method by which the computing system 1000 detects a domain-specific model in conjunction with the master model (P) in an embodiment is omitted by applying mutatis mutandis the descriptions regarding the router (RT) and the orchestrator (OCT) disclosed in the aforementioned "AI Agent Specialized Model (AIAM)".
[0380] Also, in an embodiment, the computing system 1000 determines the at least one detected domain-specific model as an application model.
[0381] Also, the computing system 1000 according to an embodiment provides output data based on the determined application model (S409).
[0382] That is, in the embodiment, the computing system 1000 can generate and provide output data (e.g., response data to specific queries and / or control signals by specific command words, etc.) for predetermined input data (e.g., text, voice, image, video, and / or specific sensor-based sensing data, etc.) based on at least one application model determined through the AI agent specialization model (AIAM) as described above.
[0383] In other words, the computing system 1000 executes a predetermined request task based on the given input data using the application model determined as described above, and provides output data by the executed data processing.
[0384] Here, in the embodiment, the computing system 1000 can provide the output data based on the on-device AI agent service described above.
[0385] As described above, in the embodiment, the computing system 1000 can effectively determine a model optimized for data processing according to the given domain even in the on-device environment based on an AI agent specialization model (AIAM) including models (as embodiments, MoELM, DMoE model, and / or specialized module model (MM), etc.) realized by applying the MoE architecture by various embodiments, and provide an output by efficient data processing through the determined model.
[0386] That is, the computing system 1000 can realize and provide an artificial intelligence model (i.e., an AI agent specialization model (AIAM)) that can better listen to, understand, execute, and answer the given task in any environment.
[0387] Accordingly, in an embodiment, the computing system 1000 can directly and significantly improve the quality and performance of various AI agent-based services (such as smartphone voice assistant services, smart camera services, fitness tracker and smartwatch services, automobile autonomous driving services, and / or home security services, etc.).
[0388] The embodiments according to the present invention described above can be realized in the form of program instruction words executable by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instruction words, data files, data structures, etc. alone or in combination. The program instruction words recorded on the computer-readable recording medium may be those specially designed or configured for the present invention or those known and usable by those skilled in the computer software field. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instruction words such as ROM, RAM, flash memory, etc. Examples of program instruction words include not only machine language code generated by a compiler but also high-level language code executable by a computer using an interpreter or the like. A hardware device can be changed into one or more software modules for performing the processing according to the present invention, and vice versa.
[0389] The specific implementation described in the present invention is one embodiment and does not limit the scope of the present invention in any way. For the sake of brevity of the specification, descriptions of conventional electronic configurations, control systems, software, and other functional aspects of the system can be omitted. Also, the connections or connection members of the lines between the components shown in the drawings exemplarily show functional connections and / or physical or circuit connections, and in actual devices, they can be alternative or shown as additional diverse functional connections, physical connections, or circuit connections. Further, without specific references such as "essential" or "importantly", components may not necessarily be required for the application of the present invention.
[0390] In addition, in the detailed description of the present invention described above, the preferred embodiments of the present invention have been described with reference thereto. However, those skilled in the art or those with ordinary knowledge in the technical field will understand that the present invention can be variously modified and changed within the scope not departing from the idea and technical field of the present invention described in the claims hereinafter. Therefore, the technical scope of the present invention should not be limited to the content described in the detailed description of the specification, but should be determined by the claims.
Explanation of Reference Numerals
[0391] 10 Dialogue Graph Generation Module 20 Dialogue Act Group Sampling Module 30 Dialogue Act Group Adjustment Module 40 Dialogue Act Group Selection Module 50 Dialogue Act Selection Module 110, 180, 181, 182, 183 User Computing Device 111, 131, 151 Processor 112, 132, 152 Memory 113, 133, 153 Data 114, 134, 154 Instruction Words 120, 140 Machine Learning Model 121 User Input Component 160 Model Trainer 161 Training Data 100, 200, 400 Computing Devices 300 Neuromorphic Circuit 310 Presynaptic Neuron Circuit 311 Presynaptic Line 320 Postsynaptic Neuron Circuit 330 Synapse Circuit
Claims
1. A task execution method based on the context of goal-oriented dialogue through an interaction model by a computing system including a memory and a processor, comprising: receiving a user interaction input; determining and providing a response interaction act for the user interaction input based on an interaction graph; analyzing a series of goal-oriented dialogue data including the user interaction input and the response interaction act to determine the context of the goal-oriented dialogue; determining the type of task requested by the user based on the context of the goal-oriented dialogue; and performing the task whose type has been determined, A task execution method based on the context of goal-oriented dialogue.
2. In the step of determining the context of the goal-oriented dialogue, extracting a plurality of keywords from the data of the goal-oriented dialogue, and analyzing at least any one of the correlation relationship between a plurality of interaction acts included in the goal-oriented dialogue, the intention of the goal-oriented dialogue, and the goal based on the plurality of keywords to determine the context. The task execution method based on the context of goal-oriented dialogue according to claim 1.
3. further comprising determining at least one task execution model optimized for the task whose type has been determined based on the context of the goal-oriented dialogue among a plurality of task execution models, In the step of performing the task, using the determined at least one task execution model to perform the task. The task execution method based on the context of goal-oriented dialogue according to claim 1.
4. In the step of determining the type of the task and the step of performing the task, the computing system determines the type of the task and supports a predetermined operation required to perform the task. The task execution method based on the context of goal-oriented dialogue according to claim 1.
5. In the step of performing the task, analyzing the programming code related to the task, generating programming code for performing the task based on the analysis result, and executing the programming code to perform the task. The task execution method based on the context of goal-oriented dialogue according to claim 1.
6. further comprising, by receiving the user interaction input, capturing a screen of an electronic device used by the user to obtain a user screen screenshot. The method for task execution based on the context of goal-oriented dialogue according to claim 1, wherein in the step of determining the type of the task, the type of the task is determined based on the information regarding the context of the determined goal-oriented dialogue and the analysis of the user screen screenshot.
7. The step of determining and providing the response dialogue act includes: generating a dialogue graph modeling at least one conditional relationship for a dialogue dataset; sampling a plurality of groups of dialogue acts for responding to the user dialogue input by using a pre-trained dialogue model; adjusting the plurality of groups of dialogue acts based on the dialogue graph; and selecting any one group of dialogue acts that satisfies a predetermined condition among the plurality of groups of dialogue acts. The method for task execution based on the context of goal-oriented dialogue according to claim 1.
8. The at least one conditional relationship includes at least any one of a first conditional relationship regarding what utterance should be made for a single utterance, a second conditional relationship regarding what utterance can be made for a single utterance, and a third conditional relationship regarding what utterance should not be made for a single utterance in the flow of the dialogue. The method for task execution based on the context of goal-oriented dialogue according to claim 7.
9. In the step of selecting any one group of dialogue acts, selecting any one group of dialogue acts that most satisfies the at least one conditional relationship among the any one group of dialogue acts. The method for task execution based on the context of goal-oriented dialogue according to claim 7.
10. Each of the plurality of groups of dialogue acts includes at least one dialogue act for the user dialogue input. The method for task execution based on the context of goal-oriented dialogue according to claim 8.
11. In the step of adjusting the plurality of groups of dialogue acts, for each of the plurality of groups of dialogue acts, adding a dialogue act that satisfies the first conditional relationship among at least one dialogue act, removing a dialogue act that does not satisfy the second conditional relationship, and removing a dialogue act that does not satisfy the third conditional relationship. The method for task execution based on the context of goal-oriented dialogue according to claim 8.
12. In the step of selecting any one group of dialogue acts, The task execution method based on the context of goal-oriented dialogue according to claim 11, wherein among the plurality of dialogue act groups, a dialogue act group that contains the most dialogue acts satisfying the at least one conditional relationship is selected.
13. The task execution method based on the context of goal-oriented dialogue according to claim 7, further comprising a response dialogue act determination step of providing any one dialogue act included in any one of the selected dialogue act groups as a response to the user dialogue input.
14. In the response dialogue act determination step, The task execution method based on the context of goal-oriented dialogue according to claim 13, wherein among at least one dialogue act included in any one of the selected dialogue act groups, a dialogue act having the highest relevance to the user dialogue input is provided as a response to the user dialogue input.
15. At least one memory, and At least one processor that reads at least one instruction word stored in the memory and performs a task execution method based on the context of goal-oriented dialogue, The at least one processor Receives user dialogue input, Determines and provides a response dialogue act to the user dialogue input based on a dialogue graph, Analyzes a series of goal-oriented dialogue data including the user dialogue input and the response dialogue act to determine the context of the goal-oriented dialogue, Determines the type of task required by the user based on the context of the goal-oriented dialogue, A task execution system based on the context of goal-oriented dialogue that performs the determined task.
16. An electronic device that receives user dialogue input, and A computing device including at least one memory and at least one processor that reads at least one instruction word stored in the at least one memory and performs a task execution method based on the context of goal-oriented dialogue, The at least one processor Receives user dialogue input, Determines and provides a response dialogue act to the user dialogue input based on a dialogue graph, Analyzes a series of goal-oriented dialogue data including the user dialogue input and the response dialogue act to determine the context of the goal-oriented dialogue, Determines the type of task required by the user based on the context of the goal-oriented dialogue, A task execution system based on the context of goal-oriented dialogue that performs the determined task.
17. The electronic device according to claim 16, which receives the user interaction input in at least one of forms of text, voice, gesture, and touch, is a task execution system based on the context of goal-oriented dialogue.
Citation Information
Patent Citations
Real scene information classification method and system for task-oriented dialogue
CN113901213A
Systems and methods for assisting agents through artificial intelligence
JP2022525362A
Entity-Level Data Augmentation in Chatbots for Robust Named Entity Recognition
JP2023530423A
System and method for generating dialogue graphs
US20200004878A1
Utilizing rule specificity in conversational ai
US20200227029A1