Task-oriented dialogue method and system
The task-oriented dialogue system addresses limitations in dialogue model diversity and resource management by using a dialogue graph to optimize response prediction and task execution, ensuring context-specific and efficient user interactions.
Patent Information
- Application Number
- JP2024212171
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-29
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-08
AI Technical Summary
Existing dialogue models struggle with limitations in generating diverse dialogue flows without human intervention and face challenges in efficiently managing large datasets and resources, particularly in on-device environments, leading to difficulties in providing context-specific AI analysis performance.
A task-oriented dialogue method and system that utilizes a dialogue graph generation model to generate a dialogue graph modeling conditional relationships for a dialogue dataset, adjusting and selecting dialogue act groups to provide context-specific responses, optimizing resource usage and enhancing response prediction performance.
The system provides more accurate and efficient task execution by selecting dialogue acts that satisfy predetermined conditional relationships, improving response prediction and resource management, thus enhancing user interaction and task completion.
Smart Images

Figure 2025102693000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a goal-oriented dialogue method and system, and more particularly, to a goal-oriented dialogue method and system that improve the response prediction performance of a dialogue model for user dialogue input by utilizing a dialogue graph modeling a predetermined conditional relationship for a dialogue dataset.
Background Art
[0002] Recently, as artificial intelligence (AI) technology has developed, various services utilizing artificial intelligence in diverse industrial fields have been commercialized. Such artificial intelligence technology enables an artificial neural network model to learn a vast amount of data and output and provide the information desired by a user.
[0003] On the other hand, there is known a dialogue-type artificial intelligence model that can interact with a user, rather than simply having an artificial intelligence model learn data patterns and provide an output value for an input value. An artificial intelligence secretary utilizing such a dialogue-type artificial intelligence model can be applied to a smart device to serve as a personal secretary, can be applied to a chatbot that answers questions from corporate customers, and can be applied to a smart home system that can control the operation of home appliances connected to the Internet of Things according to user requests. As a result, the dialogue-type artificial intelligence model can provide a more convenient and user-friendly experience to users in various technical fields.
[0004] In addition, research on artificial intelligence technology capable of performing tasks requested by a user has been actively conducted, and for this purpose, applications to dialogue-type artificial intelligence models of an analysis method based on a task-oriented dialogue graph have been attempted many times.
[0005] However, most of the research related to task-oriented dialogue graphs has remained at the level of trying to generate dialogue models using dialogue graphs directly drawn by humans or leveraging rule-based systems that automatically infer dialogue graphs based on already learned dialogue policies. As a result, it has not been possible to provide a dialogue graph-based dialogue model automatically constructed based solely on dialogue datasets without human intervention, or there are limitations in modeling diverse dialogue flows.
[0006] On the other hand, generally, artificial intelligence is realized through a large number of AI models and deep learning based on them.
[0007] Such artificial intelligence has been developed to provide diverse services by considering the user's context (such as context, environment, and / or intention, etc.).
[0008] However, when trying to process a specific task based on large-capacity data, there are limitations in that the required computing costs and time are considerable.
[0009] As a result, there are also certain restrictions in the use of AI models in the recently spotlighted on-device environment.
[0010] To solve this, conventionally, model architectures such as MoE (Mixture of Experts) have been utilized.
[0011] Here, MoE refers to the architecture of a machine learning model that combines a number of expert models to solve complex problems.
[0012] Such a MoE may include an expert model, which is a plurality of small networks designed to learn different parts and / or different characteristics of predetermined data and perform corresponding data processing operations, and a gating network that evaluates the performance of each expert model and determines which expert model is most suitable for assigning a specific task according to the predetermined data based on this evaluation.
[0013] Therefore, according to the MoE architecture, the gating network that obtains the predetermined input data determines a probabilistic or deterministic task assignment for each expert model, and the selected expert models perform their respective tasks and return the results to perform data processing for a specific task.
[0014] By utilizing such a MoE, the AI model can enhance the overall efficiency and performance by activating only specific parts and concentrating computing resources in cases such as handling complex tasks or large datasets.
[0015] However, in the case of conventional MoEs, not only is a high level of VRAM required, but there are also quite a few issues that need to be solved in the fine-tuning process.
[0016] In addition to this, the conventional MoE method is for efficiently managing large-sized models and has limitations in supporting the efficiency of the remaining resources that are not activated according to a given task.
[0017] Also, in the conventional art in this technical field, services are mostly provided using generally realized AI models, but there is a problem that it is difficult to quickly and easily ensure the AI analysis performance most suitable for a given context.
Summary of the Invention
Problems to be Solved by the Invention
[0018] According to various embodiments of the present disclosure, there is provided a task-oriented dialogue method and system in which response prediction performance is improved by utilizing a task-oriented dialogue graph generated by a dialogue graph generation model based on a dialogue dataset.
[0019] According to various embodiments of the present disclosure, there is provided a task-oriented dialogue method and system that can respond to various task requests of a user and execute the corresponding task by utilizing a dialogue model based on a task-oriented dialogue graph generated by a dialogue graph model.
[0020] However, the technical problems to be achieved by various embodiments of the present disclosure are not limited to the above technical problems, and other technical problems may exist.
Means for Solving the Problems
[0021] One embodiment is a task-oriented dialogue method provided by a computing system including a memory and a processor by utilizing a dialogue model, the method including: generating a dialogue graph modeling at least one conditional relationship for a dialogue dataset; receiving a user dialogue input; sampling a plurality of dialogue act groups for responding to the user dialogue input by utilizing a pre-trained dialogue model; adjusting the plurality of dialogue act groups based on the dialogue graph; and selecting any one dialogue act group satisfying a predetermined condition from the plurality of dialogue act groups.
[0022] In another aspect, the at least one conditional relationship may include at least any one of a first conditional relationship regarding what utterance should be made for a single utterance in the flow of a dialogue, a second conditional relationship regarding what utterance can be made for a single utterance, and a third conditional relationship regarding what utterance should not be made for a single utterance.
[0023] On the other hand, in the step of selecting the plurality of dialogue act groups, any one dialogue act group that most satisfies the at least one conditional relationship among the plurality of dialogue act groups may be selected.
[0024] On the other hand, each of the plurality of dialogue act groups may include at least one dialogue act with respect to the user dialogue input.
[0025] On the other hand, in the step of adjusting the plurality of dialogue act groups, for each of the plurality of dialogue act groups, a dialogue act that satisfies the first conditional relationship among at least one dialogue act may be added, a dialogue act that does not satisfy the second conditional relationship may be removed, and a dialogue act that does not satisfy the third conditional relationship may be removed.
[0026] On the other hand, in the step of selecting any one of the dialogue act groups, among the plurality of dialogue act groups, a dialogue act group that includes the most dialogue acts that satisfy the at least one conditional relationship may be selected.
[0027] On the other hand, the goal-oriented dialogue method may further include a response dialogue act determination step of providing, as a response to the user dialogue input, any one dialogue act included in any one of the selected dialogue act groups.
[0028] On the other hand, in the response dialogue act determination step, a dialogue act having the highest relevance to the user dialogue input among at least one dialogue act included in any one of the already selected dialogue act groups may be provided as a response to the user dialogue input.
[0029] On the other hand, in the step of generating the dialogue graph, a dialogue graph generation model learned so that the expected value is maximized when the n-th dialogue act in the dialogue context satisfies the first conditional relationship based on various dialogue acts so that the first conditional relationship is modeled for the dialogue dataset may be utilized to generate the dialogue graph.
[0030] On the other hand, in the step of generating the dialogue graph, a dialogue graph generation model learned so that the expected value is maximized when the n-th dialogue act in the dialogue context satisfies the second conditional relationship and does not satisfy the third conditional relationship at the same time based on various dialogue acts so that the second conditional relationship and the third conditional relationship are modeled for the dialogue dataset may be utilized to generate the dialogue graph.
[0031] On the other hand, in the step of generating the dialogue graph, a first dialogue graph may be generated based on a first type of dialogue dataset, and a second dialogue graph may be generated based on a second type of dialogue dataset.
[0032] On the other hand, the goal-oriented dialogue method further includes steps of determining the type of the user dialogue input, and determining a dialogue graph corresponding to the type of the user dialogue input among the first dialogue graph and the second dialogue graph, and in the step of adjusting the plurality of dialogue act groups, the plurality of dialogue act groups may be adjusted based on the determined dialogue graph.
[0033] On the other hand, the goal-oriented dialogue method may further include steps of analyzing a series of goal-oriented dialogue data including the user dialogue input and the selected one of the dialogue act groups to determine the context of the goal-oriented dialogue, determining the type of task required by the user based on the context of the goal-oriented dialogue, and performing the task whose type has been determined.
[0034] On the other hand, in the step of determining the context of the goal-oriented dialogue, a plurality of keywords may be extracted from the data of the goal-oriented dialogue, and at least one of the correlation relationships between a plurality of dialogue acts included in the goal-oriented dialogue, the intention of the goal-oriented dialogue, and the goal may be analyzed based on the plurality of keywords to determine the context.
[0035] One embodiment includes at least one memory and at least one processor that reads at least one instruction stored in the memory to execute a goal-oriented dialogue method. The at least one processor generates a dialogue graph modeling at least one conditional relationship for a dialogue dataset, receives a user dialogue input, samples a plurality of dialogue act groups for responding to the user dialogue input using a pre-learned dialogue model, adjusts the plurality of dialogue act groups based on the dialogue graph, and selects any one dialogue group that meets a predetermined condition among the plurality of dialogue act groups, providing a goal-oriented dialogue system.
[0036] One embodiment includes an electronic device that receives a user dialogue input and at least one processor that generates a dialogue graph modeling at least one conditional relationship for a dialogue dataset and samples a plurality of dialogue act groups for responding to the user dialogue input using a pre-learned dialogue model. The at least one processor adjusts the plurality of dialogue act groups based on the dialogue graph and selects any one dialogue act group that meets a predetermined condition among the plurality of dialogue act groups, providing a goal-oriented dialogue system.
[0037] On the other hand, the at least one processor may determine the type of task required by the user based on the user dialogue input and any one selected dialogue act group, and perform the determined task.
[0038] On the other hand, the electronic device may receive the user interaction input in at least one of the forms of text, voice, gesture, and touch.
Advantages of the Invention
[0039] According to various embodiments of the present disclosure, by verifying the response dialogue act sampled by the already learned dialogue model based on the dialogue graph in which the dialogue graph generation model models a predetermined conditional relationship with respect to the dialogue dataset, the sampled response dialogue act is adjusted to satisfy the predetermined conditional relationship included in the dialogue graph, and an object-oriented dialogue method and system capable of generating a more appropriate response dialogue act for the user interaction input and providing it to the user can be provided.
[0040] According to various embodiments of the present disclosure, by utilizing a dialogue model based on the object-oriented dialogue graph generated by the dialogue graph model, a more appropriate response dialogue act can be provided for various task requests of the user, and an object-oriented dialogue method and system capable of efficiently executing the task requested by the user of the type determined based on the user interaction input and the response dialogue act can be provided.
[0041] However, the effects obtained by various embodiments of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood from the following description.
Brief Description of the Drawings
[0042]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying Out the Invention
[0043] Since the present invention can be subjected to various transformations and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. The effects and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described in detail later together with the drawings. However, the present invention is not limited to the embodiments disclosed below and can be realized in various forms. In the following embodiments, terms such as first and second are not used in a limiting sense and are used to distinguish one component from another. Also, the singular expression includes plural expressions unless the context clearly has a different meaning. Also, terms such as "including" or "having" mean that the features or components described in the specification exist, and do not preclude in advance the possibility that one or more other features or components may be added. Also, in the drawings, the sizes of the components may be exaggerated or reduced for convenience of explanation. For example, since the sizes and thicknesses of each configuration in the drawings are arbitrarily shown for convenience of explanation, the present invention is not necessarily limited to the illustrated content.
[0044] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. When explaining with reference to the drawings, the same or corresponding components are given the same reference numerals, and redundant explanations thereof are omitted.
[0045] - System 1000 for providing an end-to-end dialogue service
[0046] The system 1000 according to one embodiment generates a task-oriented dialogue graph based on the analysis of the dialogue dataset, and selects the final response dialogue act by verifying and adjusting various response dialogue acts sampled by the already learned dialogue model for the user dialogue input based on the generated dialogue graph, and provides this as a response to the user dialogue input.
[0047] In this case, the system 1000 models a predetermined conditional relationship for the dialogue dataset to generate a dialogue graph, and can provide a more reliable response dialogue act to the user by verifying and adjusting the response dialogue acts sampled based on the predetermined conditional relationship included in the dialogue graph.
[0048] Also, the system 1000 determines the type of task required by the user based on the user dialogue input and the finally generated response dialogue act, and can provide a more convenient task processing experience to the user by efficiently executing the task based on the determined task type.
[0049] FIG. 1 shows an example of a block diagram of a computing system 1000 that realizes a task-oriented dialogue service and a task execution service according to one embodiment.
[0050] As shown in FIG. 1, a computing system 1000 that realizes a task-oriented dialogue service and a task execution service according to one embodiment includes a user computing device 110, a server computing system 130, and a training computing system 150, and the devices can communicate via a network 170.
[0051] According to one embodiment, the goal-oriented dialogue method may be realized and provided locally by the user computing device 110, may be realized and provided in the form of a web service by the server computing system 130 communicating with the user computing device 110, or may be realized and provided by the user computing device 110 and the server computing system 130 in cooperation with each other.
[0052] Here, in the embodiment, the user computing device 110 and / or the server computing system 130 can learn the machine learning model 120 and / or 140 through interaction with the training computing system 150 communicatively connected via the network 170. The training computing system 150 may be separate from the server computing system 130 or may be a part of the server computing system 130.
[0053] And here, the artificial intelligence model can be directly learned locally by the user computing device 110, can be learned through interaction between the server computing system 130 and the user computing device 110 via the network 170, or can be learned by a separate training computing system 150 using various training techniques and learning techniques. And it can also be realized in such a way that the artificial intelligence model learned by the training computing system 150 is transmitted to and provided / updated to the user computing device 110 and / or the server computing system 130 via the network 170.
[0054] In some embodiments, the training computing system 150 may be a part of the server computing system 130 or a part of the user computing device 110.
[0055] FIG. 2 is a conceptual diagram for explaining a method by which a computing system 1000 according to an embodiment performs a task requested by a user based on a goal-oriented dialogue service.
[0056] As shown in FIG. 2, the computing system 1000 can receive various forms of user dialogue input, provide appropriate response dialogue actions for the received user dialogue input, and further perform a task requested by the user based on the user dialogue input and the response dialogue actions. Here, the task may include various types of tasks determined based on a goal-oriented dialogue including the user dialogue input and the response dialogue actions.
[0057] The computing system 1000 grasps the context of a series of goal-oriented dialogues consisting of user dialogue input and response dialogue actions thereto.
[0058] The computing system 1000 can analyze the pattern of a goal-oriented dialogue composed of various types of user dialogue input and response dialogue actions thereto, and based on this, determine the context related to the intention, purpose, etc. of the corresponding dialogue.
[0059] The computing system 1000 can extract a plurality of keywords included in the data of the goal-oriented dialogue, and analyze the correlation between the plurality of dialogue actions included in the dialogue, the intention, purpose, etc. of the dialogue based on the extracted plurality of keywords, and finally determine the context of the goal-oriented dialogue.
[0060] In addition, the computing system 1000 can determine the type of task requested by the user based on the information related to the context of the goal-oriented dialogue.
[0061] For example, a plurality of keywords (such as hotel, reservation, date, number of people, five-star, etc.) are extracted from the data of a series of dialogues including user dialogue input requesting a hotel reservation and response dialogue actions requesting information related to the hotel reservation (such as reservation date, number of overnight guests, hotel grade, etc.).
[0062] Also, based on the plurality of extracted keywords, at least one of the correlation relationship between the user dialogue behavior included in the corresponding dialogue and the response dialogue behavior provided by the computing system 1000, the intention of the corresponding dialogue, and the purpose can be analyzed to determine the context of the corresponding dialogue.
[0063] For example, according to the keyword-based context determination process, the context of a dialogue including a plurality of dialogue behaviors such as "Please make a hotel reservation", "What is the check-in date?", "Check-in on December 31, 2024, and check-out on January 3, 2025", "How many people will stay?", and "Three people" is determined that the user requests a hotel reservation and thus an automatic hotel reservation service is required.
[0064] Furthermore, based on the determined context of the dialogue, the type of task requested by the user can be determined.
[0065] For example, based on the context of the dialogue determined that an automatic hotel reservation service is required in response to the user's hotel reservation request, the task requested by the user is determined to be "hotel reservation".
[0066] On the other hand, the number of types of tasks determined based on the context of goal-oriented dialogue may be plural. The goal-oriented dialogue may include various types of user dialogue inputs and various response dialogue behaviors corresponding thereto, and the context of such goal-oriented dialogue may be related to various tasks. Accordingly, based on the context of the goal-oriented dialogue according to one embodiment, the types of a plurality of tasks can be determined.
[0067] The system 1000 executes the task whose type is determined based on the goal-oriented dialogue including the user dialogue input and the response dialogue behavior, and provides the execution result to the user.
[0068] For example, when the task type is determined to be "hotel reservation", the system 1000 can complete a hotel reservation task that responds to the user's request based on data related to a series of goal-oriented dialogues.
[0069] In this case, the system 1000 determines the task type and supports the predetermined operations required to perform the task with the determined type.
[0070] For example, the operating system of the system 1000 can control the processors 111 and 131 to directly determine the task type and execute the predetermined operations required to perform the task with the determined type.
[0071] Also, for example, the operating system of the system 1000 can provide a predetermined API (Application Programming Interface) and / or SDK (Software Development Kit) to determine the task type and support the predetermined operations required to perform the task with the determined type.
[0072] The system 1000 generates programming code required to perform the task with the determined type based on the context of the goal-oriented dialogue, and executes this to perform the task. Here, the programming code may be code created using a programming language (for example, Java, Python, JavaScript, etc.) to perform a specific task.
[0073] On the other hand, the system 1000 can capture the screen of the electronic device used by the user by receiving user dialogue input and obtain a user screen screenshot.
[0074] Thereafter, when determining the task type, the system 1000 can determine the task type based on the information regarding the determined context of the goal-oriented dialogue and the analysis of the user screen screenshot.
[0075] For example, when the context of the conversation is determined to be "hotel reservation" and the user screen screenshot includes the screen of the homepage that provides the hotel reservation service, the system 1000 can determine the type of the task as "hotel reservation via the corresponding hotel reservation service-providing homepage".
[0076] In this case, the system 1000 can automatically perform a series of actions (such as cursor movement, clicking, text input, etc.) necessary to perform the task determined on the user screen.
[0077] However, without being limited thereto, the tasks determined based on the goal-oriented conversation may be determined to be of countless types according to the content of the user interaction input and the response interaction actions. For example, the tasks may be determined to be of various types such as product purchase, email sending, information search, document creation, etc.
[0078] Furthermore, when the type of the task is determined based on the context of the goal-oriented conversation according to an embodiment, the system 1000 determines at least one task execution model optimized for the task whose type is determined based on the context of the goal-oriented conversation among a plurality of task execution models, and performs the task using the determined at least one task execution model.
[0079] In this case, the method by which the system 1000 determines at least one task execution model optimized for the task and performs the task using the same is substantially the same as the "MoE-based model identification method" described later, and the description thereof is omitted.
[0080] Also, for example, the system 1000 may include a first type of first user computing device 180, a second type of second user computing device 181, a third type of third user computing device 182, and a fourth type of fourth user computing device 183 that can receive various forms of user interaction inputs from the user.
[0081] Here, the user interaction input may have at least one of the forms of text, voice, gesture, and touch. However, it is not limited thereto, and the user interaction input can have various forms other than the above examples.
[0082] The first user computing device 180 of the first type may be a virtual reality electronic device, the second user computing device 181 of the second type may be a mobile electronic device, the third user computing device 182 of the third type may be an augmented reality electronic device, and the fourth user computing device 183 of the fourth type may be a desktop.
[0083] However, it is not limited thereto, and the system 1000 may include user computing devices in various forms other than the above examples that can receive user interaction input.
[0084] - User Computing Device (110:User Computing Device)
[0085] The user computing device 110 may include all other types of computing devices such as smart phones, mobile phones, digital broadcast devices, PDAs (personal digital assistants), PMPs (portable multimedia players), desktops, wearable devices, embedded computing devices, tablet PCs, augmented reality (VR) devices, and / or virtual reality (AR) devices.
[0086] Such a user computing device 110 may include at least one or more processors 111 and a memory 112. Here, the processor 111 may be composed of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, micro-controllers, microprocessors and / or at least one of other electrical units for executing functions or a plurality of electrically connected processors.
[0087] For example, the ASICs may have a structure of an array-like neuromorphic circuit including a plurality of neuron circuits.
[0088] As shown in FIG. 3, for example, the neuromorphic circuit 300 may include a plurality of presynaptic neuron circuits 310, a plurality of presynaptic lines 311 extending laterally from the plurality of presynaptic neuron circuits 310, a plurality of postsynaptic neuron circuits 320, a plurality of postsynaptic lines 321 extending longitudinally from the plurality of postsynaptic neuron circuits 320, and synapse circuits 330 provided at intersections of the plurality of presynaptic lines 311 and the plurality of postsynaptic lines 321.
[0089] The plurality of presynaptic neuron circuits 310 can transmit an externally input signal in the form of an electrical signal to the plurality of synapse circuits 330 via the plurality of presynaptic lines 311.
[0090] In addition, a plurality of postsynaptic neuron circuits 320 can receive electrical signals from a plurality of synaptic circuits 330 via a plurality of postsynaptic lines 321.
[0091] Furthermore, a plurality of postsynaptic neuron circuits 320 can also transmit electrical signals to a plurality of synaptic circuits 330 via a plurality of postsynaptic lines 321.
[0092] The plurality of synaptic circuits 330 store weight values included in a layer constituting a neural network system realized by the neuromorphic circuit 300, and can perform a predetermined operation based on the weight values and input data.
[0093] For example, each of the plurality of synaptic circuits 330 may include a resistive memory cell having a variable resistor. In this case, the resistance value of the plurality of synaptic circuits 330 changes due to the voltage applied via the plurality of presynaptic neuron circuits 310 or the plurality of postsynaptic neuron circuits 320, and weight value data due to such a resistance change can be stored.
[0094] The neuromorphic circuit 300 is formed by mimicking the neuron and synapse structures, which are essential elements of the human brain. When realizing a deep neural network (DNN) using the neuromorphic circuit 300, the data processing speed can be improved and the power consumption can be reduced compared to the case of utilizing the existing von Neumann architecture.
[0095] Memory 112 may include one or more non-transitory / temporary computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof, and may also include web storage of a server that performs the storage function of the memory on the Internet. Such memory 112 can store data 113 and instruction words 114 necessary for the at least one or more processors 111 to perform functional operations such as learning an artificial intelligence model or performing vision inspection through an artificial intelligence model.
[0096] In one embodiment, the user computing device 110 can store at least one or more machine learning models 120.
[0097] For example, the machine learning model 120 may be various machine learning models such as multiple neural networks (e.g., deep neural networks) for performing goal-oriented dialogue methods and task execution methods or non-linear models and / or linear models, or other types of machine learning models including combinations thereof.
[0098] For example, linear regression, decision trees, random forests, gradient boosting, pre-trained language models, or / and deep learning models, etc. can be stored in the machine learning model. And the neural network may include at least one or more of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or / and other forms of neural networks.
[0099] Moreover, according to various embodiments, the user computing device 110 can also store, for each process utilized in performing at least a part of the processes for the goal-oriented conversation method and task execution method, a model and a prompt template that serves as the basis for the input to the model to be executed through a large language model (LLM).
[0100] In one embodiment, the user computing device 110 receives at least one or more machine learning models 120 from the server computing system 130 via the network 170, stores them in the memory 112, and then executes the stored machine learning models 120 by the processor 111 to perform tasks such as analyzing the dialogue dataset.
[0101] In other embodiments, the server computing system 130 includes at least one or more machine learning models 140, executes operations through the machine learning models 140, and can provide goal-oriented conversation services and task execution services to the user in a manner that communicates with the user computing device 110 and the data related thereto in conjunction with the user computing device 110.
[0102] For example, the user computing device 110 can perform goal-oriented conversation services in such a way that the server computing system 130 uses the machine learning model 140 to provide an output for the user's input via the web.
[0103] Also, at least a part of the machine learning models 120 and / or 140 can be executed on the user computing device 110, and the remaining part can be executed on the server computing system 130, so that the artificial intelligence model can be realized.
[0104] In addition, the user computing device 110 may include at least one or more input components 121 that sense user input. For example, the user input component 121 may include a touch sensor (such as a touch screen and / or a touch pad, etc.) that senses the touch of a user input medium (such as a finger or a stylus), an image sensor that senses user motion input, a microphone that senses user voice input, buttons, a mouse, and / or a keyboard, etc. Further, when the user input component 121 receives input to an external controller (such as a mouse and / or a keyboard, etc.) via an interface, it may include the interface and the external controller.
[0105] - Server Computing System (130:Server Computing System)
[0106] The server computing system 130 performs a series of processes to provide a goal-oriented dialogue service.
[0107] In addition, the server computing system 130 may further perform a series of processes to provide a task execution service based on the context of goal-oriented dialogue.
[0108] Specifically, in an embodiment, the server computing system 130 can provide a goal-oriented dialogue service by exchanging data necessary for driving the goal-oriented dialogue service and task execution service processes in an external device such as the user computing device 110 with the external device.
[0109] More specifically, in an embodiment, the server computing system 130 can provide an environment in which an application for providing a goal-oriented dialogue service and a task execution service can operate in the user computing device 110.
[0110] For this purpose, the server computing system 130 may include application programs, data, and / or instruction words for the application to operate, and can transmit and receive various data based on this to the external device.
[0111] The server computing system 130 may include at least one or more processors 131 and a memory 132. Here, the processor 131 may be a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or at least one of other electrical units for executing functions or a plurality of electrically connected processors.
[0112] For example, the ASICs may have a structure of an array-shaped neuromorphic circuit including a plurality of neuron circuits (see FIG. 3).
[0113] And the memory 132 may include one or more non-temporary / temporary computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. Such a memory 132 can store the data 133 and instruction words 134 necessary for the processor 131 to perform functional operations such as learning an artificial intelligence model or performing an object-oriented dialogue method and a task execution method through the artificial intelligence model.
[0114] In one embodiment, the server computing system 130 may be implemented to include at least one or more computing devices. For example, the server computing system 130 may be implemented to operate with a plurality of computing devices in a sequential computing architecture, a parallel computing architecture, or a combination thereof. Also, the server computing system 130 may include a plurality of computing devices connected by a network 170.
[0115] Further, the server computing system 130 can store at least one or more machine learning models 140. For example, the server computing system 130 may include a neural network and / or other multi-layer non-linear models as the machine learning model 140. Exemplarily, the neural network may include a feedforward neural network, a deep neural network, a recurrent neural network, and a convolutional neural network.
[0116] In an embodiment, the server computing system 130 may further include a data store computing system (hereinafter referred to as a data store), which is a storage for continuously storing and managing the raw data underlying the goal-oriented dialogue service.
[0117] Such a data store may include various forms of data storage, ranging from file systems to cloud storage. For example, the data store may include a relational database that uses Structured Query Language (SQL) to define and manipulate data, a NoSQL database designed for flexibility and scalability to process unstructured and semi-structured data, a data warehouse used for reporting and data analysis that centralizes large volumes of data from multiple sources and optimizes it for querying and analysis, a data warehouse that stores large amounts of raw data in structured, semi-structured, and unstructured data in its basic form, and at least one database, such as a local storage device or a Network Attached Storage (NAS) that stores data in a form accessible in a general computer operating system in a file.
[0118] -Training Computing System (150:Training Computing System)
[0119] The training computing system 150 may include at least one or more processors 151 and a memory 152. Here, the processor 151 may be composed of a central processing unit (CPU), a graphics processing unit (GPU), application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, and / or at least one of other electrical units for performing functions or a plurality of electrically connected processors.
[0120] For example, the ASICs may have a structure of an array-like neuromorphic circuit including a plurality of neuron circuits (see FIG. 3).
[0121] And the memory 152 may include one or more non-transitory / temporary computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, and combinations thereof. Such a memory 152 can store the data 153 and instruction words 154 necessary for the processor 151 to execute tasks such as learning of the artificial intelligence model.
[0122] For example, the training computing system 150 may include a model trainer 160 that uses various training or learning techniques such as backpropagation of errors to train the machine learning models 120 and / or 140 stored in the user computing device 110 and / or the server computing system 130.
[0123] Exemplarily, such a model trainer 160 can perform updates to one or more parameters of the machine learning models 120 and / or 140 for the purpose-oriented dialogue service in a backpropagation manner based on a defined loss function.
[0124] In some embodiments, performing backpropagation of errors may include performing truncated back propagation through time. The model trainer 160 can execute a number of generalization techniques (e.g., weight decay, dropout, and / or knowledge distillation, etc.) to improve the generalization ability of the machine learning models 120 and / or 140 to be trained.
[0125] For example, the model trainer 160 can train the machine learning model 120 and / or 140 based on a series of training data 161. Here, the training data 161 may include data in different formats such as, for example, images, audio samples, and / or text.
[0126] Also, the training data 161 may include, for example, various types of dialogue data. In this case, the dialogue data may be data related to task-oriented dialogue that requests a specific task and provides a response to the requested task.
[0127] Examples of image types that can be used may include video frames, LiDAR point clouds, X-ray images, computed tomography scans, hyperspectral images, and / or other diverse forms of images.
[0128] Such training data 161 can be provided by the user computing device 110 and / or the server computing system 130. When the training computing device makes the machine learning models 120 and / or 140 learn based on specific data of the user computing device 110, the machine learning models 120 and / or 140 can be characterized as personalized models.
[0129] And the model trainer 160 includes computer logic utilized to provide the desired functionality.
[0130] Also, the model trainer 160 can be implemented by hardware, firmware, and / or software that controls a general-purpose processor. In one implementation example, the model trainer 160 includes program files stored in a storage device, is loaded into the memory 152, and can be executed by one or more processors 151. In other implementation examples, the model trainer 160 includes one or more sets of computer-executable data 153 and instruction words 154 stored in a computer-readable storage medium of a type such as a RAM hard disk or an optical or magnetic medium.
[0131] The network 170 includes, but is not limited to, 3GPP (Registered Trademark) (3rd Generation Partnership Project) network, LTE (Long Term Evolution) network, WIMAX (World Interoperability for Microwave Access) network, Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth (Registered Trademark) network, satellite broadcast network, analog broadcast network, and / or DMB (Digital Multimedia Broadcasting) network, etc.
[0132] Generally, communication via the network 170 may be performed using any type of wired and / or wireless connection, via various communication protocols (e.g., TCP / IP, HTTP, SMTP, and / or FTP, etc.), encoding or formats (e.g., HTML and / or XML, etc.), and / or protection schemes (e.g., VPN, secure HTTP, and / or SSL, etc.).
[0133] Figure 4 is a block diagram of a computing device 100 that implements an object-oriented dialogue service and a task execution service according to an embodiment.
[0134] As shown in Figure 4, the computing device 100 included in the user computing device 110, the server computing system 130, and the training computing system 150 includes a number of applications (for example, Application 1 to Application N). Each application may include a machine learning library and one or more machine learning models. For example, the application may include an image processing (such as Detection, Classification, and / or Segmentation, etc.) application, a text messaging application, an email application, a writing application, a virtual keyboard application, a browser application, and / or a chat-bot application, etc.
[0135] In an embodiment, the computing device 100 may include a model trainer 160 for training an artificial intelligence model, and by storing and operating the trained artificial intelligence model, it can provide output data corresponding to predetermined input data (as an embodiment, dialogue dataset, etc.).
[0136] Each application of the computing device 100 can communicate with a number of other components of the computing device 100, such as, for example, at least one or more sensors, a context manager, a device state component, and / or additional components. In one embodiment, each application can communicate with each device component using an API (for example, a public API). In one embodiment, the API used by each application may be specific to that application.
[0137] FIG. 5 is a block diagram of a computing device 200 that implements an intent-based dialogue service and a task execution service according to another embodiment.
[0138] As shown in FIG. 5, the computing device 200 includes a number of applications (e.g., Application 1 to Application N). Each application can communicate with the central intelligence layer. For example, the applications may include an image processing application, a text messaging application, an email application, a writing application, a virtual keyboard application, and / or a browser application, etc. In one embodiment, each application can use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0139] The central intelligence layer may include a number of machine learning models. For example, as shown in FIG. 5, at least a part of each machine learning model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model for all applications. In some implementations, the central intelligence layer may be included within the operating system of the computing device 200 or implemented differently therefrom.
[0140] The Central Intelligence Layer can communicate with the Central Device Data Layer. The Central Device Data Layer can be a central centralized data storage for the computing device 200. As shown in FIG. 5, the Central Device Data Layer can communicate with a number of other components of the computing device 200, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the Central Device Data Layer can communicate with each device component using an API (e.g., a private API).
[0141] The techniques described herein can refer to actions taken and information sent to or from the systems, not only servers, databases, software applications, and other computer-based systems. One will recognize the inherent flexibility of computer-based systems to allow for a wide variety of possible configurations, combinations, and divisions of work and functionality among and between components. For example, the processes described herein can be implemented using a single device or component or multiple devices or components operating in combination. The database and application may be implemented in a distributed system spanning a single system or multiple systems. The distributed components can operate sequentially or in parallel.
[0142] FIG. 6 is a block diagram for explaining the functions of a computing device 400 that implements a goal-oriented dialogue service according to an embodiment.
[0143] As shown in FIG. 6, the computing device 400 may include a dialogue graph generation module 10, a dialogue act group sampling module 20, a dialogue act group adjustment module 30, a dialogue act group selection module 40, and a dialogue act selection module 50.
[0144] The dialogue graph generation module 10 can generate a dialogue graph based on the dialogue dataset received from the outside.
[0145] The dialogue dataset may include data related to conversations between various speakers conducted in various environments. For example, the dialogue dataset may include data related to a series of conversations of various types, such as conversations between a customer and a reservation agent in a hotel reservation situation, conversations between a buyer and a seller in a shopping situation, and conversations between a tutor and a tutee regarding academic tasks.
[0146] A dialogue graph is a graph representation that models the relationships between various types of utterances included in a dialogue dataset. It may include a plurality of nodes corresponding to a plurality of dialogue acts indicating the functions of a specific utterance, and a plurality of edges indicating various information regarding the relationships between the plurality of dialogue acts.
[0147] Here, for example, dialogue acts may include the functions of various types of utterances such as questions, information provision, requests, confirmations, and the like.
[0148] In addition, the relationship between a plurality of dialogue acts may include a predetermined conditional relationship between two utterances. For example, the predetermined conditional relationship may include at least any one of a first conditional relationship (Should relationship) regarding what utterance should be made in response to an utterance, a second conditional relationship (Can relationship) regarding what utterance can be made in response to an utterance in the flow of the dialogue, and a third conditional relationship (Should-not relationship) regarding what utterance should not be made in response to an utterance.
[0149] For example, as shown in FIG. 7, the dialogue graph generation module 10 generates a task-oriented dialogue flow (TOD-Flow) graph including a plurality of nodes including A, B, C, D and a plurality of edges corresponding to the relationships between the plurality of nodes (S10).
[0150] The dialogue graph generation module 10 can vectorize the dialogue dataset by performing embedding on the words and sentences included in the dialogue dataset. The dialogue graph generation module 10 can analyze the vectorized dialogue dataset to learn relevant information of the dialogue dataset, such as the intent of the dialogue utterance, dialogue acts, slots, values, context information, information on the relationships between utterances, patterns of sequential dialogue flows, etc.
[0151] For example, the dialogue graph generation module 10 can learn relevant information of the dialogue dataset by using an artificial intelligence model based on at least one of a Transformer-based model, a recurrent neural network (RNN), and a long short-term memory (LSTM).
[0152] Also, the dialogue graph generation module 10 can generate a dialogue graph by modeling at least one conditional relationship for the dialogue dataset based on the learned relevant information of the dialogue dataset.
[0153] In this case, the dialogue graph generation module 10 models at least one relationship for the dialogue dataset by maximizing a loss function optimized for each at least one conditional relationship.
[0154] For example, when the n-th dialogue act satisfies the first conditional relationship (Should relationship) in the dialogue context, the dialogue graph generation module 10 generates a dialogue graph by utilizing a dialogue graph generation model learned to maximize the expected value.
[0155] Specifically, in a situation where the first conditional relationship (Should relationship) in the dialogue context must be satisfied, the dialogue graph generation module 10 can generate a dialogue graph by utilizing a dialogue graph generation model learned to maximize the expected value when the n-th dialogue act satisfies the first conditional relationship (Should relationship).
[0156] For example, by maximizing the loss function defined by the following [Equation 1], the dialogue graph generation module 10 can model the first conditional relationship (Should relationship) regarding what utterance should be made for a single utterance in the dialogue flow.
[0157] JPEG2025102693000002.jpg20151
[0158] JPEG2025102693000003.jpg45151
[0159] Also, for example, when the n-th dialogue act in the dialogue context satisfies the second conditional relationship (Can relationship) and at the same time does not satisfy the third conditional relationship (Should-not relationship), the dialogue graph generation module 10 can generate a dialogue graph by utilizing a dialogue graph generation model learned to maximize the expected value.
[0160] Specifically, in the situation where the n-th dialogue act occurs, when the n-th dialogue act in the dialogue context satisfies the second conditional relationship (Can relationship) and does not satisfy the third conditional relationship (Should-not relationship), the first expected value, and in the situation where the n-th dialogue act does not occur, when the n-th dialogue act in the dialogue context does not satisfy the second conditional relationship (Can relationship) and satisfies the third conditional relationship (Should-not relationship), the dialogue graph generation module 10 generates a dialogue graph by utilizing a dialogue graph generation model learned to maximize the sum of the second expected values. For example, the dialogue graph generation module 10 can model the second conditional relationship (Can relationship) regarding what utterances can be made for a single utterance in the flow of dialogue and the third conditional relationship (Should-not) regarding what utterances should not be made for a single utterance by maximizing the loss function defined by the following [Equation 2].
[0161] JPEG2025102693000004.jpg29151
[0162] JPEG2025102693000005.jpg58151
[0163] The dialogue act group sampling module 20 can sample a plurality of dialogue act groups for responding to user dialogue inputs received from the outside.
[0164] For example, the dialogue act group sampling module 20 can sample a plurality of dialogue act groups appropriate as responses to user dialogue inputs received via an electronic device. As shown in FIG. 7, the next dialogue act prediction can be made for the user dialogue input. For example, for a user dialogue input related to "hotel reservation", the dialogue act group sampling module 20 samples a dialogue act group including various types of dialogue acts such as "Confirm Information" and "Book Hotel" as the dialogue act group that is a response following the user dialogue input (S20).
[0165] The dialogue act group sampling module 20 may include an artificial neural network structure that can extract the characteristics of various types of dialogue datasets and has already learned to provide appropriate output data for the input data. For example, the dialogue act group sampling module 20 may include a Transformer-based neural network architecture (e.g., GPT-3, GPT-4, BERT-based models, etc.).
[0166] The dialogue act group adjustment module 30 can adjust a plurality of dialogue act groups sampled based on the generated dialogue graph.
[0167] Each of the plurality of dialogue act groups generated by the dialogue act group sampling module 20 may include at least one dialogue act related to the user dialogue input.
[0168] The dialogue act group adjustment module 30 can adjust a plurality of dialogue act groups based on whether each of the plurality of dialogue act groups generated by the dialogue act group sampling module 20 satisfies at least one conditional relationship included in the dialogue graph for the user dialogue input.
[0169] For example, the dialogue act group adjustment module 30 determines whether each of the plurality of dialogue act groups satisfies a first conditional relationship (Should relationship) with respect to the user dialogue input, and can add a pair of dialogue acts that satisfy the first conditional relationship (Should relationship) with respect to the user dialogue input to a pair of dialogue act groups that do not satisfy it.
[0170] In addition, the dialogue act group adjustment module 30 determines whether each of a plurality of dialogue act groups satisfies a second conditional relationship (Can relationship) and a contra-third conditional relationship (not should not) with respect to the user dialogue input, and removes (removes) a dialogue act that does not satisfy the second conditional relationship (Can relationship) and the contra-third conditional relationship (not should not) among at least one dialogue act included in a pair of dialogue act groups that do not satisfy the condition.
[0171] In this way, by adjusting the plurality of dialogue act groups sampled by the dialogue act group adjustment module 30 so as to conform to the dialogue graph (S20), the reliability of the goal-oriented dialogue service provided by the system 1000 and the control power over the dialogue model are improved.
[0172] The dialogue act group selection module 40 can select any one dialogue act group that satisfies a predetermined condition from among the plurality of adjusted dialogue act groups.
[0173] For example, the dialogue act group selection module 40 determines the ranking of a plurality of dialogue act groups in descending order of the number of dialogue acts that satisfy at least one conditional relationship of the dialogue graph among the plurality of adjusted dialogue act groups (S30).
[0174] For example, the first dialogue act group, the second dialogue act group, and the third dialogue act group can be adjusted by the dialogue act group adjustment module 30. After the adjustment operation is completed, the first dialogue act group includes three dialogue acts, the second dialogue act group includes two dialogue acts, and the third dialogue act group includes four dialogue acts.
[0175] In this case, the dialogue act group selection module 40 can select the third dialogue act group including the most dialogue acts as the first rank among the first to third dialogue act groups based on at least one conditional relationship of the dialogue graph after adjustment.
[0176] The dialogue act selection module 50 can select any one of at least one dialogue act included in one dialogue act group selected from a plurality of dialogue act groups and provide a response output to the user dialogue input.
[0177] For example, the dialogue act selection module 50 can calculate the probability of the relevance of at least one dialogue act included in the selected dialogue act group to the user dialogue input. Thereafter, the dialogue act selection module 50 can select any one dialogue act with the highest calculated probability corresponding to the relevance to the user dialogue input from the selected dialogue act group and provide it to the user as a response.
[0178] -AI Agent Specialization Model (AIAM)
[0179] In other aspects, the computing system 1000 as described above may include an AI Agent Specialization Model (AIAM) according to an embodiment.
[0180] Here, the AI Agent Specialization Model (AIAM) according to an embodiment is an AI agent model that applies the MoE (Mixture of Experts) architecture realized by an embodiment, and can be an artificial intelligence model including a data processing algorithm that can autonomously act in a specific environment, solve tasks, and achieve goals. Here, MoE means an architecture of a machine learning model that combines a number of expert models to solve complex problems.
[0181] Such an AI Agent Specialization Model (AIAM) may include data processing algorithms for realizing cognitive abilities such as collecting and interpreting data from a given environment, decision mechanisms for determining optimal actions based on the collected data, execution abilities for executing the determined actions, and learning abilities for improving actions through experience.
[0182] As an embodiment, the AI Agent Specialization Model (AIAM) can obtain predetermined input data (such as text, voice, image, video, and / or specific sensor-based sensing data, etc.), and provide output data (such as response data to specific queries and / or control signals by specific command words, etc.) by performing predetermined tasks based on the obtained input data.
[0183] FIG. 8 shows an internal block diagram of an AI agent model according to an embodiment.
[0184] Specifically, as shown in FIG. 8, an AI agent model according to an embodiment may include at least one router (RT: Router, Gating Network), an orchestrator (OCT: Orchestrator), a sLLM (small Large Language Model), a general MoE model (NM: Normal MoE Model), an external model (EM: External Model), and / or a specialized model (SM: Specialized Model).
[0185] Here, in FIG. 8, in order to prevent the features according to various embodiments from being unclear, it is described that the AI agent model includes the above-mentioned components.
[0186] However, it is obvious that a person of ordinary skill in the art can understand that, in addition to the components shown in FIG. 8 according to the embodiment, other general-purpose components may be further included, or some of the components shown in FIG. 8 may be omitted.
[0187] More specifically, a router (RT: Router, Gating Network) according to an embodiment may be an artificial intelligence module that performs task assignment and / or traffic adjustment for a plurality of models in the MoE architecture.
[0188] Specifically, the router (RT) can analyze the given input data and / or requested tasks, etc., and determine which model is most suitable for the data processing.
[0189] At this time, as an embodiment, the router (RT) can determine an optimized model for the given data processing based on the performance, expertise, and / or previous experience of each model, etc.
[0190] In addition, the router (RT) can distribute the given task to at least one or more models in consideration of the system load to support efficient data processing.
[0191] Furthermore, the router (RT) can adjust the tasks assigned to a specific model to dynamically respond to changes in the real-time system.
[0192] In an embodiment, such a router (RT) may be an artificial intelligence module that selectively determines a model (hereinafter, domain-specific model) that executes a data processing operation optimized for a predetermined domain.
[0193] That is, in an embodiment, the router (RT) may be an artificial intelligence module that selects a model (i.e., a domain specialization model) that is determined to perform data processing (such as deep learning in an embodiment) most suitable for a given domain among a plurality of models included in the AI agent specialization model (AIAM).
[0194] For reference, the domain according to an embodiment means data, rules, terms, problem definitions, and / or processes used by a given AI system to perform a specific task, and the like.
[0195] As an embodiment, the router (RT) performs data analysis based on characteristics of given input data (e.g., user input and / or specific sensing data, etc.) and / or a required task, etc., grasps data processing characteristics optimized for the corresponding task based on this, and can detect a predetermined model that realizes this to determine a domain specialization model.
[0196] In other words, the router (RT) according to an embodiment may be an artificial intelligence module that detects a model that can most effectively execute data processing for a given domain, assigns / distributes the corresponding task processing work, and manages this.
[0197] Here, the router (RT) according to an embodiment may include a router (RT) that has already been learned by a disclosed predetermined algorithm, a router (RT) that has been additionally learned according to an embodiment, and / or a router (RT) that has been newly learned in a new manner, and the like.
[0198] On the other hand, the orchestrator (OCT: Orchestrator) according to an embodiment may be an artificial intelligence module that overall controls and manages the overall configuration of the AI agent specialization model (AIAM).
[0199] Specifically, in the embodiment, the orchestrator (OCT) allocates various tasks generated from the overall system to appropriate resources (such as routers (RT) and / or a predetermined model in the embodiment).
[0200] In addition, the orchestrator (OCT) manages to efficiently use resources such as available models and hardware resources (e.g., CPUs and / or GPUs).
[0201] In addition, the orchestrator (OCT) monitors the performance of the overall system and adjusts specific parameters as needed, or optimizes the network configuration, etc.
[0202] In addition, the orchestrator (OCT) manages the coordination between multiple routers (RT) and / or models and controls the data flow and processing process.
[0203] That is, in the embodiment, the orchestrator (OCT) controls and manages the entire system of the AI agent specialization model (AIAM) and performs the role of the main router (RT) that controls at least one router (RT).
[0204] Here, the orchestrator (OCT) and the router (RT) according to the embodiment can support the efficient operation of the MoE system by closely coordinating with each other.
[0205] Specifically, the orchestration (OCT) monitors the performance of the router (RT) as the administrator of the entire system and adjusts the strategy of the router (RT) as needed.
[0206] On the other hand, the router (RT) can substantially perform the allocation of data processing tasks according to the instructions of the orchestrator (OCT) and / or its own algorithms to achieve efficient system control.
[0207] In an embodiment, the aforementioned orchestrator (OCT) and / or router (RT) may be a master model (P: Master Model) that can control and manage the rest of the overall system and / or the AI agent model configuration (i.e., sLLM, general MoE model (NM), external model (EM), and / or specialized model (SM), etc.).
[0208] On the other hand, an sLLM (small Large Language Model) according to an embodiment is an artificial intelligence module realized as a lightweight version of an LLM (Large Language Model).
[0209] That is, the sLLM is an artificial intelligence module constructed to achieve performance similar to that of a large model such as an LLM with fewer resources.
[0210] In an embodiment, such an sLLM may include a plurality of specialized models (SM) according to an embodiment and a MoE model based on the combination of routers (RT) (in the embodiment, MoELM). Further, the sLLM may include a MoE model based on a domain-specific specialized model (in the embodiment, DMoE model) according to an embodiment.
[0211] Also, the general MoE model (NM: Normal MoE Model) according to an embodiment of the present invention may mean a predetermined MoE model realized by a disclosed universal method.
[0212] Exemplarily, the general MoE model (NM) may include Switch Transformer, Conditional Computation in Neural Networks, Sparse Mixture of Experts, and / or Megatron-LM, etc.
[0213] Also, an external model (EM) according to an embodiment may refer to a predetermined artificial intelligence model implemented by various disclosed algorithms.
[0214] For example, the external model (EM) may include ChatGPT, Gemini, and / or Llama, etc.
[0215] In an embodiment, such an external model (EM) can selectively be used as needed to support the given task processing.
[0216] Also, a specialized model (SM) according to an embodiment is an artificial intelligence model in which optimal learning for a specific purpose has been performed, and may refer to an artificial intelligence model learned by training data and methods specialized for the relevant purpose.
[0217] That is, the specialized model (SM) can be an artificial intelligence model learned by training data and methods specialized to achieve a predetermined purpose.
[0218] In an embodiment, such a specialized model (SM) may include a learned predetermined sLLM (including MoELM and / or DMoE models), a general MoE model (NM), and / or an external model (EM), etc. Further, the specialized model (SM) may include a specialized module model according to an embodiment disclosed in the "MoE-based model identification method" described later. A detailed description thereof will be given in the "MoE-based model identification method".
[0219] In an embodiment, the aforementioned sLLM, general MoE model (NM), external model (EM), and / or specialized model (SM) can be secondary models (S) that can execute specific tasks under the control and management of the master model (P) of the AI agent model (i.e., the orchestrator (OCT) and / or router (RT), etc.).
[0220] - Goal - Oriented Dialogue Method (S100)
[0221] The goal - oriented dialogue method (S100) will be described in detail below. By extracting and learning the characteristics of various types of dialogue datasets, generating a dialogue graph that models a predetermined conditional relationship for the dialogue dataset based on this, and selecting the optimal dialogue act from among multiple dialogue acts sampled by a dialogue model already learned for a user dialogue input based on the dialogue graph and providing it as a response, a more accurate response can be provided for the user dialogue input.
[0222] The dialogue dataset may include data related to dialogues among various speakers conducted in various environments. Therefore, the dialogue dataset includes various types of data according to the nature of the dialogues conducted among the speakers.
[0223] For example, the first dialogue dataset and the second dialogue dataset related to different types of tasks may include different types of data, and the data structures of the first dialogue graph generated based on the first dialogue dataset and the second dialogue graph generated based on the second dialogue dataset may be different from each other.
[0224] The dialogue graph models the relationships among various types of utterances included in the dialogue dataset and represents them in graph form. It can be data having a structured form of data related to multiple dialogue acts corresponding to the functions, intentions, etc. of multiple utterances and data related to the conditional relationships among multiple dialogue acts.
[0225] The goal - oriented dialogue method (S100) can select and provide the optimal dialogue act group as a response to the user dialogue input from among multiple dialogue act groups sampled by the dialogue model based on the dialogue graph.
[0226] Also, based on the user dialogue input and the dialogue action groups selected as responses, the task required by the user is determined, and the computing system 1000 according to one embodiment can perform the determined task to provide a task execution service to the user.
[0227] In the following, a goal-oriented dialogue method (S100) will be described in detail, in which the dialogue model according to one embodiment provides an appropriate response to the user dialogue input based on a dialogue graph that models at least one conditional relationship for the dialogue dataset, and can execute the task required by the user.
[0228] FIG. 9 is a flowchart of a goal-oriented dialogue method (S100) according to one embodiment.
[0229] As shown in FIG. 9, the goal-oriented dialogue method (S100) according to one embodiment may include a step (S101) of generating a dialogue graph that models at least one conditional relationship for the dialogue dataset, a step (S103) of receiving a user dialogue input, a step (S105) of sampling a plurality of dialogue action groups for responding to the user dialogue input using a previously learned dialogue model, a step (S107) of adjusting the plurality of dialogue action groups based on the dialogue graph, and a step (S109) of selecting any one dialogue group that satisfies a predetermined condition among the plurality of dialogue action groups.
[0230] In step (S101), the processors 111, 131 of the system 1000 generate a dialogue graph based on various types of dialogue datasets.
[0231] For example, the processors 111, 131 generate a first dialogue graph that models at least one conditional relationship for a first type of dialogue dataset related to the first task. Also, the processors 111, 131 generate a second dialogue graph that models at least one conditional relationship for a second type of dialogue dataset related to the first task and another second task.
[0232] Here, at least one conditional relationship for the dialogue dataset may include at least any one of a first conditional relationship (Should relationship) regarding what utterance should be made for a single utterance included in the dialogue dataset, a second conditional relationship (Can) regarding what utterance can be made for a single utterance, and a third conditional relationship regarding what utterance should not be made for a single utterance.
[0233] In step (S103), the processors 111 and 131 of the system 1000 receive user dialogue input.
[0234] For example, the processors 111 and 131 can be transmitted data of user dialogue input received by the user computing device 110, which can be realized by various types of electronic devices, via the user input component 121.
[0235] In step (S105), the processors 111 and 131 of the system 1000 sample a plurality of dialogue act groups for providing as responses to the user dialogue input by using the already learned dialogue model.
[0236] For example, as shown in FIG. 10, the processors 111 and 131 can sample a plurality of dialogue act groups a1, a2,... by using the dialogue model (π) already learned by the model trainer 160 of the training computing system 150. In this case, the number of the plurality of dialogue act groups a1, a2,... sampled by the dialogue model (π) may be several tens, but is not limited thereto.
[0237] For example, the sampled first dialogue act group a1 may include four dialogue acts A, B, C, F, and the second dialogue act group a2 may include three dialogue acts A, C, G.
[0238] In step (S107), the processors 111 and 131 of the system 1000 adjust a plurality of dialogue act groups sampled based on the dialogue graph.
[0239] The processors 111 and 131 can adjust the plurality of dialogue act groups based on whether each of the plurality of dialogue act groups satisfies at least one conditional relationship included in the dialogue graph for the user dialogue input.
[0240] First, the processors 111 and 131 determine any one of the plurality of dialogue graphs generated for various dialogue datasets that corresponds to the type of the user dialogue input.
[0241] Thereafter, the processors 111 and 131 adjust the plurality of dialogue act groups based on the determined dialogue graph.
[0242] As shown in FIG. 10, the determined dialogue graph includes G as a dialogue act that satisfies the first conditional relationship (Should relationship) for the user dialogue input, includes A, C, F, and G as dialogue acts that satisfy the second conditional relationship (Can relationship), and includes A as a dialogue act that satisfies the third conditional relationship (Should-not relationship).
[0243] The processors 111 and 131 remove B from the first dialogue act group a1 so that the first dialogue act group a1 satisfies the second conditional relationship (Can relationship) based on the determined dialogue graph.
[0244] Also, the processors 111 and 131 add G to the first dialogue act group a1 so that the first dialogue act group a1 satisfies the first conditional relationship (Should relationship) based on the determined dialogue graph.
[0245] Furthermore, the processors 111 and 131 remove A from the first dialogue act group a1 so that the first dialogue act group a1 satisfies the third conditional relationship (Should-not relationship) based on the determined dialogue graph.
[0246] Similarly, the processors 111 and 131 remove A from the second dialogue act group a2 so that the second dialogue act group a2 satisfies the first to third conditional relationships based on the determined dialogue graph.
[0247] In this case, after the adjustment operations for the plurality of dialogue act groups are completed, the first dialogue act group a1 includes three dialogue acts C, F, and G, and the second dialogue act group a2 includes two dialogue acts C and G.
[0248] By adjusting the plurality of sampled dialogue act groups to conform to the dialogue graph in this way, the reliability of the goal-oriented dialogue service provided by the system 1000 and the control power over the dialogue model can be improved.
[0249] In step (S109), the processors 111 and 131 of the system 1000 select any one dialogue act group that satisfies a predetermined condition among the adjusted plurality of dialogue act groups and provide it as a response to the user dialogue input.
[0250] The processors 111 and 131 select the dialogue act group that includes the most dialogue acts satisfying at least one conditional relationship of the dialogue graph among the adjusted plurality of dialogue act groups.
[0251] For example, through the adjustment process, the processors 111 and 131 can select the first dialogue act group a1 that includes three dialogue acts C, F, and G and the second dialogue act group (a2) that includes two dialogue acts C and G, and select the first dialogue act group a1 that includes the most dialogue acts satisfying at least one conditional relationship of the dialogue graph and provide it as a response to the user dialogue input.
[0252] FIG. 11 shows the response prediction performance index (F-1 score) of various dialogue models (FLAN-T5, GPT-turbo) for user dialogue inputs with respect to the goal-oriented dialogue method (S100) based on the dialogue graph according to an embodiment of the present disclosure for various dialogue datasets (SGD, MultiWOZ).
[0253] In this case, by changing the method of selecting, as a response, any one of a plurality of dialogue act groups generated from various dialogue models (FLAN-T5, GPT-turbo) and adjusted based on the dialogue graph, the response prediction performance index of the dialogue model changes.
[0254] Also, when selecting any one of the plurality of dialogue act groups, the response prediction performance index of the dialogue model also changes depending on whether to perform an adjustment operation based on at least one conditional relationship of the dialogue graph.
[0255] For example, the processors 111, 131 select the dialogue act group determined by the dialogue model to have the highest response probability among the plurality of dialogue act groups and provide it as a response to the user dialogue input (Greedy).
[0256] Also, the processors 111, 131 select the dialogue act group that contains the most dialogue acts satisfying at least one conditional relationship of the dialogue graph (Compliance), or select the dialogue act group that is adjusted the least based on at least one conditional relationship and provide it as a response to the user dialogue input (Violation).
[0257] Furthermore, the processors 111, 131 select the most sampled dialogue act group among the plurality of dialogue act groups sampled by the dialogue model and provide it as a response to the user dialogue input (Majority).
[0258] As shown in FIG. 11, when following the method (Compliance) of adjusting a plurality of dialogue act groups based on all of the first conditional relationship (Should relationship), the second conditional relationship (Can relationship), and the third conditional relationship (Should-not relationship) of the dialogue graph, and selecting the dialogue act group that contains the most dialogue acts satisfying at least one conditional relationship of the dialogue graph among the plurality of dialogue act groups, the response prediction performance index of the dialogue model is the highest.
[0259] In this way, by appropriately adjusting a plurality of dialogue act groups based on the dialogue graph, and selecting any one dialogue act group that most satisfies the conditional relationship of the dialogue graph among the adjusted plurality of dialogue act groups and providing it as a response to the user dialogue input, a more reliable response can be provided.
[0260] On the other hand, the method (S100) may further include a response dialogue act determination step of providing any one dialogue act included in any one selected dialogue act group as a response to the user dialogue input.
[0261] For example, when the user dialogue input is text data obtained by converting the voice "Please make a hotel reservation" into text, the selected first dialogue act group a1 among the adjusted plurality of dialogue act groups includes three dialogue acts: "How many people will be staying? (C)", "What date do you wish to make a reservation? (F)", and "What grade of hotel do you prefer? (G)".
[0262] In the response dialogue act determination step, the processors 111, 131 select any one of the plurality of dialogue acts (C, F, G) included in the first dialogue act group a1 and provide it as a response to the user dialogue input. In this case, the processors 111, 131 can calculate the relevance of each of the plurality of dialogue acts (C, F, G) included in the first dialogue act group a1 to the user dialogue input, and select and provide as a response to the user dialogue input any one dialogue act with the highest relevance.
[0263] Furthermore, the processors 111 and 131 can determine the type of task requested by the user based on the user interaction input and the dialogue act group finally selected as the response, and perform the determined task.
[0264] For example, when the type of task is determined to be "hotel reservation", based on various information related to the hotel reservation determined in the process of a series of dialogues including the user interaction input and the selected dialogue act group, the processors 111 and 131 can complete a hotel reservation that meets the user's requirements.
[0265] In this way, the system 1000 can adjust the dialogue act sampled by the dialogue model based on the dialogue graph in which the predetermined conditional relationship is structured and modeled based on the dialogue dataset, and provide the adjusted dialogue act as a response to the user interaction input, thereby providing a more reliable response to the user.
[0266] Furthermore, based on a series of dialogues including the user interaction input and the dialogue act adjusted by the dialogue graph and determined as the final response according to a predetermined condition, the type of task requested by the user is accurately determined, and by the system 1000 performing the task determined in this way, a task execution service that satisfies the user can be provided.
[0267] FIG. 12 is a flowchart of a task execution method (S200) based on the context of goal-oriented dialogue through a dialogue model according to an embodiment.
[0268] As shown in FIG. 12, a task execution method (S200) according to an embodiment may include a step (S201) of receiving user interaction input, a step (S203) of determining and providing a response interaction act for the user interaction input based on an interaction graph, a step (S205) of analyzing data of a series of goal-oriented dialogs including the user interaction input and the response interaction act to determine the context of the goal-oriented dialog, a step (S207) of determining the type of task requested by the user based on the context of the goal-oriented dialog, and a step (S209) of performing the determined task.
[0269] In step (S201), processors 111 and 131 of system 1000 receive user interaction input.
[0270] For example, processors 111 and 131 can be transmitted data of user interaction input received by user computing device 110, which can be realized by various types of electronic devices, via user input component 121.
[0271] In step (S203), processors 111 and 131 of system 1000 determine and provide a response interaction act for the user interaction input based on an interaction graph related to the user interaction input.
[0272] Step (S203) is substantially the same as the goal-oriented dialog method (S100) described with reference to FIGS. 9 and 10, and the description thereof is omitted.
[0273] In step (S205), processors 111 and 131 of system 1000 analyze data of a series of goal-oriented dialogs including the user interaction input and the response interaction act to determine the context of the goal-oriented dialog.
[0274] Processors 111 and 131 of system 1000 can grasp the context of a series of goal-oriented dialogs consisting of the user interaction input and the response interaction act thereto.
[0275] For example, the processors 111 and 131 of the system 1000 extract a plurality of keywords included in the data of the goal-oriented dialogue, and analyze the correlation between a plurality of dialogue acts included in the dialogue, the intention and purpose of the dialogue, etc. based on the extracted plurality of keywords, and finally determine the context of the goal-oriented dialogue.
[0276] In step (S207), the processors 111 and 131 of the system 1000 determine the type of task required by the user based on the context of the goal-oriented dialogue.
[0277] The processors 111 and 131 of the system 1000 determine the type of task corresponding to the context of the goal-oriented dialogue determined based on the data related to the task corresponding to the context of the goal-oriented dialogue.
[0278] In this case, the data related to the task corresponding to the context of the goal-oriented dialogue may already be stored in the memories 112 and 132 of the system 1000, or may already be learned by the machine learning models 120 and 140.
[0279] For example, when the processors 111 and 131 of the system 1000 determine that the context of the goal-oriented dialogue requires an automatic hotel reservation service in response to the user's hotel reservation request, the task required by the user can be determined as "hotel reservation".
[0280] On the other hand, the processors 111 and 131 of the system 1000 can determine a plurality of types of tasks based on the context of the goal-oriented dialogue. The goal-oriented dialogue includes various types of user dialogue inputs and various response dialogue acts corresponding thereto, and such a context of the goal-oriented dialogue is related to various tasks. Accordingly, a plurality of types of tasks can be determined based on the context of the goal-oriented dialogue according to an embodiment.
[0281] In step (S209), the processors 111 and 131 of the system 1000 perform the task whose type has been determined.
[0282] The processors 111 and 131 of the system 1000 execute tasks whose types are determined based on goal-oriented dialogues including user interaction inputs and response interaction acts, and provide the execution results to the user.
[0283] In this case, the processors 111 and 131 of the system 1000 generate programming code necessary to perform tasks whose types are determined based on the context of the goal-oriented dialogue, and execute this to perform the tasks.
[0284] Also, according to one embodiment, the processors 111 and 131 of the system 1000 can capture the screen of the electronic device used by the user by receiving a user interaction input, and obtain a user screen screenshot.
[0285] After that, when determining the type of task, the processors 111 and 131 of the system 1000 determine the type of task based on information regarding the determined context of the goal-oriented dialogue and the analysis of the user screen screenshot.
[0286] In this case, when performing the task, the processors 111 and 131 of the system 1000 can automatically execute a series of acts (for example, cursor movement, click, text input, etc.) necessary to perform the determined task on the user screen.
[0287] Furthermore, when the type of task is determined based on the context of the goal-oriented dialogue according to one embodiment, the processors 111 and 131 of the system 1000 determine at least one task execution model optimized for the task whose type is determined based on the context of the goal-oriented dialogue among a plurality of task execution models, and perform the task using the determined at least one task execution model.
[0288] In this case, the method by which the processors 111 and 131 of the system 1000 determine at least one task execution model optimized for the task and use it to perform the task is substantially the same as the "MoE-based model identification method" described later, and the description thereof is omitted here.
[0289] And when the types of tasks are determined to be plural, the processors 111 and 131 of the system 1000 can determine a plurality of task execution models optimized for each of the plurality of tasks and use them to perform the plurality of tasks.
[0290] -MoE-based model identification method
[0291] Hereinafter, a method for realizing a model providing service based on a MoE architecture by which a computing system 1000 according to an embodiment realizes modularization for a predetermined specialized model (SM) in a MoE (Mixture of Experts) model will be described in detail with reference to the accompanying drawings.
[0292] FIG. 13 is a flowchart for explaining a MoE-based model identification method according to an embodiment, and FIG. 14 is a conceptual diagram for explaining a MoE-based model identification method according to an embodiment.
[0293] As shown in FIGS. 13 and 14, a method for realizing a MoE architecture-based model providing service in which a computing system 1000 according to an embodiment performs modularization of a specialized model (SM) including a MoE model may include steps of performing MoE learning based on MoELM (S301), obtaining specialized model (SM) characteristic information by the MoE learning (S303), generating a specialized module model based on the obtained specialized model (SM) characteristic information (S305), obtaining predetermined domain information (S307), determining a domain-specialized specialized model based on the obtained domain information (S309), constructing a MoE model based on the determined domain-specialized specialized model (S311), and providing output data based on the constructed MoE model (S313).
[0294] Specifically, in many cases, it is difficult to distinguish or grasp for a general already-trained specialized model (SM) what domain the model is specialized in.
[0295] As a result, there may be certain constraints in selecting and utilizing a specialized model (SM) optimized for a specific task.
[0296] To solve this, in one embodiment, the computing system 1000 can execute the following process of specifying and modularizing the role and / or function of each specialized model (SM) and effectively selecting and utilizing a customized specialized model (SM) optimized for a specific domain based on this.
[0297] Specifically, the computing system 1000 according to an embodiment performs MoE learning based on MoELM (S301).
[0298] That is, in the embodiment, the computing system 1000 can perform MoE learning based on the above-described plurality of specialized models (SMs) and the router (RT)-coupled MoELM.
[0299] At this time, by performing learning, the computing system 1000 can realize learning for each of the plurality of specialized models (SMs) included in the MoELM.
[0300] In other words, by performing the above-described learning, the plurality of specialized models (SMs) in the MoELM can be respectively learned.
[0301] Here, in other words, the specialized model (SM) according to the embodiment is an artificial intelligence model for which optimal learning for a specific purpose has been performed, and can mean an artificial intelligence model learned by training data and a method specialized for the corresponding purpose.
[0302] In the embodiment, such a specialized model (SM) may include a learned predetermined sLLM (including MoELM and / or DMoE model), a general MoE model (NM), an external model (EM), and / or a specialized module model (MM) according to the embodiments disclosed below.
[0303] Also, the computing system 1000 according to an embodiment acquires specialized model characteristic information (SMFI) by MoE learning (S303).
[0304] Here, the specialized model characteristic information (SMFI) according to the embodiment may mean information for specifying the role and / or function of a predetermined specialized model (SM).
[0305] Specifically, referring further to FIG. 8, in the embodiment, the computing system 1000 may further include a model specialization module (MSM) according to an embodiment.
[0306] And the computing system 1000 can obtain the aforementioned specialized model feature information (SMFI) in conjunction with the model specialization module (MSM).
[0307] Here, the model specialization module (MSM) according to an embodiment may be an artificial intelligence module that generates and outputs specialized model feature information (SMFI) corresponding to a predetermined specialized model (SM) based on MoE learning.
[0308] Specifically, in an embodiment, the model specialization module (MSM) can monitor and track the task assignment status of the router (RT) for each specialized model (SM) when the aforementioned MoE learning is performed.
[0309] That is, in an embodiment, the model specialization module (MSM) can grasp how the router (RT) distributes and assigns tasks to which specialized models (SM) when the MoELM is learned and operated.
[0310] According to an embodiment, the model specialization module (MSM) can also generate a tag for identifying each tracked task assignment status and perform matching management.
[0311] Thereby, in an embodiment, the model specialization module (MSM) can determine the specialization for each of the plurality of specialized models (SM).
[0312] In addition, in an embodiment, the model specialization module (MSM) generates specialized model feature information (SMFI) corresponding to each specialized model (SM) based on the specialization for each determined specialized model (SM).
[0313] FIG. 15 is a diagram showing an example of specialized model feature information (SMFI) according to an embodiment.
[0314] Here, as shown in FIG. 15, as an embodiment, the model specification module (MSM) can generate the aforementioned specialized model feature information (SMFI) in at least one of the following forms.
[0315] [First form] Specialized model feature information (SMFI) in a form of selecting any one category of a specialized model (SM) role and / or function specification category (for example, question and answer or device control, etc.) that has already been set according to user input
[0316] [Second form] Specialized model feature information (SMFI) in a form of specifying the role and / or function of a specialized model (SM) in natural language form
[0317] [Third form] Specialized model feature information (SMFI) in a form of specifying the role and / or function of a specialized model (SM) in at least one of the first form and the second form, and further defining the input data and output data of the specialized model (SM)
[0318] Subsequently, in an embodiment, the model specification module (MSM) can provide the specialized model feature information (SMFI) generated as described above to the computing system 1000 as output data.
[0319] Therefore, in an embodiment, the computing system 1000 can obtain characteristic information for each specialized model (SM) in conjunction with the model specification module (MSM).
[0320] In addition, the computing system 1000 according to an embodiment generates a specialized module model (MM: Specialized Module Module) based on the obtained specialized model feature information (SMFI) (S305).
[0321] Here, a specialized module model (MM) according to an embodiment may mean a specialized model (SM) in which predetermined specialized model characteristic information (SMFI) is matched and independently separated.
[0322] Specifically, in an embodiment, the computing system 1000 matches the specialized model characteristic information (SMFI) obtained as described above to a corresponding specialized model (SM).
[0323] Also, in an embodiment, the computing system 1000 independently separates and databaseizes the specialized models (SM) to which the specialized model characteristic information (SMFI) is matched.
[0324] That is, in an embodiment, the computing system 1000 matches the corresponding specialized model characteristic information (SMFI) to each specialized model (SM), and executes modularization for separately classifying, storing, and managing them.
[0325] Therefore, the computing system 1000 can generate a specialized module model (MM) that is a specialized model (SM) in which the specialized model characteristic information (SMFI) is matched and independently separated at the same time.
[0326] In this way, in an embodiment, the computing system 1000 grasps the different characteristics of the specialized models (SM) within a given MoE model (in the embodiment, MoELM), and modularizes each specialized model (SM) to a small size at a level where it can be reused and shared, reflecting this.
[0327] Thereby, the computing system 1000 can quickly and efficiently select and choose specialized models (SM) that realize a data processing process optimized for a specific domain with higher accuracy, and can easily support flexible expansion or contraction of the MoE model based on this.
[0328] Also, a computing system 1000 according to an embodiment can acquire predetermined domain information (S307).
[0329] Here, the domain information according to the embodiment can be information that defines a domain that specifies data, rules, terms, problem definitions, and / or processes, etc. that a predetermined AI system uses to perform a predetermined task.
[0330] Specifically, in the embodiment, the computing system 1000 acquires predetermined input data (for example, text, voice, image, video, and / or specific sensor-based sensing data, etc.).
[0331] Also, in the embodiment, the computing system 1000 determines a domain corresponding to the acquired input data.
[0332] Here, in the embodiment, the method by which the computing system 1000 determines the domain for the input data may be performed based on various disclosed algorithms capable of executing this, and in the embodiments of the present invention, the algorithms themselves are not limited or restricted.
[0333] Therefore, the implementation system 1000 can acquire domain information for the task to be processed.
[0334] Also, a computing system 1000 according to an embodiment determines a domain specialization specialized model based on the acquired domain information (S309).
[0335] Here, the domain specialization specialized model according to the embodiment can mean a specialized model (SM) that executes data processing (as an embodiment, deep learning, etc.) operations optimized for a predetermined domain.
[0336] More specifically, referring further to FIG. 14, in an embodiment, the computing system 1000 determines at least one domain-specific specialized model based on the domain information and specialized model characteristic information (SMFI) obtained as described above.
[0337] More specifically, in an embodiment, the computing system 1000 can detect at least one specialized model characteristic information (SMFI) having characteristics corresponding to the obtained domain information.
[0338] For example, when the computing system 1000 confirms the "characteristics of the task of outputting response data for predetermined interrogation data" based on the first domain information, it can detect at least one specialized model characteristic information (SMFI) specified as the "role and / or function specialized for interrogation response" from among the plurality of specialized model characteristic information (SMFI) stored in a database.
[0339] Here, according to an embodiment, the computing system 1000 can detect at least one specialized model characteristic information (SMFI) corresponding to the domain information based on a plurality of tags generated by a model specialization module (MSM) for each task assignment state of a router (RT) for a plurality of specialized models (SM) during learning based on the aforementioned MoE architecture.
[0340] That is, according to an embodiment, the computing system 1000 can detect at least one specialized model characteristic information (SMFI) corresponding to the relevant domain information by mutually comparing the plurality of tags generated as described above and the domain information.
[0341] Here, according to an embodiment, the computing system 1000 filters the tags to be compared according to the generation time of each tag.
[0342] Specifically, the computing system 1000 sets at least one tag generated at a specific task assignment time as a comparison target tag according to user input and / or a preset unique process.
[0343] Exemplarily, the computing system 1000 can set at least one tag generated for the task assignment state performed after the preset time during the overall learning time as a comparison target tag, focusing on the fact that the accuracy of task assignment improves as the learning rate increases.
[0344] Therefore, the computing system 1000 can detect at least one specialized model characteristic information (SMFI) corresponding to the relevant domain information by comparing between at least one filtered tag that guarantees higher accuracy and the domain information.
[0345] Also, in an embodiment, the computing system 1000 extracts a specialized model (SM) (i.e., a specialized module model (MM)) that respectively matches the detected at least one specialized model characteristic information (SMFI).
[0346] And, in an embodiment, the computing system 1000 determines the extracted at least one specialized module model (MM) as a domain-specific specialized model.
[0347] Also, the computing system 1000 according to an embodiment of the present invention constructs a MoE model based on the determined domain-specific specialized model (S311).
[0348] Referring further to FIG. 14, that is, in an embodiment, the computing system 1000 can construct a model (hereinafter, DMoE model) that operates like a MoE architecture based on at least one domain-specific specialized model determined as described above.
[0349] In other words, the computing system 1000 can construct a MoE model (i.e., a DMoE model) that realizes data processing optimized for a specific domain by using at least some of a plurality of specialized models (SMs) that are modularized with a small size (i.e., domain-specific specialized models).
[0350] Specifically, in an embodiment, the computing system 1000 can construct the aforementioned DMoE model by combining at least one domain-specific specialized model and a predetermined router (RT).
[0351] Therefore, in an embodiment, the computing system 1000 can construct a DMoE model including a domain-specific specialized model and a router (RT).
[0352] Here, according to an embodiment, the DMoE model may be included in the sLLM according to an embodiment of the present invention.
[0353] In other words, the sLLM according to an embodiment may include a DMoE model constructed according to an embodiment.
[0354] Further, the computing system 1000 according to an embodiment of the present invention provides output data based on the constructed MoE model (S313).
[0355] That is, in an embodiment, the computing system 1000 uses the DMoE model constructed as described above to provide output data (e.g., response data to a specific query and / or a control signal by a specific command word, etc.) for predetermined input data (e.g., text, voice, image, video, and / or specific sensor-based sensing data, etc.).
[0356] As described above, in the embodiment, the computing system 1000 can identify the roles and / or functions of each specialized model (SM), and at the same time, separate and modularize them to levels where they can be reused and shared respectively. By leveraging this, a customized MoE model optimized for a specific domain (i.e., DMoE model) can be rapidly and flexibly constructed, and predetermined output data can be provided through efficient task processing using the constructed model.
[0357] In other words, in the embodiment, the computing system 1000 can realize and provide a MoE model with further improved data processing (and / or computing) speed and inference performance, and can support various services thereby, so that its performance and quality can be effectively improved.
[0358] -Method for providing an AI agent based on MoE applied LLM
[0359] Hereinafter, a method for realizing a model providing service based on a MoE architecture that provides an on-device specialized AI agent (Artificial Intelligence Agent) that determines an application model optimized for a domain by an external environment based on an LLM (Large Language Model) applying MoE (Mixture of Experts) in an embodiment of the present invention and executes an output based on the determined application model will be described in detail with reference to the accompanying drawings.
[0360] FIG. 16 is a flowchart for explaining a method for providing an AI agent based on MoE applied LLM according to an embodiment, and FIG. 17 is a conceptual diagram for explaining a method for providing an AI agent based on MoE applied LLM according to an embodiment.
[0361] As shown in FIGS. 16 and 17, a method for implementing a model providing service based on a MoE architecture that provides an on-device specialized AI agent specialization model (AIAM) that determines an application model optimized for a domain by an external environment based on an LLM applying MoE and outputs based on the determined application model may include steps of executing an on-device AI agent service (S401), obtaining predetermined input data (S403), determining a domain based on the obtained input data (S405), determining an application model based on the determined domain (S407), and providing output data based on the determined application model (S409).
[0362] Specifically, a computing system 1000 according to an embodiment of the present invention executes an on-device AI agent service (S401).
[0363] Here, for reference, on-device AI may mean a technology that directly executes artificial intelligence-based data processing inside a user's device rather than in the cloud and / or an external server. Since all processing is completed inside the device without sending data externally, it can provide advantages such as personal information protection, real-time processing, and reduced dependence on an Internet connection.
[0364] Therefore, in this regard, the on-device AI agent service may mean various services realized by utilizing on-device AI.
[0365] Exemplarily, the on-device AI agent service may include a voice assistant service for a smartphone (e.g., Google Assistant, Apple Siri, Samsung Bixby, etc.), a smart camera service (e.g., HDR+ of Google Pixel, Deep Fusion of Apple, etc.), a fitness tracker and smartwatch service (e.g., Apple Watch, Fitbit, etc.), an automotive autonomous driving service (e.g., Autopilot of Tesla, etc.) and / or a home security service (e.g., Nest Secure, Ring, etc.).
[0366] In an embodiment, the computing system 1000 can execute a predetermined on-device AI agent service based on the interlock with an AI agent specialization model (AIAM) according to an embodiment and / or a predetermined application, etc.
[0367] Also, the computing system 1000 according to an embodiment of the present invention acquires predetermined input data (S403).
[0368] Specifically, in an embodiment, the computing system 1000 can acquire at least one input data (e.g., predetermined text, voice, image, video and / or sensing data, etc.) based on user input based on the on-device AI agent service executed as described above and / or the interlock with an external device (e.g., a predetermined sensor, etc.).
[0369] In an embodiment, the input data acquired as described above may include predetermined data that can identify the target task of data processing.
[0370] Also, the computing system 1000 according to an embodiment determines the domain by the acquired input data (S405).
[0371] Here, in other words, the domain according to the embodiment may mean data, rules, terms, problem definitions, and / or processes, etc. that a predetermined AI system uses to perform a specific task.
[0372] Specifically, in the embodiment, the computing system 1000 determines the domain corresponding to the acquired input data.
[0373] Here, in the embodiment, the method by which the computing system 1000 determines the domain for the input data may be performed based on various disclosed algorithms capable of doing so. In one embodiment, the corresponding algorithm itself is not limited or restricted.
[0374] Therefore, in the embodiment, the computing system 1000 can acquire the domain information corresponding to the task to be processed.
[0375] Also, the computing system 1000 according to one embodiment determines an application model according to the determined domain (S407).
[0376] Here, the application model according to the embodiment may mean a model that performs a predetermined task process with the given input data.
[0377] In the embodiment, such an application model may be at least one model of the secondary model (S) as described above.
[0378] Here, in other words, the secondary model (S) according to the embodiment may mean a model that can execute a specific task by the control and management of a master model (P) (that is, an orchestrator (OCT) and / or a router (RT), etc.) responsible for the control and management of the operation of a predetermined AI system.
[0379] In an embodiment, such a secondary model (S) may include at least one model of a small language model (sLLM) (including a MoELM and / or a DMoE model), a general MoE model (NM), an external model (EM), and / or a specialized model (SM) (including a specialized module model (MM)).
[0380] Specifically, in an embodiment, the computing system 1000 determines at least one application model based on the domain information obtained as described above.
[0381] More specifically, as an embodiment, the computing system 1000, in conjunction with a master model (P) according to an embodiment (i.e., an orchestrator (OCT) and / or a router (RT), etc.), detects at least one or more models that execute data processing (such as deep learning, etc.) operations optimized for the given domain information among the aforementioned secondary models (S) (i.e., domain-specific models).
[0382] Here, the specific method by which the computing system 1000 detects a domain-specific model in conjunction with the master model (P) in an embodiment is omitted by applying mutatis mutandis the descriptions regarding the router (RT) and the orchestrator (OCT) disclosed in the aforementioned "AI Agent Specialized Model (AIAM)".
[0383] Also, in an embodiment, the computing system 1000 determines the at least one detected domain-specific model as an application model.
[0384] Also, the computing system 1000 according to an embodiment provides output data based on the determined application model (S409).
[0385] That is, in the embodiment, the computing system 1000 can generate and provide output data (e.g., response data to a specific query and / or a control signal by a specific command word, etc.) for predetermined input data (e.g., text, voice, image, video, and / or specific sensor-based sensing data, etc.) based on at least one application model determined through the aforementioned AI agent specialization model (AIAM).
[0386] In other words, the computing system 1000 executes a predetermined request task based on given input data by using the application model determined as described above, and provides output data by the executed data processing.
[0387] Here, in the embodiment, the computing system 1000 can provide the output data based on the aforementioned on-device AI agent service.
[0388] As described above, in the embodiment, the computing system 1000 effectively determines a model optimized for data processing according to a given domain even in an on-device environment based on an AI agent specialization model (AIAM) including models (as embodiments, MoELM, DMoE model, and / or specialized module model (MM), etc.) realized by applying the MoE architecture in various embodiments, and provides an output by efficient data processing through the determined model.
[0389] That is, the computing system 1000 can realize and provide an artificial intelligence model (i.e., an AI agent specialization model (AIAM)) that can better listen to, understand, execute, and answer a given task in any environment.
[0390] Accordingly, in an embodiment, the computing system 1000 can directly and significantly improve the quality and performance of various AI agent-based services (such as smartphone voice assistant services, smart camera services, fitness tracker and smartwatch services, automobile autonomous driving services, and / or home security services, etc.).
[0391] The foregoing embodiments according to the present invention can be realized in the form of program instruction words executable by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instruction words, data files, data structures, etc. alone or in combination. The program instruction words recorded on the computer-readable recording medium may be those specially designed or configured for the present invention or those known and usable to those skilled in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instruction words such as ROM, RAM, flash memory, etc. Examples of program instruction words include not only machine language codes made by compilers but also high-level language codes executable by a computer using an interpreter, etc. The hardware device can be changed into one or more software modules for performing the processing according to the present invention, and vice versa.
[0392] The specific implementation described in the present invention is one embodiment and does not limit the scope of the present invention in any way. For the sake of brevity of the specification, descriptions of conventional electronic configurations, control systems, software, and other functional aspects of the system can be omitted. Also, the connections or connecting members of the lines between the components shown in the drawings exemplarily show functional connections and / or physical or circuit connections, and in actual devices, they can be replaceable or shown as additional various functional connections, physical connections, or circuit connections. Further, without specific references such as "essential" or "importantly", components may not necessarily be required for the application of the present invention.
[0393] In the detailed description of the present invention described above, the preferred embodiments of the present invention have been described. However, those skilled in the art or those with ordinary knowledge in the technical field will understand that the present invention can be variously modified and changed within the scope not departing from the idea and technical field of the present invention described in the claims below. Therefore, the technical scope of the present invention should not be limited to the content described in the detailed description of the specification, but should be determined by the claims.
Explanation of Reference Numerals
[0394] 10 Dialogue graph generation module 20 Dialogue act group sampling module 30 Dialogue act group adjustment module 40 Dialogue act group selection module 50 Dialogue act selection module 110, 180, 181, 182, 183 User computing device 111, 131, 151 Processor 112, 132, 152 Memory 113, 133, 153 Data 114, 134, 154 Instruction words 120, 140 Machine learning model 121 User input component 160 Model Trainer 161 Training Data 100, 200, 400 Computing Devices 300 Neuromorphic Circuit 310 Presynaptic Neuron Circuit 311 Presynaptic Line 320 Postsynaptic Neuron Circuit 330 Synapse Circuit
Claims
1. A goal-oriented dialogue method provided by a computing system including a memory and a processor by utilizing a dialogue model, comprising: generating a dialogue graph modeling at least one conditional relationship for a dialogue dataset; receiving a user dialogue input; sampling a plurality of dialogue act groups for responding to the user dialogue input by using a pre-trained dialogue model; adjusting the plurality of dialogue act groups based on the dialogue graph; and selecting any one dialogue act group that meets a predetermined condition from the plurality of dialogue act groups.
2. The at least one conditional relationship includes at least any one of a first conditional relationship regarding what utterance should be made for a single utterance in the flow of a dialogue, a second conditional relationship regarding what utterance can be made for a single utterance, and a third conditional relationship regarding what utterance should not be made for a single utterance. The goal-oriented dialogue method according to Claim 1.
3. In the step of selecting any one dialogue act group, selecting any one dialogue act group that most meets the at least one conditional relationship from the plurality of dialogue act groups. The goal-oriented dialogue method according to Claim 1.
4. Each of the plurality of dialogue act groups includes at least one dialogue act for the user dialogue input. The goal-oriented dialogue method according to Claim 2.
5. In the step of adjusting the plurality of dialogue act groups, for each of the plurality of dialogue act groups, adding a dialogue act that meets the first conditional relationship among at least one dialogue act, removing a dialogue act that does not meet the second conditional relationship, and removing a dialogue act that does not meet the third conditional relationship. The goal-oriented dialogue method according to Claim 4.
6. In the step of selecting any one dialogue act group, selecting a dialogue act group that includes the most dialogue acts that meet the at least one conditional relationship from the plurality of dialogue act groups. The goal-oriented dialogue method according to Claim 5.
7. A response dialogue act determination step of providing, as a response to the user dialogue input, any one dialogue act included in any one of the selected dialogue act groups, the goal-oriented dialogue method according to claim 1, further comprising.
8. In the response dialogue act determination step, The goal-oriented dialogue method according to claim 7, wherein, among at least one dialogue act included in any one of the selected dialogue act groups, the dialogue act having the highest degree of relevance to the user dialogue input is provided as a response to the user dialogue input.
9. In the step of generating the dialogue graph, The goal-oriented dialogue method according to claim 2, wherein a dialogue graph generation model learned so that an expected value is maximized when the nth dialogue act in the dialogue context satisfies the first conditional relationship based on various dialogue acts is utilized to generate the dialogue graph so that the first conditional relationship is modeled for the dialogue dataset.
10. In the step of generating the dialogue graph, The goal-oriented dialogue method according to claim 2, wherein a dialogue graph generation model learned so that an expected value is maximized when the nth dialogue act in the dialogue context satisfies the second conditional relationship and does not satisfy the third conditional relationship at the same time based on various dialogue acts is utilized to generate the dialogue graph so that the second conditional relationship and the third conditional relationship are modeled for the dialogue dataset.
11. In the step of generating the dialogue graph, The goal-oriented dialogue method according to claim 1, wherein a first dialogue graph is generated based on a first type of dialogue dataset, and a second dialogue graph is generated based on a second type of dialogue dataset.
12. The step of determining the type of the user dialogue input, and The step of determining a dialogue graph corresponding to the type of the user dialogue input among the first dialogue graph and the second dialogue graph, further comprising In the step of adjusting the plurality of dialogue act groups, The goal-oriented dialogue method according to claim 11, wherein the plurality of dialogue act groups are adjusted based on the determined dialogue graph.
13. The step of analyzing data of a series of goal-oriented dialogues including the user dialogue input and the selected any one dialogue act group to determine the context of the goal-oriented dialogue, The step of determining the type of task required by the user based on the context of the goal-oriented dialogue, and The method for goal-oriented dialogue according to claim 1, further comprising the step of performing the task whose type has been determined.
14. In the step of determining the context of the goal-oriented dialogue, extracting a plurality of keywords from the data of the goal-oriented dialogue, and analyzing at least any one of the correlation relationship between a plurality of dialogue acts included in the goal-oriented dialogue, the intention of the goal-oriented dialogue, and the goal based on the plurality of keywords to determine the context, the method for goal-oriented dialogue according to claim 13.
15. at least one memory, and at least one processor that reads at least one instruction word stored in the at least one memory and executes a method for goal-oriented dialogue, wherein the at least one processor generates a dialogue graph modeling at least one conditional relationship for a dialogue dataset, receives a user dialogue input, samples a plurality of dialogue act groups for responding to the user dialogue input using a pre-learned dialogue model, adjusts the plurality of dialogue act groups based on the dialogue graph, and selects any one dialogue act group that satisfies a predetermined condition among the plurality of dialogue act groups, a goal-oriented dialogue system.
16. an electronic device that receives a user dialogue input, and a computing device including at least one processor that generates a dialogue graph modeling at least one conditional relationship for a dialogue dataset and samples a plurality of dialogue act groups for responding to the user dialogue input using a pre-learned dialogue model, wherein the at least one processor adjusts the plurality of dialogue act groups based on the dialogue graph, and selects any one dialogue act group that satisfies a predetermined condition among the plurality of dialogue act groups, a goal-oriented dialogue system.
17. wherein the at least one processor determines the type of a task (Task) requested by the user based on the user dialogue input and any one selected dialogue act group, and performs the determined task, the goal-oriented dialogue system according to claim 16.
18. The electronic device according to claim 16 receives the user dialogue input in at least any one form of text, voice, gesture, and touch.
Citation Information
Patent Citations
Interactive device, learning device, interactive method, learning method, and program
JP2018055548A
System and method for generating dialogue graphs
US20200004878A1
Information processing device and information processing method
WO2021235225A1