Training method and system for intent classification model using intent descriptions

The intent classification model improves accuracy by using intent descriptions and a ranker to differentiate intent names and descriptions, enhancing chatbot performance in understanding user queries.

WO2025249808A1PCT designated stage Publication Date: 2025-12-04LG MANAGEMENT DEV INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006612
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2025-05-15
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing intent classification models struggle to accurately map user utterances to predefined intents due to indistinguishable intent names and descriptions, leading to suboptimal performance in chatbots and conversational systems.

Method used

A learning method and system for an intent classification model that utilizes intent descriptions, including independent and dependent prompts to generate a data set, and employs a ranker to calculate similarity between user queries and intent descriptions, allowing for retraining to improve accuracy.

Benefits of technology

Enhances the accuracy of intent classification by distinguishing intent names and descriptions, thereby improving the performance of chatbots in understanding user queries and providing appropriate responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025006612_04122025_PF_FP_ABST
    Figure KR2025006612_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A training system for an intent classification model using intent descriptions of the present invention generates a data set including intent descriptions of an intent classification model from a language model, causes the intent classification model to perform intent classification using at least a part of the data set as an input prompt of the intent classification model, and determines the performance of the intent classification model using at least a part of the data set as an input prompt of the intent classification model. The data set may include first data including independent intent descriptions in which intent names and intent descriptions for the intent names are mutually indistinguishable, and second data including dependent intent descriptions in which intent names and intent descriptions for the intent names are mutually distinguishable.
Need to check novelty before this filing date? Find Prior Art

Description

Learning method and system for an intention classification model using intention description

[0001] The present invention relates to a learning method and system for an intention classification model, and more particularly, to a learning method and system for an intention classification model using an intention description, which can improve the accuracy of intention classification by learning an intention classification model based on input data including an intention description.

[0002] Artificial intelligence (AI) technology has recently been garnering attention across society as it demonstrates cutting-edge developments. AI encompasses "a computer brain capable of executing tasks previously reserved for human intelligence," "the engineering and science of creating intelligent machines," and "a set of algorithmic systems designed to think, perceive, and act like humans," enabling computers to perform highly advanced, human-like intellectual abilities.

[0003] AI, combined with augmented reality, the Internet of Things, edge computing, and digital twins, is being touted as a key new technology that will drive the Fourth Industrial Revolution, promising highly integrated smart spaces. Furthermore, AI is gaining traction as a next-generation growth engine capable of evolving industrial ecosystems beyond simply solving standardized problems. It is actively being applied not only to IT, healthcare, agriculture, energy, automobiles, and robotics, but also to knowledge service industries such as distribution, finance, law, education, real estate, advertising, and communications. In other words, AI is integrating with all existing systems, not just those that seek to improve the convenience and quality of daily life, but also across the entire spectrum of our society's culture and arts, preparing for a new era.

[0004] Chatbots, task-oriented conversational systems that engage users through voice or text-based input and perform specific tasks, have recently been introduced and are being used in a variety of fields. Chatbots utilize conversational AI technologies like Natural Language Processing (NLP) to understand users' queries and automatically display responses.

[0005] Meanwhile, in these task-oriented conversational systems, intent classification technology, which accurately identifies the user's intent from the query, may be essential to accurately understand the user's query and provide appropriate services in response. Research on intent classification, which accurately identifies user intent from the query and maps the user's utterance to a predefined set of intents, is actively underway.

[0006] As a related technology, Effectiveness of pre-training for few-shot intent classification.(https: / aclanthology.org / 2021.findings-emnlp.96 / )(Zhang et al., Findings 2021) was published.

[0007] One embodiment of the present invention aims to provide a method and system for learning an intent classification model using intent descriptions, which can improve the accuracy of intent classification that maps a user utterance to one of a predefined set of intents by learning an intent classification model of a large-scale language model including intent descriptions in input data.

[0008] A learning system for an intent classification model using intent descriptions according to one embodiment of the present invention comprises: at least one processor; and at least one memory for storing commands or information executed by the at least one processor; wherein operations performed by the commands or information executed by the at least one processor include: generating a data set including intent descriptions of an intent classification model using a language model; using at least a portion of the data set as an input prompt of the intent classification model to perform intent classification; and using at least a portion of the data set as an input prompt of the intent classification model to determine performance of the intent classification model; wherein the data set may include first data including independent intent descriptions in which intent names and intent descriptions for the intent names are mutually indistinguishable, and second data including dependent intent descriptions in which intent names and intent descriptions for the intent names are mutually distinguishable.

[0009] Here, the command prompt input to the language model to output the data set from the language model may include an independent prompt to obtain the first data including a plurality of intent descriptions for a single intent; and a dependent prompt to obtain the second data including a plurality of intent descriptions for a plurality of intents, but such that the plurality of intents and the plurality of intent descriptions are each distinct from each other.

[0010] Additionally, the input prompt may be comprised of an instruction and user query section and an intent options section including an intent name and intent descriptions for the intent name.

[0011] Additionally, the learning system of one embodiment further includes a ranker that adjusts the number of intent options, wherein the ranker can calculate a similarity between the user query and the intent description.

[0012] Additionally, the intent classification model may be retrained using at least a portion of the first data, or may be retrained using at least a portion of the second data.

[0013] Additionally, at least a portion of the first data may be provided as an input prompt when determining the performance of the intent classification model, or at least a portion of the second data may be provided as an input prompt when determining the performance of the intent classification model.

[0014] Additionally, when the intent classification model is retrained with at least a portion of the second data, the accuracy performance of intent classification can be improved.

[0015] A method for learning an intent classification model using intent descriptions according to one embodiment of the present invention is executed by at least one processor, and includes the steps of: generating a data set including intent descriptions of an intent classification model using a language model; using at least a portion of the data set as an input prompt of the intent classification model to perform intent classification by the intent classification model; and determining performance of the intent classification model using at least a portion of the data set as an input prompt of the intent classification model; wherein the data set may include first data including independent intent descriptions in which intent names and intent descriptions for the intent names are mutually indistinguishable, and second data including dependent intent descriptions in which intent names and intent descriptions for the intent names are mutually distinguishable.

[0016] In addition, in the step of generating the data set, a command prompt input to the language model to generate the data set from the language model is configured, and the command prompt may include an independent prompt to obtain the first data including a plurality of intent descriptions for a single intent; and a dependent prompt to obtain the second data including a plurality of intent descriptions for a plurality of intents, but such that each of the plurality of intents and the plurality of intent descriptions is distinguished from each other.

[0017] Additionally, in the step of performing intent classification by the intent classification model or the step of determining the performance of the intent classification model, the input prompt may be composed of an instruction and user query section and an intent option section including an intent name and intent descriptions for the intent name.

[0018] Additionally, the ranker that controls the number of intent options can calculate the similarity between the user query and the intent description and sort the intent descriptions in descending order according to the similarity.

[0019] Additionally, in the step where the intention classification model performs intention classification, the intention classification model may be retrained using at least a portion of the first data, or may be retrained using at least a portion of the second data.

[0020] Additionally, in the step of determining the performance of the intent classification model, at least a portion of the first data may be provided as an input prompt when determining the performance of the intent classification model, or at least a portion of the second data may be provided as an input prompt when determining the performance of the intent classification model.

[0021] Additionally, when the intent classification model is retrained with at least a portion of the second data, the accuracy performance of intent classification can be improved.

[0022] According to one embodiment of the present invention, a method and system for learning an intent classification model using intent descriptions can be provided, which can improve the accuracy of intent classification that maps a user utterance to one of a predefined set of intents by learning an intent classification model by including intent descriptions in input data.

[0023] Figure 1 is a schematic diagram of an electronic device according to one embodiment of the present invention.

[0024] Figure 2 is a schematic diagram showing an example of a command prompt input to a language model.

[0025] Figure 3 is a schematic diagram showing an example of an input prompt input into an intent classification model.

[0026] FIG. 4 is a schematic diagram of a learning method for an intention classification model using intention description according to one embodiment of the present invention.

[0027] Figure 5 is a schematic diagram showing the accuracy performance of an intent classification model according to the number of intent options.

[0028] In order to clarify the technical idea of ​​the present disclosure, embodiments of the present disclosure will be described in detail with reference to the attached drawings. In describing the present disclosure, if a detailed description of a related known function or component is determined to unnecessarily obscure the gist of the present disclosure, the detailed description will be omitted. In the drawings, components having substantially the same functional configuration are given the same reference numbers and symbols as possible even if they are shown in different drawings. For convenience of explanation, devices and methods are described together when necessary. Each operation of the present disclosure does not necessarily have to be performed in the described order and may be performed in parallel, selectively, or individually.

[0029] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this specification should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0030] Throughout this disclosure, singular expressions may include plural expressions unless the context clearly dictates otherwise. Terms such as "comprise" or "have" should be understood to indicate the presence of a feature, number, step, operation, component, part, or combination thereof, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. In other words, when it is said throughout this disclosure that a part "comprises" a certain component, unless specifically stated otherwise, this does not mean that other components may be included, but rather that other components may be excluded.

[0031] Expressions such as "at least one" modify the entire list of elements, not individual elements of the list. For example, "at least one of A, B, and C" and "at least one of A, B, or C" refer to only A, only B, only C, both A and B, both B and C, both A and C, all of A, B, and C, or any combination thereof.

[0032] In addition, terms such as “...unit”, “...module”, etc. described in the present disclosure mean a unit that processes at least one function or operation, which may be implemented as hardware or software, or a combination of hardware and software.

[0033] Throughout this disclosure, when a part is said to be "connected" to another part, this includes not only cases where the parts are "directly connected," but also cases where the parts are "electrically connected" with other elements intervening. Furthermore, when a part is said to "include" a component, this does not exclude other components, but rather includes other components, unless otherwise specifically stated.

[0034] The expression “configured to” as used throughout this disclosure can be used interchangeably with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” does not necessarily mean something that is “specifically designed to” in terms of hardware. Instead, in some contexts, the expression “a system configured to” can mean that the system, together with other devices or components, is “capable of.” For example, the phrase “a processor configured (or set) to perform A, B, and C” may mean a dedicated processor (e.g., an embedded processor) for performing those operations, or a generic-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in memory.

[0035] Artificial intelligence (AI) is a field of computer engineering and information technology that studies how to enable computers to perform human-like tasks, such as thinking, learning, and self-improvement. It aims to enable computers to mimic human intelligent behavior. Furthermore, AI does not exist in isolation; rather, it is closely linked, both directly and indirectly, to other fields of computer science. In particular, efforts are actively underway to incorporate AI elements into various fields of information technology and utilize them to solve problems in those fields.

[0036] Machine learning is a branch of artificial intelligence that empowers computers to learn without explicit programming. Specifically, machine learning is the study and development of algorithms and systems that learn from empirical data, make predictions, and improve their own performance. Rather than executing strictly defined, static program instructions, machine learning algorithms build specific models based on input data to derive predictions or decisions. The term "machine learning" can be used interchangeably with "machine learning."

[0037] Many machine learning algorithms have been developed to classify data. Representative examples include decision trees, Bayesian networks, support vector machines (SVMs), and artificial neural networks (ANNs). Decision trees are an analytical method that performs classification and prediction by diagramming decision rules in a tree-like structure. Bayesian networks are models that represent probabilistic relationships (conditional independence) between multiple variables in a graph structure. Bayesian networks are suitable for data mining through unsupervised learning. Support vector machines are supervised learning models for pattern recognition and data analysis, primarily used for classification and regression analysis. Artificial neural networks model the operating principles and interconnected relationships of biological neurons. They are information processing systems in which numerous neurons, called nodes or processing elements, are connected in layers.

[0038] An artificial neural network (ANN) is a model used in machine learning. It is a statistical learning algorithm inspired by biological neural networks (especially the brain, the central nervous system of animals) in machine learning and cognitive science. Specifically, an ANN can refer to a general model in which artificial neurons (nodes) form a network through the connection of synapses, which change the strength of the synaptic connections through learning, thereby achieving problem-solving capabilities. The term "ANN" can be used interchangeably with the term "neural network."

[0039] An artificial neural network can include multiple layers, each of which can include multiple neurons. Furthermore, an artificial neural network can include synapses, which connect neurons. An artificial neural network can generally be defined by three factors: a) the connection pattern between neurons in different layers, b) a learning process that updates the weights of the connections, and c) an activation function that generates an output value from a weighted sum of the inputs received from the previous layer.

[0040] Artificial neural networks may include, but are not limited to, network models such as Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), Multilayer Perceptrons (MLPs), and Convolutional Neural Networks (CNNs). In this specification, the term "layer" may be used interchangeably with the term "layer."

[0041] Artificial neural networks are categorized into single-layer neural networks and multi-layer neural networks based on the number of layers. A typical single-layer neural network consists of an input layer and an output layer. A typical multi-layer neural network consists of an input layer, one or more hidden layers, and an output layer.

[0042] The input layer is the layer that receives external data, and the number of neurons in the input layer is the same as the number of input variables, and the hidden layer is located between the input layer and the output layer. It receives signals from the input layer, extracts characteristics, and transmits them to the output layer. The output layer receives signals from the hidden layer and outputs output values ​​based on the received signals. The input signals between neurons are multiplied by each connection strength (weight) and then added, and if this sum is greater than the threshold of the neuron, the neuron is activated and outputs the output value obtained through the activation function.

[0043] Meanwhile, deep neural networks, which include multiple hidden layers between the input and output layers, are representative artificial neural networks that implement deep learning, a type of machine learning technique. The term "deep learning" can be used interchangeably with "deep learning," and the term "learning" can be used interchangeably with "training."

[0044] The workflow of machine learning consists of a series of steps: collecting data for learning and validation, modeling, and then training the model. This can include the processes of collecting training data, inspecting and exploring the data, preprocessing and cleaning the data, modeling, and training.

[0045] 1. Collect training data

[0046] The training data applied to the learning model of this specification can be generated using data collected from multiple samples. In this specification, at least one or more different types of training data sets can be used to train the learning model, and each training data set can further include one or more experimental results used as feature labels. At least a portion of the training data set can be used to train the learning model, and another portion can be used to validate the learned learning model.

[0047] 2. Checking and exploring data

[0048] Once training data for learning a learning model is collected, the collected training data can be inspected and explored for data structure, noisy data, and data cleaning methods for applying machine learning.

[0049] This data review and exploration phase is called Exploratory Data Analysis (EDA), and EDA can be defined as the process of observing and understanding collected data from various perspectives. Before data learning, visualizations such as graphs and statistical tests are used to examine independent and dependent variables, variable types, and their data types, allowing for preliminary identification of data characteristics and inherent structural relationships. Through EDA, data distributions and values ​​can be examined to better understand the phenomena expressed by the data and identify potential problems. Furthermore, through the process of examining data from various perspectives, various patterns that might not have been detected during the problem definition phase can be discovered, allowing for modification of existing hypotheses or the development of new ones. Exploratory data analysis can broadly include the process of searching for outliers and analyzing the relationships between data attributes.

[0050] The process of detecting outliers involves determining whether data contains outliers. This process can involve sampling, statistical methods, and visualization methods. Sampling methods extract a random sample from the data to identify overall trends and anomalies in the data values. Statistical methods can utilize summary statistics such as the mean, median, and mode to determine the center of the data, or the range and variance to determine the distribution of the data. Visualization methods can utilize probability density functions, histograms, dotplots, word clouds, time series charts, and maps to determine which statistical indicators are appropriate for each attribute of the collected data. However, when using statistical indicators, it is important to note that the mean reflects all data values ​​in the set, so outliers can affect the value, whereas the median uses the single value in the middle, so it can produce representative results even with outliers.

[0051] The process of analyzing the relationship between data attributes is to find combinations of attributes that have meaningful correlations within the data. The relationship analysis can be performed differently depending on the combination of attributes between qualitative attributes (Categorical Variable; Qualitative) that cannot be expressed numerically but can be arbitrarily quantified and quantitative attributes (Numeric Variable; Quantitative). The qualitative-qualitative relationship (Categorical - Categorical) can be displayed by using cross tables and mosaic plots to count the number of values ​​corresponding to each pair of attribute values. The quantitative-qualitative relationship (Numeric-Categorical) can be visually expressed by observing statistical values ​​(mean, median, etc.) by category or using box plots. The quantitative-quantitative relationship (Numeric-Numeric) can be analyzed for the association between two attributes using correlation coefficients. A correlation coefficient of -1 indicates a negative correlation where the two attributes change in opposite directions, 0 indicates no correlation, and 1 indicates a positive correlation where the two attributes always change in the same direction. The relationship between two attributes with a correlation coefficient can take many forms, and this can be visually represented using a scatter plot.

[0052] 3. Data preprocessing and cleaning

[0053] Once the data has been inspected and explored, data preprocessing is performed to transform it into a format suitable for machine learning training models. Data preprocessing involves refining data and transforming it into a form understandable by the model. Data preprocessing typically includes handling missing data, removing outliers, scaling, categorical data encoding, feature selection and extraction, and data transformation. The detailed data preprocessing steps can be performed in whole or in part, and a separate machine learning model may be used for data preprocessing.

[0054] Handling missing data involves handling missing values ​​in data. Missing values ​​can be displayed as NaN (Not a Number) or blank, or they can be deleted. Filling in or deleting missing values ​​within the data improves data completeness. When filling in missing values, values ​​such as the mean, median, or mode can be used.

[0055] Outlier removal is the process of removing outliers, values ​​that deviate from the normal data pattern. Outliers can degrade model performance and should therefore be removed or replaced. Identifying outliers can be accomplished by deleting the corresponding rows or columns or replacing them with different values.

[0056] Data scaling is the process of adjusting the size of data. Through data scaling, the range of data can be adjusted, and the performance of the model or the convergence speed can be improved. Through data scaling, the characteristics of the data can be adjusted to a similar range, and data scaling can generally be applied with standardization and normalization. Standardization is a method of converting data into a distribution with a mean of 0 and a standard deviation of 1, and is mainly converted using the mean and standard deviation, and the standardized value z is It can be expressed as (x is the original value, μ is the mean, σ is the standard deviation). Normalization is a method to convert the range of data to [0,1] or [-1,1], and mainly converts data using the minimum and maximum values, and the normalized value x norm silver can be expressed as (x is the original value, x min is the minimum, x max is the maximum value).

[0057] Categorical data encoding is the process of converting categorical variables, represented as strings or integers that cannot be directly input into a model, into numerical data types that can be input into the model. Typically, one-hot encoding or label encoding is used to convert categorical variables into numerical data types.

[0058] Feature selection and extraction is a process to improve model performance by selecting the most useful features for model learning or extracting new features. This process can reduce model complexity and prevent overfitting.

[0059] Data transformation is the process of transforming data to extract new information or to improve model understanding. This can include tokenizing text data or preprocessing image data. Data transformation can extract useful features from source data or transform data into an appropriate format, improving model performance.

[0060] Through data preprocessing as described above, the performance of machine learning models can be improved and stability can be secured.

[0061] Meanwhile, when training a learning model according to one embodiment of the present invention, a process of preprocessing information expressed in natural language and a process of learning a language model based on the preprocessed data may be performed.

[0062] 3-1. Text Preprocessing for Large-Scale Language Models

[0063] If the collected data has not been preprocessed to suit your needs, tokenization, cleaning, and normalization can be performed to suit the intended use of the data.

[0064] Tokenization refers to the process of dividing given data into units called tokens, which can be broadly defined as meaningful units. Tokenization can broadly include word tokenization and sentence tokenization.

[0065] Tokenization refers to the process of dividing given data into units called tokens, which can be broadly defined as meaningful units. Tokenization can broadly include word tokenization and sentence tokenization.

[0066] Word tokenization refers to cases where tokens are based on words, and in this case, words can include not only individual words but also word phrases and meaningful strings. Word tokenization separates words based on spaces or punctuation marks, such as periods, commas, question marks, semicolons, and exclamation marks. However, removing all punctuation or special characters during tokenization can sometimes result in tokens losing their meaning, necessitating a more precise tokenization algorithm. For example, if a word itself contains punctuation or uses special characters with meaning, simply removing them may not be enough. Therefore, tokenization rules such as the Penn Treebank Tokenization Rules can be applied during tokenization.

[0067] Sentence tokenization refers to the process of dividing text into sentences. Typically, if the data is unrefined, the corpus is not segmented into sentences, and thus sentence tokenization may be necessary to meet the intended use. Various rules for sentence tokenization can be defined depending on the language being used and how special characters are used within the corpus.

[0068] Tokenization is the process of classifying tokens according to their intended use. Before and after tokenization, cleaning and normalization are performed on text data to suit the intended use. Cleaning removes noise, while normalization integrates words with different representations and transforms them into a single, consistent word.

[0069] Refinement can occur before tokenization to eliminate any interference and facilitate tokenization. However, it can also be performed continuously and iteratively after tokenization to remove any remaining noise. The noise data removed during refinement are meaningless characters. Methods for removing unnecessary words include removing stopwords, low-frequency words, and short words.

[0070] Normalization work includes unifying words with different spellings based on rules, unifying uppercase and lowercase letters, etc. Unifying uppercase and lowercase letters is a normalization method that can reduce the number of words in English-speaking languages. In English-speaking languages, uppercase letters are only used in certain situations such as the beginning of a sentence, and most texts are written in lowercase letters, so unifying uppercase and lowercase letters can mostly be done by converting uppercase letters to lowercase letters.

[0071] Processing natural language in computing systems requires preprocessing, which involves digitizing text. This involves mapping each word in the text to a unique integer. This mapping process can utilize techniques such as integer encoding, padding, and one-hot encoding.

[0072] Integer encoding is a method of assigning integers to words. It creates a vocabulary by sorting words in order of frequency, and assigns integers in order of frequency, starting with the lowest number. Integer encoding performs sentence tokenization on text data containing multiple sentences, and performs word tokenization through parallel refinement and normalization. During this process, words are lowercase to unify the number of words, and stopwords and word length can be deleted. Through this, words can be recorded as keys and the frequency of each word as values. Integer encoding can be performed by sorting the text in order of frequency and assigning integers to words with high frequencies.

[0073] Padding is the process of randomly adjusting the length of sentences of different lengths within a text to the same length. Computing systems can perform parallel computations by grouping sentences of the same length into a single matrix. Specifically, to perform parallel computations, the lengths of sentences of different lengths within a text can be randomly padded with "0" to equalize the integer encoding results. Specifically, the longest sentence in a set of integer-encoded words can be identified, and a "0" can be added to the integer matrix corresponding to the length of the longest sentence. The computing system can then process sentences of the same length as a single matrix, allowing it to perform parallel processing. At this point, the computing system can ignore the "0" word, which is perceived as meaningless. This process of adjusting the size (shape) of data by filling it with a specific value is called padding. Using the number "0" to adjust the length is called zero padding.

[0074] One-hot encoding is a vector representation of words that uses the size of the vector as the dimension of the word set, assigning a value of 1 to the index of the word to be expressed, and 0 to all other indices. The vector expressed in this way is called a one-hot vector. One-hot encoding consists of integer encoding and an index assignment process. After integer encoding is performed and a unique integer is assigned to each word, the unique integer of the word to be expressed is regarded as an index, and a "1" is assigned to the corresponding position, and a "0" is assigned to the index positions of other words. However, one-hot encoding has the disadvantage that the space required to store the vector increases as the number of words increases (the dimensionality of the vector increases), and the similarity between words cannot be identified. To address these shortcomings, techniques that reflect the latent meaning of words and vectorize them into a multidimensional space include LSA (Latent Semantic Analysis), a count-based vectorization method; NNLM, RNNLM, Word2Vec, and FastText, which vectorize based on prediction; and the GloVe method, which uses both count-based and prediction-based methods.

[0075] Meanwhile, for computers to understand and process text, it must be appropriately converted into numbers. Because the performance of natural language processing can vary significantly depending on how words are represented, numerous techniques have been proposed to quantify words. Currently, the most widely used method is word embedding, which vectorizes each word through artificial neural network training.

[0076] Word embedding is a method of representing words as vectors, converting them into dense representations. The resulting word embedding is called a dense vector, or embedding vector. Word embedding methods include LSA, Word2Vec, FastText, and Glove.

[0077] 4. Modeling and Training

[0078] Artificial neural networks can be trained using training data. Here, "training" refers to the process of determining the parameters of an artificial neural network using training data to achieve objectives such as classification, regression analysis, or clustering of input data. Representative examples of artificial neural network parameters include the weights assigned to synapses and the biases applied to neurons.

[0079] An artificial neural network trained using training data can classify or cluster input data based on its patterns. Meanwhile, an artificial neural network trained using training data is referred to herein as a "trained model."

[0080] The following explains the learning methods of artificial neural networks. Learning methods of artificial neural networks can be broadly categorized into supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.

[0081] Supervised learning is a machine learning method that infers a function from training data. Among these inferred functions, regression analysis is the process of outputting continuous values, while classification is the process of predicting and outputting the class of an input vector.

[0082] In supervised learning, an artificial neural network is trained with labels for training data. Here, the label can mean the correct answer (or result value) that the artificial neural network should infer when training data is input to the artificial neural network. In this specification, the correct answer (or result value) that the artificial neural network should infer when training data is input is called a label or labeling data. In addition, in this specification, setting a label on training data for learning of the artificial neural network is called labeling the training data. In this case, the training data and the label corresponding to the training data constitute a single training set, and can be input to the artificial neural network in the form of a training set.

[0083] Meanwhile, training data represents multiple features, and labeling the training data can mean that the features represented by the training data are labeled. In this case, the training data can represent the features of the input object in vector form. An artificial neural network can use the training data and labeled data to infer a function regarding the relationship between the training data and the labeled data. Furthermore, the parameters of the artificial neural network can be determined (optimized) by evaluating the function inferred by the artificial neural network.

[0084] Unsupervised learning is a type of machine learning in which training data is not labeled. Specifically, unsupervised learning can be a learning method that trains an artificial neural network to find and classify patterns in the training data itself, rather than the relationship between the training data and the corresponding labels. Examples of unsupervised learning include clustering and independent component analysis (ICA). In this specification, the term "clustering" may be used interchangeably with the term "clustering."

[0085] Examples of artificial neural networks that utilize unsupervised learning include generative adversarial networks (GANs) and autoencoders (AEs).

[0086] Generative adversarial networks (GANs) are a machine learning method in which two different AI components, a generator and a discriminator, compete to improve performance. In this case, the generator is a model that creates new data, capable of generating new data based on original data. The discriminator, a model that recognizes data patterns, can determine whether the input data is original or new data generated by the generator. The generator learns from data that fails to fool the discriminator, while the discriminator learns from data that the generator deceives. Accordingly, the generator can evolve to fool the discriminator as effectively as possible, while the discriminator can evolve to effectively distinguish between original data and data generated by the generator.

[0087] An autoencoder is a neural network that aims to reproduce the input itself as an output. An autoencoder comprises an input layer, at least one hidden layer, and an output layer. In this case, since the number of nodes in the hidden layer is smaller than that in the input layer, the data dimensionality is reduced, leading to compression or encoding. Furthermore, data output from the hidden layer is fed into the output layer. In this case, since the number of nodes in the output layer is larger than that in the hidden layer, the data dimensionality increases, leading to decompression or decoding.

[0088] Meanwhile, autoencoders express input data as hidden layer data by adjusting the connection strengths of neurons through learning. The hidden layer expresses information with a smaller number of neurons than the input layer. The ability to reproduce input data as output implies that the hidden layer has discovered and expressed hidden patterns in the input data.

[0089] Semi-supervised learning is a type of machine learning that utilizes both labeled and unlabeled training data. One technique for semi-supervised learning is to infer labels for unlabeled training data and then use these inferred labels to perform training. This technique can be useful in situations where labeling is expensive.

[0090] Reinforcement learning is the theory that, if an agent is given an environment in which it can determine the optimal action at any given moment, it can find the optimal path through experience without data. Reinforcement learning is primarily implemented using a Markov Decision Process (MDP). A Markov Decision Process is described as follows: first, an environment containing the information necessary for the agent to take the next action is provided; second, how the agent will act in that environment is defined; third, what rewards the agent will receive for performing well and what penalties will be imposed for performing poorly is defined; and fourth, the optimal policy is derived through repeated experience until the future reward reaches its maximum.

[0091] The structure of an artificial neural network is specified by the model configuration, activation function, loss function or cost function, learning algorithm, optimization algorithm, etc., and the hyperparameters are set in advance before learning, and the model parameters are set through learning afterwards, so that the content can be specified.

[0092] For example, factors that determine the structure of an artificial neural network may include the number of hidden layers, the number of hidden nodes included in each hidden layer, the input feature vector, and the target feature vector.

[0093] Hyperparameters include various parameters that must be initially set for learning, such as initial values ​​for model parameters. Furthermore, model parameters include various parameters to be determined through learning. For example, hyperparameters may include initial values ​​for inter-node weights, initial values ​​for inter-node biases, mini-batch size, number of learning iterations, and learning rates. Furthermore, model parameters may include inter-node weights, inter-node biases, and more.

[0094] The loss function can be used as an indicator (standard) to determine the optimal model parameters during the learning process of an artificial neural network. In an artificial neural network, learning refers to the process of manipulating model parameters to reduce the loss function, and the purpose of learning can be seen as determining the model parameters that minimize the loss function. The loss function can mainly use the mean squared error (MSE) or the cross entropy error (CEE), but the present invention is not limited thereto. The cross entropy error can be used when the correct answer label is one-hot encoded. One-hot encoding is an encoding method that sets the correct answer label value to 1 only for neurons corresponding to the correct answer, and sets the correct answer label value to 0 for neurons that are not the correct answer.

[0095] In machine learning or deep learning, learning optimization algorithms can be used to minimize the loss function. Learning optimization algorithms include gradient descent (GD), stochastic gradient descent (SGD), momentum, Nesterov Accelerate Gradient (NAG), Adagrad, AdaDelta, RMSProp, Adam, and Nadam.

[0096] Gradient descent is a technique that adjusts model parameters in a direction that reduces the loss function value by considering the gradient of the loss function at the current state. The direction of model parameter adjustment is called the step direction, and the size of the adjustment is called the step size. Here, the step size can represent the learning rate. Gradient descent obtains the gradient by partially differentiating the loss function with respect to each model parameter, and updates the model parameters by changing the learning rate in the direction of the obtained gradient.

[0097] Stochastic gradient descent is a technique that divides learning data into mini-batches and performs gradient descent on each mini-batch to increase the frequency of gradient descent.

[0098] Adagrad, AdaDelta, and RMSProp are techniques for improving optimization accuracy by adjusting the step size in SGD. In SGD, momentum and NAG are techniques for improving optimization accuracy by adjusting the step direction. Adam combines momentum and RMSProp to improve optimization accuracy by adjusting the step size and step direction. Nadam combines NAG and RMSProp to improve optimization accuracy by adjusting the step size and step direction.

[0099] The learning speed and accuracy of artificial neural networks are significantly influenced by not only the network structure and the type of learning optimization algorithm, but also hyperparameters. Therefore, to obtain a good learning model, it is crucial not only to determine an appropriate artificial neural network structure and learning algorithm, but also to set appropriate hyperparameters.

[0100] Typically, hyperparameters are experimentally set to various values ​​while training an artificial neural network, and the learning results are set to the optimal values ​​that provide stable learning speed and accuracy.

[0101] Embodiments of the present invention are applied to a chatbot that converses with a user through voice or text-based input and performs specific tasks, understanding the speaker's questions and outputting appropriate answers. The chatbot's structure may include, depending on its core functions, question intent classification, named entity recognition, core keyword extraction, and answer search. The question intent classification function identifies the speaker's question intent and predicts the intent class for the question using an intent classification model. The named entity recognition function recognizes named entities by word token in the speaker's question, and the core keyword extraction function extracts nouns or verbs that are core to the meaning of the speaker's question using a morphological analyzer. The answer search function searches for and displays appropriate answers corresponding to the question intent, recognized named entities, and extracted core keywords from a learning database. In other words, when a user's query sentence is entered, the chatbot preprocesses the sentence, extracts keywords (word tokens) through a morphological analyzer, extracts only necessary keywords such as nouns and verbs, and removes stop words. Afterwards, the chatbot performs intent analysis and entity recognition on the extracted keywords to derive a corresponding response. For this purpose, chatbots generally use deep learning models such as intent analysis (classification) models and entity recognition models for natural language processing.

[0102] This embodiment may be related to fine tuning or prompt engineering for improving the accuracy of an intent classification model, which is one of the main components of such a chatbot. Fine tuning or prompt engineering is a technique for improving the accuracy and usability of a deep learning model. Fine tuning is a method that retrains a pre-trained model to fit a specific task or data set, and is a method that requires domain-specific data. Prompt engineering optimizes the input prompts of a model to obtain a desired output result, and can be usefully applied to various tasks in a relatively short period of time using a relatively small amount of data compared to fine tuning. A method and system for learning an intent classification model using intent descriptions according to an embodiment of the present invention may include both an intent classification model that retrains a pre-trained model based on learning data including intent descriptions, and an intent classification model that performs inference from the intent classification model based on input data (input prompts) including intent descriptions.

[0103] The method and system for learning an intent classification model using intent description according to one embodiment of the present invention can be applied to various fields using chatbots based on large-scale language models. The intent classification model optimized according to this embodiment can be utilized in customer support chatbots based on corporate customer response data, domain models optimized for specific domains (work fields) such as medicine, law, and finance, various natural language processing tasks such as text generation / summarization / translation, market trend analysis reflecting the latest information and trends, and personalized recommendation systems based on user service usage data.

[0104] Figure 1 is a schematic diagram of an electronic device according to one embodiment of the present invention.

[0105] As illustrated in FIG. 1, an electronic device (100) (hereinafter referred to as an electronic device) according to an embodiment of the present invention may include at least one processor (110), a memory (120), and a communication unit (130). The electronic device (100) is a basic configuration for performing a computing environment, and in other embodiments, the electronic device (100) may be implemented by additionally or alternatively including some other components, may be implemented as a single or multiple entities, or may be implemented as only some of the disclosed components. Components or at least some of the components inside or outside the electronic device (100) may be connected to each other through a BUS, a GPIO (General Purpose Input / Output), an SPI (Serial Peripheral Interface), or a MIPI (Mobile Industry Processor Interface), thereby transmitting and receiving data or signals.

[0106] Unless the context clearly indicates otherwise, the processor (110) may refer to a set of one or more processors, and may control components of the processor (110) and the electronic device (100) by executing software (e.g., commands, programs, etc.) stored in at least the memory (120). In addition, the processor (110) may perform various operations such as calculations, processing, data generation or processing, and may read data from or store data in the memory (120). The processor (110) may be composed of at least one core and may include a processor for data analysis, machine learning (ML), or deep learning (DL), such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or a tensor processing unit (TPU). The processor (110) may read software stored in the memory (120) to perform data processing for machine learning (or deep learning) of the present invention. According to one embodiment of the present disclosure, the processor (110) can perform operations for learning a neural network. The processor (110) can perform calculations for learning a neural network, such as processing input data for learning in deep learning, extracting features from the input data, calculating errors, and updating weights of the neural network using backpropagation. At least one of the CPU, GPGPU, and TPU of the processor (110) can process learning of a neural network model. For example, the CPU and GPGPU can together process learning of a neural network model and data classification using a neural network model. In addition, in one embodiment of the present disclosure, at least one processor (110) of the electronic device (100) can be used together to process learning of a neural network model and data classification using a neural network model.

[0107] The memory (120) is for storing various data, and the data is data acquired, processed, or used by at least one component of the electronic device (100), and may include software (e.g., commands, programs, etc.). Unless explicitly expressed otherwise in the context, the memory (120) may refer to a set of one or more memories, and may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., an SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, an optical disk, and a web storage that performs a storage function on the Internet. The instructions or programs or software stored in the memory (120) may be used to refer to an operating system, an application for controlling components of the electronic device (100), or middleware that provides various functions to the application so that the application can utilize the components of the electronic device (100). In one embodiment, when the processor (110) performs a specific operation, the memory (120) may store instructions that are performed by the processor (110) and correspond to the specific operation.

[0108] The communication unit (130) performs wireless or wired communication between the electronic device (100) and another device (e.g., a user terminal or another server), and the communication unit (130) can use wireless communication systems according to methods such as eMBB, URLLC, MMTC, LTE, LTE-A, NR, UMTS, GSM, CDMA, WCDMA, TDMA, FDMA, OFDMA, SCFDMA, WiBro, WiFi, Bluetooth, NFC, GPS, or GNSS. In addition, the communication unit (130) can use various wired communication systems such as USB, HDMI, RS-232 (Recommended Standard-232), POTS (Plain Old Telephone Service), PSTN (Public Switched Telephone Network), xDSL (x Digital Subscriber Line), RADSL (Rate Adaptive DSL), MDSL (Multi Rate DSL), VDSL (Very High Speed ​​DSL), UADSL (Universal Asymmetric DSL), HDSL (High Bit Rate DSL), and local area network (LAN). In one embodiment of the present invention, the communication unit (130) can be configured regardless of the communication mode such as wired or wireless, and can be configured with various communication networks such as a personal area network (PAN), a wide area network (WAN), etc. Additionally, the network may be the well-known World Wide Web (WWW), or may utilize a wireless transmission technology used for short-range communication, such as Infrared Data Association (IrDA) or Bluetooth. The technologies described in one embodiment of the present invention can also be used in other networks mentioned above.

[0109] An electronic device (100) according to an embodiment of the present invention can execute software that configures a learning system for an intention classification model using an intention description or configures a learning method for an intention classification model using an intention description.

[0110] The learning system of the intent classification model according to an embodiment of the present invention can be applied to a task-oriented dialogue system such as a chatbot to improve the accuracy of the intent classification model for understanding the intent of a user query. In general, an intent classification model can predict the domain (or class) used for learning. Therefore, even when unseen data is input to such a pre-trained intent classification model, it is necessary to smoothly perform prediction or classification of the unseen data. In this case, additional work (fine-tuning or prompt engineering) may be required for the intent classification model. In other words, even when unseen data is input to the intent classification model, the intent classification model may be modified to predict or classify an unseen domain (class; unseen domain or unseen class) not included in the learning data set, rather than classifying it into one of the learned domains (classes; seen domain or seen class) by its internal mechanism. When modifying such an intent classification model, the learning system and method of the intent classification model of the present embodiment can be applied.

[0111] In one embodiment, independent intent descriptions, dependent (or "dependent") intent descriptions, and refined intent descriptions can be generated. An independent intent description can be obtained by providing a description for only one intent and three user query examples for that intent. By omitting information about other intents, the intent can be described individually without comparison. Consequently, independent intent descriptions may not sufficiently reflect the relative differences between other intents, resulting in relatively low quality. Dependent intent descriptions can be generated by providing a prompt that includes a list of all intents, allowing each intent description to be distinguished from other intents. Therefore, generated dependent intent descriptions are likely to have higher quality than independent intent descriptions because they clearly distinguish specific intents from others. Considering the possibility that automatically generated intent descriptions, such as independent and dependent intent descriptions, may not clearly reflect the differences between intents, a refined intent description, manually refined by a human, can also be generated. A refined intent description is a manual description created by a human to clearly distinguish intents from other intents, resulting in higher quality than automatically generated descriptions.

[0112] A learning system for an intent classification model using intent descriptions according to one embodiment of the present invention outputs a data set including intent descriptions of an intent classification model from a language model, uses at least a portion of the data set as an input prompt of the intent classification model to perform intent classification, and determines the performance of the intent classification model using at least a portion of the data set as an input prompt of the intent classification model. Here, the data set may include first data including independent intent descriptions and second data including dependent intent descriptions. That is, the intent classification model may be retrained or tested with the first data or the second data.

[0113] In one embodiment, the learning system may obtain or generate a data set for retraining a pre-trained intent classification model from a separate language model. The language model of one embodiment of the present invention may be an AI chatbot that generates output in response to an input prompt, such as ChatGPT. The learning system of one embodiment may obtain two data sets from the language model, the first data including independent intent descriptions and the second data including dependent intent descriptions.

[0114] That is, as illustrated in FIG. 2, the learning system of one embodiment can obtain first data and second data by inputting a command prompt (independent prompt) to include an independent intent description in a language model such as ChatCPT and a command prompt (dependent prompt) to include a dependent intent description, respectively.

[0115] Referring to FIG. 2, first data can be acquired through an independent prompt to include an intent description for a single intent. That is, the first data acquired through the independent prompt can include an intent name and an intent description for the intent name. The first data acquired through the independent prompt includes multiple intent names and intent descriptions for the intent names, and the intent classification model makes it difficult or impossible to determine the differences between the multiple intent names included in the first data (indistinguishability). That is, the independent prompt includes a single intent and three user queries for the intent, and other intents are excluded from the prompt. Accordingly, the collected intent description lacks a comparative context with other intents, resulting in a low quality of the intent description. In one embodiment, the independent prompt can be invoked multiple times to acquire multiple independent intent descriptions for a single intent name.

[0116] Referring to FIG. 2, second data can be acquired through a dependent prompt to include intent descriptions for multiple intents. That is, the second data acquired through the dependent prompt can include multiple intent names and intent descriptions for the corresponding intent names. This second data includes multiple intent names and intent descriptions for the corresponding intent names, and the differences between the multiple intent names included in the second data can be determined by an intent classification model. That is, a dependent prompt includes multiple intents and three user queries for each intent in a single prompt. Multiple intents are included in a single dependent prompt, and the prompt's instructions acquire intent descriptions for the multiple intents, but each intent description can be set so as not to include a superordinate or subordinate concept of another intent description. Accordingly, the generated intent descriptions including all possible intents within the prompt can be uniquely distinguished from each other, and the quality of the intent descriptions in the second data can be higher than that of the first data. In one embodiment, the dependent prompt may be called multiple times to obtain multiple dependent intent descriptions for a single intent name.

[0117] In this way, the data set obtained from the language model can be used as retraining data for fine-tuning the intent classification model and as test data for testing or determining the performance of the intent classification model after fine-tuning has been completed.

[0118] Meanwhile, in one embodiment, in addition to the first or second data, third data may be included, in which the intent and the intent description for the intent are manually reviewed and adjusted by the operator. The third data may include refined intent descriptions. That is, the third data is the highest quality data for determining the accuracy of the intent classification model. The third data may be used solely for validating the intent classification model, or may be selectively included in the data set to improve the accuracy of the intent classification model.

[0119] In addition, the data set secured as above can be set in the form of an input prompt for fine-tuning an intention classification model, and an example of an input prompt based on the data set is illustrated in Fig. 3.

[0120] Referring to FIG. 3, the input prompt inputted into the intent classification model can be divided into a section including an instruction and the user query for classifying the intent given to the model, and a section including intent options. Here, the intent options can include an intent name and an intent description for the intent name. In other words, the input prompt for fine-tuning the intent classification model can be generated based on the first data, the second data, and additionally the third data, and the formats of the input prompts based on the first to third data can be generated identically.

[0121] In one embodiment, the learning system can perform intent classification of an intent classification model by inputting an input prompt based on a data set of a language model as described above to improve the performance of the intent classification model. In this case, the input prompt input to the intent classification model may be based on the first data, the second data, or the third data, or may include all of these.

[0122] Additionally, the learning system of one embodiment can determine the performance of an intent classification model by inputting a test data set into a retrained intent classification model based on a retraining data set.

[0123] In this way, an intention classification model retrained based on a data set including intention descriptions can have improved classification accuracy compared to a model that has not been retrained, and the evaluation of the model will be described below.

[0124] Meanwhile, the learning system of one embodiment may further include a ranker to adjust the number of intent options included in the input prompt for retraining. If the input prompt for model retraining includes intent descriptions, the length of the input prompt increases proportionally with the number of intent options.

[0125] For example, assuming that the model can receive up to 1024 tokens as input, providing a 10-word description for each of 100 intents will likely exceed the input length limit. To address this issue, in one embodiment, the ranker calculates the similarity between the user query and the intent description, sorts all intent options in descending order of similarity, and then passes only the top k intents to the intent classification model. In one embodiment, by selecting the k intent options with the highest similarity, the input length can be optimized instead of including the entire intent in the model input. The performance evaluation of the intent classification model by increasing or decreasing the number of intent options will be described below.

[0126] As illustrated in FIG. 4, a method for learning an intention classification model using a language model implemented in a learning system for an intention classification model using a language model according to an embodiment of the present invention may include a step of generating a data set including an intent description of the intention classification model using the language model (S110), a step of using at least a part of the data set as an input prompt of the intention classification model to perform intent classification by the intention classification model (S120), and a step of determining the performance of the intention classification model using at least a part of the data set as an input prompt of the intention classification model (S130). Here, the data set may include first data including an independent intent description and second data including a dependent intent description.

[0127] In the step (S110) of generating a data set including the intent description of the intent classification model using the language model, the learning system can obtain a data set for the pre-trained intent classification model from a separate language model. That is, the learning system of the present embodiment can obtain a data set for fine-tuning the intent classification model from a separate language model, for example, ChatGPT. To obtain the data set from the language model, the learning system inputs a pre-written command prompt into the language model to obtain the data set. At this time, the command prompt input into the language model can include two command prompts to determine the quality of the data. The independent prompt can be a command prompt for obtaining first data including the independent intent description from the language model, and the dependent prompt can be a command prompt for obtaining second data including the dependent intent description.

[0128] The first data obtained through independent prompts includes multiple intent names and intent descriptions for those intent names. However, the intent classification model makes it difficult or impossible to determine the differences between the multiple intent names included in the first data. In other words, the intent descriptions in the first data lack comparative context, resulting in low quality intent descriptions.

[0129] The second data obtained by the dependent prompt includes multiple intent names and intent descriptions for the corresponding intent names. This second data includes multiple intent names and intent descriptions for the corresponding intent names, and the intent classification model allows the intent names included in the second data to be distinguished from each other. In other words, since the intent descriptions in the second data can be uniquely distinguished from each other, the quality of the intent descriptions in the second data can be higher than that of the first data.

[0130] In this way, the data set obtained from the language model can be used as retraining data for fine-tuning the intent classification model and as test data for testing the intent classification model after fine-tuning has been completed.

[0131] In the step (S120) where the intent classification model performs intent classification by using at least a part of the data set as an input prompt of the intent classification model, the input prompt inputted to the intent classification model may have a form of a section including an instruction and a user query for classifying the intent given to the model (instruction and the user query) and a section including an intent option (intent option). Accordingly, the data set acquired from the language model is reset to the form of an input prompt of the intent classification model. A first input prompt based on first data configured to include an independent intent description or a second input prompt based on second data configured to include a dependent intent description may be inputted to the intent classification model to fine-tune the intent classification model. An intent classification model retrained with the first input prompt and an intent classification model retrained with the second input prompt may differ in classification accuracy and performance, and a performance comparison between these models will be described below.

[0132] In the step S130 of determining the performance of the intent classification model by using at least a portion of the data set as an input prompt of the intent classification model, the learning system of one embodiment may perform a test using a first input prompt based on first data including an independent intent description obtained from a language model or a second input prompt based on second data including a dependent intent description. At this time, the input prompts for the test are distinct from the input prompts for training the intent classification model, and the input prompts for the test are input to a baseline model that has not been retrained, i.e., a vanilla intent classification model, an intent classification model retrained with the first input prompt, and an intent classification model retrained with the second input prompt, respectively, so as to test each model and evaluate the performance of the model.

[0133] Meanwhile, a method for learning an intent classification model using a language model of one embodiment may further include a step of adjusting the number of intent options included in a first input prompt based on first data, a second input prompt based on second data, and a third input prompt based on third data, i.e., adjusting the number of intent options of an input prompt of an intent classification model.

[0134] To this end, the learning system of one embodiment can select only the top k intents by sorting all intent options in descending order of similarity in a ranker that calculates the similarity between the user query and the intent description included in the input prompt.

[0135] The performance of an intention classification model can be evaluated according to a learning method and system for an intention classification model using an intention description of an embodiment of the present invention having a configuration as described above.

[0136] The intent classification performance of each model can be evaluated by testing a first intent classification model retrained with a first input prompt including independent intent descriptions derived from a language model, a second intent classification model retrained with a second input prompt including dependent intent descriptions, and a baseline model trained without intent descriptions.

[0137] Types of Descriptions Used in Testingwithout descriptionsindependent descriptiondependent descriptionTypes of Descriptions Used in Trainingwithout descriptions84.28%±3.95%84.15%±3.67%90.55%±2.16%independent description81.93%±4.05%85.64%±3.53%90.97%±1.45%dependent description82.1%±3.56%86.99%±3.09%91.75%±1.91%

[0138] As shown in [Table 1], the first intent classification model (independent descriptions), the second intent classification model (dependent descriptions), and the baseline model (without descriptions) were tested on a test data set that did not include intent descriptions. As a result, the baseline model showed the highest performance of 84.28%, the first intent classification model showed a lower performance of 81.93%, and the second intent classification model showed a performance of 82.1%. However, when each model was tested with the first input prompt (independent descriptions) that included independent intent descriptions, the first intent classification model and the second intent classification model showed performances of 85.64% and 86.99%, respectively, which were 1.49% and 2.84% better than the baseline model. When testing each model with a second input prompt containing dependent intent descriptions, the first and second intent classification models achieved accuracy rates of 90.97% and 91.75%, respectively, representing a 0.42% and 1.2% performance improvement over the baseline model. This demonstrates that when testing based on input prompts containing intent descriptions, intent descriptions are more effective in interpreting the detailed meaning of the intent of a user's utterance, and that adjusting the model based on intent descriptions can improve performance in testing.

[0139] In addition, the second intent classification model retrained with dependent intent descriptions shows a performance of 86.99% when tested with the first input prompt including independent intent descriptions, whereas it shows a performance of 91.75% when tested with the second input prompt including dependent intent descriptions. This confirms that improving the quality of input prompts including intent descriptions during testing has a greater impact on improving model performance than model training.

[0140] In addition, the second intent classification model retrained based on the second input prompt including the dependent intent description was shown to have improved performance during testing compared to the first intent classification model, which confirms that the performance of the model can be improved by improving the quality of the input prompt for retraining.

[0141] Types of Descriptions Used in Testing1 dependent description1 cleansed descriptionTypes of Descriptions Used in Training5 dependent descriptions82.05%±3.15%85.28%±1.69%1 dependent descriptions82.52%±2.44%85.71%±1.61%1 cleansed descriptions83.4%±3.71%85.79%±1.69%

[0142] Meanwhile, in one embodiment, performance can be evaluated based on the number of intent descriptions (the number of intent options) within the input prompts used during retraining and testing of an intent classification model. As shown in [Table 2], testing was conducted using the trained model by varying the number of intent descriptions in a second input prompt containing dependent intent descriptions. At this time, performance was also evaluated based on the quality of the intent descriptions using a third input prompt based on third data qualitatively adjusted by a human operator, including data containing dependent intent descriptions.

[0143] After preparing an intent classification model retrained using input prompts containing 5 dependent intent descriptions per intent (5 dependent descriptions), a retrained intent classification model retrained using input prompts containing 1 dependent intent description per intent (1 dependent description), and a retrained intent classification model retrained using input prompts containing 1 dependent intent description per intent that was manually adjusted (refined) by a human operator (1 cleansed description), each model was tested based on input prompts containing 1 dependent intent description per intent and 1 dependent intent description per intent that was manually adjusted by a human operator. In this case, the input prompts manually adjusted by a human operator had the highest quality intent descriptions. As shown in [Table 2], the performance of the model trained with a single intent description and the model trained with multiple intent descriptions did not differ significantly, but the model trained with higher quality intent descriptions (manually adjusted by a human operator) was confirmed to have the highest performance. This confirmed that the intent classification model is more affected by the quality of the intent descriptions than the quantity of the intent descriptions.

[0144] Meanwhile, referring to FIG. 5, the number of intent options (number of intent descriptions) included in the input prompt for retraining the intent classification model can be optimized. In one embodiment, the learning system includes a ranker for controlling the number of intent options included in the input prompt for retraining. The ranker can calculate the similarity between the user query and the intent description of the input prompt and sort all intent options in descending order according to the similarity. In other words, the ranker performs the function of passing only the top k intents among the descendingly sorted intent options to the intent classification model. Referring to FIG. 5, it can be confirmed that the performance of the intent classification model improves as the number of top intent options k increases. For example, in the CLINC data set, the intent classification model exhibits an accuracy (performance) of approximately 44.21% when k is 1, and its performance improves as k increases, reaching an accuracy of approximately 90% when k reaches 13. Based on these results, it is confirmed that the optimal performance is achieved when retraining or testing the intent classification model using input prompts having the top 10 intent options. For example, if the CLINC data set contains 75 intent options and each intent description is 10 tokens long, approximately 1,200 to 1,300 tokens are required, but if only the top 10 intent options are used using the ranker, the length of the input prompt can be reduced to 300 to 400 tokens. In other words, the learning system of this embodiment can obtain the same similar performance in the intent classification model while reducing the input size by approximately 75% by setting the number of intent descriptions in the input prompt to 10 using the ranker.

[0145] As described above, according to the learning method and system for an intention classification model using an intention description of an embodiment of the present invention, and the evaluation results of the learning system of an embodiment, the intention classification model retrained based on the intention description of the intention shows higher performance when a prompt based on the intention description is used during testing, and it can be confirmed that a higher quality of the intention description during training and testing of the model affects the improvement of the performance of the model, and the quality of the intention description is more effective in improving the performance of the model than the number of the intention descriptions.

[0146] In one embodiment, the processor may receive a user's utterance as input, refer to a database including a plurality of intent candidates and a natural language-based description of the intent, evaluate a similarity between the user's utterance and the natural language-based description, determine a top k intent candidate having a high relevance to the user's utterance based on the similarity evaluation, and determine an intent appropriate for the user's utterance using the determined k intent candidates and the natural language-based description as input.

[0147] In one embodiment, the natural language-based description may include at least one of an independent intent description, a dependent intent description, and a refined intent description.

[0148] Meanwhile, one embodiment of the present invention may be implemented as an application-specific integrated circuit (ASIC) manufactured to suit the special functions of a specific application field and device.

[0149] An application-specific integrated circuit is also called an application-specific semiconductor. Unlike standard semiconductors that have set specifications and can be applied to any electronic product or application as long as certain requirements are met, an application-specific semiconductor is an integrated circuit that a semiconductor manufacturer manufactures according to a specific order for a specific product or function. In other words, an application-specific semiconductor is designed and manufactured to perform only the functions required for a specific device or specific function. Depending on the design method, application-specific semiconductors are largely divided into full custom ICs, which design and manufacture the circuit from scratch according to the user's needs, and semi-custom ICs, which design and manufacture the circuit using some standardized designs.

[0150] Application-specific semiconductors are primarily used in communications systems, high-performance computing systems, consumer electronics, automobiles, industrial automation, medical devices, military, and aerospace industries. Recently, they are being applied to AI semiconductors that perform large-scale calculations required for AI implementation with high performance and power efficiency.

[0151] Application-specific integrated circuits (ASICs) are core components of network routers, switches, and modems in communication systems, performing data packet processing, protocol conversion, and signal processing to deliver high throughput and low latency. In high-performance computing systems, ASICs are key components for high-speed and parallel processing. In consumer electronics such as digital cameras, smartphones, tablets, and game consoles, ASICs provide high-performance and low-power solutions required to perform specific functions. In the automotive industry, ASICs control various electronic systems within vehicles, and in industrial automation systems, ASICs provide solutions for high-precision control and high-performance processing.

[0152] An application-specific integrated circuit to which one embodiment of the present invention is applied includes a memory in which an individual memory interface (I / F) is implemented, and may include a plurality of functional blocks that request memory access. Each functional block may be a direct memory access (DMA) functional block, a processor, a video processor, a cache controller, a decompression block, or a data path block. The basic configuration of the application-specific integrated circuit may include a transistor that amplifies or switches an electrical signal, a logic gate that is a circuit that performs a logical function by combining transistors, a memory cell that stores data, an analog circuit that is a circuit that processes a continuous voltage or current by combining transistors, and an IP core (Intellectual Property Core) such as a microprocessor, DSP, or graphic core that is pre-designed to perform a specific function.

[0153] The ASIC may also include a separate memory I / F interfacing with individual memories and an embedded memory I / F interfacing with embedded memories. The separate memory I / F is connected to each functional block, receives memory access signals (e.g., control signals, address signals, and data signals), and generates signals for controlling the individual memories based on these input signals. The embedded memory I / F is connected to each functional block, receives memory access signals (e.g., control signals, address signals, and data signals), and generates modified memory access signals for controlling the embedded memories based on these input signals. The separate memory I / F and the embedded memory I / F may be designed within the memory control block of the ASIC to provide a memory control structure that can be flexibly applied to both the individual memories and the embedded memories.

[0154] Additionally, an application-specific integrated circuit (ASIC) for an artificial neural network (ANN) may be configured to include a plurality of neurons arranged in an array and a plurality of synaptic circuits, each neuron including a register, a microprocessor, and at least one input, and each synaptic circuit including a memory for storing synaptic weights. Each neuron of the ASIC may be connected to at least one other neuron through one of the plurality of synaptic circuits.

[0155] Although the present disclosure has been described above as being generally implemented by a computing device, those skilled in the art will appreciate that the present disclosure may also be implemented in combination with computer-executable instructions and / or other program modules that may be executed on one or more computers and / or as a combination of hardware and software.

[0156] Those skilled in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips referenced in the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0157] Those skilled in the art will appreciate that the various illustrative logical blocks, modules, processors, means, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, various forms of programs or design code (referred to herein, for convenience, as software), or a combination of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0158] The various embodiments presented herein can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term article of manufacture includes a computer program, carrier, or media accessible from any computer-readable storage device. For example, computer-readable storage media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., CDs, DVDs, etc.), smart cards, and flash memory devices (e.g., EEPROMs, cards, sticks, key drives, etc.). Furthermore, various storage media presented herein include one or more devices and / or other machine-readable media for storing information.

[0159] It should be understood that the specific order or hierarchy of steps in the presented processes is merely an example of exemplary approaches. It should be understood that the specific order or hierarchy of steps in the processes may be rearranged within the scope of the present disclosure based on design priorities. The appended method claims provide elements of various steps in a sample order, but are not intended to be limited to the specific order or hierarchy presented.

[0160] The description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments disclosed herein, but is to be construed in the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. As a learning system for an intention classification model using intention explanation, at least one processor; and At least one memory for storing instructions or information executed by at least one processor; The operation performed by the above command or information executed by at least one processor is: The act of generating a data set containing the intent description of the intent classification model using a language model; An operation of performing intent classification by using at least a portion of the above data set as an input prompt of the intent classification model; and An operation of determining the performance of the intent classification model by using at least a portion of the data set as an input prompt of the intent classification model; The above data set includes first data including independent intent descriptions in which intent names and intent descriptions for the intent names are mutually indistinguishable, and second data including dependent intent descriptions in which intent names and intent descriptions for the intent names are mutually distinguishable. A learning system for an intention classification model using intention descriptions.

2. In claim 1, The command prompt input to the language model to output the data set from the language model is: An independent prompt for obtaining the first data including multiple intent descriptions for a single intent; and a dependent prompt for obtaining the second data including multiple intent descriptions for multiple intents, wherein each of the multiple intents and each of the multiple intent descriptions is distinct from each other. A learning system for an intention classification model using intention descriptions.

3. In claim 1, The above input prompt consists of an instruction and user query section and an intent options section containing an intent name and intent descriptions for the intent name. A learning system for an intention classification model using intention descriptions.

4. In claim 3, It further includes a ranker that controls the number of the above intention options, The above ranker calculates the similarity between the user query and the intent description. A learning system for an intention classification model using intention descriptions.

5. In claim 1, The above intention classification model is retrained using at least a portion of the first data, or is retrained using at least a portion of the second data. A learning system for an intention classification model using intention descriptions.

6. In claim 1, At least a portion of the first data is provided as an input prompt when determining the performance of the intent classification model, or at least a portion of the second data is provided as an input prompt when determining the performance of the intent classification model. A learning system for an intention classification model using intention descriptions.

7. In claim 1, The above intention classification model, when retrained with at least a portion of the second data, improves the accuracy performance of intention classification. A learning system for an intention classification model using intention descriptions.

8. A learning method for an intention classification model using an intention description, wherein the learning method for the intention classification model is executed by at least one processor, A step of generating a data set containing an intent description of an intent classification model using a language model; A step of performing intent classification by using at least a part of the data set as an input prompt of the intent classification model; and A step of determining the performance of the intent classification model by using at least a portion of the data set as an input prompt of the intent classification model; The above data set includes first data including independent intent descriptions in which intent names and intent descriptions for the intent names are mutually indistinguishable, and second data including dependent intent descriptions in which intent names and intent descriptions for the intent names are mutually distinguishable. A learning method for an intention classification model using intention descriptions.

9. In claim 8, In the step of generating the data set, a command prompt input to the language model to generate the data set from the language model is configured, the command prompt including an independent prompt to obtain the first data including a plurality of intent descriptions for a single intent; and a dependent prompt to obtain the second data including a plurality of intent descriptions for a plurality of intents, but such that each of the plurality of intents and the plurality of intent descriptions is distinguished from each other. A learning method for an intention classification model using intention descriptions.

10. In claim 8, In the step of performing intent classification by the above intent classification model or in the step of determining the performance of the above intent classification model, The above input prompt consists of an instruction and user query section and an intent options section containing an intent name and intent descriptions for the intent name. A learning method for an intention classification model using intention descriptions.

11. In claim 10, The ranker that adjusts the number of the above intent options calculates the similarity between the user query and the intent description and sorts the intent descriptions in descending order according to the similarity. A learning method for an intention classification model using intention descriptions.

12. In claim 8, In the step where the intention classification model performs intention classification, the intention classification model is retrained using at least a part of the first data, or is retrained using at least a part of the second data. A learning method for an intention classification model using intention descriptions.

13. In claim 8, In the step of determining the performance of the intention classification model, at least a part of the first data is provided as an input prompt when determining the performance of the intention classification model, or at least a part of the second data is provided as an input prompt when determining the performance of the intention classification model. A learning method for an intention classification model using intention descriptions.

14. In claim 8, The above intention classification model, when retrained with at least a portion of the second data, improves the accuracy performance of intention classification. A learning method for an intention classification model using intention descriptions.

15. A program stored in a computer-readable recording medium that causes a computer to execute the method of any one of claims 8 to 14.

Citation Information

Patent Citations

  • Dust solidification apparatus

    KR1020220011578A

  • Method for predicting the remaining life of industrial equipment using generative adversarial networks, and apparatus thereof

    KR1020240154282A

  • Content addressable memory for large search words

    KR102690877B1

  • Intent re-ranker

    US20200279555A1

  • Multiple semantic hypotheses for search query intent understanding

    US20230004568A1