Logit enhanced for natural language processing

Enhanced logit values in chatbot systems address the overconfidence issue of deep neural networks, enhancing the accuracy of out-of-domain utterance classification.

JP2025166050AActive Publication Date: 2025-11-05ORACLE INT CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025130505
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-29
Filing Date
2025-08-05
Publication Date
2025-11-05
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Existing chatbot systems face challenges in accurately classifying out-of-scope or out-of-domain utterances due to the overconfidence of deep neural networks, leading to inaccurate classification and increased frequency of unpredictable results.

Method used

The use of enhanced logit values, determined through statistical, bounded, centroid-weighted, hyperparameter-tuned, or learned values, to improve the classification of utterances as either resolvable or unresolvable classes, addressing the overconfidence issue in neural networks.

Benefits of technology

Enhanced logit values enhance the accuracy of chatbot systems in identifying out-of-domain utterances, reducing incorrect classifications and improving overall classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025166050000001_ABST
    Figure 2025166050000001_ABST
Patent Text Reader

Abstract

To provide a chatbot system which can use a logit value enhanced for classifying speech and messages input to a chatbot system in natural language processing.SOLUTION: A chatbot system receives speech of a user, and inputs the speech into a machine learning model including a series of network layers. A final network layer of the series of the network layers. The machine learning model calculates a first probability for a class which can be solved and a second probability for a class which cannot be solved to use a logit function, maps the first probability for the class which can be solved to a first logit value and maps the second probability for the class which cannot be solved to an enhanced logit value, and, based on the first logit value and the enhanced logit value, classifies speech into the class which can be solved and the class which cannot be solved.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 119,449, filed November 30, 2020, and U.S. Non-Provisional Patent Application No. 17 / 456,687, filed November 29, 2021. Each of the above-referenced applications is incorporated herein by reference in its entirety for all purposes.

[0002] Technical Field The present disclosure relates generally to chatbot systems, and more particularly to techniques for using enhanced logit values ​​to classify utterances and messages input to chatbot systems in natural language processing. [Background technology]

[0003] background Many users around the world are on instant messaging or chat platforms to get immediate responses. Organizations often use these instant messaging or chat platforms to engage in live conversations with customers (or end users). However, hiring service personnel to engage in live communication with customers or end users can be very costly for organizations. Chatbots, or bots, have begun to be developed to simulate conversations with end users, especially over the Internet. End users can communicate with bots through messaging apps that the end users already have installed and are using. Intelligent bots, generally driven by artificial intelligence (AI), can communicate more intelligently and contextually in live conversations, thus enabling more natural conversations between bots and end users for an improved conversational experience. Instead of end users learning a fixed set of keywords or commands that the bot knows how to respond to, intelligent bots can understand the end user's intent based on end user utterances in natural language and respond accordingly. Summary of the Invention

[0004] overview A technique is provided for using enhanced logit values ​​to classify utterances and messages input to a chatbot system in natural language processing. One method includes a chatbot system receiving an utterance generated by a user interacting with the chatbot system and inputting the utterance to a machine learning model including a series of network layers. The utterance may include text data converted from speech input by the user. A final network layer in the series of network layers may include a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value.

[0005] The method may also include the machine learning model determining a first probability for the solvable class and a second probability for the unresolvable class. The machine learning model may use a logit function to map the first probability for the solvable class to a first logit value. The logit function for mapping the first probability may be the logarithm of the odds corresponding to the first probability for the solvable class, where the logarithm of the odds is weighted by the centroid of a distribution associated with the solvable class.

[0006] The method may also include the machine learning model mapping the second probability for the unresolvable class to an enhanced logit value. The enhanced logit value may be a third real number determined independently of the logit function used to map the first probability. The enhanced logit value may be selected from a range of values ​​defined by (i) a statistical value determined based on a set of logit values ​​generated from the training dataset and (ii) the logarithm of the first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being bounded within the range by a bounding function. The logit value may be a bounded value constrained to a range of values ​​by a bounding function, scaled by a scaling factor, and weighted by a centroid of a distribution associated with the unresolvable class, (iii) a weighted value generated by the logarithm of second odds corresponding to a second probability for the unresolvable class, the logarithm of the second odds being constrained to the range of values ​​by a bounding function, scaled by a scaling factor, and weighted by a centroid of a distribution associated with the unresolvable class, (iv) a hyperparameter optimized value generated based on hyperparameter tuning of a machine learning model, or (v) a learned value adjusted during training of the machine learning model. The method may also include the chatbot system classifying the utterance as a solvable class or an unresolvable class based on the first logit value and the enriched logit value.

[0007] Techniques are also provided for training a machine learning model that uses the enhanced logit values ​​to classify utterances and messages. A method may include a training subsystem receiving a training dataset. The training dataset may include a plurality of utterances generated by a user interacting with a chatbot system. At least one utterance of the plurality of utterances may include text data converted from the user's voice input. The method may also include the training subsystem accessing a machine learning model including a series of network layers. A final network layer in the series of network layers includes a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value.

[0008] The method may also include the training subsystem training the machine learning model with the training dataset, such that the machine learning model (i) determines a first probability for the solvable class and a second probability for the unresolvable class, and (ii) maps the first probability for the solvable class to a first logit value using a logit function. The logit function for mapping the first probability may be the logarithm of the odds corresponding to the first probability for the solvable class, where the logarithm of the odds is weighted by the center of gravity of a distribution associated with the solvable class.

[0009] The method may also include the training subsystem replacing the logit function with an enhanced logit value such that the second probability for the unresolvable class is mapped to the enhanced logit value. The enhanced logit value may be a third real number determined independently from the logit function used to map the first probability. The enhanced logit value may be selected from a range of values ​​defined by (i) a statistical value determined based on a set of logit values ​​generated from the training dataset, (ii) a bounded value selected from a range of values ​​defined by the logarithm of the first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function and weighted by a centroid of a distribution associated with the unresolvable class, and (iii) a bounded value generated by the logarithm of the second odds corresponding to the second probability for the unresolvable class, the logarithm of the second odds being constrained to a range of values ​​by a bounding function, scaled by a scaling factor, and weighted by a centroid of a distribution associated with the unresolvable class. The logit values ​​may include (iv) weighted values ​​generated based on hyperparameter tuning of the machine learning model, or (v) learning values ​​adjusted during training of the machine learning model. The method may also include the training subsystem deploying the trained machine learning model using the enriched logit values.

[0010] In various embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.

[0011] In various embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0012] The techniques described above and below can be implemented in several ways and in several contexts. Some example implementations and contexts are provided, as described in more detail below and with reference to the following drawings. However, the following implementations and contexts are only a few of many. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a simplified block diagram of a distributed environment incorporating an illustrative embodiment; [Figure 2] FIG. 1 is a simplified block diagram of a computing system implementing a Masterbot, according to one embodiment. [Figure 3] FIG. 1 is a simplified block diagram of a computing system implementing a skillbot, according to one embodiment. [Figure 4] FIG. 1 is a simplified block diagram of a chatbot training and deployment system according to various embodiments. [Figure 5] FIG. 1 illustrates a schematic diagram of an exemplary neural network, according to some embodiments. [Figure 6] 1 shows a flowchart illustrating an exemplary process for determining a statistic representing an enriched logit value for predicting whether an utterance corresponds to an unresolvable class, according to some embodiments. [Figure 7] 1 shows a flowchart illustrating an exemplary process for modifying a logit function to determine an enhanced logit value within a specified range for predicting whether an utterance corresponds to an unresolvable class, according to some embodiments. [Figure 8] 1 shows a flowchart illustrating an exemplary process for adding a scaling factor to a logit function to obtain an enhanced logit value for predicting whether an utterance corresponds to an unresolvable class, according to some embodiments. [Figure 9]FIG. 1 shows a flowchart illustrating an example process for using hyperparameter tuning to determine an enhanced logit value for predicting whether an utterance corresponds to an unresolvable class, according to some embodiments. [Figure 10] 1 shows a flowchart illustrating an exemplary process for using learned values ​​as enriched logit values ​​to predict whether an utterance corresponds to an unresolvable class, according to some embodiments. [Figure 11] 1 is a flowchart illustrating a process for training a machine learning model that achieves enhanced logit values ​​for unresolvable classes, according to some embodiments. [Figure 12] 10 is a flowchart illustrating a process for using enhanced logit values ​​to classify an utterance as an unresolvable class, according to some embodiments. [Figure 13] 1 is a simplified diagram of a distributed system for implementing various embodiments. [Figure 14] FIG. 1 is a simplified block diagram of one or more components of a system environment in which services provided by one or more components of an embodiment system may be offered as cloud services, according to various embodiments. [Figure 15] FIG. 1 illustrates an exemplary computer system that may be used to implement various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0014] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments. It will be apparent, however, that various embodiments may be practiced without these specific details. The figures and description are not intended to be limiting. The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.

[0015] A. Overview 1. Intent A digital assistant is an artificial intelligence-driven interface that helps users accomplish various tasks in natural language conversation. For each digital assistant, customers can assemble one or more skills. Skills (also referred to herein as chatbots, bots, or skillbots) are individual bots that focus on specific types of tasks, such as tracking inventory, submitting time cards, and creating expense reports. When an end user engages with a digital assistant, the digital assistant evaluates the end-user input and routes the conversation to and from the appropriate chatbot. Digital assistants can be made available to end users through various channels, such as Facebook Messenger, SKYPE MOBILE Messenger, or short message service (SMS). Channels carry chats back and forth from the end user to the digital assistant and its various chatbots over various messaging platforms. Channels may also support user-agent escalation, event-triggered conversations, and testing.

[0016] Intents enable a chatbot to understand what a user wants it to do. Intents consist of a sequence of typical user requests and statements, also referred to as utterances (e.g., get account balance, make a purchase, etc.). As used herein, an utterance or message may refer to a set of words (e.g., one or more sentences) exchanged during a conversation with a chatbot. An utterance may be text converted from voice input entered by a user through a user interface. An intent may be created by providing a name that denotes some user action (e.g., ordering a pizza) and compiling a set of real-life user statements or utterances commonly associated with triggering that action. Because the chatbot's cognition is derived from these intents, each intent is created from a dataset that is robust (from one to a few dozen utterances) and may vary to enable the chatbot to interpret ambiguous user input. A rich set of utterances enables the chatbot to understand what the user wants when it receives messages that mean the same thing but are expressed differently, such as "Ignore this order" or "Cancel the delivery!" Collectively, the intents and their associated utterances constitute a training corpus for chat. By training a model with a corpus, customers essentially train the model as a reference for resolving end-user input into a single intent. It can be transformed into a tool that allows customers to improve chat cognitive acuity through cycles of intent testing and intent training.

[0017] However, building a chatbot that can determine an end user's intent based on user utterances is a challenging task, in part due to the subtleties and ambiguities of natural language, as well as the dimensionality of the input space (e.g., possible user utterances) and the size of the output space (number of intents). Illustrative examples of this difficulty arise from characteristics of natural language, such as employing euphemisms, synonyms, or ungrammatical linguistic operations to express intents. For example, an utterance may express the intent to order a pizza without explicitly mentioning pizza, ordering, or delivery. For example, in the vernacular of a particular region, "pizza" is called "pie." These tendencies in natural language, such as imprecision or variability, create uncertainty and introduce reliability as a parameter for intent prediction, for example, through the inclusion of keywords, as opposed to explicit indication of intent. Therefore, chatbots may need to be trained, monitored, debugged, and retrained to improve their performance and the user experience with them. In conventional systems, training systems are provided to train and retrain machine learning models for digital assistants or chatbots in spoken language understanding (SLU) and natural language processing (NLP).

[0018] 2. Determine intent using machine learning models In one or more respects, a chatbot system can provide an utterance as input to a neural network model that uses a logistic regression function to map outputs to a probability distribution. For classification, for example, ordering a set of outputs in a probability distribution allows for prediction of the intent invoked in the utterance. Accurate predictions, in turn, allow the chatbot to accurately interact with the end user. In this sense, accuracy depends at least in part on mapping the output of the neural network classifier to a probability distribution.

[0019] To map the output of a neural network machine learning model to a probability distribution, a logit value is calculated based on the input. Logit values, also known as "logits," are the values ​​output by the logit function in the network layer of a machine learning model. Logit values ​​can represent the odds that an utterance corresponds to a particular class. The logit function is the logarithm of the odds for a particular class (e.g., the order_pizza intent class, the unresolved class). The output of the machine learning model is converted into a corresponding logit value that fits within a probability distribution. The logit value ranges between (-∞, +∞). The logit value may then be provided as input to an activation function (e.g., a softmax function) to generate a predicted likelihood that an input (e.g., an utterance) corresponds to a particular class among a set of classes. Here, the predicted likelihood lies within the probability distribution of the set of classes. The predicted likelihood can range between [0, 1]. For example, A numerical output is generated by processing an input utterance (e.g., "I want to grab a pie") through one or more hidden layers of a polynomial machine learning model. The output may be processed by a logit function for a particular class to generate a logit value of 9.4. An activation function may then be applied to the logit value to determine a probability value ranging between 0 and 1 (e.g., 0.974), which indicates that the input utterance corresponds to the order_pizza class. Utterances that do not invoke any intent that the classifier is also trained to identify will correspond to out-of-scope or out-of-domain utterances. For example, the utterance "how is the weather today?" indicates that the utterance is not intended to be specific to a food This may be considered out of scope for a classifier trained to predict whether to specify an order for an item.

[0020] Classification accuracy can be further improved by weighting one or more parameters of the logit function. For example, each intent class may be associated with a logit function that may be weighted by the centroid of the intent class. As used herein, the term "centroid" refers to the location of the center of gravity of a cluster used to classify an utterance, where the cluster identifies data corresponding to a particular end-user intent class. In some examples, the centroid is determined using data from a corresponding dataset of the utterance (e.g., a training dataset). Weighting the logit function by the centroid of the distribution allows the logit function to more accurately predict the classification of a given utterance, especially when the utterance is within the domain or scope (e.g., the utterance the system is trained to recognize).

[0021] 3. Determine the intent of out-of-scope utterances Neural networks suffer from the problem of overconfidence. Overconfidence can occur because confidence scores generated by a trained neural network (e.g., a trained NLP algorithm) for a class can be uncorrelated from the actual confidence scores. This problem typically occurs in deeper neural networks (i.e., neural network models with a larger number of layers). Deep neural network models are generally more accurate in their output predictions than shallower neural network models, but can generate erroneous high-confidence classification predictions if the actual inputs are not well represented by the training data used to train the neural network model. Thus, while the use of deep neural networks is desirable due to their high accuracy, the overconfidence problem associated with deep neural networks must be addressed to avoid performance issues for the neural network.

[0022] Conventional techniques involving the use of centroid-weighted logit functions do not effectively address the overconfidence problem. In general, for out-of-domain or out-of-scope utterances, the use of centroid-weighted logit functions can be said to treat the classifier as determining whether the utterance fits into a cluster of unresolved utterances. However, because out-of-domain or out-of-scope utterances are more likely to be dispersed rather than clustered due to their different natural language definitions and syntax, the clusters of unresolved utterances may be inaccurate. In other words, using a centroid-weighted logit function for out-of-scope utterances would mean applying centroids based on highly dispersed clusters. Applying such centroids introduces inaccuracies into the classification of utterances, for example, by underrepresenting the probability that an intent is outside the chatbot's domain.

[0023] 4. Enhanced Logit for Classifying Utterances as Having Unresolved Intents To overcome the above drawbacks, the present technology includes a system and method for accurately predicting that an out-of-scope utterance corresponds to an unresolvable class using enhanced logits in a machine learning model. This can increase the accuracy of the utterance classifier, for example, providing an improved utterance classifier that can more accurately identify out-of-domain or out-of-scope utterances that cannot be reliably classified by the classifier. This increased ability to identify out-of-domain / out-of-scope utterances can reduce the frequency with which the classifier attempts to classify utterances that it is not trained to classify, reduce the number of unpredictable or incorrect classification results that may result from attempting to classify utterances that are outside the classifier's capabilities, and thus improve overall utterance classification. The present technology can include a machine learning model (e.g., a neural network) trained to predict whether an utterance or message represents a resolvable class (e.g., the type of task the skillbot is configured to perform, the intent associated with the skillbot) or an unresolvable class. The machine learning model can predict whether an utterance corresponds to a class of resolvable intents. The system may be configured to apply a logit function to generate a first logit value that predicts whether an utterance corresponds to a class of unresolvable intents, and use an enhanced logit value that predicts whether an utterance corresponds to a class of unresolvable intents. In some examples, the enhanced logit value replaces the logit function or is determined independently from the logit function.

[0024] In some examples, the enhanced logit value for the unresolvable class includes one of: (i) a statistical value determined based on a set of logit values ​​generated from a training dataset; (ii) a bounded value selected from a range of values ​​defined by the logarithm of first odds corresponding to the probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function and weighted by the centroid of a distribution associated with the unresolvable class; (iii) a weighted value generated by the logarithm of second odds corresponding to the probability for the unresolvable class, the logarithm of the second odds being scaled by a scaling factor, bounded by the bounding function, and weighted by the centroid of the distribution associated with the unresolvable class; (iv) a hyperparameter optimized value generated based on hyperparameter tuning of a machine learning model; or (v) a learned value dynamically adjusted during training of the machine learning model.

[0025] B. Bots and Analytics Systems A bot (also referred to as a skill, chatbot, chatterbot, or talkbot) is a computer program that can conduct a conversation with an end user. Bots are generally capable of responding to natural language messages (e.g., questions or comments) through messaging applications using natural language messages. Businesses may use one or more bot systems to communicate with end users through messaging applications. The messaging application, sometimes referred to as a channel, may be the end user's preferred messaging application that the end user already has installed and is familiar with. Thus, end users do not need to download and install a new application to chat with a bot system. Messaging applications may be, for example, over-the-top (OTT) messaging channels (e.g., Facebook Messenger, Facebook WhatsApp, WeChat, Line, Kik, Telegram, Talk, Skype, Slack, or SMS), virtual private assistants (e.g., Amazon Dot, Echo, or Show, Google® Home, Apple HomePod, etc.), or messaging services. ), mobile and web app extensions that extend native or hybrid / responsive mobile or web applications with chat functionality, or voice-based input (e.g., Siri, Cortana, Google Voice, or other voice input for interaction) This may include a device or app with an interface for using the device.

[0026] In some examples, a bot system may be associated with a Uniform Resource Identifier (URI). The URI may identify the bot system using a string of characters. The URI may be used as a webhook for one or more messaging application systems. The URI may include, for example, a Uniform Resource Locator (URL) or a Uniform Resource Name (URN). The bot system may be designed to receive a message (e.g., a HyperText Transfer Protocol (HTTP) post call message) from the messaging application system. The HTTP post call message may be directed to the URI from the messaging application system. In some embodiments, the message may differ from the HTTP post call message. For example, the bot system may receive a message via Short Message Service (SMS). While the discussion herein may refer to a communication received by the bot system as a message, it should be understood that the message may be an HTTP post call message, an SMS message, or any other type of communication between two systems.

[0027] End users can interact with bot systems through conversational interactions (sometimes called conversational user interfaces (UIs)), much like interactions between people. In some cases, interactions begin with the end user saying "Hello" to the bot. , the bot may respond with "Hi" and ask the end user how it can assist them. In some cases, the interaction may also be a transactional interaction with a banking bot, e.g., transferring money from one account to another; an informational interaction with an HR bot, e.g., checking a vacation balance; or an interaction with a retail bot, e.g., discussing returning a purchased item or seeking technical support.

[0028] In some embodiments, the bot system can intelligently handle end-user interactions without interaction with an administrator or developer of the bot system. For example, an end-user may send one or more messages to the bot system to achieve a desired goal. The messages may include content such as text, emojis, audio, images, video, or other methods of conveying a message. In some embodiments, the bot system converts the content into a standardized format (e.g., a representational state transfer (REST) ​​call to an enterprise service with appropriate parameters) and provides a natural language response. The bot system may generate an answer. The bot system may also prompt the end user for additional input parameters or request other additional information. In some embodiments, the bot system may also initiate communication with the end user rather than passively responding to the end user utterance. Various techniques are described herein for identifying explicit invocations of the bot system and determining input for the invoked bot system. In one embodiment, explicit invocation analysis is performed by the master bot based on detecting a call name in the utterance. In response to detecting the call name, the utterance may be refined for input to a skill bot associated with the call name.

[0029] A conversation with a bot can follow a specific conversational flow that includes multiple states. The flow can define what happens next based on input. In some embodiments, a bot system can be implemented using a state machine that includes user-defined states (e.g., end user intents) and actions to be taken in and from states. A conversation can take different paths based on end user input, which can affect the decisions the bot makes about the flow. For example, at each state, based on the end user input or utterances, the bot can determine the end user's intent and decide the appropriate action to take next. Here, and in the context of utterances, the term "intent" refers to the intent of the user who gave the utterance. For example, a user intends to engage a bot in a conversation to order a pizza, and the user's intent may be expressed by the utterance "order a pizza." A user's intent can be directed to a specific task the user wants the chatbot to perform on their behalf. Thus, an utterance can be expressed as a question, command, request, etc. that reflects the user's intent. An intent may include a goal the end user wants to achieve.

[0030] In the context of chat configuration, the term "intent" is used herein to refer to configuration information for mapping user utterances to specific tasks / actions or categories of tasks / actions that a chatbot can perform. To distinguish between utterance intents (i.e., user intents) and chatbot intents, the latter may be referred to herein as "bot intents." A bot intent may include a set of one or more utterances associated with that intent. For example, an intent for ordering pizza may have various permutations of utterances expressing a desire to place a pizza order. These associated utterances may be used to train the chatbot's intent classifier, which can then determine whether an input utterance from a user corresponds to a specific task / action or a category of tasks / actions that a chatbot can perform. This enables the bot to determine whether a bot matches the ordering intent. A bot intent can be associated with one or more dialog flows for initiating a conversation with a user in a certain state. For example, the first message for a pizza ordering intent can be the question, "What kind of pizza would you like?" In addition to the associated utterance, a bot intent can also include named entities related to the intent. For example, a pizza ordering intent can include variables or parameters used to perform the task of ordering a pizza, such as topping 1, topping 2, pizza type, pizza size, pizza quantity, etc. The values ​​of the entities are typically obtained through a conversation with the user.

[0031] In one example, the utterance is analyzed to determine whether the utterance includes a call name for the skill bot. If a call name is not found, the utterance is considered to be an implicit call, and the process proceeds to an intent classifier, such as a trained model. If a call name is determined to be present, the utterance is considered to be an explicit call, and the process proceeds to determine which portions of the utterance are associated with the call name. When a trained model is invoked, the entire received utterance is provided as input to the intent classifier.

[0032] The intent classifier receiving the utterance may be the intent classifier of the masterbot (e.g., intent classifier 242 of FIG. 2). The intent classifier may be a machine learning-based or rule-based classifier trained to determine whether the intent of the utterance matches a system intent (e.g., exit, help) or a specific skillbot. As described herein, the intent analysis performed by the masterbot may be limited to matching against a specific skillbot without determining which intent within the skillbot is the best match for the utterance. Thus, the intent classifier receiving the utterance may identify a specific skillbot to be invoked. Alternatively, if the utterance expresses a specific system intent (e.g., the utterance includes the word “exit” or “help”), the intent classifier receiving the utterance may identify that specific system intent to trigger a conversation between the masterbot and the user based on the dialog flow configured for that specific system intent.

[0033] If an invocation name is present, one or more explicit invocation rules are applied to determine which portions of the utterance are associated with the invocation name. This determination may be based on an analysis of the sentence structure of the utterance using POS tags, dependency information, and / or other extracted information received with the utterance. For example, a portion associated with the invocation name may be a noun phrase containing the invocation name or an object with a preposition corresponding to the invocation name. Any portion associated with the invocation name, as determined based on the processing, is removed. Other portions of the utterance that are not required to convey the meaning of the utterance (e.g., prepositional words) may also be removed. The removal of certain portions of the utterance generates input for the skillbot associated with the invocation name. If any portion of the received utterance remains after the removal, the remaining portion forms a new utterance for input to the skillbot, for example, as a text string. Alternatively, if the received utterance is completely removed, the input may be an empty string. The skillbot associated with the invocation name is then invoked and provided with the generated input.

[0034] Upon receiving the generated input, the invoked skill bot processes the input by, for example, performing intent analysis using a trained skill bot intent classifier to identify a bot intent that matches the user intent expressed in the input. As a result of identifying a matching bot intent, the skill bot may perform a specific action according to the dialog flow associated with the matching bot intent. or may initiate a conversation with the user. For example, if the input is an empty string, the conversation may start in a default state defined for the dialog flow, such as a welcome message. Alternatively, if the input is not an empty string, the conversation may start in some intermediate state, such as because the input contains a value for an entity or some other information that the skill bot received as part of the input and therefore no longer needs to prompt the user. As another example, the skill bot may determine that it cannot process the input (e.g., because the confidence scores of all bot intents configured for the skill bot are below a certain threshold). In this situation, the skill bot may pass the input back to the master bot for processing (e.g., intent analysis using the master bot's intent classifier), or the skill bot may prompt the user for clarification.

[0035] 1. Overall Environment FIG. 1 is a simplified block diagram of an environment 100 incorporating a chatbot system according to one embodiment. The environment 100 includes a Digital Assistant Builder Platform (DABP) 102, which enables users of the DABP 102 to create and deploy digital assistant or chatbot systems. The DABP 102 can be used to create one or more digital assistant (or DA) or chatbot systems. For example, as shown in FIG. 1, a user 104 representing a particular business can use the DABP 102 to create and deploy a digital assistant 106 for users of the particular business. For example, a bank can use the DABP 102 to create one or more digital assistants for use by the bank's customers. Multiple businesses can use the same DABP 102 platform to create digital assistants. As another example, a restaurant (e.g., a pizza shop) owner can use the DABP 102 to create and deploy a digital assistant that enables customers of the restaurant to order food (e.g., order pizza).

[0036] For purposes of this disclosure, a "digital assistant" is an entity that helps a user of the digital assistant accomplish various tasks through natural language conversation. A digital assistant may be implemented solely using software (e.g., a digital assistant is a digital entity implemented using programs, codes, or instructions executable by one or more processors), using hardware, or using a combination of hardware and software. A digital assistant may be embodied or implemented in various physical systems or devices, such as a computer, a mobile phone, a watch, an appliance, a vehicle, etc. A digital assistant may also be referred to as a chatbot system. Thus, for purposes of this disclosure, the terms digital assistant and chatbot system are interchangeable.

[0037] A digital assistant, such as digital assistant 106 built using DABP 102, can be used to perform various tasks through natural language-based conversations between the digital assistant and its user 108. As part of the conversation, the user may provide one or more user inputs 110 to the digital assistant 106 and obtain responses 112 from the digital assistant 106. A conversation can include one or more of the inputs 110 and the responses 112. Through these conversations, the user can request one or more tasks to be performed by the digital assistant 106, and in response, the digital assistant 106 is configured to perform the user-requested task and respond to the user with an appropriate response.

[0038] User input 110 is generally in the form of natural language and is called an utterance. User utterance 110 can be in the form of text, such as when a user types a sentence, a question, a piece of text, or even a single word and provides it as input to the digital assistant 106. In some embodiments, user utterance 110 can be in the form of voice input or speech, such as when a user says or speaks something that is provided as input to digital assistant 106. The speech is typically in the language spoken by user 108. For example, the speech may be in English or some other language. If the speech is in voice form, the voice input is converted into textual speech in that particular language, and the textual speech is then processed by digital assistant 106. Various speech-to-text processing techniques may be used to convert the voice or auditory input into textual speech, which is then processed by digital assistant 106. In some embodiments, the speech-to-text conversion may be performed by digital assistant 106 itself.

[0039] The utterance, which may be a text utterance or a speech utterance, may be a fragment, a sentence, multiple sentences, one or more words, one or more questions, a combination of the aforementioned types, etc. The digital assistant 106 is configured to apply natural language understanding (NLU) techniques to the utterance to understand the meaning of the user input. As part of the NLU processing on the utterance, the digital assistant 106 is configured to perform processing to understand the meaning of the utterance, which involves identifying one or more intents and one or more entities corresponding to the utterance. Upon understanding the meaning of the utterance, the digital assistant 106 can perform one or more actions or operations in response to the understood meaning or intent. For purposes of this disclosure, it is assumed that the utterance is a text utterance provided directly by a user 108 of the digital assistant 106 or is the result of converting an input utterance into text form. However, this is not intended to be limiting or restrictive in any way.

[0040] For example, user 108's input may request that a pizza be ordered by providing an utterance such as, "I want to order a pizza." Upon receiving such an utterance, digital assistant 106 is configured to understand the meaning of the utterance and take appropriate action. The appropriate action may include responding to the user with a question requesting user input regarding, for example, the type of pizza the user wants to order, the size of the pizza, any toppings on the pizza, etc. The responses provided by digital assistant 106 may also be in natural language format, typically in the same language as the input utterance. As part of generating these responses, digital assistant 106 may perform natural language generation (NLG). Through a conversation between the user and digital assistant 106, for the user to order a pizza, the digital assistant may guide the user to provide all necessary information to order the pizza and then, at the end of the conversation, have the user order the pizza. Digital assistant 106 may end the conversation by outputting information to the user indicating that the pizza has been ordered.

[0041] At a conceptual level, digital assistant 106 performs various processes in response to utterances received from a user. In some embodiments, this processing involves a series or pipeline of processing steps, including, for example, understanding the meaning of the input utterance (sometimes referred to as natural language understanding (NLU)), determining an action to be performed in response to the utterance, causing the action to be performed if appropriate, generating a response to be output to the user in response to the user utterance, outputting the response to the user, etc. NLU processing can include parsing the received input utterance to understand the structure and meaning of the utterance, and refining and restructuring the utterance to develop a more understandable form (e.g., logical form) or structure for the utterance. Generating a response may include using NLG technology.

[0042] The NLU processing performed by a digital assistant, such as digital assistant 106, includes sentence analysis (e.g., tokenization, reordering, identifying part-of-speech tags for a sentence, and The NLU processing may include various NLP-related processes such as identifying named entities, generating a dependency tree to represent the sentence structure, splitting the sentence into clauses, analyzing the individual clauses, resolving anaphora, performing chunking, etc. In one embodiment, the NLU processing, or portions thereof, is performed by the digital assistant 106 itself. In some other embodiments, the digital assistant 106 can use other resources to perform portions of the NLU processing. For example, the syntax and structure of an input spoken sentence may be identified by processing the sentence using syntactic parsing, part-of-speech tagging, and / or named entity recognition. In one implementation, for English, syntactic parsing, part-of-speech tagging, and named entity recognition, such as those provided by the Stanford Natural Language Processing (NLP) Group, are used to analyze sentence structure and syntax. These are provided as part of the Stanford CoreNLP toolkit.

[0043] Although the various examples provided in this disclosure show English utterances, this is meant as an example only. In some embodiments, the digital assistant 106 can also process utterances in languages ​​other than English. The digital assistant 106 may provide subsystems (e.g., components that implement NLU functionality) configured to perform processing for different languages. These subsystems may be implemented as pluggable units that can be invoked using service calls from the NLU core server. This makes NLU processing flexible and extensible for each language, including allowing for different orders of processing. Language packs may be provided for individual languages, and the language packs can register a list of subsystems that can be served from the NLU core server.

[0044] 1 can be made available or accessible to its user 108 through a variety of different channels, such as, but not limited to, through an application, through a social media platform, through various messaging services and applications, and other applications or channels. A single digital assistant can have several channels configured for it, so that it can run on and be accessed by different services simultaneously.

[0045] A digital assistant or chatbot system typically includes or is associated with one or more skills. In one embodiment, these skills are individual chatbots (referred to as skillbots) that interact with a user and are configured to fulfill specific types of tasks, such as tracking inventory, submitting a timecard, creating an expense report, ordering food, checking a bank account, making a reservation, or purchasing a widget. For example, in the embodiment shown in FIG. 1 , the digital assistant or chatbot system 106 includes skills 116-1, 116-2, etc. For purposes of this disclosure, the term "skill" is used synonymously with the term "skillbot."

[0046] Each skill associated with a digital assistant helps a user of the digital assistant complete a task through a conversation with the user, where the conversation can include a combination of text or auditory input provided by the user and responses provided by the skill bot. These responses can be in the form of text or auditory messages to the user and / or with simple user interface elements (e.g., selection lists) presented to the user for the user to make a selection.

[0047] There are various ways in which a skill or skillbot can be associated with or added to a digital assistant. In one example, a skillbot can be developed by a company and then added to a digital assistant using DABP 102. In another example, a skillbot can be developed and created using DABP 102 and then added to a digital assistant created using DABP 102. In yet another example, DABP 102 DABP 102 provides an online digital store (referred to as a "skill store") that offers multiple skills aimed at a wide range of tasks. Skills offered through the skill store may also be exposed to various cloud services. To add skills to a digital assistant created using DABP 102, a user of DABP 102 can access the skill store via DABP 102, select a desired skill, and indicate that the selected skill be added to the digital assistant created using DABP 102. Skills from the skill store can be added to the digital assistant as is or in modified form (e.g., a user of DABP 102 may select and clone a particular skill bot offered by the skill store, customize or modify the selected skill bot, and then add the modified skill bot to a digital assistant created using DABP 102).

[0048] A variety of different architectures may be used to implement a digital assistant or chatbot system. For example, in one embodiment, a digital assistant created and deployed using DABP 102 may be implemented using a masterbot / child (or sub)bot paradigm or architecture. According to this paradigm, the digital assistant is implemented as a masterbot that interacts with one or more child bots, which are skillbots. For example, in the embodiment shown in FIG. 1, the digital assistant 106 includes a masterbot 114 and skillbots 116-1, 116-2, etc., that are child bots of the masterbot 114. In one embodiment, the digital assistant 106 itself may act as the masterbot.

[0049] A digital assistant implemented according to the master-child bot architecture allows users of the digital assistant to interact with multiple skills through a unified user interface, i.e., through a master bot. When a user engages with the digital assistant, user input is received by the master bot. The master bot then performs processing to determine the meaning of the user input utterance. The master bot then determines whether the task requested by the user in the utterance can be handled by the master bot itself. If not, the master bot selects an appropriate skill bot to handle the user request and routes the conversation to the selected skill bot. This allows users to converse with the digital assistant through a common, single interface while still providing the ability to use several skill bots configured to perform specific tasks. For example, in a digital assistant developed for an enterprise, the digital assistant's master bot can interface with skill bots with specific functions, such as a CRM bot for performing functions related to customer relationship management (CRM), an ERP bot for performing functions related to enterprise resource planning (ERP), and an HCM bot for performing functions related to human capital management (HCM). In this way, the end user or consumer of the digital assistant only needs to know how to access the digital assistant through a common master bot interface, and behind the scenes, multiple skill bots are provided to handle user requests.

[0050] In one embodiment, in a masterbot / childbot infrastructure, the masterbot is configured to be aware of an available list of skillbots. The masterbot may have access to various available skillbots and, for each skillbot, metadata identifying each skillbot's capabilities, including the tasks that can be performed by each skillbot. Upon receiving a user request in the form of an utterance, the masterbot is configured to identify or predict a particular skillbot from multiple available skillbots that can best accommodate or process the user request. The masterbot then routes the utterance (or a portion of the utterance) to that particular skillbot for further processing. Thus, control is transferred from the masterbot to the skillbot. A master bot can support multiple input and output channels.

[0051] 1 illustrates a digital assistant 106 with a masterbot 114 and skillbots 116-1, 116-2, and 116-3, but this is not intended to be limiting. A digital assistant can include various other components (e.g., other systems and subsystems) that provide the functionality of the digital assistant. These systems and subsystems may be realized solely in software (e.g., code, instructions stored on a computer-readable medium and executable by one or more processors), solely in hardware, or in an implementation using a combination of software and hardware.

[0052] DABP 102 provides infrastructure and various services and features that enable users of DABP 102 to create digital assistants, including one or more skillbots associated with the digital assistant. In some cases, a skillbot can be created by cloning an existing skillbot, for example, by cloning a skillbot provided by a skill store. As described above, DABP 102 provides a skill store or skill catalog that offers multiple skillbots for performing various tasks. Users of DABP 102 can clone a skillbot from the skill store. They may modify or customize the cloned skillbot as needed. In some other cases, users of DABP 102 created a skillbot from scratch using tools and services provided by DABP 102. As described above, the skill store or skill catalog provided by DABP 102 may offer multiple skillbots for performing various tasks.

[0053] In one embodiment, at one high level, creating or customizing a skillbot includes the following steps: (1) Set up the settings for the new skill bot (2) Configure one or more intents for the skill bot (3) Set one or more entities for one or more intents. (4) Training the SkillBot (5) Create a dialog flow for your skill bot (6) Add custom components to your skill bot as needed (7) Test and deploy the skill bot. Each step will be briefly described below.

[0054] (1) Set Settings for a New Skillbot—Various settings may be set for a skillbot. For example, a skillbot designer can specify one or more call names for the skillbot being created. These call names can then be used by a user of the digital assistant to explicitly call the skillbot. For example, a user can enter a call name in the user's utterance to explicitly call the corresponding skillbot.

[0055] (2) Configure one or more intents and associated example utterances for the skillbot—The skillbot designer specifies one or more intents (also called bot intents) for the skillbot being created. The skillbot is then trained based on these specified intents. These intents represent categories or classes that the skillbot is trained to reason about input utterances. Upon receiving an utterance, the trained skillbot infers the intent of the utterance. The inferred intent is selected from a predefined set of intents used to train the skill bot. The skill bot then takes an appropriate action in response to the utterance based on the intent inferred for the utterance. In some cases, the intents for a skill bot represent tasks that the skill bot can perform for a user of the digital assistant. Each intent is given an intent identifier or intent name. For example, for a skill bot trained for banking, the intents specified for the skill bot might be "CheckBalance," "TransferMoney," "DepositCheck," etc. It may include.

[0056] For each intent defined for a skillbot, the skillbot designer may also provide one or more example utterances that are representative of and demonstrate that intent. These example utterances are meant to represent utterances a user may input to the skillbot for that intent. For example, for a balance inquiry intent, example utterances might include, "What's my savings account balance?", "How much is in my checking account?", "How much money do I have in my account?", etc. Thus, various permutations of typical user utterances may be specified as example utterances for an intent.

[0057] The intents and their associated example utterances are used as training data to train the skill bot. A variety of different training techniques may be used. As a result of this training, a machine learning model is generated that is configured to take an utterance as input and output an intent inferred for the utterance by the machine learning model. In some cases, the input utterance is provided to an intent analysis engine that is configured to predict or infer an intent for the input utterance using the trained model. The skill bot may then take one or more actions based on the inferred intent.

[0058] (3) Configuring One or More Entities for One or More Intents - In some instances, additional context may be required to enable the skill bot to respond appropriately to a user utterance. For example, there may be situations where user input utterances resolve to the same intent in the skill bot. For example, in the example above, the utterances "What's my savings account balance?" and "How much is in my checking account?" both resolve to the same balance check. Although each utterance resolves to the same intent, these utterances are different requests that want different things. To clarify such requests, one or more entities are added to the intent. Using the banking skill bot example, an entity called AccountType that defines values ​​called "checking" and "saving" could be used to resolve the skill. This may enable the robot to parse the user request and respond appropriately. In the example above, the utterance resolves to the same intent, but the value associated with the AccountType entity is different. , are different for the two utterances. This allows the skill bot to potentially perform different actions for the two utterances even though they resolve to the same intent. One or more entities may be specified for a particular intent configured for the skill bot. Thus, entities are used to add context to the intent itself. Entities help to more fully describe the intent and enable the skill bot to complete the user request.

[0059] In one embodiment, there are two types of entities: (a) built-in entities provided by DABP 102, and (2) entities specified by the SkillBot designer. There are custom entities that can be used. Built-in entities are general-purpose entities that can be used with a wide variety of bots. Examples of built-in entities include, but are not limited to, entities related to time, date, address, number, email address, duration, cycle period, currency, phone number, URL, etc. Custom entities are used for more customized uses. For example, for a banking skill, the AccountType entity might be Entities may be defined by the skillbot designer to enable various banking transactions by checking user input for keywords such as checking, savings, and credit card.

[0060] (4) Train the Skillbot—The skillbot is configured to receive user input in the form of utterances, parse or otherwise process the received input, and identify or select an intent associated with the received user input. As described above, the skillbot must be trained for this. In one embodiment, the skillbot is trained based on the intents configured for the skillbot and example utterances associated with those intents (collectively, training data), thereby enabling the skillbot to resolve user input utterances to one of the skillbot's configured intents. In one embodiment, the skillbot is trained using the training data and uses a machine learning model that enables the skillbot to identify what the user is saying (or, in some cases, what they are trying to say). DABP 102 provides a variety of different training techniques that can be used by the skillbot designer to train the skillbot, including various machine learning-based training techniques, rule-based training techniques, and / or combinations thereof. In one embodiment, a portion of the training data (e.g., 80%) is used to train the skillbot model, and another portion (e.g., the remaining 20%) is used to test or validate the model. Once trained, the trained model (sometimes called a trained skill bot) can then be used to process and respond to user utterances. In some cases, a user utterance may be a question that requires only a single answer and no further conversation. To address such situations, a Q&A (Question and Answer) intent may be defined for the skill bot. This allows the skill bot to output a response to a user request without having to update the dialog definition. A Q&A intent is created similarly to a regular intent. The dialog flow for a Q&A intent may differ from the dialog flow for a regular intent.

[0061] (5) Create a Dialog Flow for the Skill Bot—The dialog flow specified for a skill bot describes how the skill bot reacts as different intents for the skill bot are resolved in response to received user input. The dialog flow defines the behavior or actions the skill bot takes, such as how the skill bot responds to user utterances, how the skill bot prompts the user for input, and how the skill bot returns data. The dialog flow is like a flowchart that the skill bot follows. Skill bot designers specify the dialog flow using a language such as Markdown. In one embodiment, a version of YAML called OBotML can be used to specify the dialog flow for a skill bot. The dialog flow definition for a skill bot serves as a model of the conversation itself, allowing skill bot designers to choreograph the interactions between the skill bot and the users it serves.

[0062] In one embodiment, a skill bot's dialog flow definition includes three sections: (a) Context Section (b) Default transition section (c) State section.

[0063] Context Section - In the context section, the skill bot designer can define variables used in the conversation flow. Other variables that can be named in the context section include, but are not limited to, variables for error handling, variables for built-in or custom entities, user variables that allow the skill bot to recognize and persist user preferences, etc.

[0064] Default Transition Section - Transitions for a skill bot can be defined in the dialog flow state section or the default transition section. Transitions defined in the default transition section act as fallbacks and are triggered when there is no applicable transition defined in a state or when the conditions required to trigger a state transition cannot be met. The default transition section can be used to define routing that allows the skill bot to gracefully handle unexpected user actions.

[0065] State Section - A dialog flow and its associated behavior are defined as a series of temporary states that govern the logic within the dialog flow. Each state node in a dialog flow definition names a component that provides the functionality needed at that point in the dialog. In this way, you build states around components. States contain component-specific characteristics and define transitions to other states that are triggered after the component executes.

[0066] Special case scenarios can be handled using the state section. For example, you might want to give a user the option to temporarily exit a first skill they're working on and do something in a second skill within the digital assistant. For example, if a user is engaged in a conversation with a shopping skill (e.g., the user has made some selections for a purchase), the user might want to jump to a banking skill (e.g., the user might want to verify that they have enough money for the purchase) and then return to the shopping skill to complete the user's order. To address this, an action in a first skill can be configured to initiate an interaction with a second, different skill in the same digital assistant and then return to the original flow.

[0067] (6) Adding Custom Components to a Skillbot—As described above, a state specified in a dialog flow for a skillbot names a component that provides the necessary functionality corresponding to that state. The component enables the skillbot to perform the function. In one embodiment, DABP 102 provides a set of pre-configured components to perform a wide range of functions. A skillbot designer can select one or more of these pre-configured components and associate them with states in the dialog flow for the skillbot. A skillbot designer can also create custom or new components using tools provided by DABP 102 and associate the custom components with one or more states in the dialog flow for the skillbot.

[0068] (7) Testing and Deploying Skillbots - DABP 102 provides several features that allow skillbot designers to test the skillbots they are developing. The skillbots can then be deployed and included in a digital assistant.

[0069] While the above description describes how to create a skillbot, similar techniques can also be used to create a digital assistant (or masterbot). At the masterbot or digital assistant level, built-in system intents can be configured for the digital assistant. These built-in system intents are used to identify common tasks that the digital assistant itself (i.e., the masterbot) can handle without invoking a skillbot associated with the digital assistant. Examples of system intents defined for a masterbot include: (1) Exit, which applies when a user signals to the digital assistant that they want to end the current conversation or context; (2) Help, which applies when a user asks for help or direction; and (3) UnresolvedIntent, which applies to user input that does not match well with the Exit and Help intents. The digital assistant also stores information about one or more skillbots associated with the digital assistant. This information allows the masterbot to select a specific skillbot to process an utterance.

[0070] At the masterbot or digital assistant level, when a user inputs a phrase or utterance into the digital assistant, the digital assistant is configured to process the utterance and the associated conversation to determine how to route the utterance. The digital assistant makes this determination using a routing model, which can be rule-based, AI-based, or a combination thereof. The digital assistant uses the routing model to determine whether the conversation corresponding to the user input utterance should be routed to a specific skill for processing, handled by the digital assistant or masterbot itself according to built-in system intents, or handled as a different state in the current conversation flow.

[0071] In one embodiment, as part of this processing, the digital assistant determines whether the user input utterance explicitly identifies a skill bot using its invocation name. If an invocation name is present in the user input, it is treated as an explicit invocation of the skill bot corresponding to the invocation name. In such a scenario, the digital assistant can route the user input to the explicitly invoked skill bot for further processing. In the absence of a specific or explicit invocation, in one embodiment, the digital assistant evaluates the received user input utterance and calculates confidence scores for system intents and skill bots associated with the digital assistant. The calculated scores for the skill bots or system intents represent the likelihood that the user input represents a task that the skill bot is configured to perform or represents a system intent. System intents or skill bots whose associated calculated confidence scores exceed a threshold (e.g., a Confidence Threshold routing parameter) are selected as candidates for further evaluation. The digital assistant then selects a specific system intent or skill bot from the identified candidates for further processing of the user input utterance. In one embodiment, after one or more skill bots are identified as candidates, the intents associated with those candidate skills are evaluated (according to the intent model for each skill), and a confidence score is determined for each intent. Generally, intents with a confidence score above a threshold (e.g., 70%) are treated as candidate intents. If a particular skill bot is selected, the user utterance is routed to that skill bot for further processing. If a system intent is selected, one or more actions are performed by the master bot itself according to the selected system intent.

[0072] 2. Components of the Masterbot System FIG. 2 is a simplified block diagram of a Masterbot (MB) system 200, according to one embodiment. 2 is a block diagram. MB System 200 can be implemented in software only, hardware only, or a combination of hardware and software. MB System 200 includes a pre-processing subsystem 210, a multiple-intent subsystem (MIS) 220, an explicit invocation subsystem (EIS) 230, a skillbot invocation unit 240, and a data store 250. MB System 200 shown in FIG. 2 is merely an example of the configuration of components in a masterbot. Those skilled in the art will recognize many possible variations, alternatives, and modifications. For example, in some implementations, MB System 200 may have more or fewer systems or components than those shown in FIG. 2, may combine two or more subsystems, or may have subsystems in a different configuration or arrangement.

[0073] The pre-processing subsystem 210 receives an utterance "A" 202 from a user and processes the utterance through a language detector 212 and a language parser 214. As described above, the utterance may be provided in a variety of ways, including as audio or text. The utterance 202 may be a fragment, a complete sentence, multiple sentences, etc. The utterance 202 may include punctuation. For example, if the utterance 202 is provided as audio, the pre-processing subsystem 210 may convert the audio to text using a speech-to-text converter (not shown), which inserts punctuation, e.g., commas, semicolons, periods, etc., into the resulting text.

[0074] The language detection unit 212 detects the language of the utterance 202 based on the text of the utterance 202. Because each language has its own grammar and semantics, the way in which the utterance 202 is processed depends on the language. Language differences are taken into account when analyzing the syntax and structure of the utterance.

[0075] The language parser 214 parses the utterance 202 to extract part-of-speech (POS) tags for individual linguistic units (e.g., words) within the utterance 202. POS tags include, for example, nouns (NN), pronouns (PN), verbs (VB), etc. The language parser 214 may also tokenize the linguistic units of the utterance 202 (e.g., to convert each word into a separate token) and lemmatize the words. A lemma is the primary form of a set of words represented in a dictionary (e.g., "run" is the lemma for run, runs, ran, running, etc.). Other types of preprocessing that the language parser 214 can perform include chunking of complex expressions, for example, combining "credit" and "card" into a single expression "credit_card". The language parser 214 may also identify relationships between words in utterance 202. For example, in some embodiments, language parser 214 generates a dependency tree that indicates which parts of the utterance (e.g., particular nouns) are direct objects, which parts of the utterance are prepositions, etc. The results of the processing performed by language parser 214 form extracted information 205, which, along with utterance 202 itself, are provided as inputs to MIS 220.

[0076] As described above, utterance 202 may include multiple sentences. For purposes of multiple-intent and explicit invocation detection, utterance 202 may be treated as a single unit even if it includes multiple sentences. However, in some embodiments, preprocessing may be performed, for example, by preprocessing subsystem 210, to identify single sentences within multiple sentences for multiple-intent analysis and explicit invocation analysis. Generally, the results produced by MIS 220 and EIS 230 are substantially the same whether utterance 202 is processed at the individual sentence level or as a single unit containing multiple sentences.

[0077] The MIS 220 determines whether the utterance 202 expresses multiple intents. While the MIS 220 can detect the presence of multiple intents in the utterance 202, the processing performed by the MIS 220 does not involve determining whether the intent of the utterance 202 matches any intent configured for the bot. Instead, the MIS 220 determines whether the utterance 202 expresses multiple intents. The processing to determine whether the intent of the utterance 202 matches a bot intent may be performed by the intent classifier 242 of the MB system 200 (e.g., as shown in the embodiment of FIG. 3) or by an intent classifier of a skill bot. The processing performed by the MIS 220 assumes that a bot (e.g., a particular skill bot or the master bot itself) exists that can process the utterance 202. Thus, the processing performed by the MIS 220 does not require knowledge of what bots are in the chatbot system (e.g., the identities of the skill bots registered with the master bot) or what intents have been set for a particular bot.

[0078] To determine that utterance 202 contains multiple intents, MIS 220 applies one or more rules from a set of rules 252 in data store 250. The rules applied to utterance 202 depend on the language of utterance 202 and may include sentence patterns that indicate the presence of multiple intents. For example, a sentence pattern may include a conjunction connecting two parts of a sentence (e.g., coordinates), both of which correspond to separate intents. If utterance 202 matches a sentence pattern, it can be inferred that utterance 202 represents multiple intents. Note that an utterance with multiple intents does not necessarily have different intents (e.g., intents directed to different bots or different intents within the same bot). Instead, an utterance may have separate instances of the same intent, such as "order a pizza using payment account X, then order a pizza using payment account Y."

[0079] As part of determining that utterance 202 represents multiple intents, MIS 220 also determines what portions of utterance 202 are associated with each intent. For each intent expressed in the multiple-intent utterance, MIS 220 constructs a new utterance for separate processing in place of the original utterance, e.g., utterance “B” 206 and utterance “C” 208, as shown in FIG. 2 . Thus, original utterance 202 may be split into two or more separate utterances that are handled one at a time. MIS 220 determines which of the two or more utterances should be processed first using extracted information 205 and / or from an analysis of utterance 202 itself. For example, MIS 220 may determine that utterance 202 includes a marker word indicating that a particular intent should be handled first. The newly formed utterance corresponding to this particular intent (e.g., utterance 206 or one of utterances 208) will be sent first for further processing by EIS 230. After the conversation triggered by the first utterance has ended (or been temporarily interrupted), the next highest priority utterance (e.g., the other of utterance 206 or utterance 208) may then be sent to EIS 230 for processing.

[0080] The EIS 230 determines whether the received utterance (e.g., utterance 206 or utterance 208) includes a call name for the skill bot. In one embodiment, each skill bot in the chatbot system is assigned a unique call name that distinguishes the skill bot from other skill bots in the chatbot system. A list of call names can be maintained in the data store 250 as part of the skill bot information 254. When the utterance includes words that match the call name, the utterance is considered to be an explicit call. If the bot is not explicitly called, the utterance received by the EIS 230 is considered an implicit call utterance 234 and is input to an intent classifier (e.g., intent classifier 242) of the master bot to determine which bot to use to process the utterance. In some examples, the intent classifier 242 determines that the master bot should process the implicit call utterance. In other examples, the intent classifier 242 determines which skill bot to route the utterance to for processing.

[0081] The explicit call feature provided by EIS 230 has several advantages. It can reduce the amount of processing that a masterbot must perform. For example, when there is an explicit call, the masterbot may not have to perform any intent classification analysis (e.g., using intent classifier 242) or may have to perform reduced intent classification analysis to select a skillbot. Thus, explicit call analysis may enable the selection of a particular skillbot without relying on intent classification analysis.

[0082] There may also be situations where there is overlap in functionality among multiple skillbots. This can occur, for example, when the intents handled by two skillbots overlap or are very close to each other. In such situations, it may be difficult for the masterbot to identify which of multiple skillbots to select based solely on intent classification analysis. In such scenarios, explicit invocation disambiguates the specific skillbot to be used.

[0083] In addition to determining that an utterance is an explicit invocation, EIS 230 is responsible for determining whether any portion of the utterance should be used as input to the explicitly invoked skillbot. In particular, EIS 230 may determine whether any portion of the utterance is not associated with an invocation. EIS 230 may make this determination through analysis of the utterance and / or analysis of extracted information 205. Instead of sending the entire utterance received by EIS 230, EIS 230 may send the portion of the utterance that is not associated with an invocation to the invoked skillbot. In some examples, the input to the invoked skillbot is formed simply by removing any portion of the utterance that is associated with an invocation. For example, "I would like to order a pizza using Pizza Bot" becomes can be shortened to "I want to order a pizza" because "Using Pizza Bot" is related to the invocation of PizzaBot, but not to any processing performed by PizzaBot. In some instances, EIS 230 may reformat the portion to be sent to the invoked bot, for example to form a complete sentence. Thus, EIS 230 determines not only that there is an explicit invocation, but also what to send to the skill bot when there is an explicit invocation. In some instances, there may be no text to input to the invoked bot. For example, if the utterance was "PizzaBot", In this case, the EIS 230 may determine that the PizzaBot is being invoked, but that there is no text to be processed by the PizzaBot. In such a scenario, the EIS 230 may indicate to the SkillBot invoker 240 that there is nothing to send.

[0084] The skillbot invoker 240 invokes a skillbot in various manners. For example, the skillbot invoker 240 can invoke the bot in response to receiving an indication 235 that a particular skillbot has been selected as a result of an explicit invoke. The indication 235 can be sent by the EIS 230 along with input for the explicitly invoked skillbot. In this scenario, the skillbot invoker 240 hands over control of the conversation to the explicitly invoked skillbot. The explicitly invoked skillbot determines an appropriate response to the input from the EIS 230 by treating the input as an independent utterance. For example, the response can be to perform a particular action or to start a new conversation in a particular state, where the initial state of the new conversation depends on the input sent from the EIS 230.

[0085] Another manner in which the skillbot invoker 240 can invoke skillbots is by implicit invocation using an intent classifier 242. The intent classifier 242 can be trained using machine learning and / or rule-based training techniques to determine the likelihood that an utterance represents a task that a particular skillbot is configured to perform. The intent classifier 242 has one intent classifier per skillbot. For example, each time a new skillbot is registered with a masterbot, a list of example utterances associated with the new skillbot can be used to train the intent classifier 242 to determine the likelihood that a particular utterance represents a task that the new skillbot can perform. Parameters (e.g., a set of values ​​for the parameters of a machine learning model) generated as a result of this training can be stored as part of the skillbot information 254.

[0086] In one embodiment, the intent classifier 242 is implemented using a machine learning model, as described in further detail herein. Training the machine learning model may include inputting at least a subset of utterances from example utterances associated with various skill bots to generate, as an output of the machine learning model, an inference about which bot is the correct bot to process any particular training utterance. For each training utterance, an indication of the correct bot to use for that training utterance may be provided as ground truth information. The behavior of the machine learning model may then be adapted (e.g., through backpropagation) to minimize the discrepancy between the generated inference and the ground truth information.

[0087] In some embodiments, the intent classifier 242 determines a confidence score for each skill bot registered with the master bot, indicating the likelihood that the skill bot can process an utterance (e.g., an implicit invocation utterance 234 received from the EIS 230). The intent classifier 242 may also determine a confidence score for each configured system-level intent (e.g., help, exit). If a particular confidence score satisfies one or more conditions, the skill bot invoker 240 will invoke the bot associated with the particular confidence score. For example, a threshold confidence score value may need to be met. Thus, the output 245 of the intent classifier 242 is either the identification of a system intent or the identification of a particular skill bot. In some embodiments, in addition to meeting the threshold confidence score value, the confidence score must exceed the next highest confidence score by a certain winning margin. Imposing such a condition enables routing to a particular skill bot when the confidence scores of multiple skill bots each exceed the threshold confidence score value.

[0088] After identifying the bot based on the evaluation of the confidence score, the skillbot invoker 240 hands over processing to the identified bot. In the case of a system intent, the identified bot is a master bot. Otherwise, the identified bot is a skill bot. Furthermore, the skillbot invoker 240 determines what to provide as input 247 to the identified bot. As described above, in the case of an explicit invoke, the input 247 can be based on a portion of the utterance not associated with the invoke, or the input 247 can be nothing (e.g., an empty string). In the case of an implicit invoke, the input 247 can be the entire utterance.

[0089] The data store 250 comprises one or more computing devices that store data used by various subsystems of the masterbot system 200. As described above, the data store 250 includes rules 252 and skillbot information 254. The rules 252 include, for example, rules for determining, by the MIS 220, when an utterance expresses multiple intents and how to split an utterance expressing multiple intents. The rules 252 further include rules for determining, by the EIS 230, which parts of an utterance that explicitly invokes a skillbot should be sent to the skillbot. The skillbot information 254 includes the call names of the skillbots in the chatbot system, for example, a list of the call names of all skillbots registered to a particular masterbot. The skillbot information 254 also includes information about each skillbot in the chatbot system, such as a list of the call names of all skillbots registered to a particular masterbot. Information used by the intent classifier 242 to determine the confidence score may include, for example, parameters of a machine learning model.

[0090] 3. Components of the SkillBot System 3 is a simplified block diagram of a Skillbot system 300 according to one embodiment. Skillbot system 300 is a computing system that may be implemented solely in software, solely in hardware, or a combination of hardware and software. In one embodiment, such as the embodiment shown in FIG. 1, Skillbot system 300 can be used to implement one or more Skillbots within a digital assistant.

[0091] Skillbot system 300 includes MIS 310, intent classifier 320, and conversation manager 330. MIS 310 is similar to MIS 220 of FIG. 2 and provides similar functionality, including being operable to determine (1) whether an utterance represents multiple intents, and if so, (2) how to split the utterance into separate utterances for each of the multiple intents using rules 352 in data store 350. In one embodiment, the rules applied by MIS 310 to detect multiple intents and split the utterance are the same as the rules applied by MIS 220. MIS 310 receives utterance 302 and extracted information 304. Extracted information 304 is similar to extracted information 205 of FIG. 1 and can be generated using language parser 214 or a language parser local to skillbot system 300.

[0092] The intent classifier 320 may be trained in a manner similar to the intent classifier 242 discussed above in connection with the embodiment of FIG. 2, and as described in more detail herein. For example, in one embodiment, the intent classifier 320 is implemented using a machine learning model. The machine learning model of the intent classifier 320 is trained for a particular skill bot using at least a subset of example utterances associated with that particular skill bot as training utterances. The ground truth for each training utterance will be the particular bot intent associated with that training utterance.

[0093] The utterance 302 may be received directly from a user or may be provided via a masterbot. When the utterance 302 is provided through a masterbot, for example, as a result of processing through the MIS 220 and the EIS 230 in the embodiment shown in FIG. 2, the MIS 310 may be bypassed to avoid repeating processing already performed by the MIS 220. However, if the utterance 302 is received directly from a user, for example, during a conversation that occurs after routing to a skillbot, the MIS 310 may process the utterance 302 to determine whether the utterance 302 represents multiple intents. If the utterance 302 represents multiple intents, the MIS 310 applies one or more rules to split the utterance 302 into separate utterances for each intent, for example, utterance “D” 306 and utterance “E” 308. If the utterance 302 does not represent multiple intents, the MIS 310 forwards the utterance 302, without segmentation, to the intent classifier 320 for intent classification.

[0094] The intent classifier 320 is configured to match a received utterance (e.g., utterance 306 or 308) with an intent associated with the skillbot system 300. As explained above, a skillbot can be configured with one or more intents, each of which includes at least one example utterance associated with that intent and used to train the classifier. In the embodiment of FIG. 2, the intent classifier 242 of the masterbot system 200 is trained to determine a confidence score for each individual skillbot and a confidence score for the system intent. Similarly, the intent classifier 242 of the masterbot system 200 is trained to determine a confidence score for each individual skillbot and a confidence score for the system intent. The intent classifier 320 can be trained to determine a confidence score for each intent associated with the skillbot system 300. While the classification performed by the intent classifier 242 is at the bot level, the classification performed by the intent classifier 320 is at the intent level and therefore at a finer granularity. The intent classifier 320 has access to intent information 354. For each intent associated with the skillbot system 300, the intent information 354 includes a list of utterances that represent the meaning of the intent and are typically associated with tasks that can be performed by the intent. The intent information 354 can further include parameters generated as a result of training on this list of utterances.

[0095] The conversation manager 330 receives as output from the intent classifier 320 an indication 322 of the particular intent identified by the intent classifier 320 as the best match for the utterance input to the intent classifier 320. In some instances, the intent classifier 320 is unable to determine any match. For example, the confidence score calculated by the intent classifier 320 may fall below a threshold confidence score value if the utterance is directed to a system intent or to the intent of a different skill bot. When this occurs, the skill bot system 300 may refer the utterance to the master bot for processing, e.g., routing to a different skill bot. However, if the intent classifier 320 successfully identifies the intent within the skill bot, the conversation manager 330 begins a conversation with the user.

[0096] The conversation initiated by the conversation manager 330 is specific to the intent identified by the intent classifier 320. For example, the conversation manager 330 may be implemented using a state machine configured to execute a dialog flow for the identified intent. The state machine may include a default starting state (e.g., for when the intent is invoked without any additional input) and one or more additional states, each having associated therewith an action to be performed by the skill bot (e.g., perform a purchase transaction) and / or a dialog to be presented to the user (e.g., question, response). Thus, the conversation manager 330 may determine an action / dialog 335 upon receiving an instruction 322 identifying an intent, and may determine the additional action or dialog in response to subsequent utterances received during the conversation.

[0097] Data store 350 comprises one or more computing devices that store data used by various subsystems of skillbot system 300. As shown in Figure 3, data store 350 includes rules 352 and intent information 354. In an embodiment, data store 350 can be integrated with a masterbot or digital assistant data store, such as data store 250 of Figure 2.

[0098] 4. Scheme for classifying utterances using a trained intent classifier 4 is a block diagram illustrating aspects of a chatbot system 400 configured to train and utilize a classifier (e.g., intent classifier 242 or 320 described with respect to FIGS. 2 and 3) based on text data 405. As shown in FIG. 4, the text classification performed by the chatbot system 400 in this example includes various stages: a machine learning model training stage 410; a skill bot invocation stage 415 for determining the likelihood that an utterance represents a task that a particular skill bot is configured to perform; and an intent prediction stage 420 for classifying the utterance as one or more intents. The machine learning model training stage 410 may individually be referred to as a machine learning model 425 and collectively referred to as a machine learning model 425) for use by the other stages. For example, machine learning models 425 may include a model for determining the likelihood that an utterance represents a task that a particular skillbot is configured to perform, another model for predicting intent from utterances for a first type of skillbot, and another model for predicting intent from utterances for a second type of skillbot. Still other types of machine learning models may be implemented in other examples consistent with this disclosure.

[0099] The machine learning model 425 may be a machine learning (“ML”) model such as a convolutional neural network (“CNN”), e.g., an inception neural network, a residual neural network (“Resnet”), or a recurrent neural network, e.g., a long short-term memory (“LSTM”) model or a gated recurrent unit (“GRU”) model, or other variants of a deep neural network (“DNN”) (e.g., a multi-level n-binary DNN classifier or a multi-class DNN classifier for single-intent classification). The machine learning model 425 may also be any other suitable ML model trained for natural language processing, such as a naive Bayes classifier, a linear classifier, a support vector machine, a bagging model such as a random forest model, a boosting model, a shallow neural network, or a combination of one or more of such techniques, e.g., a CNN-HMM or an MCNN (multiscale convolutional neural network). The chatbot system 400 may employ the same or different types of machine learning models to determine the likelihood of a task that a particular skillbot is configured to perform, to predict intents from utterances for a first type of skillbot, and to predict intents from utterances for a second type of skillbot. Still other types of machine learning models may be implemented in other examples according to this disclosure.

[0100] To train the various machine learning models 425, the training stage 410 consists of three main components: dataset preparation 430, feature engineering 435, and model training 440. Dataset preparation 430 includes the process of loading data assets 445, splitting the data assets 445 into training and validation sets 445a-n, and performing basic preprocessing so that the system can train and test the machine learning models 425. The data assets 445 may include at least a subset of utterances from example utterances associated with the various skill bots. As described above, the utterances may be provided in various ways, including as audio or text. The utterances may be fragments, complete sentences, multiple sentences, etc. For example, if the utterances are provided as audio, data preparation 430 may convert the audio to text using a speech-to-text converter (not shown), which inserts punctuation marks, such as commas, semicolons, periods, etc., into the resulting text. In some examples, the example utterances are provided by a client or customer. In other examples, example utterances are automatically generated from a library of prior utterances (e.g., identifying utterances from the library specific to the skill the chatbot is to learn). Data assets 445 for machine learning model 425 may include input text or speech (or input features of text or speech frames) and labels 450 corresponding to the input text or speech (or input features) as a matrix or table of values. For example, for each training utterance, an indication of the correct bot to use for that training utterance may be provided as ground truth information for label 450. The behavior of machine learning model 425 can then be adapted (e.g., through backpropagation) to minimize the difference between the generated inferences and the ground truth information. Alternatively, machine learning model 425 may be trained for a particular skill bot using at least a subset of example utterances associated with that skill bot as training utterances.The ground truth for the label 450 for each training utterance will be the particular bot intent associated with that training utterance.

[0101] In various embodiments, feature engineering 435 includes converting data assets 445 into feature vectors and / or creating new features using data assets 445. Feature vectors may include count vectors as features, word frequency-inverse document frequency (TF-IDF) vectors as features, such as word-level, n-gram-level, or character-level features, word embeddings as features, text / NLP as features, topic models as features, or combinations thereof. A count vector is a matrix representation of data assets 445, where each row represents an utterance, each column represents a word from the utterance, and each cell represents the frequency count of a particular word within the utterance. The TF-IDF score represents the relative importance of a word in the utterance. Word embedding is a form of representing words and utterances using dense vector representations. The location of a word in the vector space is learned from the text and is based on the words surrounding that word when it is used. Text / NLP-based features may include the number of words in the utterance, the number of characters in the utterance, the average word density, the number of punctuation marks, the number of capital letters, the number of lemmas, the frequency distribution of part-of-speech tags (e.g., nouns and verbs), or any combination thereof.Topic modeling is the technique of identifying, from a collection of utterances, groups of words (called topics) that contain the best information in the collection.

[0102] In various embodiments, model training 440 includes training a classifier using the feature vectors and / or new features created in feature engineering 435. In some examples, the training process includes iterative operations to find a set of parameters for the machine learning model 425 that minimizes or maximizes a cost function (e.g., minimizes a loss function or error function) for the machine learning model 425. Each iteration may involve finding a set of parameters for the machine learning model 425 such that the value of the cost function using the set of parameters for the machine learning model 425 is smaller or larger than the value of the cost function using another set of parameters in the previous iteration. The cost function may be constructed to measure the difference between the output predicted using the machine learning model 425 and the labels 450 included in the data asset 445. Once the set of parameters is identified, the machine learning model 425 is trained and can be utilized for prediction as designed.

[0103] In various embodiments, the training stage 410 further includes a logit decision 455. The logit decision 455 is configured to apply a logit function to the probability that an utterance is associated with a resolvable class. To determine the probability that an utterance is associated with an unresolvable class, the logit decision 455 can use enhanced logit values ​​instead of the logit function. The enhanced logit values ​​can be processed (e.g., by an activation function) to determine the classification of the utterance, including a classification that the utterance corresponds to having an unresolved intent. The use of enhanced logit values ​​by the logit decision 455 can increase the performance of the model and the chatbot, particularly for predicting whether an utterance is out-of-scope or out-of-domain. While the logit decision 455 is shown in FIG. 4 as a subprocess of the training stage, in some embodiments, the logit decision 455 may form part of the model training 440, such as when enhanced logit values ​​corresponding to unresolvable classes are learned values ​​that can be dynamically adjusted during training of the intent classifier. Additionally and / or alternatively, the enhanced logit value can be tuned by hyperparameter optimization of the intent classifier. Details of the implementation of the logit decision 455 are described further herein below.

[0104] In addition to data assets 445, labels 450, feature vectors and / or new features, other techniques and information can be employed to improve the training process of the machine learning model 425. For example, the feature vectors and / or new features can be used to refine the training process of the machine learning model 425. The machine learning models 425 may be combined with each other to help improve the accuracy of the classifier or model. Additionally or alternatively, hyperparameters may be adjusted or optimized; for example, certain parameters, such as tree length, leaf, and network parameters, may be fine-tuned to obtain the best-fit model. However, the training mechanisms described herein primarily focus on training machine learning models 425. These training mechanisms may also be utilized to fine-tune existing machine learning models 425 trained from other data assets. For example, in some cases, the machine learning model 425 may have been pre-trained using utterances specific to another skill bot. In such cases, the machine learning model 425 may be retrained using data assets 445, as discussed herein.

[0105] The machine learning model training stage 410 outputs trained machine learning models 425, including a task machine learning model 460 and an intent machine learning model 465. The task machine learning model 460 may be used in the skillbot invocation stage 415 to determine the likelihood that an utterance represents a task that a particular skillbot is configured to perform (470), and the intent machine learning model 465 may be used in the intent prediction stage 420 to classify the utterance as one or more intents (475). In some examples, the skillbot invocation stage 415 and the intent prediction stage 420 may proceed independently, using separate models. For example, the trained intent machine learning model 465 may be used to predict an intent for a skillbot in the intent prediction stage 420 without first identifying the skillbot in the skillbot invocation stage 415. Similarly, the task machine learning model 460 may be used to predict the task or skillbot to be used for an utterance in the skillbot invocation stage 415 without identifying the intent of the utterance in the intent prediction stage 420.

[0106] Alternatively, the skillbot invocation stage 415 and the intent prediction stage 420 may occur sequentially, with one stage using the output of the other stage as input, or with one stage being invoked in a specific manner for a particular skillbot based on the output of the other stage. For example, for given text data 405, the skillbot invocation unit can invoke a skillbot through implicit invocation using the skillbot invocation stage 415 and the task machine learning model 460. The task machine learning model 460 can be trained using machine learning and / or rule-based training techniques to determine the likelihood that an utterance represents a task that a particular skillbot 470 is configured to perform. Then, for an identified or invoked skillbot and given text data 405, the intent prediction stage 420 and the intent machine learning model 465 can be used to match the received utterance (e.g., an utterance in a given data asset 445) to an intent 475 associated with the skillbot. As described herein, a skillbot can be configured with one or more intents, each of which includes at least one example utterance associated with that intent and used to train a classifier. In some embodiments, the skillbot invocation stage 415 and task machine learning model 460 used in the masterbot system are trained to determine confidence scores for individual skillbots and for system intents. Similarly, the intent prediction stage 420 and intent machine learning model 465 can be trained to determine confidence scores for each intent associated with the skillbot system. While the classification performed by the skillbot invocation stage 415 and task machine learning model 460 is at the bot level, the classification performed by the intent prediction stage 420 and intent machine learning model 465 is at the intent level and therefore has finer granularity.

[0107] C. Logit function 5 shows a schematic diagram of an exemplary neural network 500, according to some embodiments. The neural network 500 can be a machine learning model trained by a training system and implemented by a chatbot system, where the machine learning model is trained to predict whether an utterance corresponds to a particular intent class. In some examples, the neural network 500 includes an input layer 502, a hidden layer 504, and an output layer 506. In some examples, the neural network 500 includes multiple hidden layers, with the hidden layer 504 corresponding to the final hidden layer of the neural network.

[0108] The input layer 502 can receive input data or a representation of input data (e.g., an n-dimensional array of values ​​representing an utterance) and apply one or more learned parameters to the input data to generate a set of outputs (e.g., a set of numerical values), which can be processed by the hidden layer 504.

[0109] The hidden layer 504 can contain one or more learned parameters that transform the set of outputs into a different feature space in which input data points from different classes are better separated.

[0110] The output layer 506 may include a classification layer that maps the output from the hidden layer to logit values, each value corresponding to one particular class (e.g., resolvable class, unresolvable class). In some examples, the output layer 506 includes an activation function to constrain the logit values ​​to a set of probability values ​​that sum to 1. Thus, the logit values ​​produced by the classification layer (logit function) for each class may then be processed by the activation function in the output layer (also referred to herein as an "activation layer") to predict a classification for the utterance.

[0111] As part of predicting a single intent from a set of output values, the machine learning model may employ a logit function (short for "logistic regression function") in a network layer of the neural network 500. The output layer 506 may include a logit function to convert intermediate outputs (e.g., probability values ​​predicting whether an utterance corresponds to a particular class) into logit values. In some cases, the logit function is the logarithm of the odds, taking inputs corresponding to probability values ​​between 0 and 1 for a particular class and outputting logit values ​​within an unbounded range between negative infinity and positive infinity. The logit function may be used to normalize each intermediate output in the set of intermediate outputs, so that the resulting set of logit values ​​can be expressed as a symmetric, unimodal probability distribution across the predicted output class. Mathematically speaking, the logit function is

[0112]

number

[0113] where p is the probability of an input corresponding to a particular class. The probability p can be set within the range between (0, 1). The output logit can correspond to a logit value ranging between (-∞, +∞). In this way, a machine learning model can employ a logit function such that the ensuing activation function can predict the most likely output (e.g., intent) from other less likely outputs of the classifier.

[0114] As part of training a machine learning model (e.g., intent classifier 320 of FIG. 3), the model may learn centroids for each class, and the centroids are part of the logit function. In particular, the centroids are used to weight the logit function, e.g., , which can be used as part of filtering the classifier output to separate possible intents from an utterance. In mathematical terms, the equation relating the logit function is logit i = f(x)*W iwhere x is the input to the model and , f(x) is a set of transformations (e.g., highway network functions), and W i is the centroid for intent "i" that acts as a weighting factor for the model output for that intent. Thus, for a set of identified intents that an intent classifier is trained to recognize, the centroids serve to classify utterances as belonging to one particular intent over another, for example, by using the centroids as locations to measure the distance between adjacent clusters.

[0115] The logit values ​​generated for each predicted output class may then be processed by an activation function to map the odds represented by the logit values ​​to a probability distribution over all predicted output classes. An example of an activation function is the softmax function. We can use Across all predicted output classes (e.g., order_pizza, unresolved intents) It may be normalized to a probability distribution.

[0116] Before applying softmax, the logit values ​​may be negative or greater than 1. and may not sum to 1. When applying softmax, each output is in the interval ( The logit values ​​will be in the range 0,1) and the output will sum to 1. Furthermore, larger input logit values ​​will correspond to larger probabilities. In functional terms, the softmax function is It is represented by:

[0117]

number

[0118] i = 1:K, and z is a set of K real numbers (z1: z K ) In some cases, the softmax function can be weighted by a basis factor b, which produces a probability distribution that is more concentrated around the location of the maximum input value. In such cases, the softmax function is

[0119]

number

[0120] where β is a real number. In the context of intent classification, z i is f(x)*w i is set to Good too.

[0121] For a set of utterances, a resolved intent may fit any of the identified intents for which an intent classifier may be trained. An unresolved intent may, by definition, describe an utterance that the classifier cannot confidently place into any of the intent categories. As an illustrative example, a skill bot trained to handle pizza orders may be trained to classify utterances into one of three intents: "order_pizza," "cancel_pizza," and "unresolvedIntent." Utterances intended for ordering a new pizza or modifying an existing order may be classified as "order_pizza," for example. In contrast, utterances intended for canceling an existing order may be classified under "cancel_pizza". Finally, utterances intended for something else may be classified under "unresolvedIntent". For example, "What's the weather like tomorrow?" is not intended to be a pizza order. They are irrelevant and should be classified as "unresolvedIntent." So in this example, "unresolvedIntent" is a negative response to "order_pizza" and "cancel_pizza." It is defined by its meaning. Utterances that do not conform to the latter conform to the former.

[0122] Thus, unresolved intents may be defined to cover utterances that do not fit the set of identified intents. In this way, for unresolved classes, the classifier output may be distributed rather than clustered. While the examples described herein focus on intent prediction, it should be understood that other classification tasks may be similarly addressed. For example, a chatbot system may include a classifier to address "unresolved" utterances at other levels, such as skillbot invocation.

[0123] D. Techniques for Classifying Utterances as Having Unresolved Intents Using Augmented Logit Applying the concept of centroids as weighting coefficients in a logit function can be useful in some instances. However, for the "unresolved intent" class, the effectiveness of using centroids is greatly reduced. For a defined intent such as "order_pizza," Stories can be clustered around a centroid: "unresolved_intent" If ,is defined as the negative classification for all utterances outside the set of classifications,,such utterances may exhibit a relative lack of clustering.,Thus, learning centroids may introduce errors into logistic,regression through inaccurate weighting coefficients.

[0124] As an example, in the case of the pizza skill bot mentioned above, utterances with intents to order or cancel may cluster around their respective centroids, while utterances ranging from weather to finance to asking about specific people may fall into an "unresolved_intent." Similarly, to direct a query to a particular skill bot, the classifier may output the probability that the utterance will invoke a particular skill bot, or that the utterance will not invoke any of the skill bots the system was trained to classify. In each of these cases, the centroid for the unresolved class may not be representative of the population characteristics that result, in large part, from the definition of the "unresolved" classification as including any utterance that does not map to a learned intent, skill, domain, or scope.

[0125] To overcome this and other problems, various embodiments are directed to techniques for using enhanced logit values ​​for unresolved intents to improve classification of utterances as unresolved intents. For example, a machine learning model (e.g., intent classifier 320 of FIG. 3 ) can implement enhanced logit values ​​for unresolved intent classes. For example, instead of calculating logit values ​​using centroids as weighting coefficients, enhanced logit values ​​for unresolvable classes may be determined by using a logit function modified to include a trainable scalar parameter. Additionally or alternatively, enhanced logit values ​​for unresolvable classes can be optimized through hyperparameter tuning of the machine learning model. In some examples, if the training data is insufficient to indicate a resolvable class, the output is classified by the intent classifier as an unresolvable class. In practice, the intent classifier may determine whether the utterance invokes or does not invoke a resolvable class, rather than determining whether the utterance invokes an unresolvable class.

[0126] The enriched logit values ​​can be determined whenever classification of an utterance (e.g., intent, scope, skill, etc.) can be undertaken using an automated process that can be integrated into a chatbot system, as described with respect to Figures 1, 2, and 3. Advantageously, the model and chatbot perform better in classifying out-of-scope utterances using the enriched logit values, at least in part because the enriched logit values This is because the calculated logit values ​​do not rely on unresolvable class centroids, which may not reflect the true clustering in the model output. Furthermore, because the augmentations are applied automatically, customers or clients are less likely to experience misdirected queries from the chatbot system.

[0127] As described further below, the enhanced logit value for classifying an utterance into the unresolvable class may be based on one of the following: (i) a statistical value determined based on a set of logit values ​​generated from a training dataset; (ii) a bounded value selected from a range of values ​​defined by the logarithm of first odds corresponding to the probability for the unresolvable class, where the logarithm of the first odds is constrained to a range of values ​​by a bounding function and weighted by the centroid of a distribution associated with the unresolvable class; (iii) a weighted value generated by the logarithm of second odds corresponding to the probability for the unresolvable class, where the logarithm of the second odds is scaled by a scaling factor, bounded by the bounding function, and weighted by the centroid of the distribution associated with the unresolvable class; (iv) a hyperparameter optimization value generated based on hyperparameter tuning of a machine learning model; or (v) a learned value dynamically adjusted during training of a machine learning model. Additionally or alternatively, a rule-based approach may be implemented as part of determining the enhanced logit value. A rule-based approach may include a logical operator that identifies an enriched logit value from a set of enriched logit values ​​for an unresolvable class. For example, the enriched logit value may be retrieved from a database or lookup table based on a query value generated by applying a set of rules to the utterance.

[0128] 1. Statistics In some embodiments, the enhanced logit value is a statistical value that replaces the logit function when determining whether an utterance corresponds to an out-of-scope class or an out-of-domain class. The logit value can include a value selected from a range of values. The function formula logit i =f(x)*W i Referring to logit i is the function f(x)*W for the unresolvable class. i The place A certain number "A" (i.e., logit i= A) Such a configuration is a fixed value

[0129]

number

[0130] In some examples, the statistical value is set as 0.3, 0.5, 0.7, 0.9, 1.0, or any higher number. The statistical value is determined such that the corresponding probability value indicates that the utterance is classified as unresolved (e.g., unresolved skill or unresolved intent).

[0131] In some embodiments, the statistics are based on the probability that an out-of-domain or out-of-scope utterance in a training set of utterances maps to an unresolvable class. Determining appropriate statistics may be part of the training phase 410 of FIG. 4, the number of which may depend in part on the training data set (e.g., data assets 445) applied in the training phase 410.

[0132] 6 shows a flowchart illustrating an example process 600 for determining a statistic representing an enhanced logit value to predict whether an utterance corresponds to an unresolvable class, according to some embodiments. The process shown in FIG. 6 may be implemented by software (e.g., a The software may be implemented in a non-transitory storage medium (e.g., code, instructions, program), hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 6 and described below is intended to be exemplary and non-limiting. While FIG. 6 depicts various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In some embodiments, the steps may be performed in some different order, or some steps may be performed in parallel. In some embodiments, such as those depicted in FIGS. 1-4, the process depicted in FIG. 6 may be performed by a training subsystem to map probability values ​​corresponding to unresolvable classes to enhanced logit values ​​of one or more machine learning models (e.g., intent classifier 242 or 320). The training subsystem can be part of a data processing system (e.g., chatbot system 400 described with respect to FIG. 4) or a component of another system configured to train and deploy machine learning models.

[0133] At 605, the training subsystem receives a training dataset. The training dataset can include a set of utterances or messages. Each utterance in the set is associated with a training label, which identifies the utterance's predicted intent class. The training dataset may also include utterances or messages that are out-of-scope from a specific task that the corresponding skill or chatbot is configured to perform. For example, if a skillbot is configured to process food orders for a pizza restaurant, an out-of-scope utterance could be a message inquiring about the weather. In some examples, each out-of-scope utterance or message in the training dataset is associated with a training label, which identifies the utterance or message as an unresolvable class.

[0134] At 610, the training subsystem uses the training dataset to train a machine learning model to predict whether an utterance or message represents a task the skillbot is configured to perform or to match the utterance or message to an intent associated with the skillbot. One or more parameters of the machine learning model may be learned to minimize the loss between the predicted output generated by the machine learning model and the expected output indicated by the training labels of the corresponding utterance. Training of the machine learning model may be performed until the loss reaches a minimum error threshold.

[0135] At 615, the training subsystem identifies a set of training logit values. Each training logit value in the set corresponds to an out-of-scope utterance in the training dataset. Each training logit value may be generated based on inputting a respective out-of-scope utterance into a machine learning model. To obtain the training logit values, one or more helper functions (e.g., a get_layer function) may be utilized to generate the training logit values. It can take in the output generated from one or more layers of the model.

[0136] At 620, the training subsystem determines a statistical value representative of the set of training logit values. For example, the statistical value may be the median of the set of training logit values. The median may be used in response to determining that the statistical distribution of the set of training logit values ​​corresponds to a skewed distribution. In some examples, a mean and median are determined for the set of training logit values. If the difference between the mean and median is statistically significant, the median is selected as the statistical value representative of the set of training logit values.

[0137] In 625, the training subsystem corresponds to a class of utterances that cannot be resolved. Set the statistic as the enhanced logit value for predicting whether the utterance corresponds to a resolvable class. The enhanced logit value can replace the existing logit function of the machine learning model that would have been used to predict that the utterance corresponds to an unresolvable class. In some instances, the existing logit function can be used to predict whether the utterance corresponds to a resolvable class (e.g., order_pizza, cancel_pizza). ) will continue to be used to predict utterances corresponding to one of the

[0138] At 630, the training subsystem deploys the trained machine learning model to the chatbot system (e.g., as part of a skillbot), where the trained machine learning model includes a statistic set as an enriched logit value for predicting whether an utterance corresponds to an unresolvable class. Process 600 then ends.

[0139] By using statistics and removing the influence of centroids from unresolvable classes, machine learning model performance can be significantly improved with respect to accurately predicting whether an utterance will invoke a resolvable or unresolvable class. In an exemplary example, the pizza skill described above includes the following three intents: "order_pizza", "cancel_pizza", and "unresolvedIntent". The utterance "What's the weather like tomorrow?" can be used to select one of the three intents: For the center of gravity weighting coefficient w i The logit value calculated based on i = F(x)*w i ) as (0.3,0.01,0.01). In this case, the logit vector The three values ​​are found using a centroid weighted relationship and are then scaled using the softmax function with a probability of (0.4 ,0.3,0.3). In this example, "order_pizza" has a 40% confidence score. This is the most likely result with a core. However, it should be clear that a query about the weather does not represent an intent to order pizza. The classification in this example is incorrect, in part, due to the relatively low logit value of the unresolved class, which is equal to 0.01 (the correct classification would be out-of-domain unresolvedIntent).

[0140] Continuing with this example, the enhanced logit value may be a statistical value equal to 0.4. The statistical value can be applied while maintaining the centroid-weighted logit values ​​for the resolvable classes. Thus, the logit value can be redefined as (0.3, 0.01, 0.4). Applying the softmax function again results in a logit value of (0.35, 0.26, 0. 39) is converted to a probability. In this case, the result of the softmax function is a 39% confidence score. This reflects the exact classification of "unresolvedIntent" that has

[0141] 2. Bounded Values Determining enhanced logit values ​​for unresolvable classes can be complicated by the fact that the standard logit function is unbounded. As mentioned above, the logit function outputs logit values ​​that range between negative infinity and positive infinity. Therefore, the unbounded range can introduce complications into determining enhanced logit values. For example, logit values ​​that span negative infinity and positive infinity can result in logit values ​​that are distributed over a very wide range of values, which can introduce difficulties in optimizing the enhanced logit values.

[0142] Thus, the enhanced logit value can be a bounded value within the range specified by the modified logit function. The modified logit function can be modified to constrain the logit value to a range that includes a lower and upper bound. In some examples, the modified logit function includes a bounded function. The bounded function can include a trigonometric function that produces an output within the range between negative one and one, such as a sine or cosine function. For example, the logit function can be expressed as: i = cosine(f(x)*W i ), with a minimum value of -1 and a maximum value of 1. The enriched logit value can then be chosen from a bounded range between -1 and 1 for the unresolvable classes. Continuing with the example, for the order_pizza class, The logit function for the classes sushi and cancel_pizza is logit i = cosine(f(x)*W i ) can be used, and a bounded value of 0.4 can be chosen for the unresolvedIntent class.

[0143] FIG. 7 shows a flowchart illustrating an example process 700 for modifying a logit function to determine an enhanced logit value within a specified range to predict whether an utterance corresponds to an unresolvable class, according to some embodiments. The process illustrated in FIG. 7 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of a respective system, hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 7 and described below is intended to be exemplary and non-limiting. While FIG. 7 shows various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In some embodiments, the steps are performed by a training subsystem to map probability values ​​corresponding to unresolvable classes to enhanced logit values ​​for one or more machine learning models (e.g., intent classifier 242 or 320). The training subsystem can be part of a data processing system (e.g., chatbot system 400 described with respect to FIG. 4) or a component of another system configured to train and deploy machine learning models.

[0144] At 705, the training subsystem initializes a machine learning model. The machine learning model may be a convolutional neural network (“CNN”), such as an Inception Neural Network, a Residual Neural Network (“Resnet”), or a recurrent neural network, such as a long short-term memory (“LSTM”) model or a gated recurrent unit (“GRU”) model, other variants of a deep neural network (“DNN”) (e.g., a multi-level n-binary DNN classifier or a multi-class DNN classifier for single-intent classification. The machine learning model 425 may also be any other suitable ML model trained for natural language processing, such as a naive Bayes classifier, a linear classifier, a support vector machine, a bagging model such as a random forest model, a boosting model, a shallow neural network, or a combination of one or more of such techniques, e.g., a CNN-HMM or an MCNN (multi-scale convolutional neural network).

[0145] Initializing a machine learning model can include defining the number of layers, the type of each layer (e.g., fully connected, convolutional neural network), and the type of activation function for each layer (e.g., sigmoid, ReLU, softmax).

[0146] At 710, the training subsystem retrieves the logit function from the final layer of the machine learning model. For example, the training system can select the final layer from a set of layers in a fully connected neural network. The logit function of the final layer can then be accessed for modification.

[0147] At 715, the training subsystem modifies the logit function by adding a bounding function. The bounding function may include a trigonometric function that produces an output within the range of negative one and one, such as a sine or cosine function. In some examples, the bounding function may be the following formula:

[0148]

number

[0149] where f(x) is the logit function. Additionally or alternatively, the logit function is selected from one of the following: i where i corresponds to each predicted class. .

[0150] At 720, the training subsystem trains the machine learning model to generate enhanced logit values ​​based on the range specified by the modified logit function. Thus, the training subsystem can train the machine learning model to generate enhanced logit values ​​by identifying statistics from the training dataset that predict whether an utterance corresponds to an unresolved class (e.g., an unresolved skill or an unresolved intent).

[0151] At 725, the training subsystem deploys the trained machine learning model to the chatbot system (e.g., as part of a skillbot), where the trained machine learning model includes the modified logit function. Thereafter, process 700 ends.

[0152] 3. Weighted Values The enhanced logit value for predicting out-of-scope utterances may be a weighted value produced by a modified logit function weighted by a scaling factor. In some examples, the scaling factor (also referred to herein as a "scaler") expressed in the function for class "i" is the logit i = scaler * cosine(f(x)*W i ) Solution The scaling factor value (e.g., 2) used for unresolved classes (e.g., unresolvedIntent) is larger than the scaling factor value (e.g., 2) used for resolvable classes (e.g., order_pizza, cancel_pizza). The value of the scaling factor (e.g., 1) used for the logit function can be different from the value of the scaling factor used for the logit function (e.g., 1). Including the scaling factor in the logit function can widen the gap between the probability of the predicted intent and the probabilities of other intents, thus allowing predicted intents in the unresolvable class to be better distinguished from other intents in the resolvable class.

[0153] FIG. 8 shows a flowchart 800 illustrating an exemplary process for adding a scaling factor to a logit function to determine an enhanced logit value for predicting whether an utterance corresponds to an unresolvable class, according to some embodiments. The process shown in FIG. 8 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of a respective system, hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 8 and described below is intended to be exemplary and non-limiting. While FIG. 8 shows various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In some embodiments, the steps are performed by a training subsystem to map probability values ​​corresponding to unresolvable classes to enhanced logit values ​​for one or more machine learning models (e.g., intent classifier 242 or 320). The training subsystem can be part of a data processing system (e.g., chatbot system 400 described with respect to FIG. 4) or a component of another system configured to train and deploy machine learning models.

[0154] At 805, the training subsystem initializes a machine learning model. The machine learning model may be a convolutional neural network ("CNN"), such as an Inception neural network, a residual neural network ("Resnet"), or a recurrent neural network, such as a long short-term memory ("LSTM") model or a gated recurrent neural network. The machine learning model 425 may be a recurrent unit ("GRU") model, other variants of a deep neural network ("DNN") (e.g., a multi-level n-binary DNN classifier or a multi-class DNN classifier for single-intent classification). The machine learning model 425 may also be any other suitable ML model trained for natural language processing, such as a naive Bayes classifier, a linear classifier, a support vector machine, a bagging model such as a random forest model, a boosting model, a shallow neural network, or a combination of one or more of such techniques, e.g., a CNN-HMM or an MCNN (multi-scale convolutional neural network).

[0155] Initializing a machine learning model can include defining the number of layers, the type of each layer (e.g., fully connected, convolutional neural network), and the type of activation function for each layer (e.g., sigmoid, ReLU, softmax).

[0156] At 810, the training subsystem retrieves the logit function from the final layer of the machine learning model. For example, the training system can select the final layer from a set of layers in a fully connected neural network. The logit function of the final layer can then be accessed for modification.

[0157] At 815, the training subsystem modifies the logit function by adding a scaling factor. In some examples, different initial values ​​of the scaling factor are assigned based on the predicted class of the corresponding utterance. For example, a scaling factor value (e.g., 2) used for an unresolvable class (e.g., unresolvedIntent) is used for a resolvable class (e.g., order_pizza, cancel_pizza). The scaling coefficients may be different from the values ​​of the scaling coefficients (e.g., 1) that are assigned to the logit functions. A first scaling coefficient of a logit function that generates a probability indicating whether an utterance corresponds to an unresolvable class is assigned a larger value than a second scaling coefficient of another logit function that generates a probability indicating whether an utterance corresponds to one of the resolvable classes. Additionally or alternatively, the logit functions may be further modified by adding a bounding function (e.g., a cosine function) to the logit function to which the scaling coefficients are added.

[0158] At 820, the training subsystem trains the machine learning model to generate enhanced logit values ​​based on the scaling coefficients assigned to the unresolvable classes. In some examples, the scaling coefficients for the unresolvable classes are learned parameters of the machine learning model. Thus, the training subsystem can train the machine learning model to adjust the scaling coefficients corresponding to the unresolvable classes. Additionally or alternatively, the scaling coefficients can be hard-coded numbers that can remain static throughout the training and deployment of the machine learning model.

[0159] At 825, the training subsystem deploys the trained machine learning model to a chatbot system (e.g., as part of a skillbot), where the trained machine learning model includes the modified logit function. Process 800 then ends.

[0160] 4. Hyperparameter Tuning The enhanced logit value for predicting whether an utterance corresponds to an unresolvable class may be a hyperparameter optimization value determined based on hyperparameter tuning of the machine learning model. In some embodiments, the hyperparameter tuning enables meta-training of the enhanced logit value as part of the logit determination 455. Tunable hyperparameters of the machine learning model include the learning rate, epoch time, and so on. The hyperparameters may include the number of blocks, momentum, regularization constant, the number of layers in the machine learning model, and the number of weights in the machine learning model. Hyperparameter tuning may include the process of determining an optimal combination of the above hyperparameters that allows for finding enhanced logit values ​​for classes that the intent classifier cannot resolve.

[0161] FIG. 9 shows a flowchart 900 illustrating an exemplary process for using hyperparameter tuning to determine an enhanced logit value to predict whether an utterance corresponds to an unresolvable class, according to some embodiments. The process shown in FIG. 9 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of a respective system, hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 9 and described below is intended to be exemplary and non-limiting. While FIG. 9 shows various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In some embodiments, the steps may be performed in some different order, or some steps may be performed in parallel. In some embodiments, such as those shown in FIGS. 1-4, the process shown in FIG. 9 may be performed by a training subsystem to map probability values ​​corresponding to unresolvable classes to enhanced logit values ​​for one or more machine learning models (e.g., intent classifier 242 or 320). The training subsystem may be a component of an intent classifier of a data processing system (e.g., the chatbot system 400 described with respect to FIG. 4) or another system configured to train and deploy machine learning models.

[0162] At 905, the training subsystem receives a training dataset. The training dataset can include a set of utterances or messages. Each utterance in the set is associated with a training label, which identifies the utterance's predicted intent class. The training dataset may also include utterances or messages that are out-of-scope from a specific task that the corresponding skill or chatbot is configured to perform. For example, if a skillbot is configured to process food orders for a pizza restaurant, an out-of-scope utterance could be a message inquiring about the weather. In some examples, each out-of-scope utterance or message in the training dataset is associated with a training label, which identifies the utterance or message as an unresolvable class.

[0163] At 910, the training subsystem uses the training dataset to train a machine learning model to predict whether an utterance or message represents a task the skill bot is configured to perform or to match the utterance or message to an intent associated with the skill bot. One or more parameters of the machine learning model may be learned to minimize a loss between a predicted output generated by the machine learning model and an expected output indicated by the training labels of the corresponding utterance. Training of the machine learning model may be performed until the loss reaches a minimum error threshold.

[0164] At 915, the training subsystem identifies a set of training logit values. Each training logit value in the set corresponds to an out-of-scope utterance in the training dataset. Each training logit value may be generated based on inputting a respective out-of-scope utterance into a machine learning model. To obtain the training logit values, one or more helper functions (e.g., a get_layer function) may be utilized to generate the training logit values. It can take in the output generated from one or more layers of the model.

[0165] At 920, the training subsystem determines a statistical value representative of the set of training logit values. For example, the statistical value may be the median of the set of training logit values. The median may be used in response to determining that the statistical distribution of the set of training logit values ​​corresponds to a skewed distribution. In some examples, a mean and median are determined for the set of training logit values. If the difference between the mean and median is statistically significant, the median is selected as the statistical value representative of the set of training logit values.

[0166] At 925, the training subsystem adjusts one or more hyperparameters of the machine learning model to optimize the statistics. The training of the machine learning model can be iterated with the adjusted hyperparameters, where the combination of hyperparameters that provided the best performance for predicting the unresolvable class can be selected. In some embodiments, the best performance corresponds to the model configuration of the machine learning model that produces the smallest loss value between the predicted output generated by the machine learning model for the out-of-scope utterances in the training dataset and the expected output indicated by the training labels of the corresponding out-of-scope utterances.

[0167] Various hyperparameter optimization techniques can be used to optimize the statistics. In some examples, a random search technique is used for hyperparameter tuning. The random search technique can include generating a grid of possible values ​​for the hyperparameters. Each iteration of the search can select a random combination of hyperparameters from this grid, where performance for each iteration can be recorded. Additionally or alternatively, a grid search technique can be used for hyperparameter tuning. The grid search technique can include generating a grid of possible values ​​for the hyperparameters. Each iteration of the search can select a combination of hyperparameters in a particular order. Performance for each iteration can be recorded. The combination of hyperparameters that provided the best performance can be selected.

[0168] Other techniques for hyperparameter tuning can be considered, including, but not limited to, Bayesian optimization algorithms, tree-based Parzen estimators (TPE), hyperbands, population-based training (PBT), and Bayesian optimization and hyperbands (BOHB).

[0169] At 930, the training subsystem sets the optimized statistic as an enhanced logit value for the unresolvable class. The enhanced logit value can replace the existing logit function of the machine learning model that would have been used to predict that the utterance corresponds to the unresolvable class. In some examples, the existing logit function may be used to predict that the utterance corresponds to one of the resolvable classes (e.g., order_pizza, cancel_pizza). The trained machine learning model continues to be used to predict corresponding utterances. At 935, the training subsystem deploys the trained machine learning model to a chatbot system (e.g., as part of a skill bot), where the trained machine learning model includes optimized statistics for predicting whether an utterance or message corresponds to the unresolvable class. Process 900 then ends.

[0170] 5. Learning Value The enhanced logit value for predicting whether an utterance corresponds to an unresolvable class can be a learned value that can be dynamically adjusted during the training of the machine learning model. The model parameters can be learned during the training phase of the machine learning model. In practice, additional training data (e.g., out-of-scope utterances) are used for training, so , the enhanced logit values ​​can be optimized.

[0171] Various techniques may be implemented to generate and adjust the learned values. FIG. 10 shows a flowchart 1000 illustrating an exemplary process for using learned values ​​as enriched logit values ​​to predict whether an utterance corresponds to an unresolvable class, according to some embodiments. The process shown in FIG. 10 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective system, hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 10 and described below is intended to be exemplary and non-limiting. While FIG. 10 shows various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In some embodiments, the steps are performed by a training subsystem to map probability values ​​corresponding to unresolvable classes to enriched logit values ​​for one or more machine learning models (e.g., intent classifier 242 or 320). The training subsystem can be part of a data processing system (e.g., chatbot system 400 described with respect to FIG. 4) or a component of another system configured to train and deploy machine learning models.

[0172] At 1005, the training subsystem receives a training dataset. The training dataset can include a set of utterances or messages. Each utterance in the set is associated with a training label, which identifies the utterance's predicted intent class. The training dataset may also include utterances or messages that are out-of-scope from a specific task that a corresponding skill or chatbot is configured to perform. For example, if a skillbot is configured to process food orders for a pizza restaurant, an out-of-scope utterance could be a message inquiring about the weather. In some examples, each out-of-scope utterance or message in the training dataset is associated with a training label, which identifies the utterance or message as an unresolvable class.

[0173] At 1010, the training subsystem performs batch balancing of the training dataset to generate an augmented training dataset that includes augmented copies of out-of-scope utterances. This step may be optional based on the amount of out-of-scope utterances available in the training dataset. For example, the training subsystem may generate duplicate or augmented copies of the training data corresponding to unresolvable classes (e.g., out-of-scope utterances). Data augmentation techniques may include reverse transformation of utterances in the training dataset, synonym substitution of one or more tokens of an utterance, random insertion of a token into an utterance, swapping between two tokens of an utterance, and random deletion of one or more tokens of an utterance.

[0174] At 1015, the training subsystem initializes a machine learning model. The machine learning model can be a convolutional neural network ("CNN"), e.g., an Inception Neural Network, a Residual Neural Network ("Resnet"), or a recurrent neural network, e.g., a long short-term memory ("LSTM") model or a gated recurrent unit ("GRU") model, other variants of a deep neural network ("DNN") (e.g., a multi-level n-binary DNN classifier or a multi-class DNN classifier for single-intent classification. The machine learning model 425 can also be a naive Bayes classifier, a linear classifier, a support vector machine, a bagging model such as a random forest model, a boosting model, a shallow neural network, or a combination of one or more of such techniques, e.g., a CNN-HMM or MCNN (multi-scheduled neural network). The model may be any other suitable ML model trained for natural language processing, such as a neural network (MLN) or a neural network (NN).

[0175] Initializing a machine learning model can include defining the number of layers, the type of each layer (e.g., fully connected, convolutional neural network), and the type of activation function for each layer (e.g., sigmoid, ReLU, softmax).

[0176] At 1020, the training subsystem trains a machine learning model using the training dataset or the augmented training dataset to predict whether an utterance or message represents a task the skill bot is configured to perform or to match the utterance or message to an intent associated with the skill bot. One or more parameters of the machine learning model may be learned to minimize a loss between a predicted output generated by the machine learning model and an expected output indicated by the training label of the corresponding utterance. Training of the machine learning model may be performed until the loss reaches a minimum error threshold. In particular, the machine learning model may be trained using the augmented training dataset so that it can generate a probability of whether a given utterance corresponds to an unresolvable intent.

[0177] In some examples, to reduce overfitting of the machine learning model with the training dataset or the augmented training dataset, the training subsystem can use regularization to adjust one or more parameter weights. For example, regularization of model parameters can include Gaussian / L2 regularization, in which a small percentage of the machine learning model's weights are removed at each training iteration. Such an approach can improve the accuracy of the enhanced logit value with a relatively small increase in the loss value, thereby reducing the overfitting problem.

[0178] In some cases, the performance of the trained machine learning model is evaluated by cross-validation. For example, the trained machine learning model can be evaluated using a k-fold cross-validation technique, which includes: (i) shuffling training data (e.g., utterances) within an augmented training dataset; (ii) dividing the training dataset into k subsets; (iii) training a respective machine learning model using each of the (k-1) subsets; (iv) evaluating the trained machine learning model using the provided k1 subsets; and (v) evaluating the trained machine learning model using a different provided k subsets. i Repeating steps (iii) and (iv) using the subset.

[0179] At 1025, the training subsystem deploys the trained machine learning model to a chatbot system (e.g., as part of a skillbot) to generate enhanced logit values ​​that predict whether an utterance or message corresponds to the unresolvable class. Because the machine learning model was trained using the expanded training dataset, the enhanced logit values ​​can be used to accurately determine probability values ​​that predict out-of-scope utterances as being associated with the unresolvable class. Process 1000 then ends.

[0180] 6. Example of finding the enhanced logit value As mentioned above, the enhanced logit values ​​for unresolvable classes can be dynamically adjusted using machine learning techniques based on coefficients including, but not limited to, data size or intent number. For example, one scaler may dynamically adjust the logit values ​​for classes (e.g., order_pizza , unresolvedIntent) to limit misoptimalities in the scalar values. Weight decay is added to reduce the logit. In another example, the logit may be dynamically changed within predetermined bounds. For unresolved intent classes, the logit value can be learned as part of the machine learning model training, but can be constrained to a range (e.g., 0.0 => 0.2).

[0181] In an illustrative example of a machine learning model using a logit value for the unresolvable intent class, a fixed logit value for the unresolvable class is applied and the trained model is tested against two standard out-of-domain test sets. These data, shown in Table 1 below, demonstrate the improvement in out-of-domain classification when applying the logit value without using a centroid-weighted approach. For example, applying a fixed logit value improved the trained classifier's accuracy by 9% (compare isMatch values ​​in rows 5 and 1) and recall across epochs by 8% (compare isMatch values ​​in rows 5 and 1) in identifying out-of-domain utterances. (Compare e2e_ood_recall value of 1).

[0182] [Table 1]

[0183] Additionally, experimental results did not reflect a commensurate or simultaneous decrease in performance with respect to accurate classification of in-domain utterances, as shown in Table 2 below. For example, the in-domain test was weighted by a cosine-bounded softmax function weighted by the centroids of the other intent classifications. There was no significant decrease in accuracy between rows 3 and 5 corresponding to applying a predetermined logit value of .5.

[0184] [Table 2]

[0185] E. Process for training a machine learning model that achieves enhanced logit values ​​for out-of-scope utterance classification FIG. 11 is a flowchart illustrating a process 1100 for training a machine learning model that achieves enhanced logit values ​​for unresolvable classes, according to some embodiments. The process illustrated in FIG. 11 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of a respective system, hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 11 and described below is intended to be exemplary and non-limiting. While FIG. 11 depicts various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In some embodiments, the steps may be performed in some different order, or some steps may be performed in parallel. In some embodiments, such as those illustrated in FIGS. 1-4, the process illustrated in FIG. 11 may be performed by a training subsystem to map probability values ​​corresponding to unresolvable classes to enhanced logit values ​​for one or more machine learning models (e.g., intent classifier 242 or 320 or machine learning model 425). The training subsystem can be part of a data processing system (e.g., the chatbot system 400 described with respect to FIG. 4) or a component of another system configured to train and deploy machine learning models.

[0186] At 1105, the training subsystem receives a training dataset. The training dataset can include a set of utterances or messages. Each utterance in the set is associated with a training label, which identifies a predicted intent class of the utterance. In some cases, the utterance includes text data converted from speech input (e.g., a voice utterance), which can be converted into a textual utterance in that particular language, and the text utterance can then be processed. The training dataset can also include utterances or messages that are out of scope from a particular task that a corresponding skill or chatbot is configured to perform. For example, if a skillbot is configured to process food orders for a pizza restaurant, an out-of-scope utterance can be a message inquiring about the weather. In some examples, each out-of-scope utterance or message in the training dataset is associated with a training label, which identifies a class to which the utterance or message cannot be resolved. and identify it.

[0187] At 1110, the training subsystem initializes a machine learning model. The machine learning model can include a series of network layers, where the final network layer in the series includes a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value. Thus, the logit function can be the logarithm of odds that converts the probability value of each class to a real number that fits within a probability distribution (e.g., a logit normal distribution). In some examples, the logit function is weighted by the center of gravity of the distribution associated with each class.

[0188] The machine learning model may be a convolutional neural network (“CNN”), e.g., an Inception Neural Network, a Residual Neural Network (“Resnet”), or a recurrent neural network, e.g., a long short-term memory (“LSTM”) model or a gated recurrent unit (“GRU”) model, other variants of a deep neural network (“DNN”) (e.g., a multi-level n-binary DNN classifier or a multi-class DNN classifier for single-intent classification. The machine learning model 425 may also be any other suitable ML model trained for natural language processing, such as a naive Bayes classifier, a linear classifier, a support vector machine, a bagging model such as a random forest model, a boosting model, a shallow neural network, or a combination of one or more of such techniques, e.g., a CNN-HMM or an MCNN (multi-scale convolutional neural network).

[0189] Initializing a machine learning model can include defining the number of layers, the type of each layer (e.g., fully connected, convolutional neural network), and the type of activation function for each layer (e.g., sigmoid, ReLU, softmax).

[0190] At 1115, the training subsystem uses the training dataset to train the machine learning model to predict whether an utterance or message represents a task the skillbot is configured to perform or to match the utterance or message to an intent associated with the skillbot. In some examples, the machine learning model is trained to determine a first probability for a solvable class and a second probability for an unsolvable class. The first probability for the solvable class may be generated by a first output channel of a final network layer of the machine learning model. The second probability for the unsolvable class may be generated by a second output channel of a final network layer of the machine learning model. In some examples, the machine learning model determines a probability for each of two or more solvable classes (e.g., order_pizza, cancel_pizza). Find the probability for .

[0191] One or more parameters of the machine learning model may be trained to minimize the loss between the predicted output generated by the machine learning model and the expected output indicated by the training label of the corresponding utterance. Training of the machine learning model may be performed until the loss reaches a minimum error threshold. In particular, the machine learning model may be trained using an expanded training dataset so that it can generate a probability of whether a given utterance corresponds to an unresolvable intent.

[0192] At 1120, the training subsystem replaces the logit functions associated with the unresolvable classes with the enhanced logit values. For the resolvable classes, the logit functions can continue to be used. Thus, the trained machine learning model applies the logit functions to the probabilities of the resolvable classes to generate the corresponding logit values. It can be achieved.

[0193] For the unresolvable class, the logit function is replaced with an enhanced logit value. In some examples, the enhanced logit value is a different real number determined independently from the logit function used for the resolvable class. The enhanced logit value can include one of: (i) a statistical value determined based on a set of logit values ​​generated from a training dataset; (ii) a bounded value selected from a range of values ​​defined by the logarithm of first odds corresponding to the probability for the unresolvable class, where the logarithm of the first odds is constrained to a range of values ​​by a bounding function and weighted by the centroid of a distribution associated with the unresolvable class; (iii) a weighted value generated by the logarithm of second odds corresponding to the probability for the unresolvable class, where the logarithm of the second odds is constrained to a range of values ​​by a bounding function, scaled by a scaling factor, and weighted by the centroid of the distribution associated with the unresolvable class; (iv) a hyperparameter optimization value generated based on hyperparameter tuning of a machine learning model; or (v) a learned value adjusted during training of a machine learning model.

[0194] At 1125, the trained machine learning model with the enriched logit values ​​may be deployed within a chatbot system (e.g., as part of a skillbot) to predict whether an utterance or message represents a task that the skillbot is configured to perform, or to match the utterance or message to an intent associated with the skillbot, or to predict whether the utterance or message corresponds to an unresolvable class. Process 1100 then ends.

[0195] F. Process for classifying out-of-scope utterances using enhanced logit values FIG. 12 is a flowchart illustrating a process 1200 for using the enhanced logit value to classify an utterance as an unresolvable class, according to some embodiments. The process illustrated in FIG. 12 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of a respective system, hardware, or a combination thereof. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The method presented in FIG. 12 and described below is intended to be exemplary and non-limiting. While FIG. 12 depicts various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In some embodiments, the steps may be performed in some different order, or some steps may be performed in parallel. In some embodiments, such as those illustrated in FIGS. 1-4, the process illustrated in FIG. 12 may be performed by a chatbot or skillbot system that can implement a machine learning model and the enhanced logit value to predict whether an utterance corresponds to an unresolvable class. In this particular exemplary process, the chatbot system performs the classification.

[0196] At 1205, the chatbot system receives an utterance generated by a user interacting with the chatbot system. The utterance may refer to a set of words (e.g., one or more sentences) exchanged during a conversation with the chatbot. In some cases, the utterance includes text data converted from a voice input (e.g., a voice utterance), which may be converted into a textual utterance in the particular language, and the text utterance may then be processed.

[0197] At 1210, the chatbot system inputs the utterance into a machine learning model. The machine learning model may include a series of network layers, with a final network The layer includes a logit function that converts a first probability for the resolvable class to a first real number representing a first logit value and a second probability for the unresolvable class to a second real number representing a second logit value.

[0198] The machine learning model may perform operations 1215, 1220, and 1225 to predict whether the utterance corresponds to an unresolvable class. At 1215, the machine learning model determines a first probability for the solvable class and a second probability for the unresolvable class. The first probability for the solvable class may be generated by a first output channel of a final network layer of the machine learning model. The second probability for the unresolvable class may be generated by a second output channel of the final network layer of the machine learning model. In some examples, the machine learning model determines a probability for each of two or more solvable classes (e.g., order_pizza, cancel_pizza).

[0199] At 1220, the machine learning model maps the first probability for the solvable class to a first logit value using a logit function. The logit function may be the logarithm of the odds corresponding to the first probability for the solvable class. The logarithm of the first odds may be weighted by the center of gravity of a distribution associated with the solvable class.

[0200] At 1225, the machine learning model maps the second probability for the unresolvable class to a second logit value, where the second logit value is determined independently from the logit function used for the solvable class. In some examples, the second logit value includes (i) a statistical value determined based on a set of logit values ​​generated from the training dataset, (ii) a bounded value selected from a range of values ​​defined by the logarithm of first odds corresponding to the probability for the unresolvable class, where the logarithm of the first odds is constrained to a range of values ​​by a bounding function and weighted by a centroid of a distribution associated with the unresolvable class, (iii) a weighted value generated by the logarithm of second odds corresponding to the probability for the unresolvable class, where the logarithm of the second odds is constrained to the range of values ​​by a bounding function, scaled by a scaling factor, and weighted by a centroid of a distribution associated with the unresolvable class, (iv) a hyperparameter optimization value generated based on hyperparameter tuning of the machine learning model, or (v) a learned value adjusted during training of the machine learning model. The first logit value and the second logit value are returned to the chatbot system for classification.

[0201] At 1230, the chatbot system classifies the utterance as a solvable class or an unsolvable class based on the first logit value and the second logit value. In some examples, the activation function (e.g., a softmax function) of the machine learning model is configured to: (i) (ii) a second logit value is applied to determine the probability that the utterance corresponds to a class that is resolvable within the multinomial distribution. Based on the determined probabilities, the chatbot system can classify the utterance. Process 1200 then ends.

[0202] G. Exemplary Systems 13 shows a simplified diagram of a distributed system 1300. In the illustrated example, the distributed system 1300 includes one or more client computing devices 1302, 1304, 1306, and 1308 coupled to a server 1312 via one or more communication networks 1310. The client computing devices 1302, 1304, 1306, and 1308 may be configured to run one or more applications.

[0203] In various examples, the server 1312 may implement one or more of the embodiments described in this disclosure. In some examples, server 1312 may also provide other services or software applications, which may include non-virtualized and virtualized environments. In some examples, these services may be provided as web-based services or cloud services, such as under a Software as a Service (SaaS) model, to users of client computing devices 1302, 1304, 1306, and / or 1308. A user operating a client computing device 1302, 1304, 1306, and / or 1308 may utilize one or more client applications to interact with server 1312 to utilize the services provided by these components.

[0204] 13, server 1312 may include one or more components 1318, 1320, and 1322 that implement the functions performed by server 1312. These components may include software components that may be executed by one or more processors, hardware components, or a combination thereof. It should be appreciated that a wide variety of system configurations are possible that may differ from distributed system 1300. Thus, the example shown in FIG. 13 is an example of a distributed system for implementing the example system and is not intended to be limiting.

[0205] Using client computing devices 1302, 1304, 1306, and / or 1308, users run one or more applications, models, or chatbots, which may generate one or more events or models, which may then be implemented or processed according to the teachings of this disclosure. A client device may provide an interface that allows a user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although FIG. 13 shows only four client computing devices, any number of client computing devices may be supported.

[0206] Client devices may include various types of computing systems, such as portable handheld devices, general-purpose computers such as personal computers and laptops, workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-like operating systems, Linux® or Linux-like operating systems), various mobile operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android®, etc.), and the like. Portable handheld devices may include cellular phones, smartphones (e.g., iPhone (registered trademark), tablet (e.g., iPad (registered trademark)), personal digital assistant (PDA) The wearable devices may include Google Glass® head-mounted displays and other devices. The gaming systems may include various handheld gaming devices and internet-connectable gaming devices (e.g., Microsoft Xbox® gaming consoles with or without Kinect® gesture input devices, Sony PlayStation® systems, various gaming systems offered by Nintendo®, etc.). The client devices may be capable of running a wide variety of applications, such as various internet-related applications, communication applications (e.g., email applications, short message service (SMS) applications), etc. Various communication protocols may be used.

[0207] Network 1310 may be any type of network known to those skilled in the art that is capable of supporting data communications using any of a variety of available protocols, including, but not limited to, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internet Packet Exchange), AppleTalk®, etc., by way of example only. As such, the network 1310 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a wireless network operating under any of the Institute of Electrical and Electronics Engineers (IEEE) 1002.11 protocol suite, Bluetooth and / or any other wireless protocol), and / or any combination of these and / or other networks.

[0208] Servers 1312 may be comprised of one or more general-purpose computers, dedicated server computers (including, by way of example, PC (personal computer) servers, UNIX servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or any other suitable configuration and / or combination. Servers 1312 may include one or more virtual machines running a virtual operating system or other computing architecture involving virtualization, such as one or more flexible pools of logical storage that can be virtualized to maintain virtual storage for the servers. In various examples, servers 1312 may be adapted to run one or more services or software applications that provide the functionality described in the above disclosure.

[0209] The computing systems within server 1312 may run one or more operating systems, including any of the operating systems described above, as well as commercially available server operating systems. Server 1312 may also run any of a variety of other server and / or middle-tier applications, including an HTTP (Hypertext Transfer Protocol) server, an FTP (File Transfer Protocol) server, a CGI (Common Gateway Interface) server, a JAVA server, a database server, etc. Exemplary database servers include, but are not limited to, those commercially available from Oracle®, Microsoft®, Sybase®, IBM® (International Business Machines), etc. It will not be done.

[0210] In some implementations, server 1312 may include one or more applications for parsing and consolidating data feeds and / or event updates received from users of client computing devices 1302, 1304, 1306, and 1308. By way of example, the data feeds and / or event updates may include, but are not limited to, Twitter® feeds, Facebook® updates, or real-time updates received from one or more third-party sources and continuous data streams that may include real-time events related to sensor data applications, financial stock tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc. Server 1312 may display the data feeds and / or real-time events on one or more display devices of client computing devices 1302, 1304, 1306, and 1308. It may also include one or more applications for display via the display.

[0211] The distributed system 1300 may also include one or more data repositories 1314, 1316. In certain examples, these data repositories may be used to store data and other information. For example, one or more of the data repositories 1314, 1316 may be used to store information, such as information related to chatbot performance or generated models for use by the chatbot used by the server 1312 when performing various functions according to various embodiments. The data repositories 1314, 1316 may reside in a variety of locations. For example, the data repository used by the server 1312 may be local to the server 1312 or may be remote from the server 1312 and communicate with the server 1312 via a network-based or dedicated connection. The data repositories 1314, 1316 may be of different types. In certain examples, the data repository used by the server 1312 may be a database, for example, a relational database such as those provided by Oracle Corporation® and other manufacturers. One or more of these databases may be adapted to allow data to be stored, updated, and retrieved from the database in response to SQL-formatted commands.

[0212] In particular examples, one or more of the data repositories 1314, 1316 may be used by an application to store application data. The data repositories used by the application may be of various types, such as, for example, a key-value store repository, an object store repository, or a general-purpose storage repository supported by a file system.

[0213] In particular examples, the functionality described in this disclosure may be provided as a service via a cloud environment. FIG. 14 is a simplified block diagram of a cloud-based system environment that may provide various services as cloud services, according to particular examples. In the example shown in FIG. 14, cloud infrastructure system 1402 may provide one or more cloud services that users may request using one or more client computing devices 1404, 1406, and 1408. Cloud infrastructure system 1402 may include one or more computers and / or servers, which may include those described above with respect to server 1312. The computers in cloud infrastructure system 1402 may be organized as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.

[0214] The network 1410 may facilitate communication and exchange of data between the clients 1404, 1406, and 1408 and the cloud infrastructure system 1402. The network 1410 may include one or more networks. The networks may be of the same type or different types. The network 1410 may support one or more communication protocols, including wired and / or wireless protocols, to facilitate communication.

[0215] The example shown in Figure 14 is merely one example of a cloud infrastructure system and is not intended to be limiting. It should be understood that in other examples, cloud infrastructure system 1402 may have more or fewer components than those shown in Figure 14, may combine two or more components, or may have components in a different configuration or arrangement. For example, while Figure 14 shows three client computing devices, in alternative examples, any number of client computing devices may be used. Multiple services may be supported.

[0216] The term cloud service is generally used to refer to services made available to users on demand via a communications network, such as the Internet, by a service provider's system (e.g., cloud infrastructure system 1402). Typically, in a public cloud environment, the servers and systems that comprise the cloud service provider's system are distinct from a customer's own on-premise servers and systems. The cloud service provider's systems are managed by the cloud service provider. Thus, customers can use cloud services offered by the cloud service provider without purchasing separate licenses, support, or hardware and software resources for the services. For example, the cloud service provider's system may host applications, and users can order and use the applications on demand via the Internet without purchasing infrastructure resources to run the applications. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, several cloud services, such as middleware services, database services, and Java cloud services, are offered by Oracle Corporation of Redwood Shores, California.

[0217] In particular examples, cloud infrastructure system 1402 may provide one or more cloud services using various models, such as a Software as a Service (SaaS) model, a Platform as a Service (PaaS) model, an Infrastructure as a Service (IaaS) model, etc., including a hybrid service model. Cloud infrastructure system 1402 may include a suite of applications, middleware, databases, and other resources that enable the provision of various cloud services.

[0218] The SaaS model allows applications or software to be delivered as a service to customers over a communications network such as the Internet, without the customer having to purchase the underlying application hardware or software. For example, the SaaS model may be used to provide customers with access to on-demand applications hosted by cloud infrastructure system 1402. Examples of SaaS services offered by Oracle Corporation® include, but are not limited to, various services for human resource / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, and social applications.

[0219] The IaaS model is commonly used to provide flexible computing and storage capabilities by providing infrastructure resources (e.g., servers, storage, hardware, and networking resources) to customers as cloud services. Various IaaS services are offered by Oracle Corporation (registered trademark).

[0220] The PaaS model is generally used to provide platform and environment resources as a service that enables customers to develop, run, and manage applications and services without having to procure, build, or manage the environment resources. Examples of PaaS services provided by Oracle Corporation (registered trademark) include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), data management cloud services, and various application development solution services. stomach.

[0221] Cloud services are generally provided on an on-demand, self-service basis, on a subscription basis, and in a flexible, scalable, reliable, highly available, and secure manner. For example, a customer may order one or more services offered by cloud infrastructure system 1402 via a subscription order. Cloud infrastructure system 1402 then performs processing to provide the services requested in the customer's subscription order. For example, a user may use utterances to request the cloud infrastructure system to take a particular action (e.g., intent), as described above, and / or to provide a service for a chatbot system, as described herein. Cloud infrastructure system 1402 may be configured to provide one cloud service or even multiple cloud services.

[0222] Cloud infrastructure system 1402 may provide cloud services through a variety of deployment models. In a public cloud model, cloud infrastructure system 1402 may be owned by a third-party cloud service provider, and cloud services are offered to general public customers. These customers may be individuals or businesses. In another example, under a private cloud model, cloud infrastructure system 1402 may function within an organization (e.g., within a corporate organization), and services are offered to customers within the organization. For example, these customers may be various departments within a company, such as human resources, payroll, or individuals within the company. In another example, under a community cloud model, cloud infrastructure system 1402 and the services it offers may be shared among various organizations within an associated community. Various other models, including hybrids of the above models, may also be used.

[0223] Client computing devices 1404, 1406, and 1408 may be of different types (e.g., client computing devices 1302, 1304, 1306, and 1308 shown in FIG. 13) and may be capable of operating one or more client applications. Users may use the client devices to interact with cloud infrastructure system 1402, such as to request services provided by cloud infrastructure system 1402. For example, users may use client devices to request information or actions from a chatbot, as described in this disclosure.

[0224] In some examples, the processing performed by cloud infrastructure system 1402 to provide services may include model training and deployment. This analysis may include using, analyzing, and processing data sets to train and deploy one or more models. This analysis may be performed by one or more processors, possibly processing the data in parallel, running simulations using the data, etc. For example, big data analysis may be performed by cloud infrastructure system 1402 to generate and train one or more models for a chatbot system. The data used in this analysis may include structured data (e.g., data stored in a database or structured according to a structured model) and / or unstructured data (e.g., data blobs (binary large objects)).

[0225] As shown in the example of FIG. 14, the cloud infrastructure system 1402 provides various cloud services. The cloud infrastructure system 1402 may include infrastructure resources 1430 utilized to facilitate the application. The infrastructure resources 1430 may include, for example, processing resources, storage or memory resources, networking resources, etc. In a particular example, a storage virtual machine available to handle storage requested by an application may be part of the cloud infrastructure system 1402. In other examples, the storage virtual machine may be part of a different system.

[0226] In certain examples, to facilitate efficient provisioning of these resources to support various cloud services offered by cloud infrastructure system 1402 to different customers, resources may be organized into resource sets or resource modules (also referred to as "pods"). Each resource module or pod may include a pre-integrated, optimized combination of one or more types of resources. In certain examples, different pods may be pre-provisioned for different types of cloud services. For example, a first set of pods may be provisioned for database services, while a second set of pods may be provisioned for Java services, etc., which may include a different combination of resources than the pods in the first set of pods. For some services, the resources allocated for provisioning these services may be shared between services.

[0227] Cloud infrastructure system 1402 itself may use services 1432 internally that are shared by different components of cloud infrastructure system 1402 and that facilitate the provisioning of services by cloud infrastructure system 1402. These internal shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelist services, high availability, backup and recovery services, services enabling cloud support, email services, notification services, file transfer services, etc.

[0228] Cloud infrastructure system 1402 may include multiple subsystems. These subsystems may be implemented in software, hardware, or a combination thereof. As shown in FIG. 14 , the subsystems may include a user interface subsystem 1412 that allows users or customers of cloud infrastructure system 1402 to interact with cloud infrastructure system 1402. User interface subsystem 1412 may include a variety of different interfaces, such as a web interface 1414, an online store interface 1416 through which cloud services offered by cloud infrastructure system 1402 are advertised and available for purchase by consumers, and other interfaces 1418. For example, a customer may use a client device to request one or more services (service request 1434) offered by cloud infrastructure system 1402 using one or more of interfaces 1414, 1416, and 1418. For example, a customer may access an online store, browse cloud services offered by cloud infrastructure system 1402, and place a subscription order for one or more services offered by cloud infrastructure system 1402 and for which the customer wishes to subscribe. The service request may include information identifying the customer and one or more services for which the customer wishes to subscribe. For example, a customer may submit an order to subscribe to services provided by cloud infrastructure system 1402. As part of the order, the customer may provide information identifying the chatbot system for which the service will be provided, and optionally one or more credentials for the chatbot system.

[0229] 14, cloud infrastructure system 1402 may include an order management subsystem (OMS) 1420 configured to process new orders. As part of this processing, OMS 1420 may be configured to create an account for the customer if not already created, receive billing and / or account information from the customer to use for billing the customer for providing the requested services to the customer, verify the customer information, and, once verified, reserve the order for the customer and prepare the order for provisioning by coordinating various workflows.

[0230] Upon proper validation, the OMS 1420 may invoke an order provisioning subsystem (OPS) 1424 configured to provision resources for the order, including processing, memory, and networking resources. Provisioning may include allocating resources for the order and configuring the resources to facilitate the service requested by the customer order. The manner in which resources are provisioned for the order and the type of resources provisioned may depend on the type of cloud service ordered by the customer. For example, following a workflow, the OPS 1424 may be configured to determine the specific cloud service being requested and identify the number of pods that will be pre-configured for this specific cloud service. The number of pods allocated for an order may depend on the size / amount / level / scope of the service requested. For example, the number of pods to allocate may be determined based on the number of users the service is to support, the duration for which the service is requested, etc. The allocated pods may then be customized to the specific requesting customer to provide the requested service.

[0231] In particular examples, the setup phase processing may be performed as part of the provisioning process, as described above, by cloud infrastructure system 1402. Cloud infrastructure system 1402 may generate an application ID and select a storage virtual machine for the application from among storage virtual machines provided by cloud infrastructure system 1402 itself or from storage virtual machines provided by other systems other than cloud infrastructure system 1402.

[0232] Cloud infrastructure system 1402 may send a response or notification 1444 to the requesting customer to indicate when the requested service will be available for use. In some examples, the customer may be sent information (e.g., a link) that enables the customer to begin using and utilizing the benefits of the requested service. In particular examples, for a customer requesting a service, the response may include a chatbot system ID generated by cloud infrastructure system 1402 and information identifying the chatbot system selected by cloud infrastructure system 1402 for the chatbot system corresponding to the chatbot system ID.

[0233] Cloud infrastructure system 1402 may provide services to multiple customers. For each customer, cloud infrastructure system 1402 manages information related to one or more subscription orders received from the customer, maintains customer data related to the orders, and is responsible for providing the requested services to the customer. Cloud infrastructure system 1402 may also collect usage statistics regarding the customer's use of the subscribed services. For example, statistics may be collected about the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and downtime. This usage information may be used to bill the customer. Billing may be on a monthly basis, for example.

[0234] Cloud infrastructure system 1402 may provide services to multiple customers in parallel. Cloud infrastructure system 1402 may store information about these customers, possibly including copyright information. In particular examples, cloud infrastructure system 1402 includes an identity management subsystem (IMS) 1428 configured to manage customer information and separate the managed information so that information about one customer is not accessed from information about another customer. IMS 1428 may be configured to provide various security-related services, such as identity services such as information access management, authentication and authorization services, services for managing customer identities and roles and associated capabilities, etc.

[0235] FIG. 15 illustrates an example of a computer system 1500. In some examples, the computer system 1500 may be used to implement any of the digital assistant or chatbot systems in a distributed environment, as well as the various servers and computer systems described above. As shown in FIG. 15, the computer system 1500 includes various subsystems, including a processing subsystem 1504 that communicates with several other subsystems via a bus subsystem 1502. These other subsystems may include a processing acceleration unit 1506, an I / O subsystem 1508, a storage subsystem 1518, and a communication subsystem 1524. The storage subsystem 1518 may include non-transitory computer-readable storage media, including a storage medium 1522 and a system memory 1510.

[0236] Bus subsystem 1502 provides a mechanism for allowing the various components and subsystems of computer system 1500 to communicate with each other as intended. While bus subsystem 1502 is shown schematically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1502 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, a local bus, etc., using any of a variety of bus architectures. For example, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a bus conforming to the IEEE P1386.1 standard. The bus may include a Peripheral Component Interconnect (PCI) bus, which may be implemented as a mezzanine bus manufactured by

[0237] The processing subsystem 1504 controls the operation of the computer system 1500 and may include one or more processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The processors may include single-core or multi-core processors. The processing resources of the computer system 1500 may be organized into one or more processing units 1532, 1534, etc. The processing units may include one or more processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some examples, the processing subsystem 1504 may include one or more dedicated coprocessors such as a graphics processor, a digital signal processor (DSP), etc. In some examples, some or all of the processing units of the processing subsystem 1504 may use customized circuitry such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0238] In some examples, processing units within processing subsystem 1504 may execute instructions stored in system memory 1510 or computer-readable storage medium 1522. In various examples, the processing units may execute various program or code instructions and maintain multiple programs or processes running simultaneously. At any given time, some or all of the program code to be executed may reside in system memory 1510 and / or computer-readable storage medium 1522, potentially including one or more storage devices. Through appropriate programming, processing subsystem 1504 may provide the various functions described above. In examples in which computer system 1500 is running one or more virtual machines, one or more processing units may be assigned to each virtual machine.

[0239] In certain examples, a processing acceleration unit 1506 may optionally be provided to accelerate the overall processing performed by the computer system 1500, to perform customized processing, or to offload some of the processing performed by the processing subsystem 1504.

[0240] I / O subsystem 1508 may include devices and mechanisms for inputting information into computer system 1500 and / or outputting information from or through computer system 1500. In general, use of the term "input device" is intended to include all conceivable types of devices and mechanisms for inputting information into computer system 1500. User interface input devices may include, for example, keyboards, pointing devices such as mice or trackballs, touchpads or touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, voice input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may also include motion-sensing and / or gesture-recognition devices, such as a Microsoft Kinect® motion sensor that allows a user to control and interact with the input device, a Microsoft Xbox® 360 game controller, or devices that provide an interface for receiving input using gestures and voice commands. The user interface input devices may also include eye gesture recognition devices, such as a Google Glass® blink detector, that detects eye movements from the user (e.g., "blinks" while taking a picture and / or making a menu selection) and translates the eye gestures as input to the input device (e.g., Google Glass®). The user interface input devices may also include voice recognition sensing devices that allow the user to interact with a voice recognition system (e.g., Siri® navigator) via voice commands.

[0241] Other examples of user interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser range finders, and eye-tracking devices. User interface input devices may also include medical imaging input devices such as, for example, computed tomography, magnetic resonance imaging, position emission tomography, and medical ultrasound devices. User interface input devices may also include audio input devices such as, for example, MIDI keyboards, digital musical instruments, and the like.

[0242] In general, the use of the term output device refers to the information transmitted from the computer system 1500 to the user. The term "user interface output devices" is intended to include all conceivable types of devices and mechanisms for outputting information to a computer or other computer. User interface output devices may include display subsystems, indicator lights, or non-visual displays such as audio output devices. Display subsystems may be flat panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, plotting devices, touch screens, etc. For example, user interface output devices may include, but are not limited to, various display devices that visually convey text, graphics, and audio / visual information, such as monitors, printers, speakers, headphones, automobile navigation systems, plotters, audio output devices, and modems.

[0243] The storage subsystem 1518 provides a repository or data store for storing information and data used by the computer system 1500. The storage subsystem 1518 provides a tangible, non-transitory, computer-readable storage medium for storing the basic programming and data constructs that provide some example functionality. Software (e.g., programs, code modules, instructions) that, when executed by the processing subsystem 1504, provide the functionality described above may be stored in the storage subsystem 1518. The software may be executed by one or more processing units of the processing subsystem 1504. The storage subsystem 1518 may also provide authentication in accordance with the teachings of the present disclosure.

[0244] The storage subsystem 1518 may include one or more non-transitory memory devices, including volatile and non-volatile memory devices. As shown in FIG. 15, the storage subsystem 1518 includes a system memory 1510 and a computer-readable storage medium 1522. The system memory 1510 may include several memories, including volatile primary random access memory (RAM) for storing instructions and data during program execution, and non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some implementations, a basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within the computer system 1500, such as during start-up, may typically be stored in ROM. Typically, RAM contains data and / or program modules currently operated on and executed by the processing subsystem 1504. In some implementations, the system memory 1510 may include multiple different types of memory such as static random access memory (SRAM), dynamic random access memory (DRAM), etc.

[0245] 15, system memory 1510 may load running application programs 1512, program data 1514, and operating system 1516, which may include various applications such as a web browser, a middle-tier application, a relational database management system (RDBMS), etc. By way of example, operating system 1516 may be a Microsoft Windows®, Apple Macintosh®, and / or Linux operating system. rating systems, various commercially available UNIX or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome OS, etc.), and / or iOS ), Windows Phone, various versions of mobile operating systems such as Android® OS, BlackBerry® OS, Palm® OS operating systems, and the like.

[0246] The computer-readable storage medium 1522 may include programming that provides functionality for some examples. and data structures. The computer-readable storage medium 1522 can provide storage of computer-readable instructions, data structures, program modules, and other data for the computer system 1500. Software (programs, code modules, instructions) that, when executed by the processing subsystem 1504, provide the above-described functionality may be stored in the storage subsystem 1518. By way of example, the computer-readable storage medium 1522 may include non-volatile memory such as a hard disk drive, a magnetic disk drive, a CD-ROM, a DVD, an optical disk drive such as a Blu-Ray® disk, or other optical media. The computer-readable storage medium 1522 may include, but is not limited to, a Zip® drive, a flash memory card, a Universal Serial Bus (USB) flash drive, a Secure Digital (SD) card, a DVD disk, a digital video tape, etc. The computer-readable storage medium 1522 may also include solid-state drives (SSDs) based on non-volatile memory such as flash memory-based SSDs, enterprise flash drives, solid-state ROM, etc., SSDs based on volatile memory such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs.

[0247] In particular examples, storage subsystem 1518 may also include a computer-readable storage medium reader 1520 that may be further connected to a computer-readable storage medium 1522. Reader 1520 may be configured to receive and read data from a memory device such as a disk, flash drive, or the like.

[0248] In certain examples, computer system 1500 may support virtualization techniques, including, but not limited to, virtualization of processing and memory resources. For example, computer system 1500 may provide support for running one or more virtual machines. In certain examples, computer system 1500 may execute a program such as a hypervisor that facilitates configuration and management of virtual machines. Each virtual machine may be assigned memory, computing (e.g., processors, cores), I / O, and networking resources. Each virtual machine typically executes independently from other virtual machines. A virtual machine typically executes its own operating system, which may be the same as or different from the operating systems executed by other virtual machines executed by computer system 1500. Thus, potentially multiple operating systems may be executed simultaneously by computer system 1500.

[0249] The communications subsystem 1524 provides an interface to other computer systems and networks. The communications subsystem 1524 serves as an interface for sending and receiving data between other systems and the computer system 1500. For example, the communications subsystem 1524 may enable the computer system 1500 to establish a communications channel to one or more client devices over the Internet to send and receive information from the one or more client devices. For example, if the computer system 1500 is used to implement the bot system 120 shown in FIG. 1, the communications subsystem may be used to communicate with a chatbot system selected for the application.

[0250] The communications subsystem 1524 may support both wired and / or wireless communications protocols. In one example, the communications subsystem 1524 may support wireless voice and / or video (e.g., using cellular telephone technology, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE 802.XX family of standards, or other mobile communications technologies, or any combination thereof). and / or may include a radio frequency (RF) transceiver component for accessing a data network, a global positioning system (GPS) receiver component, and / or other components. In some examples, the communications subsystem 1524 may provide a wired network connection (e.g., Ethernet) in addition to or instead of a wireless interface.

[0251] The communications subsystem 1524 may receive and transmit data in a variety of formats. In some examples, the communications subsystem 1524 may receive incoming communications in the form of structured and / or unstructured data feeds 1526, event streams 1528, event updates 1530, and the like, among other formats. For example, the communications subsystem 1524 may receive incoming communications from social media networks and / or Twitter (registered users). (trademark) Feeds, Facebook® Updates, Rich Site Summary (RSS) Feeds The network may be configured to receive (or transmit) data feeds 1526 in real time from users of other communications services, such as web feeds, such as Yahoo! News, and / or real-time updates from one or more third-party sources.

[0252] In particular examples, the communications subsystem 1524 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1528 of real-time events and / or event updates 1530 that may be continuous or infinite in nature without a clear end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial stock tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc.

[0253] Communications subsystem 1524 may be configured to communicate data from computer system 1500 to other computer systems or networks. This data may be communicated in a variety of different formats, such as structured and / or unstructured data feeds 1526, event streams 1528, event updates 1530, etc., to one or more databases that may communicate with one or more streaming data source computers coupled to computer system 1500.

[0254] Computer system 1500 may be one of a variety of types, including a handheld portable device (e.g., an iPhone® cellular phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head-mounted display), a personal computer, a workstation, a mainframe, a kiosk, a server rack, or other data processing system. Due to the ever-changing nature of computers and networks, the description of computer system 1500 shown in FIG. 15 is intended only as a specific example. Many other configurations are possible, having more or fewer components than the system shown in FIG. 15. It should be recognized, based on the disclosure and teachings herein, that there are other aspects and / or methods for implementing the various examples.

[0255] While specific examples have been described, various variations, modifications, alternative configurations, and equivalents are possible. The examples are not limited to operation in a particular data processing environment, but may freely operate in multiple data processing environments. Furthermore, while the examples have been described using a particular sequence of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. While some flowcharts describe operations as sequential processes, many of these operations may be performed in parallel or simultaneously. Additionally, operations The order of steps may be rearranged. The process may have additional steps not included in the figures. The various features and aspects of the above examples may be used individually or together.

[0256] Additionally, while particular examples have been described using particular combinations of hardware and software, it should be understood that other combinations of hardware and software are possible. Particular examples may be implemented exclusively in hardware, exclusively in software, or using a combination thereof. The various processes described herein may be implemented on the same processor or any combination of different processors.

[0257] Where a device, system, component, or module is described as being configured to perform a particular operation or function, such configuration may be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, by executing, for example, computer instructions or code, or a processor or core programmed to execute code or instructions stored on a non-transitory memory medium, or any combination thereof. Processes may communicate using a variety of techniques, including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different techniques, and the same pair of processes may use different techniques at different times.

[0258] In this disclosure, specific details are provided to ensure a thorough understanding of the examples. However, the examples may be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary detail so as not to obscure the examples. This specification provides illustrative examples only and is not intended to limit the scope, applicability, or configuration of other examples. Rather, the above description of the examples provides one skilled in the art with an enabling description for implementing various examples. Various changes are possible within the function and configuration of elements.

[0259] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will be apparent, however, that additions, subtractions, deletions, and other modifications and alterations may be made thereto without departing from the broader spirit and scope as set forth in the claims. Thus, while specific examples have been described, they are not intended to be limiting. Various modifications and equivalents are within the scope of the appended claims.

[0260] While the foregoing specification describes aspects of the disclosure with reference to specific examples thereof, those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above disclosure may be used individually or together. Moreover, the examples can be utilized in a variety of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings should be regarded as illustrative rather than restrictive.

[0261] In the above description, for purposes of illustration, the methods have been described in a particular order. It should be understood that in alternative examples, the methods may be performed in an order different from that described. It should also be understood that the methods described above may be performed by hardware components or embodied in a sequence of machine-executable instructions that, when used, may cause a machine, such as a general-purpose or special-purpose processor or logic circuitry programmed with such instructions, to perform the method. These machine-executable instructions may be stored in a variety of media, such as a CD-ROM or other type of optical disk, floppy disk, ROM, RAM, EPROM, EEPROM, magnetic or optical card, flash memory, or the like. The methods may be stored on one or more machine-readable media, or other types of machine-readable media suitable for storing electronic instructions. Alternatively, these methods may be performed by a combination of hardware and software.

[0262] Where a component is described as being configured to perform particular operations, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the particular operations, by programming a programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the particular operations, or any combination thereof.

[0263] While illustrative examples of the present application have been described in detail herein, it is to be understood that the concepts of the present invention may be variously embodied and employed in other forms, and that the claims are intended to be construed to include such variations except insofar as limited by the prior art.

Claims

1. 1. A method comprising: The method further comprises: a chatbot system receiving an utterance generated by a user interacting with the chatbot system, the utterance including text data converted from a voice input of the user; The chatbot system includes inputting the utterance to a machine learning model including a series of network layers, a final network layer of the series including a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value, and the method further includes: the machine learning model determining the first probability for the resolvable class and the second probability for the unresolvable class; the machine learning model includes mapping the first probability for the solvable class to the first logit value using the logit function, wherein the logit function for mapping the first probability is the logarithm of the odds corresponding to the first probability for the solvable class, the logarithm of the odds being weighted by a centroid of a distribution associated with the solvable class, and the method further comprises: The machine learning model includes mapping the second probability for the unresolvable class to an enhanced logit value, the enhanced logit value being a third real number determined independently from the logit function used to map the first probability, the enhanced logit value being selected from a range of values ​​defined by (i) a statistical value determined based on a set of logit values ​​generated from a training dataset, and (ii) a logarithm of first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function, and (iii) a weighted value generated by the logarithm of second odds corresponding to the second probability for the unresolvable class, the logarithm of the second odds being constrained to the range of values ​​by the bounding function, scaled by a scaling factor, and weighted by the centroid of the distribution associated with the unresolvable class; (iv) a hyperparameter optimization value generated based on hyperparameter tuning of the machine learning model; or (v) a learned value adjusted during training of the machine learning model, wherein the method further comprises: The method includes the chatbot system classifying the utterance into the resolvable class or the unresolvable class based on the first logit value and the enriched logit value.

2. The method of claim 1 , further comprising the chatbot system responding to the user based on the classification of the utterance as the resolvable class or the unresolvable class.

3. The method of claim 1 or claim 2, wherein the resolved classes are in-domain and in-scope skills or intents, and the unresolved classes are out-of-domain or out-of-scope skills or intents.

4. The enhanced logit value is the statistical value determined based on a set of the logit values ​​generated from the training data set, and determining the statistical value includes: accessing a subset of the training data set, the subset of the training data set comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and determining the statistics further comprises: 、 generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and determining the statistical value further comprises: determining the statistical value, the statistical value representing the set of training logit values, and determining the statistical value further comprises:

10. A method according to any one of the preceding claims, comprising setting the statistical value as the enhanced logit value.

5. The method of claim 4 , wherein the statistic is the median of the set of training logit values.

6. The method of claim 4 , wherein the statistical value is the mean of the set of training logit values.

7. The method of any one of claims 1 to 3, wherein the enhanced logit values ​​are the bounded values, and the logit function is constrained to the range of values ​​by the bounding function.

8. 4. The method of claim 1, wherein the enhanced logit value is the weighted value, and the logit function is constrained to the range of values ​​by the bounding function and scaled by the scaling factor.

9. 9. The method of claim 8, wherein a first value is assigned to the scaling factor of the logit function and a second value is assigned to the scaling factor of the logarithm of the second odds corresponding to the second probability for the unresolvable class, the second value being greater than the first value.

10. The enhanced logit values ​​are the hyperparameter optimization values, and determining the hyperparameter optimization values ​​includes: accessing a subset of the training dataset, the subset of the training dataset comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and determining the hyperparameter optimization values ​​further comprises: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances; and determining the hyperparameter optimization values ​​further comprising: determining the statistical values, the statistical values ​​representing the set of training logit values, and determining the hyperparameter optimization values ​​further comprises: tuning one or more hyperparameters of the machine learning model to generate optimized statistics; and setting the optimized statistical value as the enhanced logit value.

11. 1. A system comprising: one or more data processors; a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations, including: receiving an utterance generated by a user interacting with the chatbot system, the utterance including text data converted from a voice input of the user, the operation further comprising: inputting the utterance to a machine learning model including a series of network layers, a final network layer of the series including a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value, the operations further comprising: the machine learning model determining the first probability for the resolvable class and the second probability for the unresolvable class; the machine learning model includes mapping the first probability for the solvable class to the first logit value using the logit function, wherein the logit function for mapping the first probability is the logarithm of the odds corresponding to the first probability for the solvable class, the logarithm of the odds being weighted by a centroid of a distribution associated with the solvable class, and the operations further include: The machine learning model includes mapping the second probability for the unresolvable class to an enhanced logit value, the enhanced logit value being a third real number determined independently from the logit function used to map the first probability, the enhanced logit value being selected from a range of values ​​defined by (i) a statistical value determined based on a set of logit values ​​generated from a training dataset, and (ii) a logarithm of first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function, and (iii) a bounded value weighted by the centroid of the distribution associated with the unresolvable class, (iv) a hyperparameter optimization value generated based on hyperparameter tuning of the machine learning model, or (v) a learned value adjusted during training of the machine learning model, wherein the operations further comprise: classifying the utterance into the resolvable class or the unresolvable class based on the first logit value and the enhanced logit value.

12. The instructions further cause the one or more data processors to perform operations, the operations including: The system of claim 11 , further comprising: responding to the user based on the classification of the utterance as the resolvable class or the unresolvable class.

13. The system of claim 11 or claim 12, wherein the resolved classes are in-domain and in-scope skills or intents, and the unresolved classes are out-of-domain or out-of-scope skills or intents.

14. The enhanced logit value is the statistical value determined based on a set of the logit values ​​generated from the training data set, and determining the statistical value includes: accessing a subset of the training data set, the subset of the training data set comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and determining the statistics further comprising: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances; and determining the statistical value further To, determining the statistical value, the statistical value representing the set of training logit values, and determining the statistical value further comprises: The system of any one of claims 11 to 13, further comprising setting the statistical value as the enhanced logit value.

15. The system of claim 14 , wherein the statistical value is the median of the set of training logit values.

16. The system of claim 14 , wherein the statistical value is the mean of the set of training logit values.

17. The system of any one of claims 11 to 13, wherein the enhanced logit values ​​are the bounded values, and the logit function is constrained to the range of values ​​by the bounding function.

18. 14. The system of claim 11, wherein the enhanced logit value is the weighted value, and the logit function is constrained to the range of values ​​by the bounding function and scaled by the scaling factor.

19. 20. The system of claim 18, wherein a first value is assigned to the scaling factor of the logit function and a second value is assigned to the scaling factor of the logarithm of the second odds corresponding to the second probability for the unresolvable class, the second value being greater than the first value.

20. The enhanced logit values ​​are the hyperparameter optimization values, and determining the hyperparameter optimization values ​​includes: accessing a subset of the training dataset, the subset of the training dataset comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and determining the hyperparameter optimization values ​​further comprises: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances; and determining the hyperparameter optimization values ​​further comprising: determining the statistical values, the statistical values ​​representing the set of training logit values, and determining the hyperparameter optimization values ​​further comprises: tuning one or more hyperparameters of the machine learning model to generate optimized statistics; and setting the optimized statistical value as the enhanced logit value.

21. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform operations, said operations including: receiving an utterance generated by a user interacting with the chatbot system, the utterance including text data converted from a voice input of the user, the operation further comprising: inputting the utterance to a machine learning model including a series of network layers, wherein a final network layer of the series of network layers provides a first probability for a resolvable class. , a logit function that converts the first probability for the unresolvable class into a first real number representing a first logit value, and converts the second probability for the unresolvable class into a second real number representing a second logit value, and the operations further include: the machine learning model determining the first probability for the resolvable class and the second probability for the unresolvable class; the machine learning model includes mapping the first probability for the solvable class to the first logit value using the logit function, wherein the logit function for mapping the first probability is the logarithm of the odds corresponding to the first probability for the solvable class, the logarithm of the odds being weighted by a centroid of a distribution associated with the solvable class, and the operations further include: The machine learning model includes mapping the second probability for the unresolvable class to an enhanced logit value, the enhanced logit value being a third real number determined independently from the logit function used to map the first probability, the enhanced logit value being selected from a range of values ​​defined by (i) a statistical value determined based on a set of logit values ​​generated from a training dataset, and (ii) a logarithm of first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function, and (iii) a bounded value weighted by the centroid of the distribution associated with the unresolvable class, (iv) a hyperparameter optimization value generated based on hyperparameter tuning of the machine learning model, or (v) a learned value adjusted during training of the machine learning model, wherein the operations further comprise: classifying the utterance into the resolvable class or the unresolvable class based on the first logit value and the enhanced logit value.

22. The instructions further cause the one or more data processors to perform operations, the operations including:

22. The computer program product of claim 21, further comprising: responding to the user based on the classification of the utterance as the resolvable class or the unresolvable class.

23. 23. The computer program product of claim 21 or claim 22, wherein the resolved classes are in-domain and in-scope skills or intents, and the unresolved classes are out-of-domain or out-of-scope skills or intents.

24. The enhanced logit value is the statistical value determined based on a set of the logit values ​​generated from the training data set, and determining the statistical value includes: accessing a subset of the training data set, the subset of the training data set comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and determining the statistics further comprising: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and determining the statistical value further comprises: determining the statistical value, the statistical value representing the set of training logit values, and determining the statistical value further comprises: A computer program product according to any one of claims 21 to 23, comprising setting the statistical value as the enhanced logit value.

25. 25. The computer program product of claim 24, wherein the statistical value is the median of the set of training logit values.

26. 25. The computer program product of claim 24, wherein the statistical value is the mean of the set of training logit values.

27. 24. A computer program product according to any one of claims 21 to 23, wherein the enhanced logit values ​​are the bounded values, and the logit function is constrained to the range of values ​​by the bounding function.

28. 24. The computer program product of claim 21, wherein the enhanced logit value is the weighted value, and the logit function is constrained to the range of values ​​by the bounding function and scaled by the scaling factor.

29. 29. The computer program product of claim 28, wherein a first value is assigned to the scaling factor of the logit function and a second value is assigned to the scaling factor of the logarithm of the second odds corresponding to the second probability for the unresolvable class, the second value being greater than the first value.

30. The enhanced logit values ​​are the hyperparameter optimization values, and determining the hyperparameter optimization values ​​includes: accessing a subset of the training dataset, the subset of the training dataset comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and determining the hyperparameter optimization values ​​further comprises: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances; and determining the hyperparameter optimization values ​​further comprising: determining the statistical values, the statistical values ​​representing the set of training logit values, and determining the hyperparameter optimization values ​​further comprises: tuning one or more hyperparameters of the machine learning model to generate optimized statistics; and setting the optimized statistical value as the enhanced logit value.

31. 1. A method comprising: a training subsystem receiving a training dataset, the training dataset including a plurality of utterances generated by a user interacting with a chatbot system, at least one utterance of the plurality of utterances including text data converted from a voice input of the user; The training subsystem includes accessing a machine learning model including a series of network layers, a final network layer of the series including a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value, and the method further includes: The training subsystem trains the machine learning model with the training dataset, and the machine learning model: determining the first probability for the resolvable class and the second probability for the unresolvable class; using the logit function to map the first probability for the solvable class to the first logit value, wherein the logit function for mapping the first probability is the logarithm of the odds corresponding to the first probability for the solvable class, the logarithm of the odds being weighted by the centroid of a distribution associated with the solvable class, the method further comprising: the training subsystem replacing the logit function with an enhanced logit value such that the second probability for the unresolvable class is mapped to the enhanced logit value; the enhanced logit value is a third real number determined independently from the logit function used to map the first probability; The enhanced logit value comprises: (i) a statistical value determined based on a set of logit values ​​generated from the training dataset; (ii) a bounded value selected from a range of values ​​defined by the logarithm of first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function and weighted by a centroid of a distribution associated with the unresolvable class; (iii) a weighted value generated by the logarithm of second odds corresponding to the second probability for the unresolvable class, the logarithm of the second odds being constrained to a range of values ​​by the bounding function, scaled by a scaling factor, and weighted by the centroid of the distribution associated with the unresolvable class; (iv) a hyperparameter optimization value generated based on hyperparameter tuning of the machine learning model; or (v) a learned value adjusted during training of the machine learning model; The method, wherein the training subsystem includes deploying the trained machine learning model with the enriched logit values.

32. The method further comprises: generating an augmented training data set from the training data set, the augmented training data set including transforming one or more copies of a particular utterance of the plurality of utterances, the particular utterance being associated with a training label that identifies the particular utterance as being associated with the unresolvable class, the method further comprising:

32. The method of claim 31 , comprising training the machine learning model using the augmented training data set.

33. 33. The method of claim 32, wherein transforming the one or more copies of the particular utterance comprises performing one or more of the following: (i) a reverse transformation of the one or more copies of the particular utterance; (ii) a synonym substitution of one or more tokens of the one or more copies of the particular utterance; (iii) a random insertion of a token into the one or more copies of the particular utterance; (iv) a swap between two tokens of the one or more copies of the particular utterance; or (v) a random deletion of one or more tokens of the one or more copies of the particular utterance.

34. The enhanced logit value is the statistical value determined based on a set of the logit values ​​generated from the training dataset, and training the machine learning model further includes: accessing a subset of the training data set, the subset of the training data set comprising a subset of the plurality of utterances, Each utterance of the subset is associated with the unresolvable class, and training the machine learning model further comprises: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and training the machine learning model further comprises: determining the statistical values, the statistical values ​​representing the set of training logit values, and training the machine learning model further comprises: A method according to any one of claims 31 to 33, comprising setting the statistical value as the enhanced logit value.

35. 35. The method of claim 34, wherein the statistic is the median of the set of training logit values.

36. 35. The method of claim 34, wherein the statistical value is the mean of the set of training logit values.

37. 34. A method according to any one of claims 31 to 33, wherein the enhanced logit values ​​are the bounded values ​​and the logit function is constrained to the range of values ​​by the bounding function.

38. 34. The method of any one of claims 31 to 33, wherein the enhanced logit value is the weighted value, and the logit function is constrained to the range of values ​​by the bounding function and scaled by the scaling factor.

39. 39. The method of claim 38, wherein a first value is assigned to the scaling factor of the logit function and a second value is assigned to the scaling factor of the logarithm of the second odds corresponding to the second probability for the unresolvable class, the second value being greater than the first value.

40. 39. The method of claim 38, wherein training the machine learning model further comprises adjusting the scaling factor for the unresolvable class.

41. the enriched logit values ​​are the hyperparameter optimization values, and training the machine learning model further comprises: accessing a subset of the training dataset, the subset of the training dataset comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and training the machine learning model further comprises: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and training the machine learning model further comprises: determining the statistical values, the statistical values ​​representing the set of training logit values, and training the machine learning model further comprises: tuning one or more hyperparameters of the machine learning model to generate optimized statistics; and setting the optimized statistical value as the enhanced logit value.

42. 1. A system comprising: one or more data processors; a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations, including: receiving a training dataset, the training dataset including a plurality of utterances generated by a user interacting with a chatbot system, at least one utterance of the plurality of utterances including text data converted from a voice input of the user, the operations further comprising: accessing a machine learning model including a series of network layers, a final network layer of the series including a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value, the operations further comprising: training the machine learning model using the training dataset, determining the first probability for the resolvable class and the second probability for the unresolvable class; using the logit function to map the first probability for the resolvable class to the first logit value, wherein the logit function for mapping the first probability is the logarithm of the odds corresponding to the first probability for the resolvable class, the logarithm of the odds being weighted by the centroid of a distribution associated with the resolvable class, and the operations further include: replacing the logit function with an enhanced logit value such that the second probability for the unresolvable class is mapped to the enhanced logit value; the enhanced logit value is a third real number determined independently from the logit function used to map the first probability; The enhanced logit value may comprise: (i) a statistical value determined based on a set of logit values ​​generated from the training dataset; (ii) a bounded value selected from a range of values ​​defined by the logarithm of first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function and weighted by a centroid of a distribution associated with the unresolvable class; (iii) a weighted value generated by the logarithm of second odds corresponding to the second probability for the unresolvable class, the logarithm of the second odds being constrained to a range of values ​​by the bounding function, scaled by a scaling factor, and weighted by the centroid of the distribution associated with the unresolvable class; (iv) a hyperparameter optimization value generated based on hyperparameter tuning of the machine learning model; or (v) a learned value adjusted during training of the machine learning model; and the operations further comprise: deploying the trained machine learning model with the enriched logit values.

43. The instructions further cause the one or more data processors to perform operations, the operations including: generating an augmented training data set from the training data set, the augmented training data set including transforming one or more copies of a particular utterance of the plurality of utterances, the particular utterance being associated with a training label that identifies the particular utterance as being associated with the unresolvable class, the operations further comprising:

43. The system of claim 42, further comprising training the machine learning model using the augmented training dataset.

44. 44. The system of claim 43, wherein transforming the one or more copies of the particular utterance comprises performing one or more of the following: (i) a reverse transformation of the one or more copies of the particular utterance; (ii) a synonym substitution of one or more tokens of the one or more copies of the particular utterance; (iii) a random insertion of a token into the one or more copies of the particular utterance; (iv) a swap between two tokens of the one or more copies of the particular utterance; or (v) a random deletion of one or more tokens of the one or more copies of the particular utterance.

45. The enhanced logit value is the statistical value determined based on a set of the logit values ​​generated from the training dataset, and training the machine learning model further includes: accessing a subset of the training dataset, the subset of the training dataset comprising a subset of the plurality of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and training the machine learning model further comprising: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and training the machine learning model further comprises: determining the statistical values, the statistical values ​​representing the set of training logit values, and training the machine learning model further comprises: The system of any one of claims 42 to 44, comprising setting the statistical value as the enhanced logit value.

46. 46. ​​The system of claim 45, wherein the statistical value is the median of the set of training logit values.

47. 46. ​​The system of claim 45, wherein the statistical value is the mean of the set of training logit values.

48. 45. The system of any one of claims 42 to 44, wherein the enhanced logit values ​​are the bounded values, and the logit function is constrained to the range of values ​​by the bounding function.

49. 45. The system of claim 42, wherein the enhanced logit value is the weighted value, and the logit function is constrained to the range of values ​​by the bounding function and scaled by the scaling factor.

50. 50. The system of claim 49, wherein a first value is assigned to the scaling factor of the logit function and a second value is assigned to the scaling factor of the logarithm of the second odds corresponding to the second probability for the unresolvable class, the second value being greater than the first value.

51. 50. The system of claim 49, wherein training the machine learning model further comprises adjusting the scaling factor for the unresolvable class.

52. the enriched logit values ​​are the hyperparameter optimization values, and training the machine learning model further comprises: accessing a subset of the training data set, the subset of the training data set comprising a subset of utterances, the subset of utterances each utterance of the generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and training the machine learning model further comprises: determining the statistical values, the statistical values ​​representing the set of training logit values, and training the machine learning model further comprises: tuning one or more hyperparameters of the machine learning model to generate optimized statistics; and setting the optimized statistical value as the enhanced logit value.

53. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform operations, said operations including: receiving a training dataset, the training dataset including a plurality of utterances generated by a user interacting with a chatbot system, at least one utterance of the plurality of utterances including text data converted from a voice input of the user, the operations further comprising: accessing a machine learning model including a series of network layers, a final network layer of the series including a logit function that converts a first probability for a resolvable class to a first real number representing a first logit value and a second probability for an unresolvable class to a second real number representing a second logit value, the operations further comprising: training the machine learning model using the training dataset, determining the first probability for the resolvable class and the second probability for the unresolvable class; using the logit function to map the first probability for the resolvable class to the first logit value, wherein the logit function for mapping the first probability is the logarithm of the odds corresponding to the first probability for the resolvable class, the logarithm of the odds being weighted by the centroid of a distribution associated with the resolvable class, and the operations further include: replacing the logit function with an enhanced logit value such that the second probability for the unresolvable class is mapped to the enhanced logit value; the enhanced logit value is a third real number determined independently from the logit function used to map the first probability; The enhanced logit value may comprise: (i) a statistical value determined based on a set of logit values ​​generated from the training dataset; (ii) a bounded value selected from a range of values ​​defined by the logarithm of first odds corresponding to the second probability for the unresolvable class, the logarithm of the first odds being constrained to a range of values ​​by a bounding function and weighted by a centroid of a distribution associated with the unresolvable class; (iii) a weighted value generated by the logarithm of second odds corresponding to the second probability for the unresolvable class, the logarithm of the second odds being constrained to a range of values ​​by the bounding function, scaled by a scaling factor, and weighted by the centroid of the distribution associated with the unresolvable class; (iv) a hyperparameter optimization value generated based on hyperparameter tuning of the machine learning model; or (v) a learned value adjusted during training of the machine learning model; and the operations further comprise: deploying the trained machine learning model with the enriched logit values.

54. The instructions further cause the one or more data processors to perform operations, the operations including: generating an augmented training data set from the training data set, the augmented training data set including transforming one or more copies of a particular utterance of the plurality of utterances, the particular utterance being associated with a training label that identifies the particular utterance as being associated with the unresolvable class, the operations further comprising:

54. The computer program product of claim 53, comprising training the machine learning model using the augmented training dataset.

55. 55. The computer program product of claim 54, wherein transforming the one or more copies of the particular utterance comprises performing one or more of the following: (i) a reverse transformation of the one or more copies of the particular utterance; (ii) a synonym substitution of one or more tokens of the one or more copies of the particular utterance; (iii) a random insertion of a token into the one or more copies of the particular utterance; (iv) a swap between two tokens of the one or more copies of the particular utterance; or (v) a random deletion of one or more tokens of the one or more copies of the particular utterance.

56. The enhanced logit value is the statistical value determined based on a set of the logit values ​​generated from the training dataset, and training the machine learning model further includes: accessing a subset of the training dataset, the subset of the training dataset comprising a subset of the plurality of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and training the machine learning model further comprising: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and training the machine learning model further comprises: determining the statistical values, the statistical values ​​representing the set of training logit values, and training the machine learning model further comprises:

56. A computer program product according to any one of claims 53 to 55, comprising setting the statistical value as the enhanced logit value.

57. 57. The computer program product of claim 56, wherein the statistical value is the median of the set of training logit values.

58. 57. The computer program product of claim 56, wherein the statistical value is the mean of the set of training logit values.

59. 56. A computer program product according to any one of claims 53 to 55, wherein the enhanced logit values ​​are the bounded values ​​and the logit function is constrained to the range of values ​​by the bounding function.

60. 56. The computer program product of claim 53, wherein the enhanced logit value is the weighted value, and the logit function is constrained to the range of values ​​by the bounding function and scaled by the scaling factor.

61. 61. The computer program product of claim 60, wherein a first value is assigned to the scaling factor of the logit function and a second value is assigned to the scaling factor of the logarithm of the second odds corresponding to the second probability for the unresolvable class, the second value being greater than the first value.

62. 61. The computer program product of claim 60, wherein training the machine learning model further comprises adjusting the scaling factor for the unresolvable class.

63. the enriched logit values ​​are the hyperparameter optimization values, and training the machine learning model further comprises: accessing a subset of the training dataset, the subset of the training dataset comprising a subset of utterances, each utterance of the subset of utterances being associated with the unresolvable class, and training the machine learning model further comprises: generating a set of training logit values, each training logit value in the set of training logit values ​​being generated by applying the machine learning model to a respective utterance in the subset of utterances, and training the machine learning model further comprises: determining the statistical values, the statistical values ​​representing the set of training logit values, and training the machine learning model further comprises: tuning one or more hyperparameters of the machine learning model to generate optimized statistics; and setting the optimized statistical value as the enhanced logit value.

Citation Information

Patent Citations

  • Data processing method and device, computer storage medium and electronic equipment

    CN111274374A

  • Systems and methods for human inspired simple question answering (HISQA)

    JP2017076403A

  • Chatbot Integrating Derived User Intent

    US20190188590A1

  • Question answering device and computer program

    WO2020004136A1