Methods and systems for overprediction in neural networks
The method addresses the overconfidence issue in deep neural networks by generating and aggregating confidence score distributions across layers, resulting in a more reliable confidence score for chatbot systems, enhancing their performance and trustworthiness.
Patent Information
- Application Number
- JP2023532791
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-16
- Filing Date
- 2021-11-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-17
AI Technical Summary
Deep neural networks used in chatbot systems for classification purposes often suffer from overconfidence issues, where the confidence score generated by the neural network becomes uncorrelated with the actual confidence score, leading to performance problems.
A method is introduced that generates a confidence score distribution for each layer of a machine learning model, determines the prediction assigned to each layer based on this distribution, and assigns an overall confidence score associated with the overall prediction by identifying a layer where the assigned prediction meets a criterion.
This approach helps mitigate the overconfidence problem in deep neural networks by providing a more accurate and reliable confidence score for the model's predictions, thereby improving the performance and trustworthiness of chatbot systems.
Smart Images

Figure 0007692482000002 
Figure 0007692482000003 
Figure 0007692482000004
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application is a non-provisional application of U.S. Provisional Application No. 63 / 119,566, filed on November 30, 2020, and U.S. Non-Provisional Application No. 17 / 455,181, filed on November 16, 2021, and claims the benefit and priority under 35 U.S.C. § 119(e). The entire contents of the above applications are hereby incorporated by reference in their entirety for all purposes.
[0002] Field The present disclosure generally relates to chatbot systems, and more particularly to techniques for addressing overconfidence issues associated with machine learning models, such as neural networks, used in chatbot systems for classification purposes.
Background Art
[0003] Background Many users around the world are on instant messaging or chat platforms to get instant responses. Organizations often engage in live conversations with customers (or end-users) using these instant messaging or chat platforms. However, hiring service staff to engage in live communication with customers or end-users can be very costly for organizations. Chatbots or bots have begun to be developed, especially on the Internet, to simulate conversations with end-users. End-users can communicate with the bot via a messaging app that the end-user already has installed and uses. Generally, intelligent bots driven by artificial intelligence (AI) can communicate more knowledgeably and contextually in a live conversation, and thus enable a more natural conversation between the bot and the end-user for an improved conversation experience. Instead of the end-user learning a fixed set of keywords or commands that the bot knows how to respond to, intelligent bots can understand the end-user's intent based on the user's utterance in natural language and respond accordingly.
[0004] However, building chatbots is difficult because these automated solutions require specific knowledge in certain domains and the application of specific technologies that may only be within the capabilities of specialized developers. As part of building such chatbots, developers may first need to understand the needs of the enterprise and end users. Then, for example, the developers may select the dataset to be used for analysis, prepare the input dataset for analysis (such as data cleansing, extraction of data before analysis, formatting and / or conversion, performing data feature engineering, etc.), identify the appropriate machine learning (ML) techniques or models for performing the analysis, and analyze and make decisions related to improving the techniques or models to improve the results / achievements based on feedback. The task of identifying the appropriate model may involve developing multiple models in parallel in some cases, and after iteratively testing and experimenting with these models, identifying a specific model for use. Further, supervised learning-based solutions typically include a training stage, a subsequent application (i.e., inference) stage, and an iterative loop between the training stage and the application stage. The developer may be responsible for carefully implementing and monitoring these stages to achieve the optimal solution.
[0005] Typically, individual bots are trained as classifiers and configured to utilize a machine learning model, such as a neural network, that predicts or infers the class or category of the input from a set of target classes or categories for a given input. Generally, deep neural networks (i.e., neural network models having a large number of layers, such as four or more layers) are more accurate in terms of output prediction than shallow neural networks (i.e., neural network models having a small number of layers). However, deep neural networks have an overconfidence problem (in the confidence score) where the confidence score generated by the neural network for a certain class may become uncorrelated with the actual confidence score.
[0006] Accordingly, while deep neural networks are desirable to use because of their high accuracy, in order to avoid performance issues of deep neural networks, the overconfidence problem associated with deep neural networks must be addressed. The embodiments described herein address these and other problems, individually and collectively. SUMMARY OF THE INVENTION
[0007] Summary Techniques (e.g., methods, systems, non-transitory computer-readable media storing code or instructions executable by one or more processors) for addressing the overconfidence problem associated with machine learning models (e.g., neural networks) used in chatbot systems for classification purposes are disclosed. Various embodiments are described herein, including methods, systems, non-transitory computer-readable storage media storing programs, code, or instructions executable by one or more processors, and the like.
[0008] One aspect of the present disclosure provides a method, which includes, for each layer of a plurality of layers of a machine learning model, generating a confidence score distribution of a plurality of predictions regarding an input utterance; determining a prediction assigned to each layer of the machine learning model based on the confidence score distribution generated for the layer; generating an overall prediction of the machine learning model based on the determination; repeatedly processing a subset of the plurality of layers of the machine learning model to identify a layer of the machine learning model for which the assigned prediction meets a criterion; and assigning a confidence score associated with the assigned prediction of the layer of the machine learning model as an overall confidence score associated with the overall prediction of the machine learning model.
[0009] According to one embodiment, a computing device is provided, the computing device includes a processor and a memory including instructions, which when executed by the processor, cause the computing device to, at least, generate a confidence score distribution of a plurality of predictions regarding an input utterance for each layer of a plurality of layers of a machine learning model; determine a prediction assigned to each layer of the machine learning model based on the confidence score distribution generated for the layer; generate an overall prediction of the machine learning model; repeatedly process a subset of the plurality of layers of the machine learning model to identify a layer of the machine learning model for which the assigned prediction meets a criterion; and assign a confidence score associated with the assigned prediction of the layer of the machine learning model as an overall confidence score associated with the overall prediction of the machine learning model.
[0010] One aspect of the present disclosure provides a method, the method comprising, for each layer of a plurality of layers of a machine learning model, generating a confidence score distribution of a plurality of predictions regarding an input utterance; for each prediction of the plurality of predictions, calculating a score based on the confidence score distribution for the plurality of layers of the machine learning model; determining one of the plurality of predictions to correspond to an overall prediction of the machine learning model; and assigning the score associated with the one of the plurality of predictions as an overall confidence score associated with the overall prediction of the machine learning model.
[0011] The above matters, together with other features and embodiments, will become more apparent with reference to the following specification, claims and accompanying drawings.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed description In the following description, various embodiments are described. For the sake of thorough understanding of these embodiments, specific configurations and details are described for the purpose of explanation. However, it will be apparent to those skilled in the art that these embodiments can be implemented without these specific details. Furthermore, well-known features may be omitted or simplified so as not to obscure the described embodiments.
[0014] Although specific embodiments have been described, various modifications, changes, alternative configurations, and equivalents are possible. Embodiments are not limited to operations within a specific data processing environment and can operate freely within multiple data processing environments. Furthermore, although specific embodiments have been described using a specific series of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. Some of the flowcharts describe operations as sequential processes, but many of these operations may be executed in parallel or simultaneously. Additionally, the order of operations may be re-specified. The process may have additional steps not included in the figures. The various features and aspects of the above embodiments may be used individually or together. The term "exemplary" as used herein is meant to mean "served as an example, instance, or illustration." None of the embodiments or designs described herein as "exemplary" should necessarily be construed as being more preferred or advantageous than other embodiments or designs.
[0015] Furthermore, while certain embodiments have been described using specific combinations of hardware and software, it should be understood that other combinations of hardware and software are possible. The particular embodiments may be implemented using only hardware, only software, or a combination thereof. The various processes described herein may be implemented on the same processor or on different processors in any combination.
[0016] If a device, system, component, or module is described as being configured to perform a particular operation or function, such a configuration can be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, by programming computer instructions or code to execute code or instructions stored in a non-transitory memory medium, or any combination thereof, or by executing a processor or core, etc. Processes can communicate using a variety of techniques including, but not limited to, conventional techniques for inter-process communication, different pairs of processes may use different techniques, and the same pair of processes may use different techniques at different times.
[0017] Introduction A machine learning model, such as a neural network, trained as a classifier is configured to predict or infer, for a given input, the class or category of the input from a set of target classes or categories. Such a classifier neural network model is generally trained to generate a probability distribution over the set of target classes, where the probabilities are generated by the neural network for each target class in the set and the sum of the generated probabilities is 1 (or 100% when expressed as a percentage). In such a neural network classifier, the output layer of the neural network can generate a probability score distribution over the set of classes using the softmax function as its activation function. These probabilities are also referred to as confidence scores. The class with the highest associated confidence score can be output as the answer to the input.
[0018] For example, in a chatbot domain, the chatbot can take an utterance as input and use a neural network model trained to predict a distribution of probabilities or confidence scores over a set of classes for which the neural network is trained. The set of classes can include, for example, intent classes that represent the intent underlying the utterance. The neural network is configured to generate a confidence score for each of the intent classes, and the intent class with the highest confidence score can be selected as the most relevant intent class for the input utterance. Also, in some embodiments, the highest confidence score must exceed a preset threshold (e.g., 70% confidence) in order to be selected as the relevant intent class for the input utterance.
[0019] The set of intent classes can include one or more in-domain classes and out-of-domain (OOD) classes. In-domain classes are classes that represent intents that a particular chatbot can handle. OOD classes are generally classes that represent unresolved intents (i.e., intents that cannot be resolved to one of the in-domain classes) that the chatbot is not configured to handle.
[0020] For example, consider a chatbot for ordering pizza ("Pizza Bot"). A user can interact with the Pizza Bot to order and / or cancel a pizza order. The Pizza Bot can be trained to obtain an input utterance and classify the utterance into a set of intent classes that includes one or more in-domain classes and an OOD class. As an example, the in-domain classes can include an "order pizza" intent class and a "cancel pizza order" intent class, and the out-of-domain class can be an "unresolved" class. Thus, if the input utterance is related to ordering a pizza, a properly trained neural network will generate the highest confidence score for the "order pizza" class. Similarly, if the input utterance is related to canceling a pizza order, a properly trained neural network will generate the highest confidence score for the "cancel pizza order" class. If the utterance is unrelated to ordering a pizza or canceling a pizza order, a properly trained neural network may generate the highest confidence score for the OOD "unresolved" class. Further processing performed by the Pizza Bot in response to the input utterance will depend on which class received the highest confidence score for that utterance. Therefore, assigning appropriate confidence scores to the set of classes for a given utterance is important for the performance of the Pizza Bot.
[0021] A neural network generally must be trained before it can be used for inference or prediction. Training can be performed using training data (sometimes referred to as labeled data), where the inputs and the labels (ground truth) associated with those inputs are known. For example, the training data can include an input x(i) and, for each input x(i), a target value or correct answer (also referred to as the ground truth) y(i) for that input. The pair (x(i), y(i)) is called a training sample, and the training data can include many such training samples. For example, the training data used to train a neural network model for a chatbot can include a set of utterances and, for each utterance in the set, the known (ground truth) class for that utterance. As an example, in the above pizza bot, the training data can include a set of utterances with the associated label class of "order pizza", a set of utterances with the associated label class of "cancel pizza order", and a set of utterances with the associated label OOD class of "unsolved".
[0022] The space of all input x(i) within the training data can be represented by X, and the space of all corresponding target y(i) can be represented by Y. The goal of training a neural network is to learn a hypothesis function "h()" that maps the training input space X to the target value space Y such that h(x) is a good predictor of the corresponding value of y. In some implementations, an objective function, such as a loss function that measures the difference between the ground truth value for an input and the value predicted for that input by the neural network, is defined as part of deriving the hypothesis function. This objective function is optimized, i.e., minimized or maximized, as part of the training. Training techniques, such as backpropagation training techniques that repeatedly modify / manipulate the weights associated with the inputs to the perceptrons within the neural network, may be used to minimize or maximize the objective function associated with the output provided by the neural network.
[0023] The depth of a neural network model is measured by the number of layers in the neural network model. A neural network generally has an input layer that receives the input provided to the neural network, an output layer that outputs the result of the input, and one or more hidden layers between the input layer and the output layer. Generally, a deep neural network (i.e., a neural network model with a large number of layers) is more accurate in terms of output prediction than a shallow neural network model (i.e., a neural network model with a small number of layers). However, deep neural networks have an overconfidence problem (with respect to the confidence score) where the confidence score generated by the neural network for a certain class may become uncorrelated with the actual confidence score. A deep neural network model may even generate highly confident misclassification predictions when the actual input is not well represented by the training data used to train the neural network model, i.e., when the actual sample is derived from outside the distribution observed during training, i.e., when the model is not well calibrated. Overconfidence makes it difficult to post-process the model output (such as setting a threshold for the prediction), which means that it needs to be addressed by the architecture (typically used for aleatoric uncertainty) and / or training (such as typically used for epistemic uncertainty).
[0024] Therefore, deep learning models such as deep neural networks are desirable to use because of their high accuracy, but in order to avoid the performance problems of neural networks, the overconfidence problem associated with deep learning models must be addressed.
[0025] Bot system A bot (also referred to as a skill, chatbot, chatterbot, or talkbot) is a computer program that can execute a conversation with an end user. Bots can generally respond to natural language messages (such as questions or comments) through a messaging application that uses natural language messages. A company can communicate with an end user through one or more bot systems via a messaging application. The messaging application, sometimes called a channel, can be the end user's preferred messaging application that the end user has already installed and is familiar with. Thus, the end user does not need to download and install a new application to chat with the bot system. The messaging application can include, for example, an over-the-top (OTT) messaging channel (such as Facebook Messenger, Facebook WhatsApp, WeChat, Line, Kik, Telegram, Talk, Skype, Slack, or SMS), a virtual private assistant (such as Amazon Dot, Echo, or Show, Google® Home, Apple HomePod, etc.), a native or hybrid / responsive mobile app or web app with a chat feature, a mobile and web app extension that extends a mobile or web app, or a voice-based input (such as a device or app with an interface that uses Siri, Cortana, Google Voice, or other voice input for conversation).
[0026] FIG. 1 is a simplified block diagram of an environment 100 incorporating a chatbot system according to a particular embodiment. The environment 100 includes a Digital Assistant Builder Platform (DABP) 102 that enables users of the DABP 102 to create and deploy digital assistants or chatbot systems. The DABP 102 can be used to create one or more digital assistants (or DAs) or chatbot systems. For example, as shown in FIG. 1, a user 104 representing a particular company can use the DABP 102 to create and deploy a digital assistant 106 for the users of the particular company. For example, a bank can use the DABP 102 to create one or more digital assistants for use by the bank's customers. Multiple companies can use the same DABP 102 platform to create digital assistants. As another example, the owner of a restaurant (e.g., a pizza shop) can use the DABP 102 to create and deploy a digital assistant that enables customers of the restaurant to order food (e.g., order pizza).
[0027] FIG. 1 is a simplified block diagram of an environment 100 incorporating a chatbot system according to a particular embodiment. The environment 100 includes a Digital Assistant Builder Platform (DABP) 102 that enables users of the DABP 102 to create and deploy digital assistants or chatbot systems. The DABP 102 can be used to create one or more digital assistants (or DAs) or chatbot systems. For example, as shown in FIG. 1, a user 104 representing a particular enterprise can use the DABP 102 to create and deploy a digital assistant 106 for users of the particular enterprise. For example, a bank can use the DABP 102 to create one or more digital assistants for use by the bank's customers. Multiple enterprises can use the same DABP 102 platform to create digital assistants. As another example, the owner of a restaurant (e.g., a pizza shop) can use the DABP 102 to create and deploy a digital assistant that enables customers of the restaurant to order food (e.g., order pizza).
[0028] Digital assistants, such as the digital assistant 106 built using the DABP 102, can be used to perform various tasks via a natural language-based conversation between the digital assistant and its user 108. As part of the conversation, the user may provide one or more user inputs 110 to the digital assistant 106 and receive a response 112 from the digital assistant 106. The conversation can include one or more of the inputs 110 and responses 112. Through these conversations, the user can request that one or more tasks be performed by the digital assistant, and in response, the digital assistant is configured to execute the user-requested task and respond to the user with an appropriate response.
[0029] User input 110 is generally in natural language form and is referred to as an utterance. The user utterance 110 can be in text form, such as when the user types a sentence, question, text snippet, or even a single word and provides the text as input to the digital assistant 106. In some embodiments, the user utterance 110 can be in voice input or spoken form, such as when the user says or speaks something that is provided as input to the digital assistant 106. The utterance is typically the language spoken by the user 108. For example, the utterance can be in English or some other language. If the utterance is in voice form, the voice input is converted to text form of the utterance in that particular language, and then the text utterance is processed by the digital assistant 106. Various voice-to-text processing techniques can be used to convert the voice or auditory input to a text utterance, and the text utterance is then processed by the digital assistant 106. In some embodiments, the conversion from voice to text can be performed by the digital assistant 106 itself.
[0030] The utterance may be a text utterance or a voice utterance, and may be a fragment, a sentence, multiple sentences, one or more words, one or more questions, a combination of the aforementioned types, etc. The digital assistant 106 is configured to apply natural language understanding (NLU) technology to the utterance to understand the meaning of the user input. As part of the NLU processing for the utterance, the digital assistant 106 is configured to perform processing to understand the meaning of the utterance, which involves identifying one or more intents and one or more entities corresponding to the utterance. Once the meaning of the utterance is understood, the digital assistant 106 can perform one or more actions or operations in response to the understood meaning or intent. For the purposes of the present disclosure, it is assumed that the utterance is either a text utterance directly provided by the user 108 of the digital assistant 106 or the result of a conversion of an input voice utterance into text form. However, this is not intended to be limiting or restrictive in any way.
[0031] For example, the input of user 108 may request that a pizza be ordered by providing an utterance such as "I want to order a pizza". Upon receiving such an utterance, digital assistant 106 is configured to understand the meaning of the utterance and take appropriate action. Appropriate actions may include, for example, asking questions that request user input regarding the type of pizza the user wants to order, the size of the pizza, any toppings for the pizza, etc., and responding to the user. The response provided by digital assistant 106 may also be in natural language form and typically may be in the same language as the input utterance. As part of generating these responses, digital assistant 106 may perform natural language generation (NLG). To order a pizza, through the conversation between the user and digital assistant 106, the digital assistant may guide the user to provide all the necessary information for ordering a pizza, and then, at the end of the conversation, may cause the pizza to be ordered. Digital assistant 106 may end the conversation by outputting to the user information indicating that the pizza has been ordered.
[0032] At the concept level, digital assistant 106 executes various processes in response to an utterance received from a user. In some embodiments, this process includes, for example, understanding the meaning of the input utterance (using NLU), determining the actions to be executed in response to the utterance, causing the actions to be executed when appropriate, generating a response to be output to the user in response to the user utterance, and outputting the response to the user, among other things, and involves a series of processing steps or a pipeline of processing steps. The NLU process can include parsing the received input utterance to understand the structure and meaning of the utterance, and refining and restructuring the utterance to develop a more understandable form (e.g., logical form) or structure for the utterance. Generating a response may include using natural language generation (NLG) techniques. Thus, the natural language processing (NLP) performed by a digital assistant can include a combination of NLU processing and NLG processing. The NLU processing performed by a digital assistant such as digital assistant 106 can include various NLU-related processes such as syntactic analysis (e.g., tokenization, rearrangement, identification of part-of-speech tags for a sentence, identification of named entities in a sentence, generation of a dependency tree representing the sentence structure, splitting of a sentence into clauses, analysis of individual clauses, resolution of anaphora, execution of chunking, etc.). In certain embodiments, the NLU processing or a part thereof is performed by digital assistant 106 itself. In some other embodiments, digital assistant 106 can use other resources to execute a part of the NLU processing. For example, the syntax and structure of the input utterance sentence may be identified by processing the sentence using syntactic analysis, part-of-speech tagging, and / or named entity recognition. In one implementation, in the case of English, syntactic analysis, part-of-speech tagging, and named entity recognition, such as those provided by the Stanford NLP Group, are used to analyze the sentence structure and syntax. These are provided as part of the Stanford CoreNLP toolkit.
[0033] Although the various examples provided in this disclosure show English utterances, this is meant as merely an example. In certain embodiments, digital assistant 106 can also process utterances in languages other than English. Digital assistant 106 may provide subsystems (e.g., components that implement NLU functionality) configured to perform processing for different languages. These subsystems may be implemented as pluggable units that can be invoked using service calls from an NLU core server. This makes the NLU processing flexible and extensible for each language, including allowing for different orders of processing. A language pack may be provided for each individual language, and the language pack can register a list of subsystems that can be serviced from the NLU core server.
[0034] A digital assistant, such as digital assistant 106 shown in FIG. 1, can be made available or accessible to its user 108 via, but not limited to, an application, via a social media platform, via various messaging services and applications (e.g., instant messaging applications), and via various different channels such as other applications or channels. Since a single digital assistant can constitute several channels for it, it can be run simultaneously on different services and accessed simultaneously by different services.
[0035] A digital assistant or chatbot system generally includes one or more skills or is associated with one or more skills. In certain embodiments, these skills are individual chatbots (referred to as skillbots) configured to interact with a user and fulfill certain types of tasks such as inventory tracking, time card submission, expense report creation, food ordering, bank account verification, reservation creation, widget purchase, and the like. For example, in the embodiment shown in FIG. 1, the digital assistant or chatbot system 106 includes skills 116-1, 116-2, and the like. For the purposes of the present disclosure, the term "skill" is used synonymously with the term "skillbot".
[0036] Each skill associated with the digital assistant helps a user of the digital assistant complete a task through a conversation with the user, and the conversation can include a combination of text or auditory input provided by the user and a response provided by the skillbot. These responses may be in the form of a text message or an auditory message to the user and / or in the form of using simple user interface elements (e.g., a selection list) presented to the user to make a selection.
[0037] There are various ways in which a skill or skill bot can be associated with or added to a digital assistant. In one example, a skill bot is developed by an enterprise and can then be added to a digital assistant using DABP102, such as via a user interface provided by DABP102 for registering the skill bot with the digital assistant. In another example, a skill bot is developed and created using DABP102 and can then be added to a digital assistant created using DABP102. In yet another example, DABP102 provides an online digital store (referred to as a "skill store") that offers a plurality of skills directed to a wide range of tasks. Skills provided through the skill store may also expose various cloud services. To add a skill to a digital assistant generated using DABP102, a user of DABP102 can access the skill store via DABP102, select a desired skill, and indicate that the selected skill is to be added to the digital assistant created using DABP102. Skills from the skill store can be added to the digital assistant as is or in a modified form (for example, a user of DABP102 can select and clone a particular skill bot provided by the skill store, customize or modify the selected skill bot, and then add the modified skill bot to a digital assistant created using DABP102).
[0038] To implement a digital assistant or chatbot system, various different architectures may be used. For example, in certain embodiments, a digital assistant created and deployed using DABP102 may be implemented using a master bot / sub (or child) bot paradigm or architecture. According to this paradigm, the digital assistant is implemented as a master bot that interacts with one or more child bots that are skill bots. For example, in the embodiment shown in FIG. 1, the digital assistant 106 includes a master bot 114 and skill bots 116-1, 116-2, etc. that are child bots of the master bot 114. In certain embodiments, the digital assistant 106 itself is considered to operate as a master bot.
[0039] A digital assistant implemented according to the master - slave bot architecture enables a user of the digital assistant to interact with multiple skills via an integrated user interface, i.e., via the master bot. When the user engages with the digital assistant, the user input is received by the master bot. Then, the master bot performs a process to determine the meaning of the user input utterance. Next, the master bot determines whether the task requested by the user in the utterance can be processed by the master bot itself. If not, the master bot selects an appropriate skill bot to process the user request and routes the conversation to the selected skill bot. Thereby, the user can converse with the digital assistant via a common single interface and still be provided with the ability to use several skill bots configured to perform specific tasks. For example, in the case of a digital assistant developed for enterprise use, the master bot of the digital assistant can interface with skill bots having specific functions such as a CRM bot for performing functions related to customer relationship management (CRM), an ERP bot for performing functions related to enterprise resource planning (ERP), and an HCM bot for performing functions related to human capital management (HCM). Thus, the end - user or consumer of the digital assistant only needs to know how to access the digital assistant via the common master bot interface, and behind it, multiple skill bots are provided to process the user requests.
[0040] In certain embodiments, in a master bot / slave bot infrastructure, the master bot is configured to recognize a list of available skill bots. The master bot may have access to various available skill bots, and for each skill bot, metadata identifying the capabilities of each skill bot, including the tasks that each skill bot can perform. When a user request is received in the form of speech, the master bot is configured to identify or predict a particular skill bot from among the plurality of available skill bots that is most likely to respond well to the user request or best process the user request. The master bot then routes that utterance (or a portion thereof) to that particular skill bot for further processing. Thus, control flows from the master bot to the skill bot. The master bot can support multiple input and output channels. In certain embodiments, routing can be performed with the assistance of processing performed by one or more available skill bots. For example, as described below, a skill bot can be trained to infer the intent of an utterance and determine whether the inferred intent matches the intent for which the skill bot is configured. Thus, the routing performed by the master bot can involve the skill bot communicating to the master bot an indication of whether the utterance is configured with an intent suitable for the skill bot to process.
[0041] The embodiment of FIG. 1 shows a digital assistant 106 including a master bot 114 and skill bots 116-1, 116-2, and 116-3, but this is not intended to be limiting. The digital assistant can include various other components (e.g., other systems and subsystems) that provide the functionality of the digital assistant. These systems and subsystems may be implemented in realizations using only software (e.g., code, instructions stored on a computer-readable medium and executable by one or more processors), only hardware, or a combination of software and hardware.
[0042] DABP102 provides an infrastructure, as well as various services and features, that enable a user of DABP102 to create a digital assistant that includes one or more skillbots associated with the digital assistant. In some cases, a skillbot can be created by cloning an existing skillbot, for example, by cloning a skillbot provided by a skill store. As described above, DABP102 can provide a skill store or skill catalog that provides multiple skillbots for performing various tasks. A user of DABP102 can clone a skillbot from the skill store. Optionally, the cloned skillbot may be modified or customized. In some other cases, a user of DABP102 creates a skillbot from scratch using the tools and services provided by DABP102.
[0043] In certain embodiments, at a high level, creating or customizing a skillbot includes the following steps: (1) Set the settings for the new skillbot (2) Set one or more intents for the skillbot (3) Set one or more entities for one or more intents (4) Train the skillbot (5) Create a dialog flow for the skillbot (6) Optionally add custom components to the skillbot (7) Test and deploy the skillbot.
[0044] Each of the above steps is briefly described below. (1) Configure settings for a new skill bot - Various settings may be configured for the skill bot. For example, a skill bot designer can specify one or more call names for the skill bot being created. These call names, which function as identifiers for the skill bot, can then be used by the user of the digital assistant to explicitly call the skill bot. For example, the user can include the call name in their utterance to explicitly call the corresponding skill bot.
[0045] (2) Configure one or more intents and associated exemplary utterances for the skill bot - The skill bot designer specifies one or more intents (also called bot intents) for the skill bot being created. The skill bot is then trained based on these specified intents. These intents represent categories or classes for which the skill bot is trained to make inferences about input utterances. When an utterance is received, the trained skill bot infers the intent of the utterance, and the inferred intent is selected from a predefined set of intents used to train the skill bot. The skill bot then takes an appropriate action in response to the utterance based on the intent inferred for the utterance. In some cases, the intents for the skill bot represent tasks that the skill bot can perform for the user of the digital assistant. Each intent is given an intent identifier or intent name. For example, in the case of a skill bot trained for a bank, the intents specified for that skill bot may include "CheckBalance", "TransferMoney", "DepositCheck", etc.
[0046] For each intent defined for the skill bot, the skill bot designer may also provide one or more exemplary utterances that represent that intent. These exemplary utterances are meant to represent utterances that a user may enter for that intent into the skill bot. For example, for the intent of balance inquiry, exemplary utterances may include "What's my savings account balance?", "How much is in my checking account?", "How much money do I have in my account?", etc. Thus, various permutations of typical user utterances may be specified as utterance examples for the intent.
[0047] Intents and their associated exemplary utterances are used as training data for training the skill bot. Various different training techniques may be used. As a result of this training, a prediction model is generated, which is configured to take an utterance as input and output the intent inferred for the utterance by the prediction model. In some cases, the input utterance is provided to an intent analysis engine (e.g., a rule-based or machine learning-based classifier executed by the skill bot) that is configured to predict or infer the intent for the input utterance using the trained model. The skill bot may then take one or more actions based on the inferred intent.
[0048] (3) Set entities for one or more intents of the skill bot - In some examples, additional context may be required to enable the skill bot to respond appropriately to user utterances. For example, there may be situations where user input utterances resolve to the same intent in the skill bot. For example, in the above example, the utterances "What's my savings account balance?" and "How much is in my checking account?" both resolve to the same balance inquiry intent, but these utterances are different requests with different desired information. To clarify such requests, one or more entities may be added to the intent. Using the example of a banking skill bot, an entity called AccountType that defines values such as "checking" and "saving" may enable the skill bot to analyze the user request and respond appropriately. In the above example, the utterances resolve to the same intent, but the values associated with the AccountType entity are different for the two utterances. This allows the skill bot to perform different actions for the two utterances in some cases, even though they resolve to the same intent. One or more entities may be specified for a particular intent set for the skill bot. Thus, entities are used to add context to the intent itself. Entities help to more fully describe the intent and enable the skill bot to complete the user request.
[0049] In certain embodiments, there are two types of entities, namely, (a) the built-in entities provided by DABP102, and (2) custom entities that can be specified by a skillbot designer. The built-in entities are general-purpose entities that can be used with a wide variety of bots. Examples of built-in entities include, but are not limited to, entities related to time, date, address, number, email address, duration, cycle period, currency, phone number, URL, etc. Custom entities are used for more customized applications. For example, for banking skills, the AccountType entity may be defined by the skillbot designer to enable various banking transactions by checking user input for keywords such as current, regular, and credit card.
[0050] (4) Training the Skill Bot - The skill bot is configured to receive user input in the form of speech, analyze or otherwise process the received input, and identify or select an intent associated with the received user input. As described above, the skill bot must be trained for this purpose. In certain embodiments, the skill bot is trained based on the intents set for the skill bot and exemplary utterances associated with those intents (collectively training data), whereby the skill bot can resolve a user input utterance to one of the set intents of the skill bot. In certain embodiments, the skill bot is trained using training data and uses a prediction model that enables the skill bot to identify what the user is saying (or in some cases, what the user is trying to say). DABP102 provides various different training techniques that can be used by skill bot designers to train the skill bot, including various machine learning-based training techniques, rule-based training techniques, and / or combinations thereof. In certain embodiments, a portion of the training data (e.g., 80%) is used to train the skill bot model and another portion (e.g., the remaining 20%) is used to test or validate the model. Once trained, the trained model (sometimes referred to as the trained skill bot) can then be used to process and respond to user utterances. In some cases, the user's utterance may be a question that requires only a single answer and no further conversation. To handle such situations, a Q&A (question and answer) intent may be defined for the skill bot. The Q&A intent is created in the same way as a normal intent. The dialog flow for the Q&A intent may differ from the dialog flow for a normal intent. For example, unlike a normal intent, the dialog flow for the Q&A intent may not include a prompt to request further information from the user (e.g., a value for a particular entity).
[0051] (5) Create a dialogue flow for the skill bot - The dialogue flow specified for the skill bot describes how the skill bot reacts when different intents for the skill bot are resolved in response to the received user input. The dialogue flow defines the actions or operations that the skill bot takes, such as how the skill bot responds to user utterances, how the skill bot prompts the user for input, and how the skill bot returns data. The dialogue flow is like a flowchart that the skill bot follows. The skill bot designer specifies the dialogue flow using a language such as the Markdown language. In a particular embodiment, a version of YAML called OBotML can be used to specify the dialogue flow for the skill bot. The dialogue flow definition for the skill bot functions as a model of the conversation itself, enabling the skill bot designer to choreograph the conversation between the skill bot and the user corresponding to the skill bot.
[0052] In a particular embodiment, the dialogue flow definition of the skill bot includes three sections: (a) Context section (b) Default transition section (c) State section.
[0053] Context section - The skill bot designer can define the variables used in the conversation flow in the context section. Other variables that can be named in the context section include, but are not limited to, variables for error handling, variables for built-in entities or custom entities, and user variables that enable the skill bot to recognize and persist user preferences.
[0054] Default Transition Section - Transitions for the skill bot can be defined in the dialog flow state section or the default transition section. Transitions defined in the default transition section act as a fallback and are triggered when there are no applicable transitions defined within the state or when the conditions required to trigger a state transition cannot be met. The default transition section can be used to define routing that enables the skill bot to handle unexpected user actions seamlessly.
[0055] State Section - The dialog flow and its related operations are defined as a series of transient states that manage the logic within the dialog flow. Each state node within the dialog flow definition names the components that provide the functionality required at that point in the dialog. In this way, states are constructed around the components. A state includes component-specific characteristics and defines transitions to other states that are triggered after the component has been executed.
[0056] Special case scenarios can be handled using the state section. For example, there may be a case where you want to give the user the option to temporarily step out of the first skill they are engaged in and do something in a second skill within the digital assistant. For instance, if the user is involved in a conversation with the shopping skill (e.g., the user has made some selection for a purchase), the user may want to jump to the banking skill (e.g., to check if they have sufficient funds for that purchase) and then return to the shopping skill to complete their order. To handle this, the state section in the dialog flow definition of the first skill can be configured to start a conversation with a second different skill within the same digital assistant and then return to the original dialog flow.
[0057] (6) Add custom components to the skill bot - As described above, the states specified in the dialog flow for the skill bot designate the components that provide the necessary functions corresponding to those states. The components enable the skill bot to execute functions. In certain embodiments, DABP102 provides a set of preconfigured components for performing a wide range of functions. The skill bot designer can select one or more of these preconfigured components and associate them with states within the dialog flow for the skill bot. The skill bot designer can also create custom or new components using the tools provided by DABP102 and associate the custom components with one or more states within the dialog flow for the skill bot.
[0058] (7) Test and deploy the skill bot - DABP102 provides several features that enable the skill bot designer to test the skill bot under development. The skill bot can then be deployed in and included in a digital assistant.
[0059] The above description explains how to create a skill bot, but using a similar technique, it is also possible to create a digital assistant (or master bot). At the master bot or digital assistant level, system intents for the digital assistant can be set. These system intents for the embedded system are used to identify common tasks that the digital assistant itself (i.e., the master bot) can handle without invoking the skill bots associated with the digital assistant. Examples of system intents defined for the master bot include the following: (1) Exit: Applicable when the user wants to end the current conversation or context in the digital assistant; (2) Help: Applicable when the user requests help or directions; (3) Unresolved intent: Applicable to user inputs that do not match well with the exit intent and the help intent. The digital assistant also stores information about one or more skill bots associated with the digital assistant. This information enables the master bot to select a specific skill bot to process the utterance.
[0060] At the master bot or digital assistant level, when the user inputs a phrase or utterance to the digital assistant, the digital assistant is configured to perform a process of determining how to route the utterance and the related conversation. The digital assistant uses a routing model, which can be rule-based, AI-based, or a combination of them, to make this determination. The digital assistant uses the routing model to determine whether the conversation corresponding to the user input utterance should be routed to a specific skill for processing, processed by the digital assistant or the master bot itself according to the embedded system intent, or processed as a different state in the current conversation flow.
[0061] In certain embodiments, as part of this process, the digital assistant determines whether the user input utterance explicitly identifies a skill bot by its invocation name. If the invocation name is present in the user input, it is treated as an explicit invocation of the skill bot corresponding to the invocation name. In such a scenario, the digital assistant can route the user input to the explicitly invoked skill bot for further processing. In the absence of a specific or explicit invocation, in certain embodiments, the digital assistant evaluates the received user input utterance and calculates a confidence score for system intents and skill bots associated with the digital assistant. The score calculated for a skill bot or system intent represents the likelihood that the user input represents a task configured for the skill bot to perform or represents a system intent. System intents or skill bots with an associated calculated confidence score that exceeds a threshold (e.g., the Confidence Threshold routing parameter) are selected as candidates for further evaluation. The digital assistant then selects a specific system intent or skill bot from the identified candidates for further processing of the user input utterance. In certain embodiments, after one or more skill bots are identified as candidates, the intents associated with those candidate skills are evaluated (using the trained model for each skill) and a confidence score is determined for each intent. Generally, an intent with a confidence score exceeding a threshold (e.g., 70%) is treated as a candidate intent. If a specific skill bot is selected, the user utterance is routed to that skill bot for further processing. If a system intent is selected, one or more actions are performed by the master bot itself according to the selected system intent.
[0062] Techniques for addressing overconfidence issues According to some embodiments, the chatbot takes an utterance as input and uses a neural network model trained to predict a probability or distribution of confidence scores for a set of classes for which the neural network is trained. The set of classes can include, for example, intent classes that represent the intent underlying the utterance. The neural network is configured to generate a confidence score for each of the intent classes, and the intent class with the highest confidence score can be selected as the most relevant intent class for the input utterance. Also, in some embodiments, the highest confidence score must exceed a preset threshold (e.g., 70% confidence) in order to be selected as the relevant intent class for the input utterance.
[0063] FIG. 2 is a diagram showing an exemplary machine learning model according to various embodiments. The machine learning model shown in FIG. 2 is a deep neural network (DNN) 210, and the DNN 210 includes an encoder 220, a plurality of layers 230A, 230B... and 230N, a plurality of prediction modules 240A, 240B... and 240N, and a confidence score processing unit 250. Each layer is associated with a corresponding prediction module. For example, layer 1 230A is associated with prediction module 240A, layer 2 230B is associated with prediction module 240B, and layer 3 230N is associated with prediction module 240N.
[0064] The utterance is an input to an encoder 220 that generates an embedding of the utterance. In some examples, the encoder 220 may be a multi-lingual universal sentence encoder (MUSE) that maps natural language elements (e.g., sentences, words, n-grams (i.e., a collection of n words or characters)) to an array of numbers, i.e., an embedding. Each of the layers 230A, 230B, and 230N of the deep neural network 210 processes the embedding sequentially. Specifically, a prediction module 240A associated with the first layer, i.e., layer 230A, generates a confidence score distribution associated with the first layer based on the embedding, while a prediction module 240B associated with the second layer, i.e., layer 230B, generates a confidence score distribution associated with the second layer based on the embedding processed by the first layer. Thereafter, each layer utilizes its corresponding prediction module to generate a distribution associated with the layer based on the processing performed by the previous layer. Each layer of the DNN 210 is configured to generate a probability (i.e., a confidence score) distribution over a set of classes, such as an intent class that represents the intent underlying the utterance. More specifically, in each layer, the corresponding prediction module generates a confidence score distribution over the set of intent classes. The output of the DNN is an overall prediction / classification and an overall confidence score assigned to this overall prediction.
[0065] It is understood that each layer of the DNN contains one or more neurons (i.e., processing entities). The neurons of a particular layer are connected to the neurons of subsequent layers. Each connection between neurons is associated with a weight, which indicates the importance of the input values for the neurons. Further, each neuron is associated with an activation function, which processes each input to the neuron. It is understood that different activation functions can be assigned to each neuron or layer of the DNN. Thus, each layer processes the input in a unique manner (i.e., based on the activation function and the weights), and the associated prediction module of each layer generates a confidence score distribution for the set of intent classes based on the processing performed by each layer.
[0066] Typically, the neural network model assigns the intent having the highest confidence score in the last layer (i.e., layer N) as the overall prediction of the model. Further, the confidence score associated with such an intent (in the last layer, such as layer 230N which is the output layer of the DNN) is assigned as the overall confidence score of the model. In doing so, the neural network model may encounter the overconfidence problem, i.e., the problem that the confidence scores generated by the neural network may be uncorrelated with the actual confidence scores. To address this overconfidence problem, the deep neural network 210 determines the overall prediction and the overall confidence score associated with the overall prediction in a manner different from the processing performed by a typical neural network, by the confidence score processing unit 250. Specifically, techniques for determining the overall prediction and the overall confidence score associated with the overall prediction of the DNN (referred to herein as iterative techniques and ensemble techniques) are described below.
[0067] The reliability score processing unit 250 obtains the reliability score distribution calculated for each layer by the corresponding prediction module. Specifically, each prediction module is trained to generate a reliability score distribution based on the process executed by the corresponding layer of the DNN 210. For example, layer 1 230A processes the embedding generated by the encoder 220. The prediction module 240A generates a reliability score distribution (associated with various intents) based on the embedding processed by layer 1 230A. Subsequently, layer 2 230B receives the processed embedding from layer 1 230A as an input and performs further processing on this embedding. The prediction module 240B associated with layer 2 230B is trained to generate a reliability score distribution (for various intents) based on the process executed by layer 2 230B.
[0068] According to one embodiment, the reliability score processing unit 250 determines the prediction assigned to each layer of the DNN 210. For each layer of the DNN 210, the reliability score processing unit 250 determines the prediction with the highest reliability score (from the corresponding reliability score distribution generated by the associated prediction module of the layer) as the prediction assigned to the layer. Further, the reliability score processing unit 250 selects the overall prediction of the model to correspond to the assigned prediction of the last layer (i.e., the output layer, layer N 240N) of the DNN 210.
[0069] To assign a comprehensive reliability score to the comprehensive prediction, in an iterative technique approach, the reliability score processing unit 250 performs iterative processing over layers i = 1 to N - 1 to compare the assigned prediction of layer i with the comprehensive prediction (i.e., the assigned prediction of the last layer). When a match is found, the reliability score processing unit 250 stops further processing and assigns the reliability score associated with the assigned prediction of the i-th layer (i.e., the layer whose prediction matches the comprehensive prediction) as the comprehensive reliability associated with the comprehensive prediction of the DNN 210. In other words, the DNN model uses the prediction of the last layer (to handle high accuracy) and the reliability score of the i-th layer (to help mitigate the overconfidence problem).
[0070] It is understood that the term last layer corresponds to the layer of the DNN that processes the input utterance last (e.g., layer 230N in FIG. 2). For example, since the layers shown in FIG. 2 are arranged in a horizontal manner (i.e., from left to right), layer 230N is considered the last layer to process the input utterance. However, note that the DNN may be arranged in different manners, such as a pyramid structure, i.e., a top-down (or bottom-up) manner. In this case as well, the last layer can be the bottom layer (or the top layer) of the pyramid structure and corresponds to the layer that processes the input utterance from the user last.
[0071] Referring to FIG. 3, an exemplary classification executed by the DNN 210 of FIG. 2 according to various embodiments of the present disclosure is shown. For illustrative purposes, consider a pizza bot that includes a set 310 of intent classes, an input utterance 320, and a DNN model having N = 4 layers. Further, for simplicity purposes, the set 310 of intent classes is assumed to include three intents, namely, Intent 1 - "order a pizza", Intent 2 - "cancel a pizza", and Intent 3 - "deliver a pizza". The input utterance 320 is assumed to be "I want a pepperoni pizza". Further, it should be understood that a deep neural network will typically have five or more layers implemented. However, for simplicity purposes, only four layers are used in this example.
[0072] In FIG. 3, table 330 shows the assigned predictions for each of the four layers of the DNN. It is understood that in each of the four layers, the assigned prediction corresponds to the prediction with the highest confidence score. For example, the prediction for layer 1 is "deliver a pizza" (Intent 3) with a confidence score of 70%, the prediction for layer 2 is "cancel a pizza" (Intent 2) with a confidence score of 50%, the prediction for layer 3 is "order a pizza" (Intent 1) with a confidence score of 70%, and the prediction for layer 4 is "order a pizza" (Intent 1) with a confidence score of 90%.
[0073] According to some embodiments, the overall prediction of the model is determined to be the assigned prediction of the last layer of the DNN. For example, referring to FIG. 3, the overall prediction of the DNN model is Intent 1, i.e., the assigned intent of layer 4 that has the highest confidence score in the last layer. To assign an overall confidence score to this overall prediction, the confidence score processing unit of the DNN performs iterative processing over layers i = 1 to N - 1 to compare the prediction of layer i with the overall prediction (i.e., the prediction of the last layer). When a match is found, further processing is aborted, and the confidence score of the i-th layer (i.e., the layer whose prediction matches the overall prediction) is assigned as the overall confidence associated with the overall prediction. For example, referring to FIG. 3, it is determined that the confidence score 340 (i.e., 70%) of layer 3 is the overall confidence score of the DNN model. It is understood that layer 3 is the first layer (within the range from layer 1 to layer 3) whose prediction matches the overall prediction of the model (i.e., the prediction of the last layer, i.e., layer 4). Thus, the DNN model uses the prediction of the last layer (to handle high precision) and the confidence score of the i-th layer (to help mitigate the overconfidence problem).
[0074] Returning to FIG. 2, according to some embodiments, in an ensemble mechanism for determining the overall prediction of the DNN and the overall confidence score associated with this overall prediction of the DNN, the confidence score processing unit 250 of the DNN 210 calculates an ensemble score (i.e., an average score) for each intent class based on the confidence score distribution generated for each layer of the DNN 210. Specifically, the confidence score processing unit 250 calculates probability(intent_i|x) = avg(probability_layer_k(intent_i|x)), where k iterates over the range 1 → N. Details regarding the ensemble calculation will be described next with reference to FIG. 4.
[0075] FIG. 4 is a diagram showing an exemplary classification performed by the DNN model 210 of FIG. 2 according to various embodiments. For illustrative purposes, consider a pizza bot that includes a set 410 of intent classes, an input utterance 420, and a DNN model having N = 4 layers. Further, for simplicity purposes, assume that the set 410 of intent classes includes three intents, namely, intent 1 - "order a pizza", intent 2 - "cancel a pizza", and intent 3 - "deliver a pizza". Assume that the input utterance 420 is "I want a pepperoni pizza". Further, it should be understood that a deep neural network will typically have five or more layers implemented. However, for simplicity purposes, only four layers are used in this example.
[0076] In FIG. 4, table 430 shows the predictions for each of the four layers of the DNN 210. For example, the predicted distributions at each layer are as follows.
[0077] Layer 1 · Intent 1: order a pizza, 20% · Intent 2: cancel a pizza, 10% · Intent 3: deliver a pizza, 70% Layer 2 · Intent 1: order a pizza, 40% · Intent 2: cancel a pizza, 50% · Intent 3: deliver a pizza, 10% Layer 3 · Intent 1: order a pizza, 70% · Intent 2: cancel a pizza, 10% · Intent 3: deliver a pizza, 20% Layer 4 · Intent 1: order a pizza, 90% · Intent 2: cancel a pizza, 5% · Intent 3: deliver a pizza, 5% In an ensemble mechanism that determines a comprehensive prediction of a DNN and a comprehensive confidence score associated with this comprehensive prediction of the DNN, the prediction distribution of each layer is calculated in a manner similar to the iterative approach. Further, similar to the iterative approach, the DNN model 210 determines the comprehensive prediction of the model to correspond to the prediction of the last layer of the DNN having the highest ensemble score. However, as will be described below, in the ensemble approach, the generation of the comprehensive confidence score is different from the iterative approach.
[0078] The confidence score processing unit 250 takes in, as input, the predictions made by the prediction modules of each layer and calculates an ensemble score (e.g., an average score) for each intent class as follows.
[0079] Ensemble score · Intent 1: Order pizza (0.2 + 0.4 + 0.7 + 0.9) / 4 = 55% · Intent 2: Cancel pizza (0.1 + 0.5 + 0.1 + 0.05) / 4 = 18.75% · Intent 3: Deliver pizza (0.7 + 0.1 + 0.2 + 0.05) / 4 = 26.25% Specifically, the confidence score processing unit 250 calculates an ensemble score for each intent so as to correspond to the average confidence score for each intent based on the confidence score distribution for each layer of the DNN model 210. For example, referring to FIG. 4, the comprehensive prediction of the DNN model is determined to be the intent of Intent 1, i.e., the last layer (i.e., layer 4) having the highest confidence score (i.e., 90%). Further, the DNN model assigns the ensemble score (corresponding to the determined comprehensive intent, i.e., Intent 1) as the comprehensive confidence score of the model. That is, in the example shown in FIG. 4, the DNN model assigns a score of 55% as the comprehensive confidence score of the model.
[0080] Figure 5A is a flowchart showing a process executed by a deep neural network (DNN) model according to various embodiments. Specifically, Figure 5A is a flowchart showing an iterative technique for determining an overall prediction and an overall confidence score of a DNN. The processes shown in Figure 5A may be implemented by software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of each system, may be implemented by hardware, or may be implemented by a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The methods shown in Figure 5A and described below are intended to be exemplary and non-limiting. Figure 5A shows that various processing steps are performed in a particular sequence or order, which is not intended to be limiting. In certain alternative embodiments, the steps may be executed in some different order, or some steps may be executed in parallel. In certain embodiments, the processes shown in Figure 5A may be executed by the confidence score processing unit 250 discussed with respect to Figures 2 and 3.
[0081] The process begins at step 510, where a confidence score distribution is generated for each layer of the DNN model with respect to the input utterance. For example, the prediction modules associated with each layer of the DNN model generate a confidence score distribution associated with each layer. At step 515, based on the generated distribution, a prediction is determined for each layer of the DNN model. Specifically, the confidence score processing unit 250 determines the prediction having the highest confidence score (from the corresponding confidence score distribution associated with the layer) as the prediction for the layer, and then assigns the determined prediction to the layer. For example, referring to Figure 4, among the confidence score distributions associated with layer 1, i.e., intent 1 (20%), intent 2 (10%), and intent 3 (70%), intent 3 has the highest confidence score, so the prediction assigned to layer 1 is intent 3.
[0082] Next, the process proceeds to step 520, where the DNN model determines the overall prediction of the model. In some embodiments, the overall prediction is the prediction assigned to the last layer of the DNN model (i.e., layer N). It is understood that the prediction assigned to the last layer corresponds to the prediction having the highest confidence score from the prediction score distribution associated with the last layer.
[0083] In step 525, the value of counter (C) is initialized to 1. Counter C is utilized to perform iterative processing through a subset of the multiple layers of the DNN. For example, in a DNN model including k = N layers, the value of counter C can iterate from layer k = 1 to layer k = N - 1. Thereafter, the process proceeds to step 530, where a query is executed to determine whether the assigned prediction of layer (C) of the DNN model is the same as the overall prediction of the model. If the response to the query is affirmative, the process proceeds to step 540; if the response to the query is negative, the process proceeds to step 535. In step 535, the value of counter (C) is incremented by 1, and the process loops back to step 530 to evaluate the assigned prediction of the next layer.
[0084] Once the layer of the DNN where the assigned prediction is the same as the overall prediction of the model is successfully identified, in step 540, the confidence score associated with the identified layer is assigned as the overall confidence score associated with the overall prediction of the DNN model 210.
[0085] FIG. 5B is a flowchart showing another process executed by a deep neural network (DNN) model according to various embodiments. Specifically, FIG. 5B is a flowchart showing an ensemble technique for determining an overall prediction and an overall confidence score of a DNN. The processes shown in FIG. 5B may be implemented by software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of each system, may be implemented by hardware, or may be implemented by a combination thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The method shown in FIG. 5B and described below is intended to be exemplary and non-limiting. Although FIG. 5B shows that various processing steps are performed in a particular sequence or order, this is not intended to be limiting. In certain alternative embodiments, the steps may be executed in some different order, or some steps may be executed in parallel. In certain embodiments, the processes shown in FIG. 5B may be executed by the confidence score processing unit 250 discussed with respect to FIGS. 2 and 4.
[0086] The process starts at step 555, where a confidence score distribution is generated for each layer of the DNN model with respect to the input utterance. For example, the prediction modules associated with each layer of the DNN model generate a confidence score distribution associated with each layer. At step 560, the process calculates an ensemble score for each prediction based on the generated confidence score distribution. For example, the confidence score processing unit calculates the ensemble score for each prediction (i.e., intent) as the average (or mean value) of the confidence scores corresponding to the predictions associated with different layers of the DNN. For example, referring to Figure 4, intent 1 has a confidence score of 20% in the first layer, 40% in the second layer, 70% in the third layer, and 90% in the fourth layer. Thus, the ensemble score corresponding to intent 1 is (0.2 + 0.4 + 0.7 + 0.9) / 4 = 55%, which is the average of the respective confidence scores of intent 1 in different layers.
[0087] Next, the process proceeds to step 565, where the overall prediction of the model is determined to be the prediction of the last layer of the model having the highest confidence score. For example, referring to FIG. 4, it can be seen that the last layer (layer 4) has a confidence score of 90% associated with intent 1, a confidence score of 5% associated with intent 2, and a confidence score of 5% associated with intent 3. Thus, intent 1 is determined to correspond to the overall prediction of the model. At step 570, the process assigns an overall confidence score to the overall prediction based on the calculated ensemble score. For example, the DNN model assigns the ensemble score corresponding to the overall prediction as the overall confidence score of the model. For example, referring to FIG. 4, the ensemble score for intent 1 (i.e., the intent determined to be the overall intent of the model) is 55%. Thus, a score of 55% (different from the 90% score associated with intent 1 in the fourth layer) is assigned as the overall confidence score associated with the overall prediction of the model.
[0088] According to some embodiments, the performance of the above-described techniques (i.e., the iterative technique and the ensemble technique) for determining the overall prediction and overall confidence score of a DNN was evaluated over 200 data sets. The evaluation was performed under two different scenarios: a) a DNN model with hyperparameter tuning and b) a DNN model without hyperparameter tuning. Hyperparameter tuning is the process of selecting the optimal set of hyperparameters for a DNN model, where hyperparameters are understood to be parameters whose values are used to control the learning process of the DNN model.
[0089] The performance results of these two cases are shown in Table 1 below. The DNN model is observed to improve the average performance (across 200 datasets) when assigning an appropriate confidence score, i.e., the overall confidence score of the model, as compared to a standard technique that simply assigns the confidence score of the last layer of the model as the overall confidence score.
[0090] [Table 1]
[0091] In the above embodiments, the prediction module is associated with each layer of the DNN, but it is understood that the prediction module may be associated with the MUSE layer, i.e., the encoder layer. Further, although the embodiments of the present disclosure are described in the context of a DNN model used in a chatbot setting, it is understood that the techniques for addressing the overconfidence problem described herein are equally applicable to any neural network in different settings.
[0092] Exemplary system FIG. 6 shows a schematic diagram of a distributed system 600. In the illustrated example, the distributed system 600 includes one or more client computing devices 602, 604, 606, and 608 coupled to a server 612 via one or more communication networks 610. The client computing devices 602, 604, 606, and 608 may be configured to execute one or more applications.
[0093] In various examples, server 612 may be adapted to execute one or more services or software applications that enable one or more embodiments described in this disclosure. In certain examples, server 612 may also provide other services or software applications that may include non-virtual and virtual environments. In some examples, these services may be provided as web-based services or cloud services to users of client computing devices 602, 604, 606, and / or 608, such as under a software as a service (SaaS) model. Users operating client computing devices 602, 604, 606, and / or 608 may utilize one or more client applications to interact with server 612 to utilize the services provided by these components.
[0094] In the configuration shown in FIG. 6, server 612 may include one or more components 618, 620, and 622 that implement the functions executed by server 612. These components may include software components that may be executed by one or more processors, hardware components, or combinations thereof. It should be recognized that a wide variety of system configurations may be possible that may differ from distributed system 600. Accordingly, the example shown in FIG. 6 is an example of a distributed system for implementing an example system and is not intended to be limiting.
[0095] The user may execute one or more applications, models, or chatbots using client computing devices 602, 604, 606, and / or 608, which may generate one or more events or models, which may then be implemented or processed in accordance with the teachings of the present disclosure. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although FIG. 6 shows only four client computing devices, any number of client computing devices may be supported.
[0096] The client device can include various types of computing systems, such as portable handheld devices, general-purpose computers such as personal computers and laptops, workstation computers, wearable devices, game systems, thin clients, various messaging devices, sensors or other sensing devices. These computing devices can run various types and versions of software applications and operating systems, including various mobile operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android®, BlackBerry®, Palm OS®), and operating systems such as Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-based operating systems, Linux® or Linux-based operating systems such as Google Chrome® OS. Portable handheld devices can include cellular phones, smartphones (e.g., iPhone®), tablets (e.g., iPad®), personal digital assistants (PDAs), etc. Wearable devices can include Google Glass® head-mounted displays and other devices. Game systems can include various handheld game devices, Internet-connected game devices (e.g., Microsoft Xbox® game consoles with / without Kinect® gesture input devices, Sony PlayStation® systems, various game systems provided by Nintendo®, etc.).The client device may be capable of executing a variety of applications such as various Internet-related applications, communication applications (such as e-mail applications, Short Message Service (SMS) applications), and may use various communication protocols.
[0097] The network 610 may be any type of network well known to those skilled in the art that can support data communication using any of a variety of available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Exchange), AppleTalk®, etc. By way of example only, the network 610 may be a Local Area Network (LAN), an Ethernet®-based network, Token Ring, a Wide-Area Network (WAN), the Internet, a virtual network, a Virtual Private Network (VPN), an intranet, an extranet, a Public Switched Telephone Network (PSTN), an infrared network, a wireless network (such as a network operating under any of the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suites, Bluetooth® and / or any other wireless protocol), and / or any combination of these and / or other networks.
[0098] Server 612 may be composed of one or more general-purpose computers, dedicated server computers (including, for example, PC (Personal Computer) servers, UNIX (registered trademark) servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or other appropriate configurations and / or combinations. Server 612 may include one or more virtual machines that execute a virtual operating system, or other computing architectures with virtualization. This can be, for example, one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server. In various examples, Server 612 may be adapted to execute one or more services or software applications that provide the functions described in the above disclosure.
[0099] The computing system within Server 612 may include one or more operating systems including any of the above operating systems, and may execute commercially available server operating systems. Also, Server 612 may execute any of various other server applications and / or middleware applications including, for example, an HTTP (Hypertext Transfer Protocol) server, an FTP (File Transfer Protocol) server, a CGI (Common Gateway Interface) server, a JAVA (registered trademark) server, a database server, etc. Exemplary database servers include, but are not limited to, those commercially available from Oracle (registered trademark), Microsoft (registered trademark), Sybase (registered trademark), IBM (registered trademark) (International Business Machines), etc.
[0100] In some implementations, server 612 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of client computing devices 602, 604, 606, and 608. As an example, the data feeds and / or event updates may include real-time events related to sensor data applications, financial stock market dashboards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc., and may include, but are not limited to, Twitter (registered trademark) feeds, Facebook (registered trademark) updates, or real-time updates received from one or more third-party information sources and continuous data streams. Server 612 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 602, 604, 606, and 608.
[0101] The distributed system 600 may also include one or more data repositories 614, 616. In certain examples, these data repositories can be used to store data and other information. For example, one or more of the data repositories 614, 616 can be used to store information related to generated models used by the chatbot performance or the chatbot used by the server 612 when executing various functions according to various embodiments. The data repositories 614, 616 can exist in various locations. For example, the data repository used by the server 612 can be at a local location of the server 612 or at a remote location from the server 612 and communicate with the server 612 via a network-based connection or a dedicated connection. The data repositories 614, 616 can be of different types. In certain examples, the data repository used by the server 612 can be a database, such as a relational database provided by Oracle Corporation (registered trademark) and other manufacturers. One or more of these databases can be adapted to enable storage, update, and retrieval of data in response to SQL format commands.
[0102] In certain examples, one or more of the data repositories 614, 616 may be used by an application to store application data. The data repository used by the application can be of various types, such as, for example, a key-value store repository, an object store repository, or a general-purpose storage repository supported by a file system.
[0103] In certain examples, the functions described in this disclosure may be provided as services via a cloud environment. FIG. 7 is a simplified block diagram of a cloud-based system environment that can provide various services as cloud services according to a particular example. In the example shown in FIG. 7, a cloud infrastructure system 702 can provide one or more cloud services that a user can request using one or more client computing devices 704, 706, and 708. The cloud infrastructure system 702 can include one or more computers and / or servers that may include what was previously described with respect to server 612. The computers within the cloud infrastructure system 702 can be organized as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0104] A network 710 can facilitate the communication and exchange of data between clients 704, 706, and 708 and the cloud infrastructure system 702. The network 710 can include one or more networks. The networks can be of the same type or different types. The network 710 can support one or more communication protocols, including wired and / or wireless protocols, to facilitate communication.
[0105] The example shown in FIG. 7 is only one example of a cloud infrastructure system and is not intended to be limiting. It should be understood that in some other examples, the cloud infrastructure system 702 may have more or fewer components than shown in FIG. 7, may combine two or more components, or may have components with different configurations or arrangements. For example, although FIG. 7 shows three client computing devices, in alternative examples, any number of client computing devices can be supported.
[0106] The term cloud service is generally used to refer to services that are made available to users on demand via a communication network such as the Internet by a service provider's system (e.g., cloud infrastructure system 702). Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premises servers and systems. The cloud service provider's system is managed by the cloud service provider. Thus, customers can utilize cloud services provided by the cloud service provider without having to separately purchase licenses, support, or hardware and software resources for the services. For example, the cloud service provider's system can host an application, and users can order and use the application on demand via the Internet without having to purchase infrastructure resources to run the application. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, several cloud services such as middleware services, database services, Java (registered trademark) cloud services, etc. are provided by Oracle Corporation (registered trademark) of Redwood Shores, California.
[0107] In certain examples, the cloud infrastructure system 702 can provide one or more cloud services using various models, including a software as a service (SaaS) model, a platform as a service (PaaS) model, an infrastructure as a service (IaaS) model, and a hybrid service model. The cloud infrastructure system 702 can include a suite of applications, middleware, databases, and other resources that enable the provisioning of various cloud services.
[0108] The SaaS model enables an application or software to be delivered to a customer as a service over a communication network such as the Internet without the customer having to purchase the underlying hardware or software for the application. For example, by using the SaaS model, a customer can be given access to on-demand applications hosted by the cloud infrastructure system 702. Examples of SaaS services provided by Oracle Corporation (registered trademark) include, but are not limited to, various services for human resources / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, social applications, and the like.
[0109] The IaaS model is generally used to provide flexible computing and storage capabilities by providing infrastructure resources (such as servers, storage, hardware, and networking resources) to a customer as a cloud service. Various IaaS services are provided by Oracle Corporation (registered trademark).
[0110] The PaaS model is generally used to provide, as services, a platform and environmental resources that enable customers to develop, run, and manage applications and services without having to procure, build, or manage environmental resources. Examples of PaaS services provided by Oracle Corporation (registered trademark) include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), data management cloud services, and various application development solution services.
[0111] Cloud services are generally provided in an on-demand self-service, subscription-based, flexible, scalable, reliable, highly available, and secure manner. For example, a customer may order one or more services provided by the cloud infrastructure system 702 via a subscription order. The cloud infrastructure system 702 then provides the services requested in the customer's subscription order by performing processing. For example, a user can use speech to cause the cloud infrastructure system to take a specific action (e.g., an intent) as described above and / or request that the cloud infrastructure system provide services for a chatbot system as described herein. The cloud infrastructure system 702 may be configured to provide one cloud service or multiple cloud services.
[0112] The cloud infrastructure system 702 can provide cloud services via various deployment models. In the public cloud model, the cloud infrastructure system 702 may be owned by a third-party cloud service provider, and the cloud services are provided to general public customers. This customer may be an individual or an enterprise. In another example, under the private cloud model, the cloud infrastructure system 702 may function within an organization (e.g., within a corporate organization), and the services are provided to customers within this organization. For example, this customer may be various departments within the enterprise such as the human resources department, the payroll department, or an individual within the enterprise. In another example, under the community cloud model, the cloud infrastructure system 702 and the provided services may be shared among several organizations within the relevant community. Other various models such as a hybrid model of the above models may also be used.
[0113] The client computing devices 704, 706, and 708 may be of different types (e.g., the client computing devices 602, 604, 606, and 608 shown in FIG. 6), and may be capable of operating one or more client applications. The user can interact with the cloud infrastructure system 702, such as by using the client device to request services provided by the cloud infrastructure system 702. For example, the user can use the client device to request information or actions from a chatbot, as described in this disclosure.
[0114] In some examples, the processing that the cloud infrastructure system 702 performs to provide services may include model training and deployment. This analysis may include training and deploying one or more models by using, analyzing, and processing a dataset. This analysis may be performed by one or more processors, optionally, by processing data in parallel, performing simulations using the data, and the like. For example, big data analysis may be performed by the cloud infrastructure system 702 to generate and train one or more models for a chatbot system. The data used in this analysis may include structured data (e.g., data stored in a database or data structured according to a structured model) and / or unstructured data (e.g., a data blob (binary large object)).
[0115] As shown in the example of FIG. 7, the cloud infrastructure system 702 may include infrastructure resources 730 that are utilized to facilitate the provisioning of the various cloud services provided by the cloud infrastructure system 702. The infrastructure resources 730 may include, for example, processing resources, storage or memory resources, networking resources, and the like. In a particular example, a storage virtual machine that is available to process the storage requested from an application may be part of the cloud infrastructure system 702. In other examples, the storage virtual machine may be part of a different system.
[0116] In certain examples, to facilitate the efficient provisioning of these resources for supporting the various cloud services provided by the cloud infrastructure system 702 to different customers, the resources may be grouped into sets of resources or resource modules (also referred to as "pods"). Each resource module or pod may include a pre-integrated and optimized combination of one or more types of resources. In certain examples, different pods may be pre-provisioned for different types of cloud services. For example, a first set of pods may be provisioned for database services, and a second set of pods, which may include a different combination of resources from the pods within the first set of pods, may be provisioned for Java services or the like. For some services, the resources allocated for provisioning these services may be shared among the services.
[0117] The cloud infrastructure system 702 itself may internally use services 732 that are shared by different components of the cloud infrastructure system 702 and that facilitate the provisioning of services by the cloud infrastructure system 702. These internal shared services may include, but are not limited to, security identity services, integration services, enterprise repository services, enterprise manager services, virus scan white list services, high availability, backup recovery services, services enabling cloud support, email services, notification services, file transfer services, and the like.
[0118] The cloud infrastructure system 702 may include a plurality of subsystems. These subsystems may be implemented in software, or hardware, or a combination thereof. As shown in FIG. 7, the subsystems may include a user interface subsystem 712 that enables a user or customer of the cloud infrastructure system 702 to interact with the cloud infrastructure system 702. The user interface subsystem 712 may include various different interfaces such as a web interface 714, an online store interface 716 where cloud services provided by the cloud infrastructure system 702 are advertised and available for purchase by consumers, and other interfaces 718. For example, a customer may use a client device to request (service request 734) one or more services provided by the cloud infrastructure system 702 using one or more of the interfaces 714, 716, and 718. For example, a customer may access an online store, browse cloud services provided by the cloud infrastructure system 702, and place a subscription order for one or more services provided by the cloud infrastructure system 702 and desired by the customer to subscribe to. This service request may include information identifying the customer and one or more services the customer desires to subscribe to. For example, a customer may place an order to subscribe to services provided by the cloud infrastructure system 702. As part of the order, the customer may provide information identifying a chatbot system through which the service is to be provided and, optionally, one or more qualification information for the chatbot system.
[0119] In certain examples, such as the example shown in FIG. 7, the cloud infrastructure system 702 may include an order management subsystem (OMS) 720 configured to process new orders. As part of this process, the OMS 720 may create a customer account if it has not already been created, receive billing and / or account information from the customer for use in billing the customer for the requested services provided to the customer, verify the customer information, and after verification, reserve the order for the customer and configure various workflows to prepare the order for provisioning.
[0120] Once properly authenticated, the OMS 720 may call an order provisioning subsystem (OPS) 724 configured to provision resources for this order, including processing, memory, and networking resources. Provisioning may include allocating resources for the order and configuring the resources to facilitate the services requested by the customer order. The way resources are provisioned for an order and the type of resources provisioned may depend on the type of cloud service the customer ordered. For example, following a certain workflow, the OPS 724 may be configured to determine the specific cloud service requested and identify the number of pods that would be pre-configured for this specific cloud service. The number of pods allocated for an order may depend on the size / volume / level / range of the service requested. For example, the number of pods to allocate may be determined based on the number of users the service should support, the period for which the service is requested, etc. Next, the allocated pods may be customized for the specific customer making the request to provide the requested services.
[0121] In certain examples, the setup phase processing can be performed by the cloud infrastructure system 702 as part of the provisioning process, as described above. The cloud infrastructure system 702 can generate an application ID and select a storage virtual machine for the application from among the storage virtual machines provided by the cloud infrastructure system 702 itself or from storage virtual machines provided by other systems outside the cloud infrastructure system 702.
[0122] The cloud infrastructure system 702 may send a response or notification 744 to the requesting customer to indicate when the requested service will be available. In some examples, the customer may be sent information (such as a link) that enables the customer to begin using and utilizing the benefits of the requested service. In certain examples, for a customer requesting a service, the response may include a chatbot system ID generated by the cloud infrastructure system 702 and information identifying the chatbot system selected by the cloud infrastructure system 702 for the chatbot system corresponding to the chatbot system ID.
[0123] The cloud infrastructure system 702 can provide services to multiple customers. For each customer, the cloud infrastructure system 702 manages information related to one or more subscription orders received from the customer, maintains customer data related to the orders, and serves to provide the requested services to the customer. Also, the cloud infrastructure system 702 may collect usage statistics regarding the use of the subscribed services by the customers. For example, the statistics may be collected regarding the amount of storage used, the amount of data transferred, the number of users, as well as the amount of system uptime and system downtime. This usage information may be used to bill the customers. The billing may be performed, for example, monthly.
[0124] The cloud infrastructure system 702 may provide services to multiple customers in parallel. The cloud infrastructure system 702 may store information about these customers, which may include proprietary information in some cases. In a particular example, the cloud infrastructure system 702 includes an Identity Management Subsystem (IMS) 728 configured to manage customer information and separate the information being managed so that information about one customer is not accessible to another customer. The IMS 728 may be configured to provide various security-related services such as identity services like information access management, authentication and authorization services, services for managing customer identities and roles and related capabilities, and the like.
[0125] FIG. 8 shows an example of a computer system 800. In some examples, the computer system 800 may be any of a digital assistant or chatbot system in a distributed environment, and may be used to implement the various servers and computer systems described above. As shown in FIG. 8, the computer system 800 includes various subsystems including a processing subsystem 804 that communicates with several other subsystems via a bus subsystem 802. These other subsystems may include a processing acceleration unit 806, an I / O subsystem 808, a storage subsystem 818, and a communication subsystem 824. The storage subsystem 818 may include a non-transitory computer-readable storage medium including a storage medium 822 and a system memory 810.
[0126] The bus subsystem 802 provides a mechanism for enabling communication among the various components and subsystems of computer system 800 as intended. Although the bus subsystem 802 is schematically shown as a single bus, alternative examples of bus subsystems may utilize multiple buses. The bus subsystem 802 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a local bus, using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus that can be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard, among others.
[0127] The processing subsystem 804 controls the operation of the computer system 800 and may include one or more processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The processor may include a single-core or multi-core processor. The processing resources of the computer system 800 can be organized into one or more processing units 832, 834, etc. The processing unit may include one or more processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some examples, the processing subsystem 804 may include one or more dedicated coprocessors such as a graphics processor, a digital signal processor (DSP), etc. In some examples, some or all of the processing units of the processing subsystem 804 may be implemented using customized circuits such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs).
[0128] In some examples, the processing units within the processing subsystem 804 may execute instructions stored in the system memory 810 or the computer-readable storage medium 822. In various examples, the processing units may execute various programs or code instructions and may maintain multiple programs or processes to be executed simultaneously. At any given point in time, some or all of the program code to be executed may reside in the system memory 810 and / or potentially in the computer-readable storage medium 822 including one or more storage devices. Through appropriate programming, the processing subsystem 804 may provide the various functions described above. In an example where the computer system 800 is executing one or more virtual machines, one or more processing units may be assigned to each virtual machine.
[0129] In certain examples, a processing acceleration unit 806 can optionally be provided to execute customized processing to accelerate the overall processing executed by the computer system 800, or to offload a portion of the processing executed by the processing subsystem 804.
[0130] The I / O subsystem 808 can include devices and mechanisms for inputting information into the computer system 800 and / or for outputting information from, or via, the computer system 800. Generally, the use of the term "input device" is intended to include all conceivable types of devices and mechanisms for inputting information into the computer system 800. User interface input devices can include, for example, a keyboard, a mouse or other pointing device such as a trackball, a touchpad or touch screen incorporated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices can also include motion sensing and / or gesture recognition devices such as a Microsoft Kinect (registered trademark) motion sensor, a Microsoft Xbox (registered trademark) 360 game controller, a device that provides an interface for receiving input using gestures and voice commands, which enable a user to control and interact with the input device. User interface input devices can also include gesture recognition devices such as a Google Glass (registered trademark) blink detector that detects eye movements (e.g., "blinks" while taking a photo and / or while making a menu selection) from a user and converts the eye gesture into an input to the input device (e.g., Google Glass (registered trademark)). Additionally, user interface input devices can include voice recognition sensing devices that enable a user to interact with a voice recognition system (e.g., a Siri (registered trademark) navigator) via voice commands.
[0131] Other examples of user interface input devices include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, game pads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. Also, the user interface input device may include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasonic examination devices. The user interface input device may also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, etc.
[0132] In general, the use of the term output device is intended to include all possible types of devices and mechanisms for outputting information from computer system 800 to a user or another computer. The user interface output device may include, for example, a display subsystem, indicator lights, or a non-visual display such as an audio output device. The display subsystem may be, for example, a flat panel device using a cathode ray tube (CRT), a liquid crystal display (LCD), or a plasma display, a projection device, a touch screen, etc. For example, the user interface output device may include, but is not limited to, various display devices for visually conveying text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, audio output devices, and modems.
[0133] Storage subsystem 818 provides a repository or data store for storing information and data used by computer system 800. Storage subsystem 818 provides a tangible non-transitory computer-readable storage medium for storing basic programming and data configurations that provide some example functionality. Software (e.g., programs, code modules, instructions) that provides the above-described functionality when executed by processing subsystem 804 may be stored in storage subsystem 818. The software may be executed by one or more processing units of processing subsystem 804. Storage subsystem 818 may also provide authentication in accordance with the teachings of the present disclosure.
[0134] Storage subsystem 818 may include one or more non-transitory memory devices including volatile and non-volatile memory devices. As shown in FIG. 8, storage subsystem 818 includes system memory 810 and computer-readable storage medium 822. System memory 810 may include several memories including volatile main random access memory (RAM) for storing instructions and data during program execution and non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some implementations, a basic input / output system (BIOS) including basic routines that assist in transferring information between elements within computer system 800, such as during startup, may typically be stored in ROM. Typically, RAM includes data and / or program modules that are currently being operated on and executed by processing subsystem 804. In some implementations, system memory 810 may include multiple different types of memories such as static random access memory (SRAM), dynamic random access memory (DRAM), and the like.
[0135] As an example, without limitation, as shown in FIG. 8, the system memory 810 may load an application program 812, program data 814, and an operating system 816 that are in execution and may include various applications such as a web browser, a middle-tier application, a relational database management system (RDBMS), etc. As an example, the operating system 816 may include various versions of Microsoft Windows (registered trademark), Apple Macintosh (registered trademark) and / or Linux operating systems, various commercially available UNIX (registered trademark) or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google Chrome (registered trademark) OS, etc.), and / or mobile operating systems such as iOS (registered trademark), Windows (registered trademark) Phone, Android (registered trademark) OS, BlackBerry (registered trademark) OS, Palm (registered trademark) OS operating systems.
[0136] The computer-readable storage medium 822 can store programming and data configurations that provide the functions of several examples. The computer-readable medium 822 can provide storage of computer-readable instructions, data structures, program modules, and other data for the computer system 800. Software (programs, code modules, instructions) that provides the above functions when executed by the processing subsystem 804 may be stored in the storage subsystem 818. As an example, the computer-readable storage medium 822 may include non-volatile memory such as a hard disk drive, a magnetic disk drive, a CD ROM, a DVD, an optical disk drive such as a Blu-Ray (registered trademark) disk, or other optical media. The computer-readable storage medium 822 may include, but is not limited to, a Zip (registered trademark) drive, a flash memory card, a Universal Serial Bus (USB) flash drive, a Secure Digital (SD) card, a DVD disk, a digital video tape, etc. The computer-readable storage medium 822 may also include solid-state drives (SSDs) based on non-volatile memory such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, SSDs based on volatile memory such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs.
[0137] In certain examples, the storage subsystem 818 may further include a computer-readable storage medium reader 820 that is connectable to the computer-readable storage medium 822. The reader 820 may be configured to receive and read data from memory devices such as disks, flash drives, etc.
[0138] In certain examples, computer system 800 may support virtualization techniques including, but not limited to, virtualization of processing and memory resources. For example, computer system 800 may provide support for running one or more virtual machines. In certain examples, computer system 800 may execute a program such as a hypervisor that facilitates the configuration and management of virtual machines. Memory, computing (e.g., processor, core), I / O, and networking resources may be allocated to each virtual machine. Each virtual machine typically runs independently of other virtual machines. A virtual machine may execute its own operating system, which may be the same as or different from the operating systems executed by other virtual machines typically run by computer system 800. Thus, potentially multiple operating systems may be executed simultaneously by computer system 800.
[0139] Communication subsystem 824 provides an interface to other computer systems and networks. Communication subsystem 824 functions as an interface for sending and receiving data between other systems and computer system 800. For example, communication subsystem 824 may enable computer system 800 to establish a communication channel to one or more client devices via the Internet for sending and receiving information between computer system 800 and the one or more client devices. For example, if computer system 800 is used to implement the bot system 120 shown in FIG. 1, the communication subsystem may be used to communicate with the chatbot system selected for the application.
[0140] Communication subsystem 824 may support both wired and / or wireless communication protocols. In certain examples, communication subsystem 824 may include, for example, radio frequency (RF) transceiver components for accessing wireless voice and / or data networks using cellular phone technology, advanced data network technologies such as 3G, 4G or EDGE (Enhanced Data rates for Global Evolution), WiFi (IEEE 802.XX family of standards, or other mobile communication technologies, or any combination thereof), a global positioning system (GPS) receiver component, and / or other components. In some examples, communication subsystem 824 may provide a wired network connection (such as Ethernet (registered trademark)) in addition to or instead of a wireless interface.
[0141] Communication subsystem 824 may receive and transmit data in various formats. In some examples, communication subsystem 824 may receive input communications in the form of, in addition to other formats, structured data feeds and / or unstructured data feeds 826, event streams 828, event updates 830, etc. For example, communication subsystem 824 may be configured to receive (or transmit) data feeds 826 in real time from users of other communication services such as social media networks and / or Twitter (registered trademark) feeds, Facebook (registered trademark) updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0142] In certain examples, the communication subsystem 824 may be configured to receive data in the form of a continuous data stream, which may include an event stream 828 and / or event updates 830 of real-time events that are inherently continuous or infinite and have no clear end. Examples of applications that generate continuous data include, for example, sensor data applications, financial stock market dashboards, network performance measurement tools (such as network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and the like.
[0143] The communication subsystem 824 may be configured to communicate data from the computer system 800 to other computer systems or networks. This data may be communicated to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 800 in various different forms such as structured and / or unstructured data feeds 826, event streams 828, event updates 830, and the like.
[0144] The computer system 800 can be of any one of various types, including a handheld portable device (such as an iPhone (registered trademark) cellular phone, an iPad (registered trademark) computing tablet, a PDA), a wearable device (such as Google Glass (registered trademark) head-mounted display), a personal computer, a workstation, a mainframe, a kiosk, a server rack, or other data processing systems. Since the nature of computers and networks is constantly changing, the description of the computer system 800 shown in FIG. 8 is only intended as a specific example. Many other configurations are possible that have more or fewer components than the system shown in FIG. 8. It should be recognized that there are other aspects and / or methods for implementing various examples based on the disclosure and teachings herein.
[0145] Although specific examples have been described, various modifications, changes, alternative configurations, and equivalents are possible. The examples are not limited to operating within a specific data processing environment and can operate freely within multiple data processing environments. Further, although specific examples have been described using a specific series of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. Some of the flowcharts describe the operations as sequential processes, but many of these operations may be performed in parallel or simultaneously. Additionally, the order of the operations may be re-specified. The process may have additional steps not included in the figures. The various features and aspects of the above examples may be used individually or together.
[0146] Furthermore, while specific examples have been described using specific combinations of hardware and software, it should be understood that other combinations of hardware and software are possible. The specific examples may be implemented using only hardware, or only software, or a combination thereof. The various processes described herein may be implemented on the same processor or on different processors in any combination.
[0147] When a device, system, component, or module is described as being configured to perform a particular operation or function, such a configuration can be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, by programming a computer instruction or code that executes code or instructions stored in a non-transitory memory medium or any combination thereof, or by executing a processor or core, etc. Processes can communicate using a variety of techniques including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different techniques, and the same pair of processes may use different techniques at different times.
[0148] In the present disclosure, examples are made to be well understood by showing specific details. However, the examples can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary details so as not to obscure the examples. This specification provides only exemplary examples and is not intended to limit the scope, applicability, or configuration of other examples. Rather, the above description of the examples provides those skilled in the art with an explanation that enables the implementation of various examples. Various changes are possible within the scope of the functions and configurations of the elements.
[0149] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it will be apparent that additions, deletions, omissions, as well as other modifications and changes may be made to them without departing from the broader spirit and scope described in the claims. Thus, while specific examples have been described, these are not intended to be limiting. Various variations and equivalents are within the scope of the appended claims.
[0150] In the foregoing specification, specific examples have been referred to in connection with aspects of the present disclosure, but those skilled in the art will recognize that the present disclosure is not limited thereto. The various features and aspects of the foregoing disclosure may be used individually or together. Further, the examples can be utilized in a variety of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
[0151] In the foregoing description, for purposes of illustration, methods were described in a particular order. In alternative examples, it should be understood that the methods may be performed in an order different from that described. Also, the foregoing methods may be performed by hardware components or may be embodied in a sequence of machine-executable instructions that, when used, cause a machine such as a general or special purpose processor or logic circuit programmed with such instructions to perform the methods. These machine-executable instructions can be stored on one or more machine-readable media such as a CD-ROM or other type of optical disk, a floppy disk, ROM, RAM, EPROM, EEPROM, magnetic or optical card, flash memory, or other type of machine-readable media suitable for storing electronic instructions. Alternatively, these methods may be performed by a combination of hardware and software.
[0152] If an element is described as being configured to perform a particular operation, such a configuration may be achieved, for example, by designing an electronic circuit or other hardware to perform the particular operation, programming a programmable electronic circuit (such as a microprocessor or other suitable electronic circuit) to perform the particular operation, or any combination thereof.
[0153] Examples for the description of the present application have been described in detail herein, but it should be understood that the concept of the present invention can be embodied and adopted in various other aspects, and the claims are intended to be construed to include such variations except when limited by the prior art.
Claims
Claim 1 A method comprising: for each layer of a plurality of layers of a machine learning model, generating a reliability score distribution of a plurality of predictions regarding an input utterance; based on the reliability score distribution generated for the layer, determining a prediction assigned to each layer of the machine learning model; based on the determination, generating an overall prediction of the machine learning model; repeatedly processing a subset of the plurality of layers of the machine learning model to identify a layer of the machine learning model for which the assigned prediction meets a criterion; assigning a reliability score associated with the assigned prediction of the identified layer of the machine learning model as an overall reliability score associated with the overall prediction of the machine learning model; and wherein the criterion corresponds to the assigned prediction of the layer being identical to the overall prediction of the machine learning model. Claim 2 The step of determining a prediction assigned to each layer of the machine learning model further comprises: assigning, as the prediction for the layer, one of the plurality of predictions having the highest reliability score among the reliability score distributions generated for the layer. The method according to claim 1. Claim 3 The step of generating an overall prediction of the machine learning model further comprises: assigning, as the overall prediction of the machine learning model, a prediction of the last layer of the machine learning model having the highest reliability score among the reliability score distributions associated with the last layer, wherein the last layer is an output layer of the machine learning model. The method according to claim 1 or 2. Claim 4 The plurality of layers of the machine learning model includes N layers, the subset of the plurality of layers corresponds to the first N - 1 layers of the machine learning model, and the machine learning model is a deep neural network model. The method according to any one of claims 1 to 3. Claim 5 The machine learning model includes an encoder configured to receive the input utterance and generate an embedding, and each layer of the plurality of layers of the machine learning model includes a prediction module configured to generate the reliability score distribution associated with the layer. The method according to any one of claims 1 to 4. Claim 6 The first prediction module associated with the first layer of the machine learning model generates a first reliability score distribution associated with the first layer based on the embedding generated by the encoder, and the second layer of the machine learning model generates a second reliability score distribution associated with the second layer based on the embedding processed by the first layer. The method according to claim 5.
7. A method comprising: for each layer of a plurality of layers of a machine learning model, generating a reliability score distribution of a plurality of predictions regarding an input utterance; for each prediction of the plurality of predictions, calculating a score based on the reliability score distribution for the plurality of layers of the machine learning model; determining one of the plurality of predictions to correspond to a comprehensive prediction of the machine learning model; and assigning the score associated with the one of the plurality of predictions as a comprehensive reliability score associated with the comprehensive prediction of the machine learning model. A method.
8. One of the plurality of predictions corresponding to the comprehensive prediction is a prediction of the last layer of the machine learning model having the highest reliability score among the reliability score distributions associated with the last layer, and the last layer is an output layer of the machine learning model. The method according to claim 7.
9. The score of the prediction is an average of the reliability scores of the prediction regarding the plurality of layers of the machine learning model. The method according to claim 8.
10. The machine learning model is a deep neural network model. The method according to any one of claims 7 to 9.
11. The machine learning model includes an encoder configured to receive the input utterance and generate an embedding, and each layer of the plurality of layers of the machine learning model includes a prediction module configured to generate the reliability score distribution associated with the layer. The method according to any one of claims 7 to 10.
12. The first prediction module associated with the first layer of the machine learning model generates a first confidence score distribution associated with the first layer based on the embedding generated by the encoder, and the second layer of the machine learning model generates a second confidence score distribution associated with the second layer based on the embedding processed by the first layer. The method according to claim 11.
13. A computing device comprising: a processor; a memory containing instructions that, when executed by the processor, cause the computing device to perform the method according to any one of claims 1 to 12.
14. A program that causes a processor to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Convolution neural network classifier system, training method for the same, classifying method, and usage
JP2014049118A
Ultrasonic imaging device, and image processor
JP2020168233A
Routing for chatbots
US20200342850A1