Target function optimization in target-based hyper-parametric tuning
By introducing multi-objective optimization and modification of objective functions in the hyperparameter tuning system, the problem of hyperparameter tuning algorithm ignoring multiple targets and instance-level regression in the prior art is solved, and more stable and efficient model performance is achieved.
Patent Information
- Application Number
- CN202380072106.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-15
- Filing Date
- 2023-07-26
- Publication Date
- 2025-06-03
AI Technical Summary
Existing hyperparameter tuning algorithms ignore multiple important goals, consider only a single domain, and fail to effectively solve the instance-level regression problem.
A target-based hyperparameter tuning system is adopted, through multi-objective optimization, the weights of multiple indicators are considered, and the hyperparameters are tuned in multiple domains of different levels of importance. At the same time, the objective function is modified to exclude unstable instances and include acceptable regression ratio parameters.
The balanced hyperparameter tuning between multiple domains and different indicators is achieved, reducing instance-level regression, and improving the stability and performance of the model.
Smart Images

Figure CN120092248A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Non - Provisional Application No. 18 / 197,224, filed on May 15, 2023, which claims the benefit and priority of U.S. Provisional Application No. 63 / 405,981, filed on Sep. 13, 2022. The entire contents of these U.S. Provisional Applications are incorporated herein by reference for all purposes. Technical Field
[0003] The present disclosure generally relates to machine learning techniques, and more particularly to objective function optimization techniques in target - based hyperparameter tuning. Background Art
[0004] Artificial intelligence has many applications. To illustrate this, many users around the world use instant messaging or chat platforms to obtain instant responses. Organizations often use these instant messaging or chat platforms to conduct real - time sessions with customers (or end - users). However, hiring service personnel to communicate with customers or end - users in real - time can be very expensive for organizations. Chatbots or bots have been developed to simulate conversations with end - users, especially via the Internet. End - users can communicate with the bot through a messaging application that the end - user has installed and uses. Intelligent bots (usually powered by artificial intelligence (AI)) can communicate more intelligently and contextually in real - time sessions, and thus can allow for a more natural conversation between the bot and the end - user to improve the conversation experience. Instead of the end - user having to learn a fixed set of keywords or commands that the bot knows how to respond to, the intelligent bot can be able to understand the end - user's intent based on natural - language user utterances and respond accordingly.
[0005] However, AI-based solutions (such as chatbots) can be difficult to build because many automated solutions require specific knowledge in certain domains and the application of certain techniques that may only be within the capabilities of professional developers. To illustrate this, as part of building such a chatbot, developers can first understand the needs of the enterprise and end users. The developers can then analyze and make decisions regarding, for example: selecting a data set to be used for analysis; preparing the input data set for analysis (e.g., cleaning data, extracting, formatting, and / or transforming data, performing data feature engineering, etc. before analysis); identifying appropriate machine learning (ML) techniques or ML models to perform the analysis; and improving the techniques or models to improve the results / effectiveness based on feedback. The task of identifying an appropriate model can include developing multiple models (possibly in parallel), iteratively testing and experimenting with these models, and then identifying a specific model (or models) for use. Further, supervised learning-based solutions typically involve a training phase, followed by an application (i.e., inference) phase, and an iterative loop between the training phase and the application phase. The developers can be responsible for carefully implementing and monitoring these phases to achieve an optimal solution. For example, to train ML techniques or models, precise training data is needed so that the algorithms can understand and learn certain patterns or features (e.g., for a chatbot - intent extraction and careful syntactic analysis are needed, not just raw language processing), and these ML techniques or models will use these patterns or features to predict the desired results (e.g., inferring intent from utterances). To ensure that the ML techniques or models correctly learn these patterns and features, the developers can be responsible for selecting, enriching, and optimizing the training data and hyperparameter sets for these ML techniques or models. SUMMARY OF THE INVENTION
[0006] The techniques disclosed herein generally relate to machine learning techniques. More specifically and without limitation, the techniques disclosed herein relate to objective function optimization in target-based hyperparameter tuning.
[0007] In various embodiments, a computer-implemented method is provided, the computer-implemented method comprising: initializing a machine learning algorithm with a set of hyperparameter values; accessing a hyperparameter objective function defined at least in part over a plurality of domains of a search space associated with the machine learning algorithm, wherein the search space includes a training data set and an evaluation data set, wherein each domain includes a subdivision of the search space, the subdivision having at least one training data set and at least one evaluation data set, and wherein the hyperparameter objective function includes a domain score for each domain, the domain score being calculated based on the number of instances correctly or incorrectly predicted by the machine learning algorithm within the at least one evaluation data set during a given trial; for each trial of a hyperparameter tuning process: training the machine learning algorithm for each domain using the at least one training data set associated with each domain and the set of hyperparameter values, wherein the training outputs a plurality of machine learning models, the plurality of machine learning models including a machine learning model for each domain; evaluating the machine learning model for each domain using the at least one evaluation data set associated with each domain and the set of hyperparameter values, wherein the evaluation includes generating the domain score for each domain; calculating a current trial objective score using the hyperparameter objective function based on the domain score for each domain and a domain weight associated with each domain; and determining whether the machine learning model has reached convergence based on the current trial objective score; and in response to determining that the machine learning model has reached convergence, providing at least one of the plurality of machine learning models.
[0008] In some embodiments, the hyperparameter objective function is formulated to normalize instance-level improvements and instance-level regressions over the total number of instances to obtain an improvement score and a regression score, and wherein the domain score for each domain is calculated based on the improvement score and the regression score.
[0009] In some embodiments, the improvement score is calculated based on: (i) the number of instances correctly predicted by the machine learning model within the at least one evaluation data set during the given trial, (ii) the number of instances incorrectly predicted by a baseline machine learning model within the at least one evaluation data set, and (iii) the total count of the number of instances within the at least one evaluation data set, and wherein the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model within the at least one evaluation data set during the given trial, (ii) the number of instances correctly predicted by a baseline machine learning model within the at least one evaluation data set, and (iii) the total count of the number of instances within the at least one evaluation data set.
[0010] In some embodiments, the hyperparameter objective function is formulated to exclude unstable instances, which are instances within the at least one evaluation dataset that are determined to produce different prediction results from each other using the same machine learning model.
[0011] In some embodiments, excluding unstable instances includes: (i) subtracting the count of excluded unstable instances from the number of instances correctly predicted by the machine learning model within the at least one evaluation dataset during the given trial, and (ii) subtracting the count of the excluded unstable instances from the number of instances incorrectly predicted by the machine learning model within the at least one evaluation dataset during the given trial.
[0012] In some embodiments, the hyperparameter objective function is formulated to include a parameter for an acceptable regression ratio m for one or more of the multiple domains.
[0013] In some embodiments, the parameter is defined based on a regression score, and if the regression score is less than the acceptable regression ratio m, the regression score is set to zero; otherwise, the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model within the at least one evaluation dataset during the given trial, (ii) the number of instances correctly predicted by a baseline machine learning model within the at least one evaluation dataset, and (iii) the total count of the number of instances within the at least one evaluation dataset.
[0014] In various embodiments, a system is provided that includes one or more processors and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the system to perform some or all of one or more methods disclosed herein.
[0015] In various embodiments, one or more non-transitory computer-readable media are provided for storing instructions that, when executed by one or more processors, cause a system to perform some or all of one or more methods disclosed herein.
[0016] The techniques described above and below can be implemented in many ways and in many contexts. As described in more detail below, several example implementations and contexts are provided with reference to the following drawings. However, the following implementations and contexts are only some of many implementations and contexts. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A simplified block diagram of a distributed environment incorporating an exemplary embodiment is depicted.
[0018] Figure 2 Depicts a simplified block diagram of a computing system implementing a master robot according to certain embodiments.
[0019] Figure 3 Depicts a simplified block diagram of a computing system implementing a skill robot according to certain embodiments.
[0020] Figure 4A Depicts exemplary types of hyperparameters according to various embodiments.
[0021] Figure 4B Depicts exemplary types of metrics according to various embodiments.
[0022] Figure 4C Depicts an exemplary set of specifications associated with metrics according to various embodiments.
[0023] Figure 5 Illustrates a hyperparameter tuning system according to various embodiments.
[0024] Figure 6 Depicts a flowchart illustrating a training process performed by a hyperparameter tuning system according to various embodiments.
[0025] Figure 7 Depicts a flowchart illustrating a validation process performed by a hyperparameter tuning system according to various embodiments.
[0026] Figure 8 Depicts a simplified diagram of a tuning workflow according to various embodiments.
[0027] Figure 9 Depicts a flowchart illustrating objective function optimization and tuning performed by a hyperparameter tuning system according to various embodiments.
[0028] Figure 10 Depicts a simplified diagram of a distributed system for implementing various embodiments.
[0029] Figure 11 Depicts a simplified block diagram of one or more components of a system environment according to various embodiments, through which services provided by one or more components of an embodiment system can be provided as cloud services.
[0030] Figure 12 Illustrates an example computer system that can be used to implement various embodiments. Detailed Description
[0031] In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments of the invention. It will be evident, however, that various embodiments may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.
[0032] Introduction
[0033] Artificial intelligence has many applications. For example, digital assistants are AI-driven interfaces that help users complete various tasks using natural language conversations. For each digital assistant, a customer can assemble one or more skills. A skill (also described herein as a chatbot, bot, or skillbot) is a separate bot that focuses on specific types of tasks such as tracking inventory, submitting time cards, and creating expense reports. When an end user engages with a digital assistant, the digital assistant evaluates the end user input and routes the conversation to the appropriate chatbot and from the appropriate chatbot routes the conversation. A digital assistant can be made available to an end user through various channels such as Messenger, SKYPE messenger, or short message service (SMS). The channels enable the chat to be transmitted back and forth between the end user and the digital assistant and its respective chatbots through various messaging platforms. The channels can also support user agent upgrades, event-initiated conversations, and testing.
[0034] When creating a model (e.g., a machine learning model) for performing various tasks for an application (e.g., a digital assistant), developers will face choices on how to define the model architecture and implement a learning process for the model. Typically, developers do not immediately know what the best model architecture or learning process for a given model should be, so developers hope to be able to explore a range of possibilities. Typically, developers will ask a computing device or subsystem (such as a hyperparameter tuner) to perform this exploration and automatically select the best model architecture and learning process. The parameters that define the model architecture and learning process are called hyperparameters, and the exploration process of finding the ideal model architecture and learning process is called hyperparameter tuning. Hyperparameters are used to improve the learning of the model, and their values are set before the learning process of the model begins. Hyperparameters are typically tuned by: defining the model and learning process (including selecting an initial set of values for the hyperparameters that define the architecture and learning process), defining the range of possible values for the hyperparameters, defining a method for sampling hyperparameter values, defining an evaluation criterion for judging the model (such as overall accuracy), running the model on training data to learn a set of model parameters, evaluating the performance of the learned model based on the evaluation criterion, and using the method for sampling hyperparameter values to adjust the values of the hyperparameters accordingly. Finding the best hyperparameters can be very cumbersome, so tuning algorithms such as grid search and random search are used to sample hyperparameter values.
[0035] The drawback of such standard hyperparameter tuning algorithms is that they ignore other important objectives (such as regression error) when performing the optimization process. Additionally, standard hyperparameter tuning mechanisms only consider a single domain (each with a training dataset and an evaluation dataset) to train / evaluate a machine learning model. Although some hyperparameter tuning algorithms consider multiple target domains when evaluating a model, the importance levels of each domain in the training of the machine learning model may not be the same. Additionally, standard hyperparameter tuning algorithms fail to adequately address the problem of instance-level regression. A single training / validation / test case within a training / validation / test dataset or other dataset is considered an instance. Instance-level regression can mean that an instance that was correctly classified by a previous version of the model is misclassified by a subsequent version of the model. Similarly, instance-level improvement can refer to an instance that was misclassified by a previous version of the model being correctly classified by a subsequent version of the model.
[0036] Standard hyperparameter tuning algorithms may not be able to detect instance-level regression or improvement. For example, both the first version of the model and the second version of the model are able to correctly classify 600 instances out of a dataset with a total of 1000 instances. However, 200 instances that were correctly classified by the first model are misclassified by the second model. Similarly, 200 instances that were misclassified by the first model are correctly classified by the second model. Therefore, there are 200 instance-level improvements and 200 instance-level regressions between the first model and the second model.
[0037] Standard hyperparameter tuning objective functions can use an accuracy score (e.g., percentage of accuracy) to evaluate a model. In this example, both the first model and the second model will have an accuracy score of 60% (e.g., 600 correct instances / 1000 instances total), and the standard algorithm cannot detect instance-level regression. Undetected instance-level regression can mean that the two models perform differently after deployment. Instance-level regression can lead to a perceived decrease in accuracy by customers using the model because, although the accuracy scores are consistent, the new model performs differently. Additionally, in some cases, instance-level regression can be caused by data points (e.g., examples from the training set) that affect unstable prediction results (i.e., predictions that are not the same or similar across runs). The reason for the unstable prediction results can be that the data points are near the decision boundary. The data points can cause the standard hyperparameter tuning model to chase random noise during later stages of the tuning process, which can lead to instance-level regression.
[0038] Accordingly, different approaches are needed to address these and other challenges. The present disclosure provides a hyperparameter tuning system and technique that optimizes an objective function while considering multiple metrics simultaneously, i.e., performs multi-objective optimization. Each metric is assigned a weight that represents the level of importance of the metric to the performance of the machine learning model. The hyperparameter tuning system and technique can also tune hyperparameters while considering different domains with different levels of importance. Specifically, a weight is assigned to each domain that specifies the importance of the domain in training the machine learning model. Additionally, the hyperparameter tuning system and technique enable the definition of one or more constraints for the machine learning model being trained. The training infrastructure (also referred to herein as the hyperparameter tuning system) employs various automated techniques to automatically identify, set, and tune the hyperparameters used to train the model such that the trained model complies with and meets the constraints specified for the model.
[0039] Additionally, the objective function in the hyperparameter tuning system can be modified to address instance-level regression. Instead of computing an accuracy score, the hyperparameter tuning system can compute instance-level regression or improvement by cross-referencing sets of correctly or incorrectly classified instances generated by different models. The objective function can be further modified to exclude unstable instances from the set of correctly or incorrectly classified instances. By excluding unstable instances, the objective function is optimized during tuning without chasing random noise caused by unstable instances. Additionally, not all domains may be equally important to a customer (e.g., based on business use). In such cases, the objective function can be further modified to include an acceptable regression ratio for each domain. Generally, the lower the acceptable regression ratio, the more important the domain. For example, if the acceptable regression ratio for a domain is set to zero, this means that the domain does not accept any regression at all. In this way, the hyperparameter tuning system of the present disclosure provides a single machine learning model that works across different domains and different metrics while minimizing instance-level regression.
[0040] In an exemplary embodiment, a computer-implemented method is provided, the method comprising: initializing a machine learning algorithm with a set of hyperparameter values; and accessing a hyperparameter objective function defined at least in part over a plurality of domains of a search space associated with the machine learning algorithm. The search space includes a training data set and an evaluation data set, wherein each domain includes a subdivision of the search space having at least one training data set and at least one evaluation data set, and the hyperparameter objective function includes a domain score for each domain, the domain score being computed based on the number of instances correctly or incorrectly predicted by the machine learning algorithm within the at least one evaluation data set during a given trial. For each trial of the hyperparameter tuning process: training the machine learning algorithm for each domain using the at least one training data set associated with each domain and the set of hyperparameter values (the training outputting a plurality of machine learning models, the plurality of machine learning models including a machine learning model for each domain); evaluating the machine learning model for each domain using the at least one evaluation data set associated with each domain and the set of hyperparameter values (the evaluation including generating the domain score for each domain); computing a current trial objective score using the hyperparameter objective function based on the domain score for each domain and a domain weight associated with each domain; and determining whether the machine learning model has reached convergence based on the current trial objective score. In response to determining that the machine learning model has reached convergence, providing at least one of the plurality of machine learning models.
[0041] Robot system
[0042] A bot (also known as a skill, chatbot, conversational bot, or talkbot) is a computer program that can perform a conversation with an end user. A bot can typically respond to natural language messages (e.g., questions or comments) via a messaging application that uses natural language messages. An enterprise can use one or more bots to communicate with end users via a messaging application. The messaging application can include, for example, over-the-top (OTT) messaging channels (such as Facebook Messenger, Facebook WhatsApp, WeChat, Line, Kik, Telegram, Talk, Skype, Slack, or SMS), virtual personal assistants (such as Amazon Dot, Echo, or Show, Google Home, Apple HomePod, etc.), native or hybrid extensions of mobile and web applications / responsive mobile applications or web applications with chat capabilities, or voice-based input (such as a device or application with the use of Siri, Microsoft Cortana, Google Voice, or other voice input for interaction).
[0043] In some examples, a bot can be associated with a Uniform Resource Identifier (URI). The URI can use a string of characters to identify the bot. The URI can be used as a webhook for one or more messaging application systems. The URI can include, for example, a Uniform Resource Locator (URL) or a Uniform Resource Name (URN). A bot can be designed to receive messages (e.g., Hypertext Transfer Protocol (HTTP) post call messages) from a messaging application system. The HTTP post call messages can relate to the URI from the messaging application system. In some examples, the messages can be different from the HTTP post call messages. For example, a bot can receive messages from Short Message Service (SMS). Although the discussion herein refers to the communications received by a bot as messages, it should be understood that the messages can be HTTP post call messages, SMS messages, or any other type of communication between two systems.
[0044] End users interact with a robot through conversational interactions (sometimes referred to as a conversational user interface (UI)), just as end users interact with other people. In some cases, the conversational interaction can include the end user saying "Hello" to the robot and the robot responding with "Hi" and asking the end user how the robot can help. End users also interact with the robot through other types of interactions, such as transactional interactions (e.g., with a banking robot trained at least to transfer funds from one account to another), informational interactions (e.g., interacting with a human resources robot trained at least to check the end user's remaining vacation time), and / or retail interactions (e.g., interacting with a retail robot trained at least to discuss returning purchased goods or seeking technical support).
[0045] In some examples, the robot can intelligently handle end user interactions without the intervention of the robot's administrator or developer. For example, the end user can send one or more messages to the robot to achieve a desired goal. The messages can include some content, such as text, emojis, audio, images, videos, or other ways of conveying messages. In some examples, the robot can automatically convert the content into a standardized form and generate a natural language response. The robot can also automatically prompt the end user for additional input parameters or request other additional information. In some examples, the robot can also initiate a conversation with the end user rather than passively responding to the end user's words.
[0046] A conversation with the robot can follow a specific conversation flow that includes multiple states. The flow can define what will happen next based on the input. In some examples, a state machine that includes user-defined states (e.g., end user intent) and actions to be taken within or between states can be used to implement the robot. The conversation can take different paths based on the end user input, which may affect the decisions made by the robot for the flow. For example, at each state, based on the end user input or utterance, the robot can determine the end user's intent in order to determine the next appropriate action to take. As used herein and in the context of an utterance, the term "intent" refers to the intent of the user providing the utterance. For example, a user may intend to engage the robot in a conversation to order pizza, where the user's intent will be expressed by the utterance "order pizza". The user intent can involve a specific task that the user wants the robot to perform on behalf of the user. Thus, an utterance reflecting the user's intent can be expressed as a question, command, request, etc.
[0047] In the context of a robot's configuration, the term "intent" as used herein also refers to configuration information used to map a user's utterance to a specific task / action or class of tasks / actions that the robot can perform. To distinguish the intent of an utterance (i.e., the user intent) from the robot's intent, the latter is sometimes referred to herein as the "robot intent". A robot intent can include a set of one or more utterances associated with the intent. For example, the intent to order a pizza can have various permutations of utterances expressing the desire to place an order for a pizza. These associated utterances can be used to train the robot's intent classifier so that the intent classifier can subsequently determine whether an input utterance from the user matches the intent to order a pizza. A robot intent can be associated with one or more conversation flows used to initiate a conversation with the user in a certain state. For example, a first message for the intent to order a pizza can be the question "What kind of pizza would you like?". In addition to the associated utterances, a robot intent can further include named entities related to the intent. For example, the intent to order a pizza can include variables or parameters (such as topping1, topping2, pizza type, pizza size, number of pizzas, etc.) used to perform the task of ordering a pizza. The values of the entities are typically obtained by talking to the user.
[0048] Figure 1 is a simplified block diagram of an environment 100 incorporating a chatbot system according to certain embodiments. Environment 100 includes a Digital Assistant Builder Platform (DABP) 102 that enables a user 104 of the DABP 102 to create and deploy a digital assistant or chatbot system. The DABP 102 can be used to create one or more digital assistants (or DAs) or chatbot systems. For example, as Figure 1 shown, a user 104 representing a particular enterprise can use the DABP 102 to create and deploy a digital assistant 106 for the users of the particular enterprise. For example, the DABP 102 can be used by a bank to create one or more digital assistants for use by the bank's customers. Multiple enterprises can use the same DABP 102 platform to create digital assistants. As another example, the owner of a restaurant (e.g., a pizza shop) can use the DABP 102 to create and deploy a digital assistant that enables the restaurant's customers to order food (e.g., order a pizza).
[0049] For the purposes of the present disclosure, a "digital assistant" is a tool that helps a user of the digital assistant complete various tasks through a natural language conversation. A digital assistant can be implemented using only software (e.g., the digital assistant is a digital tool implemented using a program, code, or instructions executable by one or more processors), using hardware, or using a combination of hardware and software. The digital assistant can be embodied or implemented in various physical systems or devices such as a computer, mobile phone, watch, appliance, vehicle, etc. A digital assistant is sometimes also referred to as a chatbot system. Thus, for the purposes of the present disclosure, the terms "digital assistant" and "chatbot system" are interchangeable.
[0050] A digital assistant (such as digital assistant 106 built using DABP 102) can be used to perform various tasks via a natural language-based conversation between the digital assistant and its user 108. As part of the conversation, the user can provide one or more user inputs 110 to the digital assistant 106 and receive a returned response 112 from the digital assistant 106. The conversation can include one or more of the inputs 110 and responses 112. Via these conversations, the user can request that the digital assistant perform one or more tasks, and in response, the digital assistant is configured to perform the tasks requested by the user and respond to the user with an appropriate response.
[0051] User inputs 110 are typically in the form of natural language and are referred to as utterances. The user utterance 110 can be in text form, such as when the user types a sentence, question, text fragment, or even a single word and provides the text as an input to the digital assistant 106. In some examples, the user utterance 110 can be in the form of an audio input or speech, such as when the user speaks or says something as an input to the digital assistant 106. The utterance is typically in the form of the language spoken by the user. For example, the utterance can be in English or some other language. When the utterance is in speech form, the speech input is converted into a text form utterance in that particular language, and then the text utterance is processed by the digital assistant 106. Various speech-to-text processing techniques can be used to convert the speech or audio input into a text utterance, which is then processed by the digital assistant 106. In some examples, the speech-to-text conversion can be done by the digital assistant 106 itself.
[0052] A discourse (which can be a text discourse or a voice discourse) can be a fragment, a sentence, multiple sentences, one or more words, one or more questions, a combination of the above types, etc. The digital assistant 106 is configured to apply natural language understanding (NLU) technology to the discourse to understand the meaning of the user input. As part of the NLU processing for the discourse, the digital assistant 106 is configured to perform a processing for understanding the meaning of the discourse, which involves identifying one or more intents and one or more entities corresponding to the discourse. After understanding the meaning of the discourse, the digital assistant 106 can perform one or more actions or operations in response to the understood meaning or intent. For the purposes of this disclosure, it is assumed that these discourses are text discourses directly provided by the user of the digital assistant 106, or the result of converting input voice discourses into text form. However, this is not intended to be limiting or restrictive in any way.
[0053] For example, the user input can request to order a pizza by providing a discourse such as "I want to order a pizza". After receiving such a discourse, the digital assistant 106 is configured to understand the meaning of the discourse and take appropriate actions. Appropriate actions can involve, for example, responding to the user with questions that request user input regarding the type of pizza the user expects to order, the size of the pizza, any toppings for the pizza, etc. The response provided by the digital assistant 106 can also be in natural language form and is typically in the same language as the input discourse. As part of generating these responses, the digital assistant 106 can perform natural language generation (NLG). For the user to be able to order a pizza via a conversation between the user and the digital assistant 106, the digital assistant can guide the user to provide all the necessary information for pizza ordering and then place the pizza order at the end of the conversation. The digital assistant 106 can end the conversation by outputting information to the user indicating that the pizza has been ordered.
[0054] At the conceptual level, the digital assistant 106 performs various processes in response to the discourse received from the user. In some examples, the process involves a series of processing steps or a processing step pipeline, including, for example, understanding the meaning of the input discourse, determining the actions to be performed in response to the discourse, causing the actions to be performed when appropriate, generating a response to be output to the user in response to the user discourse, outputting the response to the user, etc. The NLU processing can include parsing the received input discourse to understand the structure and meaning of the discourse, refining and reformulating the discourse to develop a better understandable form (e.g., logical form) or structure of the discourse. Generating the response can include using NLG technology.
[0055] NLU processing performed by a digital assistant (such as digital assistant 106) may include various NLP-related tasks such as sentence parsing (e.g., tokenization, lemmatization, identifying part-of-speech tags for the sentence, identifying named entities in the sentence, generating a dependency tree to represent the sentence structure, splitting the sentence into clauses, analyzing individual clauses, resolving anaphora, performing chunking, etc.). In some examples, the NLU processing is performed by digital assistant 106 itself. In some other examples, digital assistant 106 may use other resources to perform parts of the NLU processing. For example, a sentence can be processed by using a parser, a part-of-speech tagger, and / or NER to identify the syntax and structure of the input utterance sentence. In one implementation, for the English language, a parser, a part-of-speech tagger, and a named entity recognizer provided by the Stanford NLP Group are used to analyze the sentence structure and syntax. These are provided as part of the Stanford CoreNLP toolkit.
[0056] Although the various examples provided in this disclosure illustrate utterances in the English language, this is only meant as an example. In some examples, digital assistant 106 is also capable of handling utterances in languages other than English. Digital assistant 106 may provide subsystems (e.g., components implementing NLU functionality) that are configured to perform processing for different languages. These subsystems can be implemented as pluggable units that can be invoked from the NLU core server using service calls. This makes the NLU processing flexible and scalable for each language, including allowing different processing sequences. Language packs can be provided for individual languages, where the language pack can register a list of subsystems that can provide services from the NLU core server.
[0057] A digital assistant (such as Figure 1 the digital assistant 106 depicted therein) can be made available or accessible to its user 108 through various different channels such as, but not limited to, via certain applications, via social media platforms, via various messaging services and applications, and other applications or channels. A single digital assistant can be configured with several channels for itself, such that a single digital assistant can run on different services simultaneously and be accessed through different services.
[0058] A digital assistant or chatbot system typically includes one or more skills or is associated with one or more skills. In some embodiments, these skills are separate chatbots (referred to as skillbots) that are configured to interact with a user and complete a specific type of task (such as tracking inventory, submitting a time card, creating an expense report, ordering food, querying a bank account, making an appointment, purchasing a widget, etc.). For example, for Figure 1In the illustrated embodiment, the digital assistant or chatbot system 106 includes skills 116-1, 116-2, 116-3, etc. For the purposes of this disclosure, the terms "skill" and "skills" are used synonymously with the terms "skill bot" and "skill bots", respectively.
[0059] Each skill associated with the digital assistant helps the user of the digital assistant complete a task through a conversation with the user, where the conversation may include a combination of text or audio input provided by the user and responses provided by the skill bot. These responses can take the following forms: text or audio messages to the user and / or simple user interface elements (e.g., selection lists) presented to the user for the user to select.
[0060] There are various ways to associate or add a skill or skill bot with a digital assistant. In some instances, a skill bot can be developed by an enterprise and then added to a digital assistant that uses DABP 102. In other instances, a skill bot can be developed and created using DABP 102 and then added to a digital assistant created using DABP 102. In still other instances, DABP 102 provides an online digital store (referred to as the "Skill Store") that offers multiple skills related to a wide variety of tasks. The skills provided through the Skill Store can also expose various cloud services. To add a skill to a digital assistant generated using DABP 102, a user of DABP 102 can access the Skill Store via DABP 102, select the desired skill, and indicate that the selected skill is to be added to the digital assistant created using DABP 102. Skills from the Skill Store can be added to the digital assistant as-is or in a modified form (e.g., a user of DABP 102 can select and copy a specific skill bot provided by the Skill Store, customize or modify the selected skill bot, and then add the modified skill bot to the digital assistant created using DABP 102).
[0061] Various different architectures can be used to implement the digital assistant or chatbot system. For example, in certain embodiments, a digital assistant created and deployed using DABP 102 can be implemented using the master bot / slave (or sub) bot paradigm or architecture. According to this paradigm, the digital assistant is implemented as a master bot that interacts with one or more slave bots that are skill bots. For example, in Figure 1In the depicted embodiment, the digital assistant 106 includes a primary robot 114 and skill robots 116-1, 116-2, etc. that are secondary robots to the primary robot 114. In some examples, the digital assistant 106 itself is considered to act as the primary robot.
[0062] A digital assistant implemented according to a primary-secondary robot architecture enables a user of the digital assistant to interact with multiple skills through a unified user interface (i.e., via the primary robot). When a user interacts with the digital assistant, the primary robot receives the user input. The primary robot then performs processing to determine the meaning of the user input utterance. The primary robot then determines whether the task requested by the user in the utterance can be handled by the primary robot itself, otherwise the primary robot selects an appropriate skill robot to handle the user request and routes the conversation to the selected skill robot. This enables the user to have a conversation with the digital assistant through a common single interface and still have the ability to use several skill robots configured to perform specific tasks. For example, for a digital assistant developed for an enterprise, the primary robot of the digital assistant can interface with skill robots having specific functions, such as a customer relationship management (CRM) robot for performing functions related to customer relationship management, an enterprise resource planning (ERP) robot for performing functions related to enterprise resource planning, a human capital management (HCM) robot for performing functions related to human capital management, etc. In this way, the end user or consumer of the digital assistant only needs to know how to access the digital assistant through the common primary robot interface, and multiple skill robots are provided in the background to handle user requests.
[0063] In some examples, in the primary robot / secondary robot infrastructure, the primary robot is configured to know a list of available skill robots. The primary robot can access metadata identifying the various available skill robots and, for each skill robot, access the capabilities of the skill robot including the tasks that can be performed by the skill robot. After receiving a user request in the form of an utterance, the primary robot is configured to identify or predict a specific skill robot from among the multiple available skill robots that can best serve or handle the user request. The primary robot then routes the utterance (or a portion of the utterance) to that specific skill robot for further handling. Thus, control flows from the primary robot to the skill robot. The primary robot can support multiple input channels and output channels. In some examples, the routing can be performed by processing executed by one or more of the available skill robots. For example, as discussed below, a skill robot can be trained to infer the intent of an utterance and determine whether the inferred intent matches the intent configured for the skill robot. Thus, the routing performed by the primary robot can involve the skill robot transmitting an indication to the primary robot indicating whether the skill robot has been configured with an intent suitable for handling the utterance.
[0064] Although Figure 1 the embodiments shown include a main robot 114 and skill robots 116-1, 116-2, and 116-3 in the digital assistant 106, this is not intended to be limiting. The digital assistant may include various other components (e.g., other systems and subsystems) that provide the functionality of the digital assistant. These systems and subsystems may be implemented in software only (e.g., code, instructions stored on a computer-readable medium and executable by one or more processors), in hardware only, or in embodiments using a combination of software and hardware.
[0065] The DABP 102 provides the infrastructure, as well as various services and features, that enable users of the DABP 102 to create digital assistants (including one or more skill robots associated with the digital assistant). In some instances, a skill robot can be created by cloning an existing skill robot, e.g., cloning a skill robot provided by a skill store. As previously mentioned, the DABP 102 provides a skill store or skill catalog that provides multiple skill robots for performing various tasks. A user of the DABP 102 can clone a skill robot from the skill store. The cloned skill robot can be modified or customized as needed. In some other instances, a user of the DABP 102 creates a skill robot from scratch using the tools and services provided by the DABP 102. As previously mentioned, the skill store or skill catalog provided by the DABP 102 can provide multiple skill robots for performing various tasks.
[0066] In certain examples, at a high level, creating or customizing a skill robot involves the following steps:
[0067] (1) Configure settings for the new skill robot
[0068] (2) Configure one or more intents for the skill robot
[0069] (3) Configure one or more entities for one or more intents
[0070] (4) Train the skill robot
[0071] (5) Create a dialogue flow for the skill robot
[0072] (6) Add custom components to the skill robot as needed
[0073] (7) Test and deploy the skill robot
[0074] Each of the above steps is briefly described below.
[0075] (1)Configure settings for the new skill bot - Various settings can be configured for the skill bot. For example, the skill bot designer can specify one or more invocation names for the skill bot being created. Then, the user of the digital assistant can use these invocation names to explicitly invoke the skill bot. For example, the user can enter the invocation name in the user's utterance to explicitly invoke the corresponding skill bot.
[0076] (2)Configure one or more intents and associated example utterances for the skill bot - The skill bot designer specifies one or more intents (also known as bot intents) for the skill bot being created. Then, the skill bot is trained based on these specified intents. These intents represent the categories or classifications that the skill bot is trained to infer for an input utterance. After receiving an utterance, the trained skill bot infers the intent of the utterance, where the inferred intent is selected from a set of predefined intents used to train the skill bot. Then, the skill bot takes an appropriate action in response to the utterance based on the intent inferred for the utterance. In some instances, the intents of the skill bot represent tasks that the skill bot can perform for the user of the digital assistant. Each intent is assigned an intent identifier or intent name. For example, for a skill bot trained for a bank, the intents specified for the skill bot can include "CheckBalance", "TransferMoney", "DepositCheck", etc.
[0077] For each intent defined for the skill bot, the skill bot designer can also provide one or more example utterances that represent and illustrate the intent. These example utterances are intended to represent the utterances that the user can enter for the intent to the skill bot. For example, for the CheckBalance intent, the example utterances can include "What's my savings account balance?", "How much is in my checking account?", "How much money do I have in my account", etc. Thus, various permutations of typical user utterances can be specified as example utterances for the intent.
[0078] These intents and their associated example utterances are used as training data for training the skill robot. A variety of different training techniques can be used. As a result of this training, a prediction model is generated that is configured to take an utterance as input and output the intent inferred by the prediction model for the utterance. In some instances, the input utterance is provided to an intent analysis engine that is configured to use the trained model to predict or infer the intent of the input utterance. The skill robot can then take one or more actions based on the inferred intent.
[0079] (3) Configure entities for one or more intents of a skill robot - In some instances, additional context may be required to enable a skill robot to respond appropriately to user utterances. For example, there may be situations where user input utterances resolve to the same intent in a skill robot. For example, in the above example, the utterances "What's my savings account balance?" and "How much is in my checking account?" both resolve to the same CheckBalance intent, but these utterances are different requests asking different things. To clarify such requests, one or more entities are added to the intent. Using the example of a banking skill robot, an entity called AccountType (which defines values called "checking" and "saving") can enable the skill robot to parse the user request and respond appropriately. In the above example, although these utterances resolve to the same intent, the values associated with the AccountType entity for the two utterances are different. This enables the skill robot to perform potentially different actions for the two utterances, even though the two utterances resolve to the same intent. One or more entities can be specified for certain intents configured for a skill robot. Therefore, entities are used to add context to the intent itself. Entities help describe the intent more fully and enable the skill bot to complete the user request.
[0080] In some examples, there are two types of entities: (a) built-in entities provided by the DABP 102; and (2) custom entities that can be specified by the skill robot designer. Built-in entities are general entities that can be used with various robots. Examples of built-in entities include, but are not limited to, entities related to time, date, address, number, email address, duration, recurring time period, currency, phone number, URL, etc. Custom entities are used for more customized applications. For example, for a banking skill, the AccountType entity can be defined by the skill robot designer to enable various banking transactions by examining keywords (such as current deposit, savings, and credit card, etc.) entered by the user.
[0081] (4) Training the skill robot - The skill robot is configured to receive user input in the form of utterances, parse or otherwise process the received input, and identify or select an intent related to the received user input. As indicated above, for this purpose, the skill robot must be trained. In some embodiments, the skill robot is trained based on the intents configured for the skill robot and the example utterances associated with the intents (collectively referred to as training data), such that the skill robot can parse the user input utterance into one of the intents it is configured with. In some examples, the skill robot uses a prediction model that is trained using the training data and allows the skill robot to discern what the user is saying (or, in some cases, trying to say). The DABP 102 provides various different training techniques that can be used by the skill robot designer to train the skill robot, including various machine learning-based training techniques, rule-based training techniques, and / or combinations thereof. In some examples, a portion (e.g., 80%) of the training data is used to train the skill robot model and another portion (e.g., the remaining 20%) is used to test or validate the model. Once trained, the trained model (sometimes also referred to as the trained skill robot) can be used to handle the user's utterances and respond to the user's utterances. In some cases, the user's utterance can be a question that only requires a single answer and no additional conversation. To handle such a situation, a Q&A (question answering) intent can be defined for the skill robot. This enables the skill robot to output a response to the user's request without having to update the conversation definition. The Q&A intent is created in a similar manner to a regular intent. The dialogue flow for the Q&A intent may be different from the dialogue flow for a regular intent.
[0082] (5) Create a dialogue flow for the skill robot - The dialogue flow specified for the skill robot describes how the skill robot reacts when parsing different intents of the skill robot in response to received user input. The dialogue flow defines the operations or actions that the skill robot will take. For example, how the skill robot responds to user utterances, how the skill robot prompts the user for input, and how the skill robot returns data. The dialogue flow is like a flowchart that the skill robot follows. Skill robot designers use languages such as the markdown language to specify the dialogue flow. In some embodiments, a YAML version called OBotML can be used to specify the dialogue flow for the skill robot. The dialogue flow definition of the skill robot serves as a model for the conversation itself, which is the model that enables skill robot designers to choreograph the interaction between the skill robot and the users it serves.
[0083] In some examples, the dialogue flow definition of the skill robot includes the following three parts:
[0084] (a) Context part
[0085] (b) Default transition part
[0086] (c) State part
[0087] Context part - Skill robot designers can define variables used in the conversation flow in the context part. Other variables that can be named in the context part include, but are not limited to: variables for error handling, variables for built-in entities or custom entities, user variables that enable the skill robot to recognize and save user preferences, etc.
[0088] Default transition part - The transitions of the skill robot can be defined in the dialogue flow state part or in the default transition part. The transitions defined in the default transition part act as a fallback and are triggered when there is no applicable transition defined within the state or the conditions required to trigger a state transition are not met. The default transition part can be used to define the routing that allows the skill robot to gracefully handle unexpected user actions.
[0089] State part - The dialogue flow and its associated operations are defined as a sequence of temporary states that manage the logic within the dialogue flow. Each state node within the dialogue flow definition names a component that provides the functionality required at that point in the conversation. Thus, the state is built around the component. The state contains properties specific to the component and defines the transitions to other states that are triggered after the component executes.
[0090] Special case scenarios can be handled using the state section. For example, sometimes it may be desirable to provide the user with an option to temporarily leave the first skill to allow them to do something in a second skill within the digital assistant. For example, if the user is engaged in a conversation with the shopping skill (e.g., the user has made some purchase selections), the user may want to jump to the banking skill (e.g., the user may want to ensure he / she has enough money for the purchase) and then return to the shopping skill to complete the user's order. To address this, the actions in the first skill can be configured to initiate an interaction with a second different skill in the same digital assistant and then return to the original flow.
[0091] (6) Adding custom components to the skill bot - As described above, the states specified in the dialog flow of the skill bot name the components that provide the functionality required for the corresponding state. The components enable the skill bot to perform functions. In some embodiments, the DABP 102 provides a set of pre-configured components for performing a variety of functions. The skill bot designer can select one or more of these pre-configured components and associate them with the states in the dialog flow of the skill bot. The skill bot designer can also use the tools provided by the DABP 102 to create custom or new components and associate the custom components with one or more states in the dialog flow of the skill bot.
[0092] (7) Testing and deploying the skill bot - The DABP 102 provides several features that enable the skill bot designer to test the skill bot being developed. The skill bot can then be deployed and included in the digital assistant.
[0093] While the above description describes how to create a skill bot, similar techniques can also be used to create a digital assistant (or master bot). At the master bot or digital assistant level, built-in system intents can be configured for the digital assistant. These built-in system intents are used to identify the general tasks that the digital assistant itself (i.e., the master bot) can handle without invoking the skill bot associated with the digital assistant. Examples of system intents defined for the master bot include: (1) Exit: Applicable when the user signals a desire to exit the current session or context in the digital assistant; (2) Help: Applicable when the user requests help or orientation; and (3) Unresolved Intent: Applicable to user inputs that do not closely match the exit intent and the help intent. The digital assistant also stores information about one or more skill bots associated with the digital assistant. This information enables the master bot to select a specific skill bot for handling the utterance.
[0094] At the main robot or digital assistant level, when a user inputs a phrase or utterance to the digital assistant, the digital assistant is configured to perform processing for determining how to route the utterance and associated conversation. The digital assistant uses a routing model to determine this, which can be rule-based, AI-based, or a combination thereof. The digital assistant uses the routing model to determine whether the conversation corresponding to the user input utterance is to be routed to a specific skill for disposition, to be disposed of by the digital assistant or the main robot itself according to built-in system intents, or to be disposed of as a different state in the current conversation flow.
[0095] In some embodiments, as part of this processing, the digital assistant determines whether the user input utterance explicitly identifies a skill bot using its invocation name. If the invocation name is present in the user input, the invocation name is considered an explicit invocation of the skill bot corresponding to the invocation name. In such a scenario, the digital assistant can route the user input to the explicitly invoked skill bot for further disposition. In some embodiments, if there is no specific or explicit invocation, the digital assistant evaluates the received user input utterance and calculates confidence scores for system intents and skill bots associated with the digital assistant. The scores calculated for the skill bots or system intents represent how likely the user input represents a task that the skill bot is configured to perform or represents a system intent. Any system intent or skill bot for which the associated calculated confidence score exceeds a threshold (e.g., a confidence threshold routing parameter) is selected as a candidate for further evaluation. Then, the digital assistant selects a specific system intent or skill bot from the identified candidates for further disposition of the user input utterance. In some embodiments, after one or more skill bots are identified as candidates, the intents associated with those candidate skills are evaluated (according to the intent model for each skill) and confidence scores are determined for each intent. Any intent for which the confidence score exceeds a threshold (e.g., 70%) is generally considered a candidate intent. If a specific skill bot is selected, the user utterance is routed to that skill bot for further processing. If a system intent is selected, the main robot itself performs one or more actions according to the selected system intent.
[0096] Figure 2 is a simplified block diagram of a main robot (MB) system 200 according to some embodiments. The MB system 200 can be implemented in software only, in hardware only, or in a combination of hardware and software. The MB system 200 includes a preprocessing subsystem 210, a multi-intent subsystem (MIS) 220, an explicit invocation subsystem (EIS) 230, a skill bot invoker 240, and a data store 250. Figure 2The MB system 200 depicted is merely an example of the component arrangement in the main robot. Those of ordinary skill in the art will recognize many possible variations, alternatives, and modifications. For example, in some embodiments, the MB system 200 may have more or fewer systems or components than Figure 2 those shown, may combine two or more subsystems, or may have different subsystem configurations or arrangements.
[0097] The preprocessing subsystem 210 receives the utterance "A" 202 from the user and processes the utterance through the language detector 212 and the language parser 214. As indicated above, the utterance can be provided in various ways including audio or text. The utterance 202 can be a sentence fragment, a complete sentence, multiple sentences, etc. The utterance 202 can include punctuation. For example, if the utterance 202 is provided as audio, the preprocessing subsystem 210 can use a speech-to-text converter (not shown) that inserts punctuation (e.g., commas, semicolons, periods, etc.) into the resulting text to convert the audio to text.
[0098] The language detector 212 detects the language of the utterance 202 based on the text of the utterance 202. The way the utterance 202 is handled depends on the language, as each language has its own grammar and semantics. Differences between languages are considered when analyzing the syntax and structure of the utterance.
[0099] The language parser 214 performs a syntactic analysis of the utterance 202 to extract the part-of-speech (POS) tags of the individual language units (e.g., words) in the utterance 202. POS tags include, for example, nouns (NN), pronouns (PN), verbs (VB), etc. The language parser 214 can also tokenize the language units of the utterance 202 (e.g., convert each word into a separate token) and classify the words by inflectional form. A lemma is the main form of a set of words as represented in a dictionary (e.g., "run" is the lemma for run, runs, ran, running, etc.). Other types of preprocessing that the language parser 214 can perform include chunking of compound expressions, e.g., combining "credit" and "card" into a single expression "credit card". The language parser 214 can also identify the relationships between the words in the utterance 202. For example, in some embodiments, the language parser 214 generates a dependency tree that indicates which part of the utterance (e.g., a particular noun) is the direct object, which part of the utterance is a preposition, etc. The result of the processing performed by the language parser 214 forms the extracted information 205 and is provided as input to the MIS 220 along with the utterance 202 itself.
[0100] As indicated above, utterance 202 may include more than one sentence. For purposes of detecting multiple intents and explicit invocations, utterance 202 may be treated as a single unit even if it includes multiple sentences. However, in some embodiments, preprocessing may be performed, for example, by preprocessing subsystem 210, to identify a single sentence within the multiple sentences for multiple intent analysis and explicit invocation analysis. Generally, whether utterance 202 is processed at the level of a single sentence or as a single unit including multiple sentences, the results produced by MIS 220 and EIS 230 are substantially the same.
[0101] MIS 220 determines whether utterance 202 represents multiple intents. Although MIS 220 may detect the presence of multiple intents in utterance 202, the processing performed by MIS 220 does not involve determining whether the intent of utterance 202 matches any of the intents that have been configured for the robot. Instead, the processing of determining whether the intent of utterance 202 matches a robot intent may be performed by intent classifier 242 of MB system 200 or an intent classifier of a skill robot (e.g., as Figure 3 shown). The processing performed by MIS 220 assumes the existence of a robot (e.g., a specific skill robot or the main robot itself) that can handle utterance 202. Thus, the processing performed by MIS 220 does not need to know which robots are in the chatbot system (e.g., the identities of the skill robots registered with the main robot) or what intents have been configured for a particular robot.
[0102] To determine that utterance 202 includes multiple intents, MIS 220 applies one or more rules from a set of rules 252 in data store 250. The rules applied to utterance 202 depend on the language of utterance 202 and may include sentence patterns indicating the presence of multiple intents. For example, a sentence pattern may include a coordinating conjunction that connects two parts of a sentence (e.g., a conjunction), where the two parts correspond to different intents. If utterance 202 matches the sentence pattern, it may be inferred that utterance 202 represents multiple intents. It should be noted that an utterance with multiple intents does not necessarily have different intents (e.g., intents related to different robots or different intents within the same robot). Instead, an utterance may have different instances of the same intent (e.g., "Place a pizza order using payment account X, then place a pizza order using payment account Y").
[0103] As part of determining that utterance 202 represents multiple intents, MIS 220 also determines which parts of utterance 202 are associated with each intent. MIS 220 constructs new utterances for separate processing to replace the original utterance for each intent represented in an utterance that contains multiple intents, e.g., utterance "B" 206 and utterance "C" 208 depicted in Figure 2 . Thus, the original utterance 202 can be split into two or more separate utterances to be processed one at a time. MIS 220 uses the extracted information 205 and / or based on an analysis of the utterance 202 itself to determine which of the two or more utterances should be processed first. For example, MIS 220 may determine that utterance 202 contains a marker word indicating that a particular intent should be processed first. The newly formed utterance corresponding to that particular intent (e.g., one of utterance 206 or utterance 208) will be sent first for further processing by EIS 230. After the session triggered by the first utterance has ended (or has been temporarily suspended), then the next highest priority utterance (e.g., the other of utterance 206 or utterance 208) can be sent to EIS 230 for processing.
[0104] EIS 230 determines whether the utterance it receives (e.g., utterance 206 or utterance 208) contains the call name of a skill bot. In some embodiments, each skill bot in the chatbot system is assigned a unique call name that differentiates the skill bot from other skill bots in the chatbot system. The list of call names can be saved in data store 250 as part of the skill bot information 254. When the utterance contains a word that matches the call name, the utterance is considered an explicit call. If the bot is not explicitly called, the utterance received by EIS 230 is considered a non-explicit call utterance 234 and is input into the main bot's intent classifier (e.g., intent classifier 242) to determine which bot to use to process the utterance. In some instances, intent classifier 242 will determine that the main bot should process the non-explicit call utterance. In other instances, intent classifier 242 will determine the skill bot to which the utterance should be routed for processing.
[0105] The explicit call feature provided by EIS 230 has several advantages. It can reduce the amount of processing that the main bot has to perform. For example, when there is an explicit call, the main bot may not have to (e.g., using intent classifier 242) perform any intent classification analysis, or may have to perform a simplified intent classification analysis to select a skill bot. Thus, explicit call analysis can enable the selection of a particular skill bot without resorting to intent classification analysis.
[0106] Moreover, there may be cases where there is a functional overlap between multiple skill robots. For example, this may occur if the intents handled by two skill robots overlap or are very close to each other. In such cases, it may be difficult for the master robot to identify which one of the multiple skill robots to select based solely on intent classification analysis. In such scenarios, an explicit call makes it unambiguous which specific skill robot is to be used.
[0107] In addition to determining that the utterance is an explicit call, the EIS230 is also responsible for determining whether any part of the utterance should be used as input to the skill robot being explicitly called. In particular, the EIS230 can determine whether a part of the utterance is unrelated to the call. The EIS230 can perform this determination by analyzing the utterance and / or analyzing the extracted information 205. The EIS230 can send the part of the utterance that is unrelated to the call to the called skill robot instead of sending the entire utterance that the EIS230 receives. In some instances, the input to the called skill robot is simply formed by deleting any part of the utterance associated with the call. For example, "I want to order pizza using Pizza Bot" can be shortened to "I want to order pizza" because "using PizzaBot" is related to the call to the pizza robot but not to any processing to be performed by the pizza robot. In some instances, the EIS230 can reformat the part to be sent to the called robot, such as to form a complete sentence. Thus, the EIS230 not only determines that there is an explicit call but also determines what to send to the skill robot when there is an explicit call. In some instances, there may be no text that can be input to the called robot. For example, if the utterance is "Pizza Bot", the EIS230 can determine that the pizza robot is being called but there is no text for the pizza robot to process. In such a scenario, the EIS230 can indicate to the skill robot invoker 240 that there is nothing to send.
[0108] The skill robot invoker 240 invokes skill robots in various ways. For example, the skill robot invoker 240 can invoke a robot in response to receiving an indication 235 that a specific skill robot has been selected as a result of an explicit invocation. The indication 235 can be sent by the EIS 230 together with the input for the skill robot for the explicit invocation. In this scenario, the skill robot invoker 240 hands over the control of the conversation to the explicitly invoked skill robot. The explicitly invoked skill robot will determine an appropriate response to the input by treating the input from the EIS 230 as an independent utterance. For example, the response can be to perform a specific action or start a new conversation in a specific state, where the initial state of the new conversation depends on the input sent from the EIS 230.
[0109] Another way the skill robot invoker 240 can invoke a skill robot is through an implicit invocation using the intent classifier 242. Machine learning and / or rule-based training techniques can be used to train the intent classifier 242 to determine the likelihood that an utterance represents a task that a specific skill robot is configured to perform. The intent classifier 242 is trained on different classifications, one for each skill robot. For example, whenever a new skill robot is registered with the main robot, a list of example utterances associated with the new skill robot can be used to train the intent classifier 242 to determine the likelihood that a specific utterance represents a task that the new skill robot can perform. The parameters (e.g., a set of parameter values of a machine learning model) resulting from this training can be stored as part of the skill robot information 254.
[0110] In some embodiments, the intent classifier 242 is implemented using a machine learning model, as described in further detail herein. The training of the machine learning model can involve at least inputting a subset of utterances from example utterances associated with various skill robots to generate an inference as the output of the machine learning model about which robot is the correct one for handling any specific training utterance. For each training utterance, an indication of the correct robot for the training utterance can be provided as ground truth information. The behavior of the machine learning model can then be adapted (e.g., through backpropagation) to minimize the difference between the generated inference and the ground truth information.
[0111] In some embodiments, the intent classifier 242 determines a confidence score indicating the likelihood that a skill robot can handle an utterance (e.g., the non-explicit call utterance 234 received from the EIS 230) for each skill robot registered with the primary robot. The intent classifier 242 can also determine a confidence score for each system-level intent that has been configured (e.g., help, exit). If a particular confidence score meets one or more conditions, the skill robot invoker 240 will invoke the robot associated with the particular confidence score. For example, it may be required to meet a threshold confidence score value. Thus, the output 245 of the intent classifier 242 is an identification of a system intent or an identification of a particular skill robot. In some embodiments, in addition to meeting the threshold confidence score value, the confidence score must also exceed the second-highest confidence score by a certain margin. Imposing such a condition will enable routing to a particular skill robot when the confidence scores of multiple skill robots all exceed the threshold confidence score value.
[0112] After the robot is identified based on the confidence score-based evaluation, the skill robot invoker 240 hands over the processing to the identified robot. In the case of a system intent, the identified robot is the primary robot. Otherwise, the identified robot is a skill robot. Further, the skill robot invoker 240 will determine what to provide as the input 247 to the identified robot. As indicated above, in the case of an explicit call, the input 247 can be based on the part of the utterance that is not related to the call, or the input 247 can be nothing (e.g., an empty string). In the case of an implicit call, the input 247 can be the entire utterance.
[0113] The data store 250 includes one or more computing devices that store data used by various subsystems of the primary robot system 200. As explained above, the data store 250 includes rules 252 and skill robot information 254. The rules 252 include, for example, rules for the MIS 220 to determine when an utterance represents multiple intents and how to split an utterance that represents multiple intents. The rules 252 further include rules for the EIS 230 to determine which parts of an utterance that explicitly calls a skill robot to send to the skill robot. The skill robot information 254 includes the invocation names of the skill robots in the chatbot system, e.g., a list of the invocation names of all skill robots registered with a particular primary robot. The skill robot information 254 can also include information that the intent classifier 242 uses to determine the confidence scores for each skill robot in the chatbot system, e.g., the parameters of a machine learning model.
[0114] Figure 3is a simplified block diagram of a skills bot system 300 according to certain embodiments. The skills bot system 300 is a computing system that can be implemented solely in software, solely in hardware, or in a combination of hardware and software. In certain embodiments, as Figure 1 depicted in the embodiments, the skills bot system 300 can be used to implement one or more skills bots within a digital assistant.
[0115] The skills bot system 300 includes a MIS 310, an intent classifier 320, and a session manager 330. The MIS 310 is similar to Figure 2 the MIS220 in and provides similar functionality, including operably using rules 352 in a data store 350 to determine: (1) whether an utterance represents multiple intents, and if so, (2) how to split the utterance into separate utterances for each of the multiple intents. In certain embodiments, the rules applied by the MIS 310 for detecting multiple intents and for splitting utterances are the same as those applied by the MIS220. The MIS 310 receives an utterance 302 and extracted information 304. The extracted information 304 is similar to Figure 1 the extracted information 205 in and can be generated using a language grammar parser 214 or a language grammar parser local to the skills bot system 300.
[0116] The intent classifier 320 can be trained in a manner similar to the intent classifier 242 discussed above in connection with the Figure 2 embodiments and is described in further detail herein. For example, in certain embodiments, the intent classifier 320 is implemented using a machine learning model. The machine learning model of the intent classifier 320 is trained using at least a subset of example utterances associated with the particular skills bot as training utterances. The ground truth for each training utterance will be the particular bot intent associated with the training utterance.
[0117] The utterance 302 can be received directly from the user or provided by the main bot. When the utterance 302 is provided by the main bot, for example, as passed through Figure 2In the embodiments depicted, the results processed by MIS220 and EIS230 can bypass MIS 310 to avoid repeating the processing already performed by MIS220. However, if utterance 302 is received directly from the user, for example, during a session that occurs after routing to the skill bot, then MIS 310 can process utterance 302 to determine whether utterance 302 represents multiple intents. If so, MIS 310 applies one or more rules to split utterance 302 into separate utterances for each intent, such as utterance "D" 306 and utterance "E" 308. If utterance 302 does not represent multiple intents, then MIS 310 forwards utterance 302 to intent classifier 320 for intent classification without splitting utterance 302.
[0118] Intent classifier 320 is configured to match the received utterance (e.g., utterance 306 or 308) with the intents associated with the skill bot system 300. As explained above, a skill bot can be configured with one or more intents, each intent including at least one example utterance associated with the intent and used to train the classifier. In Figure 2 the embodiment, the intent classifier 242 of the main robot system 200 is trained to determine the confidence scores for the individual skill bots and the confidence score for the system intent. Similarly, intent classifier 320 can be trained to determine the confidence score for each intent associated with the skill bot system 300. The classification performed by intent classifier 242 is at the robot level, while the classification performed by intent classifier 320 is at the intent level and is thus more fine-grained. Intent classifier 320 can access intent information 354. For each intent associated with the skill bot system 300, intent information 354 includes a list of utterances that represent the intent, explain the meaning of the intent, and are generally associated with the tasks that can be performed by the intent. Intent information 354 can further include parameters that result from training on the list of utterances.
[0119] Session manager 330 receives an indication 322 of a specific intent as the output of intent classifier 320, the intent being identified by intent classifier 320 as the best match for the utterance input to intent classifier 320. In some instances, intent classifier 320 cannot determine any match. For example, if the utterance relates to a system intent or the intent of a different skill bot, the confidence score calculated by intent classifier 320 may be lower than a threshold confidence score value. When this occurs, the skill bot system 300 can submit the utterance to the main robot for disposition, e.g., to route to a different skill bot. However, if intent classifier 320 successfully identifies an intent within the skill bot, then session manager 330 will initiate a session with the user.
[0120] The session initiated by the session manager 330 is a session specific to the intent recognized by the intent classifier 320. For example, the session manager 330 can be implemented using a state machine configured to execute a dialogue flow for the recognized intent. The state machine can include a default starting state (e.g., when the intent is invoked without any additional input) and one or more additional states, where each state is associated with an action to be performed by the skill bot (e.g., performing a purchase transaction) and / or a dialogue to be presented to the user (e.g., a question, a response). Thus, the session manager 330 can determine the action / dialogue 335 upon receiving the indication 322 that an intent has been recognized, and can determine additional actions or dialogues in response to subsequent utterances received during the session.
[0121] The data store 350 includes one or more computing devices that store data used by the various subsystems of the skill bot system 300. As Figure 3 depicted, the data store 350 includes rules 352 and intent information 354. In some embodiments, the data store 350 can be integrated into the data store of the main robot or digital assistant, such as Figure 2 the data store 250 in
[0122] Goal - Based Hyperparameter Tuning
[0123] As previously mentioned, AI-based solutions can use one or more machine learning models to perform various functions. For example, a chatbot can use a machine learning model configured to take an utterance as input and infer or predict the intent of each utterance. The chatbot can then use the intent of the utterance inferred by the model to determine how to respond to the utterance. Implementing a machine learning model (also referred to as a model) typically occurs in two phases: (1) a training phase, where training data is run on one or more algorithms to create a trained model; and (2) an inference phase, where the trained model is used to make predictions based on new data. Training infrastructure is typically provided to implement the training phase to train the model. The training infrastructure can be provided by tools, applications, or software for performing training. The training infrastructure is configured to run training data on one or more algorithms to train or learn the algorithms and create a model. The training infrastructure typically provides control over hyperparameters that manage the training process. A set of hyperparameter values determines the network structure of the algorithm (e.g., the number of input layers, the number of hidden layers, activation functions, etc.), the way the algorithm is trained (e.g., learning rate, number of epochs, etc.), and any other hyperparameters (e.g., data augmentation settings, batch balancing settings, cache settings, etc.). The set of hyperparameter values is determined by the training infrastructure using a hyperparameter tuning process or algorithm.
[0124] Hyperparameter tuning can involve optimizing the performance of a model across multiple domains by trying different values for different hyperparameters. At the core of hyperparameter tuning is a tuning objective function, the value of which can be optimized (e.g., maximized or minimized) during hyperparameter tuning. For hyperparameter tuning optimized for N domains, e.g., D = {D_0, D_1, ..., D_N}, the objective function can be defined as follows:
[0125] F_{tuning} = w_0 * f(D_0) + w_1 * f(D_i) +... + w_N * f(D_N)
[0126] where f(D_i) (i = 0, 1,..., N) is the domain score calculated for each domain, and w_0, w_1,..., w_N are the domain weights to be applied to each domain score. The domain score is a measure of the performance of the model for a given domain D_0, D_i, D_N, etc. (e.g., model accuracy or F1). The domain score can be defined as the evaluation score of the trained model or a target-based score, in which a baseline performance is set to understand the improvement or regression of the trained model relative to the baseline. For example:
[0127] · The evaluation score of the trained model, e.g., the accuracy of intent classification, the Bilingual Evaluation Understudy (BLEU) score for machine translation, the F1 score for image classification, etc.
[0128] · Target-based score - improvement and / or regression compared to the baseline score of the domain
[0129] Each hyperparameter tuning can run T trials, where for any trial t, the objective value F_{tuning}_t can be calculated, where t = 0, 1,..., T; and the purpose of hyperparameter tuning can be to find such a trial: where a set of assignments to the hyperparameters maximizes or minimizes the objective function value, e.g., H_best = argmax_{t} F_{tuning}_t, where t = 0, 1,..., T. The domain weights in the objective function, e.g., W = {w_0, w_1,..., w_N}, can indicate the importance of each domain; and domains with higher weights can receive more attention, so domains with higher domain weights are more likely to have better performance. The domain score with a larger domain weight will have a greater impact on the final value of the objective function than the same domain score with a smaller domain weight, because the objective score is scaled by the weighted value.
[0130] Figure 4Adepicts exemplary types of hyperparameters according to various embodiments. The hyperparameters may include: the number of layers in the model, the type of learning algorithm used to train the model, the learning rate, the number of training epochs, the number of hidden units in each layer, and the width (i.e., the number of units in each hidden layer). In some instances, a user using the training infrastructure would manually set the values of the hyperparameters. However, this can be a very difficult task that requires a very in-depth understanding of the training process. As will be described next with reference to Figure 5 it is provided a hyperparameter tuning system according to various embodiments, which is configured to perform objective optimization in an automated manner, i.e., optimize a function of one or more metrics (e.g., a loss function). As Figure 4B shown, the metrics may include a stability metric, a regression error metric, a confidence score metric, a model size metric, a training time or execution time metric, an accuracy metric, or any combination thereof.
[0131] The stability metric ensures that the training process is stable, i.e., when making a small change to the training data (e.g., when adding or removing a training example), the predictions made by the model do not change fundamentally. The regression error metric minimizes the number of regressions of the model, i.e., the regression error corresponds to the model misclassifying an input for which a previous version of the model had made a correct classification. The confidence score metric ensures that the model predicts certain examples with high confidence (i.e., the model not only makes a correct prediction, but it makes the correct prediction with high confidence). The model size metric corresponds to the size of the trained model being within a user-defined threshold (e.g., 130 megabytes), while the training time metric corresponds to the amount of time used to train the model. The accuracy metric ensures that the trained model reaches a user-defined level of accuracy, e.g., an accuracy of 125% on certain validation datasets. In other words, when a machine learning model is trained on a specific training dataset, it reaches a specific target accuracy on a specific validation dataset. It can be understood that the selection of the training dataset and the validation dataset can span all use cases of the chatbot, i.e., in a range of applications, the size of the datasets ranges from very small datasets to very large datasets.
[0132] The hyperparameter tuning system of the present disclosure is configured to train a machine learning model for multiple domains (e.g., each domain consists of a training dataset and an evaluation dataset) and evaluate the performance of the machine learning model for one or more metrics. According to some embodiments, each dataset used to train and evaluate the machine learning model is assigned a domain weight, which indicates the importance of the dataset in training and evaluating the machine learning model. In other words, the weight assigned to the dataset corresponds to the degree of influence of the dataset on the training and evaluation of the machine learning model.
[0133] Further, the hyperparameter tuning system can be configured to assign weights to each metric used in the target optimization. Specifically, the weight assigned to a metric indicates the importance of that metric to the performance of the machine learning model. As will be described in detail below, the assignment of weights to metrics and different domains can be performed according to one or more policies governing the hyperparameter tuning system.
[0134] In some instances, the hyperparameter tuning system also enables the specification of one or more constraints when training a machine learning model. A constraint can be a requirement imposed on the machine learning model, i.e., a quality or characteristic that the user desires to achieve in the trained model. Constraints may be related to the training process itself. Constraints can be specified before starting model training. Thus, for a given set of constraints, the training infrastructure employs various automated techniques for the automatic identification, value setting, and tuning of hyperparameter values to train a machine learning model such that the trained machine learning model complies with and meets the set of constraints.
[0135] In some embodiments, the hyperparameter tuning system allows the validation of a machine learning model on a wide range of validation / test datasets. The hyperparameter tuning system associates specific target values with one or more metrics used to evaluate the performance of the machine learning model. A hyperparameter tuning objective function (e.g., a loss function) is constructed based on the target values of the one or more metrics. In certain instances, as will be referenced Figure 5 and described in detail below, the hyperparameter tuning system utilizes an asymmetric loss mechanism (i.e., not achieving the target is severely penalized compared to the case where reaching or exceeding the target is rewarded) to assign weights to different domains and / or metrics when validating the machine learning model.
[0136] Hyperparameter tuning system
[0137] Go to Figure 5 , which depicts a hyperparameter tuning system according to various embodiments. The hyperparameter tuning system 500 includes a dataset weight assignment unit 510, a metric selection and weight assignment unit 520, a constraint establishment unit 530, and a hyperparameter tuner 550. The hyperparameter tuner 550 includes an optimizer 551 (also referred to herein as a tuning unit) and a hyperparameter set 555.
[0138] The hyperparameter tuning system 500 is configured to train a machine learning model (e.g., with respect to Figures 1 to 3a model associated with the chatbot described), and evaluates the performance of the machine learning model based on one or more metrics. The dataset weight assignment unit 510 captures one or more domains, where each domain represents a knowledge base that a skill bot can use to conduct natural language conversations on a specific topic (e.g., PizzaBot can use the first domain to discuss ordering pizza, FlightBookingBot can use the second domain to help users book flights, FinancialBot can use the third domain to answer financial questions, InsuranceBot can use the fourth domain to provide insurance quotes), such as domain 1 505A, domain 2 505B, and domain K 505C, and assigns weights to each domain according to one or more policies 515. The weight assigned to a domain corresponds to the importance of the domain in training the machine learning model. Different weights are assigned to different domains so that the domains have an appropriate impact on the training of the machine learning model (i.e., according to their respective weights).
[0139] As an example, one of the policies in policy 515 may require the machine learning model to achieve a low regression relative to the baseline performance. Therefore, the dataset weight assignment unit 510 assigns a higher weight to the regression dataset (e.g., dataset 1 505A) than to another type of domain. As another example, domain 1 505A may correspond to a domain obtained from a first client of the hyperparameter tuning system 500, while domain 2 505B may correspond to a domain obtained from a second client different from the first client. The training data included in domain 1 and domain 2 may correspond to different types of user utterances related to the context (provided by their respective clients). Assuming that one of the policies 515 indicates that the first client is more important than the second client (e.g., the service level agreement (SLA) between the first client and the hyperparameter tuning system 500 is higher), the dataset corresponding to the first client (i.e., domain 1 505A) can be assigned a higher weight than the weight of domain 2 505B. Additionally, it can be understood that the system administrator of the hyperparameter tuning system 500 determines the policy 515 before training the machine learning model. The weighted domain 505 can be provided as a first input to the hyperparameter tuner 550. Additionally, it can be understood that the system administrator of the hyperparameter tuning system 500 can define a default or general policy, thus initializing all domain weights to the same value or level, e.g., all domain weights = 1.
[0140] The metric selection and weight assignment unit 520 selects one or more of the metrics 540, such as metric 1 540A, metric 5 540B, and metric M 540C, for evaluating the performance of the machine learning model for one or more domains 505. It should be noted that the metrics 540 correspond to metrics such as stability, accuracy, model size, regression error, etc., as described with respect to Figure 4BAs depicted and described. In one embodiment, the metric selection and weight assignment unit 520 selects one or more metrics from the set of available metrics 540 based on certain criteria, such as metric 1 540A, metric 5 540B, etc. For example, a first client of the hyperparameter tuning system 500 may want the machine learning model to focus on the accuracy parameter, while another client of the hyperparameter tuning system 500 may want the machine learning model to focus on another metric, such as the regression error metric. The requirements of different clients can be stored as one of the policies 515, and the metric selection and weight assignment unit 520 selects multiple metrics based on the policy to evaluate the performance of the machine learning model.
[0141] The metric selection and weight assignment unit 520 is further configured to assign weights to each selected metric. The weight assigned to a particular metric indicates the importance of that metric to the performance of the machine learning model. In one embodiment, the metric selection and weight assignment unit 520 assigns weights to metrics based on the importance level of the clients of the hyperparameter tuning system 500. For example, if the first client is more important than the second client (i.e., the first client has a higher SLA than the second client), then the metrics requested by the first client are assigned higher weights than the metrics requested by the second client. As another example, consider a machine learning model trained for domain detection. In this case, there are two metrics: in-domain recall and out-of-domain recall, which are used to evaluate the performance of the machine learning model. In such a domain detection model, it is generally desired that the model perform better in in-domain detection than in out-of-domain detection, i.e., above a certain threshold level. In this case, the weight assignment unit 520 assigns a higher weight to the in-domain recall metric than to the out-of-domain recall metric. The multiple weighted metrics are provided as a second input to the hyperparameter tuner 550.
[0142] Each metric 540 is associated with a corresponding set of specifications. The set of specifications includes multiple specification parameters that define or characterize the metric. As Figure 5 shown, metric 1 540A is associated with specification set 1 542A, metric 2 540B is associated with specification set 5 542B, and metric M 540C is associated with specification set M 542C. It can be understood that the set of specifications associated with a particular metric can be configured independently of the sets of specifications associated with other metrics. Referring to Figure 4C , the set of specifications for a metric can include: (1) a training data set, (2) a validation data set, (3) a metric definition, which defines a measure of how well the model meets the goal on the data set (e.g., the training data set and the validation data set), (4) the target score of the metric (i.e., the metric score that the model is expected to meet), (5) a penalty factor for the metric, and (6) a reward factor for the metric.
[0143] According to some embodiments, it may be allowed to share the training dataset and / or the validation dataset for metrics. However, if there are diverse training datasets and validation datasets among different metrics, the machine learning model is expected to produce more robust results. The specification parameters of each of the specification sets 542A, 542B, and 542C can be set to specific values to achieve the desired results. For example, for the metric of regression error, the corresponding specification set can be configured as follows:
[0144] 1. Model the training dataset based on a specific customer set (e.g., important customers). The validation dataset includes examples that were correctly classified by a previous machine learning model and are expected to be widely used by customers.
[0145] 2. The metric definition of the regression error is set to accuracy.
[0146] 3. The target score for accuracy is set to 95%.
[0147] 4. The penalty factor is set to 130, and the reward factor is set to 1. In this way, each percentage point below 95% is penalized 130 times that of a percentage point above 95%.
[0148] In one embodiment, for the stability metric, the corresponding specification set can be configured as follows:
[0149] 1. The training dataset is set to a smaller-sized dataset because a small dataset is expected to produce higher instability. The size of the validation dataset can be set to be much larger than the training dataset.
[0150] 2. The metric definition of stability is set to the standard deviation of the accuracy scores of the machine learning model on the validation dataset, e.g., when the machine learning model is trained 13 times.
[0151] 3. The target score is set to 8%, i.e., it is desired that the change in the accuracy scores of the machine learning model is at most 8%.
[0152] 4. The penalty factor is set to 13, and the reward factor is set to 1. In this way, the loss caused by each percentage point that does not reach the target score is 13 times the improvement of a percentage point that exceeds the target score.
[0153] In one embodiment, for the confidence score metric, the corresponding specification set can be configured as follows:
[0154] 1. The size range of the training dataset can vary from small to large and depends on the domain. The validation dataset can include in-domain examples belonging to its intended classification label.
[0155] 2. The metric definition of the confidence score metric is set to the proportion of sentences where the model confidence threshold is greater than 55% (i.e., the prediction is correct and the confidence is at least 55% higher than any other prediction).
[0156] 3. The target score is set to 90%, i.e., at least 90% of the examples need to be marked confidently.
[0157] 4. The penalty factor is set to 13 and the reward factor is set to 1.
[0158] It can be understood that the configuration of the above specification set is intended to be illustrative rather than restrictive. The system administrator can configure each of the specification sets in any other way based on different requirements. Further, different specification sets used in validating a machine learning model will be described below with reference to the hyperparameter tuner 550.
[0159] In some embodiments, the hyperparameter tuning system 500 enables a user (e.g., a system administrator) to specify one or more constraints for hyperparameter tuning of a machine learning model. The one or more constraints are specified via the constraint establishment unit 530. The one or more constraints are provided as a third input to the hyperparameter tuner 550. Each constraint is a requirement imposed on the hyperparameter tuning of the machine learning model, i.e., each constraint is a requirement that the trained machine learning model should satisfy. Given the one or more constraints, the hyperparameter tuner 550 is configured to identify the set of hyperparameters that affect each constraint, specify values for the identified hyperparameters, and iteratively tune the hyperparameters until the trained machine learning model satisfies each of the one or more constraints, as described below.
[0160] In one embodiment, for each of the one or more constraints, the hyperparameter tuner 550 identifies one or more hyperparameters from the hyperparameter set 555 that affect each constraint. The hyperparameter tuner 550 identifies one or more hyperparameters that affect a constraint by changing the value of the hyperparameter and determining whether the change in the value of the hyperparameter affects the value associated with the constraint. Further, it can be understood that the first set of hyperparameters that affect the first constraint may be different from the second set of hyperparameters that affect the second constraint.
[0161] After identifying one or more hyperparameters that affect each constraint, the hyperparameter tuner 550 specifies values for the hyperparameter set 555 and iteratively tunes the hyperparameters until each constraint is satisfied. As an example, consider the hyperparameter set 555 includes five hyperparameters: H = [h1, h2, h3, h4, h5]. Further, for the sake of illustration, consider the user specifies two constraints C1 and C2, where the hyperparameter tuner 550 has identified that hyperparameters h1 and h3 affect constraint C1, while hyperparameters h2, h3, and h5 affect constraint C2.
[0162] The optimizer 551 of the hyperparameter tuner 550 (also referred to herein as the tuning unit) assigns values to the hyperparameter set H, i.e., V(H) = [v(h1), v(h2), v(h3), v(h4), v(h5)], which is referred to herein as the configuration of the hyperparameters. It should be noted that the optimizer 551 assigns an initial configuration of the hyperparameters in a random manner. Further, for the constraint C1, the optimizer iteratively changes the values of the hyperparameters h1 and / or h3 (while keeping the values of h2, h4, and h5) until the constraint C1 is satisfied. Note that the constraint is satisfied when the values of the hyperparameters affecting the constraint meet the requirements imposed by the constraint. For the constraint C2, the optimizer 551 iteratively changes / modifies the values of the hyperparameters h2, h4, and / or h5 (while keeping the values of h1 and h3) until the constraint C2 is satisfied. Note that the optimizer performs the above iterations while training the machine learning model for one or more data sets. Specifically, as described below, the optimizer 551 determines the best configuration of the hyperparameters that satisfy each constraint while optimizing the objective function (e.g., cost function or loss function) of the machine learning model for multiple metrics. In one embodiment, examples of constraints specified by the user include the following constraints: -
[0163] - The inference latency for a given batch should be less than a specific threshold. For example, for a batch size of one, the latency should be less than 80 milliseconds.
[0164] - The maximum size of the trained model is required to be below a certain specified threshold (e.g., 13 MB).
[0165] - The training time of the model should be less than a certain user-specified time threshold. For example, the training time should be less than or equal to five minutes.
[0166] Further, in some embodiments, a priority order is assigned to multiple constraints specified by the user. For example, an importance level is assigned to each constraint to reflect the satisfaction of the constraint. In some embodiments, constraints can be specified such that the trained machine learning model must satisfy certain constraints, while satisfying other constraints is desirable but optional.
[0167] The optimizer 551 of the hyperparameter tuner 550 constructs / formulates the objective function to be optimized. The objective function can be a loss function or a cost function, which serves as a performance index for training / validating a machine learning model using one or more training / validation data sets. In one embodiment, the independent variable of the objective function is a set of hyperparameters associated with the machine learning model, which is optimized by the optimizer 551. The value of the objective function is a weighted combination of the differences between the actual values of each metric and the target values configured for each metric. The weight of each metric in the weighted combination depends on whether the metric exceeds the target value. In some instances, an asymmetric loss technique is utilized, where higher weights are assigned to metrics that fail to reach the target value (rather than exceed the target value).
[0168] For example, according to one embodiment, the domain score can be expressed as follows: If v is a vector of hyperparameter values, m i (v) represents the model performance of the i-th metric in the case of v, t i represents the target performance for the i-th metric, and p i and b i represent the penalty factor and the reward factor for the i-th metric respectively, then the score (i.e., the domain score) of a single domain (L(v)) can be expressed as:
[0169] L(v) = Σib i max(mi(v) - t i , 0) - p i max(t i - m i (v), 0) - And the objective function or loss function (Ftuning) for calculating the scores of multiple domains (i.e., the objective score) can be expressed as:
[0170] Ftuning = w_0 * f(D_0) + w_1 * f(D_i) +... + w_N * f(D_N)
[0171] where f(D_i) (i = 1, 2,..., N) is the domain score calculated for each domain using the objective function (L(v)), and w_0, w_1,... w_N are the domain weights to be applied to each domain score. In this example, the goal of hyperparameter tuning is to maximize Ftuning. Other objective functions are envisioned, and for example, as described in further detail herein, the objective function can be modified to include: scores that include regression scores (e.g., regression scores), or scores that include improvement scores (e.g., improvement scores). The scores can be formulated to capture instance-level regression or improvement. An instance can be an input classified by a machine learning model, and the scores formulated to capture instance-level regression or improvement can capture changes in model performance that are not reflected in the accuracy score. In
[0172] The optimizer 551 tunes a set of hyperparameters 555 associated with a machine learning model to optimize an objective function on one or more metrics (e.g., to obtain a minimum of a loss function). The optimizer can utilize one or more hyperparameter tuning methods to tune the set of hyperparameters, such as grid-based methods (e.g., constructing models for each possible combination of all provided hyperparameter values, evaluating each model, and selecting the model that produces the best result), gradient search methods (e.g., computing gradients for the hyperparameters and then using gradient descent to optimize the hyperparameters), and Bayesian methods (e.g., defining a model constructed with hyperparameters λ, which is scored as v according to certain evaluation metrics after training; then using the previously evaluated hyperparameter values to compute the posterior expectation of the hyperparameter space; selecting the best hyperparameter value as the next model candidate according to this posterior expectation; and iteratively repeating this process until convergence to an optimal value).
[0173] References in this article Figure 6 and Figure 7 describe details regarding hyperparameter tuning. In this way, the hyperparameter tuner 550 trains / validates a machine learning model by tuning a set of hyperparameters to achieve optimal performance for different weighted metrics. In some instances, this process is performed while ensuring that each of one or more constraints is satisfied. In other words, the hyperparameter tuning system 500 performs optimization of multiple weighted metrics by training a machine learning model on different weighted domains, while optionally supporting one or more user-specified constraints. After optimizing the machine learning model for multiple metrics, the hyperparameter tuning system 500 outputs the trained / validated ML model and the configuration of the hyperparameters that optimize the objective function and, in some instances, satisfy one or more constraints. In the above embodiment, the optimizer 551 constructs and optimizes the objective function. It should be noted that the configuration of the hyperparameter tuner 550 as described above is not intended to limit the scope of the present disclosure. For example, the hyperparameter tuner 550 can include an objective function formulation unit (not shown) that formulates the objective function optimized by the optimizer 551.
[0174] Hyperparameter tuning techniques
[0175] Figure 6 Depicts a simplified flowchart 600 that depicts a training process performed by a hyperparameter tuning system (e.g., regarding Figure 5 the hyperparameter tuning system 500 described). Figure 6 The depicted processing can be implemented in software (e.g., code, instructions, programs), hardware, or a combination thereof executed by one or more processing units (e.g., processors, cores) of the corresponding system. The software can be stored on a non-transitory storage medium (e.g., on a memory device). Figure 6The methods presented and described below are intended to be illustrative and not restrictive. Although Figure 6 various processing steps are depicted as occurring in a particular order or sequence, this is not intended to be restrictive. In certain alternative embodiments, the steps may be performed in a different order or some steps may also be performed in parallel.
[0176] At block 610, a domain for training a machine learning model is obtained. For example, a user or a subsystem may provide a domain for training a machine learning model. Each domain may be associated with one or more training data sets. At block 620, the hyperparameter tuning system assigns weights to each obtained domain according to a policy. It is noted that different weights are assigned to different domains such that the domains have an appropriate impact on the training of the machine learning model (i.e., according to their respective weights). In some instances, the policy specifies that all domain weights are initialized to the same value, e.g., domain weight for each domain = X.
[0177] At block 630, one or more metrics are selected to evaluate the performance of the machine learning model on the obtained domains. For example, a user (e.g., a system administrator) selects one or more metrics as Figure 4B depicted to evaluate the performance of the machine learning model. For each selected metric, a weight may be assigned to the metric according to another policy to indicate the importance of the metric to the performance of the machine learning model (block 640). At optional block 650, a user establishes one or more constraints. Each constraint is a quality or characteristic that the user desires to achieve in the trained machine learning model. In other words, each constraint is a requirement imposed on the machine learning model.
[0178] At block 660, the process formulates / constructs a function (i.e., an objective function) based on the input weighted metrics and hyperparameter set. In one embodiment, the objective function is a loss function or a cost function that serves as a performance metric for training the machine learning model within one or more domains.
[0179] At block 670, the process iteratively tunes the hyperparameter set associated with the machine learning model in order to optimize the machine learning model for one or more metrics (e.g., to obtain the best value of the objective function). For example, when training the machine learning model on the weighted domains, one or more hyperparameters that affect one or more constraints and / or functions are identified by changing the values of the hyperparameters and determining whether the change in the values of the hyperparameters affects the values associated with the function and / or constraint.
[0180] In some embodiments, during the process of tuning hyperparameters, the process evaluates the current configuration of the hyperparameters (i.e., the values of the hyperparameters), the value of the function, and determines whether the model converges and / or whether the current configuration satisfies each of one or more constraints. Convergence is the point in model training after which the change in the learning rate becomes lower and the error produced by the model in training reaches a minimum value (e.g., the model reaches a state during training - the loss stabilizes within the error range around the final value). When the model converges, further training does not significantly reduce the model error. Convergence can be divided into two types: global convergence or local convergence. Thus, convergence can be the point in model training after which the error or performance is within the error range around the local / global minimum. Mathematically, convergence can be regarded as the study of series and sequences. When the series is a convergent series, the model can be considered to be convergent. If at least one of the constraints is violated and / or the value of the function is not optimal, the tuning process modifies the values of one or more hyperparameters to obtain a new configuration of the hyperparameters and continues to train the machine learning model based on the new configuration. Note that it can be determined whether the model converges and / or whether the current configuration violates a specific constraint by determining whether the values of one or more hyperparameters affecting the model and the constraints satisfy the requirements imposed by regression analysis (e.g., no regression across all domains) and / or the constraints. Further, the value of the function (indicating the performance of the machine learning model with respect to the current configuration) is determined to be optimal by comparing the value of the function with the new values of the function obtained via different configurations of the hyperparameters.
[0181] In this way, the tuning process iterates through the hyperparameter value space until a configuration is achieved that produces the optimal value of the function (and optionally does not violate any constraints). Additionally, it can be understood that the process of tuning hyperparameters can start from an initial configuration of the hyperparameters assigned in a random manner. Further, when searching the hyperparameter value space to obtain a new hyperparameter configuration, the tuning process can implement one of a random search method, a Bayesian search method, a branch and bound method, a grid search method, a genetic algorithm, etc. When the machine learning model is optimized, the machine learning model (and the values of the hyperparameters that achieve the optimized machine learning model) is output to the user as the trained machine learning model.
[0182] Figure 7 Depicts a simplified flowchart 700 that depicts a verification process performed by a hyperparameter tuning system (e.g., with respect to Figure 5 the hyperparameter tuning system 500 described). Figure 7 The depicted processing can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., on a memory device). Figure 7The methods presented and described below are intended to be illustrative and not restrictive. Although Figure 7 various processing steps are depicted as occurring in a particular order or sequence, this is not intended to be restrictive. In certain alternative embodiments, the steps may be performed in a different order or some steps may also be performed in parallel.
[0183] At block 710, one or more metrics are selected to evaluate the performance of a machine learning model, and hyperparameters of the machine learning model are tuned with respect to these metrics. At block 720, a set of specifications associated with each selected metric is configured according to criteria. Specifically, values are assigned to the specification parameters included in the set of specifications based on the criteria. For example, if it is desired to reduce the regression error, the set of specifications associated with the regression error metric is configured as follows: a low value is set for the target score, and a high value is set for the penalty factor corresponding to the regression error metric. As an example, the target accuracy improvement may be set to 120%, i.e., it is expected that at least 120% of the previously correctly labeled training examples are correctly labeled by the current version of the machine learning model. Further, setting a higher penalty factor, e.g., the penalty factor is 130, means that each percentage point below 120% (i.e., the performance of the machine learning model) is penalized 130 times as much as the percentage points above 120%. When configuring the set of specifications for the regression error metric in this way, the regression error dominates the objective function (i.e., the loss function), and Figure 5 the hyperparameter tuner performs hyperparameter tuning that minimizes the regression error.
[0184] At block 730, a metric score is calculated for each metric. Specifically, one or more selected metrics are evaluated on a validation dataset to produce a metric score for each metric. At block 740, the metric score calculated for each metric is compared with the corresponding target score of the metric. In one embodiment, the difference between the metric score and the target score of the metric is calculated. If the metric score is higher than the target score, the difference is multiplied by the reward factor associated with the metric. However, if the metric score is lower than the target score, the difference is multiplied by the penalty factor. At block 750, an objective function (e.g., a loss function) is formulated based on the processing performed at block 740. Specifically, the loss function is determined as the sum of the differences (between the metric score and the target score) multiplied by the corresponding reward factor or penalty factor. Thereafter, the process moves to block 760 to optimize the formulated objective function.
[0185] At block 760, the hyperparameter tuner iteratively tunes hyperparameters associated with a machine learning model to optimize a formulated objective function (e.g., obtain a minimum of a loss function) formulated at block 750. In one embodiment, when tuning hyperparameters, the hyperparameter tuner evaluates the result of the loss function (e.g., a first objective score) for the current configuration (i.e., the values of the hyperparameters). The hyperparameter tuner further determines whether the result of the loss function is the best by comparing the result of the loss function (e.g., a first objective score) with the result of the loss function (e.g., a second objective score) obtained via different configurations of the hyperparameters.
[0186] In this way, the tuning process iterates through the hyperparameter value space until a configuration is achieved that yields the best value of the formulated objective function. The process of tuning hyperparameters can start with an initial configuration of hyperparameters assigned in a random manner. Additionally, the tuning process can implement a search algorithm to explore the hyperparameter space to obtain new hyperparameter configurations. The hyperparameter tuner can utilize one of a random search method, a Bayesian search method, a branch and bound method, etc. to search the hyperparameter space. When the objective function is optimized (i.e., the minimum of the loss function is achieved), the hyperparameter tuner provides the validated machine learning model and the values of the hyperparameters that achieve the optimized objective function as output to the user.
[0187] Objective Function Optimization in Goal - Based Hyperparameter Tuning
[0188] Techniques for optimizing an objective function in goal-based hyperparameter tuning are disclosed herein. The techniques include modifying the objective function to: minimize instance-level regression, handle data points with unstable predictions, incorporate a regression acceptance level for one or more domains, or any combination thereof.
[0189] Objective Function Optimization
[0190] Attempting to find a unified hyperparameter configuration for thousands of different domains or experiments is a difficult hyperparameter tuning task. To perform this task, it is important to consider which domains are more important and which are less important to achieve a specified goal, e.g., the accuracy performance of a model. Domain weights can be defined as part of the tuning objective function, indicating the importance of a domain; and, domains with higher weights can receive more attention during hyperparameter tuning than domains with lower weights. As discussed herein, for hyperparameter tuning that optimizes for N domains, e.g., D = {D_0, D_1,..., D_N}, the objective function can be defined as follows:
[0191] F_{tuning} = w_0 * f(D_0) + w_1 * f(D_1) + … + w_N * f(D_N)
[0192] Among them, f(D_i) (i = 0, 1,..., N) are domain scores calculated for each domain, and w_0, w_1,..., w_N are domain weights to be applied to each domain score.
[0193] The domain score is a metric for measuring the model performance, and is calculated based on the improvement and regression of the model M_t compared to the baseline score of a certain domain in trial t. For example, f(D_i, t) =
[0194] ·improvement_weight * [improvement_score] = (score_{trial_t}(M_t, D_i) - score_{baseline}(M_t, D_i)) --------- if score_{troal_t}(M_t, D_i) >= score_{baseline}(M_t, D_i)
[0195] -(negative)
[0196] ·regression_weight * [regression_score] = (score_{trial_t}(M_t, D_i) - score_{baseline}(M_t, D_i)) --------- if score_{trial_t}(M_t, D_i) < score_{baseline}(M_t, D_i)
[0197] As discussed in this article, score_{trial_t}(M_t, D_i) for each domain can be calculated as different evaluation metrics, such as accuracy, F1, precision, recall, etc. Hyperparameter tuning largely depends on the domain scores calculated for each domain to optimize the hyperparameters. For example, the ratio between regression_weight and improvement_weight indicates whether hyperparameter tuning focuses on tuning the hyperparameters to improve performance (while allowing a certain degree of regression); or whether hyperparameter tuning focuses on tuning the hyperparameters to minimize regression (while allowing less overall improvement).
[0198] However, there are still some challenges with objective-based hyperparameter tuning because tuning heavily relies on domain scores computed for each domain to optimize hyperparameters. First, it has been found that dataset-level domain evaluation scores do not always reflect true improvements and regressions in some cases. More specifically, domain scores indicate improvements and regressions at the dataset level, but not necessarily at the instance level. For example, the overall accuracy of the model in the current trial may be the same as the baseline accuracy; however, there are more than 1000 instances with correct predictions and more than 1000 instances with incorrect predictions. Since instance-level improvements and instance-level regressions cancel each other out, the overall accuracy remains unchanged. However, more than 1000 instances with incorrect predictions may pose problems for model deployment and application. To overcome this challenge, the objective function can be modified to include the number of correctly predicted instances and incorrectly predicted instances when computing the domain score. This allows hyperparameter tuning to be performed in a way that takes into account not only dataset-level improvements and regressions but also instance-level improvements and regressions. The advantage of this tuning technique is that instance-level regressions can be minimized during objective-based hyperparameter tuning, improving the overall performance of the model and the performance of the computing system running the model (e.g., increasing the speed or efficiency of the underlying computing device and / or reducing the processing requirements or memory usage of the underlying computing device).
[0199] To modify the objective function to include the number of correctly predicted instances and incorrectly predicted instances when computing the domain score, the objective function is formulated to normalize instance-level improvements and instance-level regressions over the total number of instances to obtain an improvement_score and a regression_score, as follows:
[0200] f(D_i, t) = improvement_weight * improvement_score - regression_weight *
[0201] .regression_score
[0202] improvement_score = count(correct{trial_t}(M_t, D_j) ∩
[0203] οincorrect{baseline}(M_b, D_j)) / count(instances)
[0204] regression_score = coumt(incorrect{trial_t}(M_t, D_i) ∩
[0205] οcorrect{baseline}(M_b, D_i)) / count(instances)
[0206] Where correct{trial_t}(M_t, D_j) is the set of instances correctly predicted by the model for domain D_j in trial t; incorrect{baseline}(M_b, D_j) is the set of instances incorrectly predicted by the baseline model for domain D_j; incorrect{trial_t}(M_t, D_i) is the instance incorrectly predicted by the model for domain D_j in trial t; correct{baseline}(M_b, D_i) is the set of instances correctly predicted by the baseline model for domain D_j; and count() calculates the total number of instances in the instance set. Correct and incorrect can be determined by comparing the predictions of the model with the ground truth established for each example or instance.
[0207] By way of non-limiting example:
[0208] 1. Given a test data set for domain A, sized 5000 utterances (i.e., instances).
[0209] 2. The accuracy performance of the baseline model can be 60%, with 3000 correct predictions and 2000 incorrect predictions;
[0210] 3. For the model in the current trial T, 2000 utterances out of [the 3000 utterances correctly predicted by the baseline model] can be correctly predicted; additionally, 1000 utterances out of [the 2000 utterances incorrectly predicted by the baseline model] can be correctly predicted; → thus, 3000 can be correctly predicted and 2000 can be incorrectly predicted, which is equivalent to an accuracy performance of 60%.
[0211] 4. However, 1000 utterances out of [the 3000 utterances correctly predicted by the baseline model] are now incorrectly predicted, and these are regarded as "instance-level regression". 1000 instance-level regressions may pose problems for the deployment and application of the model. However, according to the "data set-level" accuracy (60% → 60%), 0% regression is observed.
[0212] 5. Correspondingly, 1000 utterances out of [the 2000 utterances incorrectly predicted by the baseline model] can now be correctly predicted; and are regarded as "instance-level improvement".
[0213] 6. The "instance-level improvement" (e.g., 1000) and "instance-level regression" (e.g., -1000) can be normalized over the total number of instances (e.g., 5000) to obtain the improvement_score and regression_score.
[0214] a. For example, improvement_score = 1000 / 5000 = 0.2; regression_score = -1000 / 5000 = -0.2.
[0215] b. If it is assumed that the importance of regression is 100 times that of improvement, e.g., improvement_weight:regression_weight = 100:1; then f(D_i, t) = 100.0 * (-0.2) + 1.0 * (0.2) = -19.8
[0216] Second, it has been found that due to the statistical nature of machine learning models or deep learning models, the domain score is partially calculated based on data points that affect unstable prediction results; and when tuning is performed based on these data points, the tuner is forced to chase random noise in the later stages of tuning, which is not ideal. For example, a 0% regression acceptance rate can be configured for all customer data sets. However, during tuning, it is observed that some customer data sets may not be perfect (i.e., outlier detection tools and manual analysis indicate that the predictions for N test cases are unstable); therefore, these N test cases should be excluded from hyperparameter tuning so that the hyperparameter tuner does not chase random noise or incorrect test cases. To overcome this challenge, the objective function can be modified to exclude unstable instances from the calculation of the domain score. This allows hyperparameter tuning to be performed in a way that takes into account unstable instances. The advantage of this tuning technique is that objective-based hyperparameter tuning does not chase random noise or incorrect test cases when optimizing hyperparameters, improves the overall performance of the model, and improves the performance of the computing system running the model (e.g., increases the speed or efficiency of the underlying computing device and / or reduces the processing requirements or memory usage of the underlying computing device).
[0217] To modify the objective function to exclude unstable instances, the objective function is formulated to subtract the number of unstable instances and normalize the instance-level improvement and instance-level regression over the total number of instances to obtain the improvement_score and regression_score, as follows:
[0218] f(D_i, t) = improvement_weight * improvement_score - regression_weight *
[0219] .regression_score
[0220] improvement_score = (count(correct{trial_t}(M_t, D_j) ∩
[0221] οincorrect{baseline}(M_b, D_j)) - count(unstable_instances)) / count(instances) regression_score = (count(incorrect{trial_t}(M_t, D_i) ∩
[0222] οcorrect{basedline}(M_b, D_i)) - coumt(umstable_instances)) / count(instances)
[0223] Unstable instances can be determined and calculated by running an instability experiment on the model. For example, the model can run a predetermined number of times (e.g., 10 times) on the same example or instance, and those instances that have unstable prediction results in the predetermined number of runs are determined and counted as unstable. Unstable prediction results are prediction results that are different from each other for the same instance run on the same model (e.g., 9 times predicted as "red" and 1 time predicted as "blue" in 10 runs, which indicates instability). In some instances, the determination of instability is thresholded such that a predetermined number of different prediction results from each other must exist to determine the existence of instability in the instance (e.g., 6 times predicted as "red" and 4 times predicted as "blue" in 10 runs, and if the threshold is set to be equal to or greater than 3, then the experiment will indicate instability).
[0224] By way of non - limiting example:
[0225] 1. Continuing with the above example, given 1000 "instance - level regressions" and 1000 "instance - level improvements" for domain A.
[0226] 2. Considering the instability experiment, 200 of the regressions are determined to potentially be due to the quality of the test examples, e.g., the utterances are very close to the decision boundary.
[0227] 3. In such an instance, 200 unstable "instance - level regressions" can be excluded from the calculation of the objective function (e.g., - 1000 - (-200)= - 800), so that hyper - tuning does not chase this random noise.
[0228] 4. Additionally, "instance-level improvement" (e.g., 1000) and "instance-level regression" (e.g., -800) can be normalized over the total number of instances (e.g., 5000) to obtain improvement_score and regression_score.
[0229] a. For example, improvement_score = 1000 / 5000 = 0.2; regression_score = -800 / 5000 = -0.16.
[0230] b. If it is assumed that the importance of regression is 100 times that of improvement, e.g., improvement_weight:regression_weight = 100:1; then f(D_i, t) = 100.0 * (-0.16) + 1.0 * (0.2) = -15.8
[0231] Finally, it has been found that an acceptable regression level (e.g., regression ratio, m) can be defined for different domains according to various goals of the user (e.g., business requirements). For example, 0% regression can be configured or accepted on domain A; while 3% regression can be configured or accepted on domain B. To achieve this enhancement, the objective function can be modified to include a parameter for the acceptable regression level in one or more domains. For example, an acceptable regression ratio level of 2% can be configured or accepted for the training-test-split dataset; but an acceptable regression ratio level of 0% can be configured or accepted for the customer dataset as the "custom set" has a greater impact on the user. This allows hyperparameter tuning to be performed in a way that takes into account the (multiple) acceptable regression levels of one or more domains. The advantage of this tuning technique is that goal-based hyperparameter tuning will allow for a more fine-grained configuration in terms of domain importance assignment during hyperparameter tuning. For example, some users may be more concerned with regression while some users may be more concerned with improvement. It incorporates the regression acceptance levels of the (multiple) domains while optimizing the hyperparameters, improving the overall performance of the model, and improving the performance of the computing system running the model (e.g., increasing the speed or efficiency of the underlying computing device and / or reducing the processing requirements or memory usage of the underlying computing device).
[0232] To modify the objective function to incorporate the regression acceptance levels of the (multiple) domains, the objective function is formulated to include a parameter for the acceptable regression ratio m for one or more domains, as follows:
[0233] .f(D_i, t) = improvement_weight * improvement_score - regression_weight *
[0234] (If - regression_score < m, then it is 0, otherwise it is regression_score)
[0235] improvement_score = count(correct{trial_t}(M_t, D_j) ∩
[0236] οincorrect{baseline}(M_b, D_j)) / count(instances)
[0237] regression_score = count(incorrect{trial_t}(M_t, D_i) ∩
[0238] οcorrect{basedline}(M_b, D_i)) / count(instances)
[0239] Note that although the above objective functions are modified to be built on each other and incorporate the number of correctly predicted instances and incorrectly predicted instances in the calculation of domain scores, exclusion of unstable instances, and / or regression acceptance levels of (multiple) domains, it should be understood that each modification of the objective functions described herein can be incorporated into the objective function individually or in any combination.
[0240] By way of non - limiting example:
[0241] 1. Continuing the above example, among 5000 test cases in domain A, given 1000 (20%) "instance - level regressions" and 1000 (20%) "instance - level improvements".
[0242] 2. Considering business requirements, domain A requires 0% regression acceptance.
[0243] a. regression_score = - 1000 / 5000 = - 0.2, because 20% > 0%
[0244] 3. Additionally, among 2000 test cases, domain B can have 30 (1.5%) "instance - level regressions" and 50 (2.5%) "instance - level improvements".
[0245] 4. Considering business requirements, domain B requires 3% regression acceptance
[0246] a. regression_score = 0, because 1.5% < 3%
[0247] 5. On the other hand, if the regression acceptance of domain B is 1%, the regression_score of domain B will be regression_score = -30 / 2000 = -0.015
[0248] Experimental Example
[0249] The systems and methods implemented in various embodiments can be better understood by referring to the following experimental examples. Consider three types of tests that a model developer wishes to ensure no regression with high priority: simple, user regression, and QA user regression.
[0250] · The simple test type includes very simple test cases that are expected to guarantee a high accuracy score so that these simple test cases do not fail during the production / deployment of the model.
[0251] · User (e.g., customer) regression and QA user regression include test cases shared by users that are expected to guarantee no regression in the production / deployment of the model. For each failure, an appropriate reason is expected.
[0252] By applying the experiments with the objective function modifications described herein in goal-based hyperparameter tuning, it can be observed that the total number of stable regressions for each test type can be minimized while maintaining or improving accuracy. As shown in Table 1: the number of stable regressions for the simple test type decreased to 4; the number of stable regressions for the customer regression test type decreased to 35; and the number of stable regressions for the QA customer regression test type decreased to 27.
[0253] Table 1:
[0254]
[0255] Objective Function Optimization and Tuning Workflow
[0256] Figure 8 is a simplified diagram of the tuning workflow 800 according to various embodiments. The objective function optimization workflow can be used to evaluate a proposed set of hyperparameter values by calculating an objective score for each proposed set. An objective score can be calculated using an objective function that includes a domain score and a domain weight for each domain in the hyperparameter search space, as described in detail with respect to Figures 4A to 7 The domain score is calculated based on the improvement and regression compared to the baseline score of a certain domain in trial t of model M_t, e.g., f(D_i,t), as described in further detail above.
[0257] More specifically, turning to the objective function optimization workflow 800, at block 805, a domain for training a machine learning model is obtained. For example, a user or subsystem may provide a domain for training a machine learning model. Each domain may be associated with one or more training data sets. At block 810, one or more metrics are selected to evaluate the performance of the machine learning model over the obtained domain. For example, a user (e.g., a system administrator) selects one or more metrics as depicted in Figure 4B to evaluate the performance of the machine learning model. For each selected metric, a weight may be assigned to the metric according to a policy to indicate the importance of the metric to the performance of the machine learning model. At optional block 820, a user establishes one or more constraints. Each constraint is a quality or characteristic that the user desires to achieve in the trained machine learning model. In other words, each constraint is a requirement imposed on the machine learning model. At block 825, the user makes a modification or enhancement to the objective function (e.g., minimizing instance-level regression, incorporating a regression acceptance level for one or more domains, handling data points with unstable predictions, or any combination thereof). Each modification or enhancement is a quality or characteristic that the user desires to achieve in the goal-based hyperparameter tuning. In other words, each modification or enhancement is a requirement imposed on the tuning process via the objective function.
[0258] At block 830, the process formulates / constructs an objective function based on the domain, metrics, optional constraints, and modifications or enhancements. As described herein, the objective function is an equation that includes domain weights and domain scores for each domain, where the domain includes a training data set and an evaluation data set. In one embodiment, the objective function is a loss function or a cost function that serves as a performance metric for training a machine learning model within one or more domains. The independent variables of the objective function are a set of hyperparameters associated with the machine learning model, which are optimized by an optimizer (e.g., the optimizer 551 described with respect to Figure 5 . The value of the objective function is a weighted combination of the differences between the actual values of each metric and the target values configured for each metric. The weight of each metric in the weighted combination depends on whether the metric exceeds the target value. Specifically, an asymmetric loss technique is utilized, where a higher weight is assigned to a metric in the case of failing to reach the target value (compared to the case of exceeding the target value).
[0259] For example, the objective function (e.g., the weighted objective function) may be expressed as follows: If v is a vector of hyperparameter values, m i (v) represents the value of the i-th metric in the case of v, t i represents the target value of the i-th metric, and p i and b i represent the regression factor and the improvement factor (e.g., weight) of the i-th metric, respectively, then the objective function or loss function (L(v)) may be expressed as Equation (1):
[0260] L(v) = ∑ i p i max(t i - mi(v), 0) - b i max(m i (v) - t i , 0) (1)
[0262] Other objective functions are envisioned and, for example, the objective function can be modified to include: a score that includes a regression score (e.g., regression score), or a score that includes an improvement score (e.g., improvement score). The score can be formulated to capture regression or improvement at the instance level. An instance is an example input into a machine learning model, and a domain score formulated to capture regression or improvement at the instance level can capture changes in model performance not reflected in the accuracy score. In some cases, the domain score can include a normalized change in correctly predicted instances by dividing the change in correctly predicted instances by the total number of instances.
[0263] Unstable instances or instances within a decision boundary threshold distance may cause the model to chase noise, and the domain score can be formulated to exclude unstable instances. Unstable instances can include unstable correct instances, correctly classified instances that are too close to the decision boundary, or unstable incorrect instances (including incorrectly classified instances that are too close to the decision boundary). By subtracting the number of unstable instances from the number of correctly classified or incorrectly classified instances and then dividing that amount by the total number of instances, a normalized difference in the number of correctly classified or incorrectly classified instances can be generated.
[0264] In some cases, different domains of a hyperparameter set may not be equally important, and the objective function can be formulated to allow for different amounts of regression in different domains. A regression acceptance threshold (e.g., percentage regression) can be established for one or more domains of the hyperparameter set. The regression acceptance threshold can be based on various objectives such as business requirements.
[0265] For example, the objective function (e.g., weighted objective function) can be expressed as follows: If v is a vector of hyperparameter values, p_score ij (v) represents the regression score of the i-th metric in the j-th domain, b_score ij (v) is the improvement score of the i-th metric in the j-th domain, and p ij and b ij represent the regression factor and improvement factor (e.g., weight) of the i-th metric in the j-th domain, respectively, then the objective function or loss function (L(v)) can be expressed as Equation (2):
[0266] L(v) = ∑ij p ij *p_score ij (v)-b ij *b_score ij (v) (2)
[0268] In another example, the cost function can be expressed as follows: If v is a vector of hyperparameter values, p_score ij (v) represents the regression score of the i-th metric in the j-th domain, b_score ij (v) is the improvement score of the i-th metric in the j-th domain, p ij and b ij represent the regression factor and improvement factor of the i-th metric in the j-th domain respectively, and r ij is the regression acceptance threshold of the i-th metric in the j-th domain, then the objective function or loss function (L(v)) can be expressed as Equation (3):
[0269]
[0270] In some examples, the improvement score can be formulated to capture instance-level regression or improvement as follows: If v is a vector of hyperparameter values, b_score ij (v) is the improvement score of the j-th domain and the i-th metric in the case of v, correct ij (v) represents the number of instances correctly classified using the i-th metric in the j-th domain in the case of v, incorrect_t ij is the target (baseline) number of incorrect instances of the j-th domain and the i-th metric, and instances j can be the number of instances in the j-th domain, then the improvement score can be expressed as Equation (4):
[0271] b_score ij (v) = [correct ij (v) ∩ incorrect_t ij / instances j (4)
[0272] In some examples, the regression score can be formulated to capture instance-level regression or improvement as follows: If v is a vector of hyperparameter values, p_score ij (v) is the regression score of the j-th domain and the i-th metric in the case of v, incorrect ij (v) represents the number of instances misclassified using the i-th metric in the j-th domain in the case of v, correct_t ijis the target (baseline) number of correct instances for the j-th domain and the i-th metric, and instances j can be the number of instances in the j-th domain, then the improvement score can be expressed as Equation (5):
[0273] p score ij (v)=[incorrect ij (v)∩correct t ij / lnstances j (5)
[0275] In some instances, the improvement score can be formulated to resist the influence of unstable instances as follows: If v is a vector of hyperparameter values, b_score ij (v) is the improvement score for the j-th domain and the i-th metric in the case of v, correct ij (v) represents the number of instances correctly classified in the j-th domain using the i-th metric in the case of v, incorrect_t ij is the target number of incorrect instances for the j-th domain and the i-th metric, instances j can be the number of instances in the j-th domain, and unstable ij can be the number of unstable instances, then the improvement score can be expressed as Equation (6):
[0276] b_score ij (v)=[(correct ij (v)∩incorrect_t ij ) - unstable ij / instances j (6)
[0278] In some instances, the regression score can be formulated to resist the influence of unstable instances as follows: If v is a vector of hyperparameter values, p_score ij (v) is the regression score for the j-th domain and the i-th metric in the case of v, incorrect ij (v) represents the number of instances misclassified in the j-th domain using the i-th metric in the case of v, correct_t ij is the target number of correct instances for the j-th domain and the i-th metric, instances j can be the number of instances in the j-th domain, and unstable ij can be the number of unstable instances, then the improvement score can be expressed as Equation (7):
[0279] p_score ij (v) = [(incorrect ij (v) ∩ correct_t ij ) - unstable ij / instances j (7)
[0281] At block 835, the values of the domain weights to be used in the objective function formulated at block 830 are initialized. During initialization or as part of a periodic domain weight update, the values of the domain weights are assigned to the objective function, and the domain score is calculated by performing the evaluation process described at block 845 using the hyperparameter values suggested at block 840. In some embodiments, all domain weights may be initialized to a uniform distribution (e.g., according to a general strategy, all domain weights may be set to 1.0); and the uniform distribution may be stored in a database (as Figure 8 shown). In a sequential tuning technique, the domain weights may be initialized in the first trial, and after initialization, the domain weights in the database may be updated every K trials. For example, the update may be performed when the trial counter reaches a threshold (e.g., determining whether trial number n % K == 0 holds). In other instances, the domain weights in the database may be updated after a given evaluation time period (e.g., a time period threshold, such as six hours).
[0282] At block 840, hyperparameter values for trial i are searched for and suggested for the model. As described herein, hyperparameters are parameters used to control the model architecture and the learning process. The search and selection of hyperparameters may be implemented using hyperparameter tuning algorithms (e.g., random search method, Bayesian search method, branch and bound method, grid search method, genetic algorithm, etc.). The hyperparameter tuning algorithm uses the objective scores calculated from previous trials and the corresponding trial hyperparameter values (e.g., based on a matrix of previous objective scores and the corresponding values of the hyperparameter set) to determine a new set of hyperparameter values for trial i, as described in detail with respect to Figures 4A to 7 this.
[0283] At block 845, trial i is evaluated. During trial evaluation, the model is trained on one or more training data sets associated with the domain using the hyperparameter values suggested for trial i, the trained model is evaluated on one or more evaluation data sets associated with the domain using the hyperparameter values suggested for trial i, and the domain score f(D_i) is calculated for the domain based on the training and / or evaluation. This process is repeated for the model of each domain. The domain score f(D_i) is calculated using the (multiple) objective function formulated / constructed at block 830.
[0284] At block 850, an objective score is calculated for trial i. The objective score is calculated using the domain scores from block 845 and the domain weights saved in the database for each domain. For example, the objective score can be calculated by retrieving the domain weights from the database and substituting the domain weights and domain scores of all domains from block 845 into a combined objective function (such as a linear combination or weighted linear combination of each domain score) to calculate the objective score for trial i.
[0285] At decision block 855, the trial counter or timer is checked to determine if a threshold has been reached. If the trial counter or timer has not reached the threshold, the tuning technique will continue with the next trial i+1 by suggesting a new set of hyperparameter values at block 850. If the trial counter or timer has reached the threshold, the tuning technique will proceed to block 860, where the domain weights are updated and the objective scores of the previous trials are recalculated based on the updated domain weights.
[0286] At block 860, the domain weights are updated and the objective scores of the previous trials are recalculated based on the updated domain weights. The domain weights of each domain can be updated if the domain scores meet specific requirements. For example, in threshold-based domain weight adjustment, for a given domain, if the domain score is less than the domain score threshold, its domain weight can be adjusted. In other embodiments, the domain weights can be adjusted for multiple domains with the M lowest domain scores (e.g., the domain weights can be adjusted for the M domains with the worst performance in the trial).
[0287] After adjusting the domain weights, the updated domain weights are saved to the database, the objective scores of the previous trials are recalculated based on the updated domain weights, and the recalculated objective scores of the previous trials are saved to the database. The objective scores of the previous trials are updated using the updated domain weights because the hyperparameter tuning algorithm uses the previous and current objective scores (and all relevant hyperparameter values) to evaluate and select new hyperparameters for each trial. Updating the objective scores of the previous and current trials using the updated domain weights effectively updates all the correlations between the previous objective scores and the evaluated hyperparameter values, enabling the hyperparameter tuning algorithm to better evaluate and select new hyperparameters for each trial. In some instances, the objective scores of all previous trials are recalculated based on the updated domain weights. In other instances, the objective scores of at least one previous trial are recalculated based on the updated domain weights.
[0288] After updating the domain weights and recalculating the objective scores of previous trials, the tuning technique can proceed to the next trial i+1 and can suggest a new set of hyperparameter values at block 840. For the next trial i+1, one or more hyper-tuning processes use the objective scores recalculated for the previous trial and the corresponding trial hyperparameter values (e.g., a matrix based on the recalculated previous objective scores and the corresponding values of the hyperparameter set) to determine a new set of hyperparameter values for the next trial i+1. This process iterates until the values of the hyperparameters converge to an optimal value or a stopping condition is met (e.g., a threshold number of trials have been executed).
[0289] Objective function optimization and tuning techniques
[0290] Figure 9 depicts a simplified flowchart 900 that depicts objective function optimization and tuning techniques performed by a hyperparameter tuning system (e.g., the hyperparameter tuning system 500 described with respect to Figure 5 ). The depicted processing can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., a memory device). Figure 9 The techniques presented and described below are intended to be illustrative and not restrictive. Although Figure 9 depicts various processing steps occurring in a particular order or sequence, this is not intended to be restrictive. In certain alternative embodiments, the steps can be performed in a different order or some steps can also be performed in parallel. Figure 9 At block 905, a machine learning algorithm is initialized using a set of hyperparameter values. Initialization includes configuring the machine learning algorithm and the training protocol of the machine learning algorithm using the set of hyperparameter values.
[0291] At block 910, a hyperparameter objective function defined at least in part over multiple domains of a search space associated with the machine learning algorithm is accessed. The search space includes a training data set and an evaluation data set, where each domain includes a subdivision of the search space that has at least one training data set and at least one evaluation data set. The hyperparameter objective function includes a domain score for each domain, which is calculated based on the number of instances correctly or incorrectly predicted by the machine learning algorithm within at least one evaluation data set during a given trial.
[0292]
[0293] In some instances, the hyperparameter objective function is formulated to normalize instance-level improvements and instance-level regressions over the total number of instances to obtain an improvement score and a regression score. A domain score for each domain can be calculated based on the improvement score and the regression score. In some cases, the improvement score is calculated based on: (i) the number of instances correctly predicted by the machine learning model during a given trial within at least one evaluation dataset, (ii) the number of instances incorrectly predicted by a baseline machine learning model within at least one evaluation dataset, and (iii) the total count of the number of instances within at least one evaluation dataset. In some cases, the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model during a given trial within at least one evaluation dataset, (ii) the number of instances correctly predicted by a baseline machine learning model within at least one evaluation dataset, and (iii) the total count of the number of instances within at least one evaluation dataset.
[0294] In some instances, the hyperparameter objective function is formulated to exclude unstable instances, which are instances within at least one evaluation dataset that are determined to produce different prediction results from each other using the same machine learning model. In some instances, excluding unstable instances includes: (i) subtracting the count of the excluded unstable instances from the number of instances correctly predicted by the machine learning model during a given trial within at least one evaluation dataset, and (ii) subtracting the count of the excluded unstable instances from the number of instances incorrectly predicted by the machine learning model during a given trial within at least one evaluation dataset.
[0295] In some instances, the hyperparameter objective function is formulated to include a parameter for an acceptable regression ratio m for one or more of multiple domains. In some instances, the parameter is defined based on the regression score, and if the regression score is less than the acceptable regression ratio m, the regression score is set to zero; otherwise, the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model during a given trial within at least one evaluation dataset, (ii) the number of instances correctly predicted by a baseline machine learning model within at least one evaluation dataset, and (iii) the total count of the number of instances within at least one evaluation dataset.
[0296] At block 915, for each trial of the hyperparameter tuning process, blocks 920 - 950 and 960 are iteratively executed.
[0297] At block 920, a machine learning algorithm for each domain is trained using at least one training dataset associated with each domain and a set of hyperparameter values. The training outputs multiple machine learning models, the multiple machine learning models including the machine learning model for each domain.
[0298] At block 925, a machine learning model for each domain is evaluated using at least one evaluation data set associated with each domain and a set of hyperparameter values. The evaluation includes generating a domain score for each domain.
[0299] At block 930, a current trial objective score is calculated using a hyperparameter objective function based on the domain score for each domain and the domain weight associated with each domain.
[0300] At block 935, the current trial objective score is stored in a database. The database includes multiple trial objective scores from current and previous trials of the hyperparameter tuning process, the domain weights associated with each domain, the domain scores for each domain from current and previous trials of the hyperparameter tuning process, and the set of hyperparameter values and all other sets of hyperparameter values from previous trials of the hyperparameter tuning process.
[0301] At block 940, it is determined whether the machine learning model has reached convergence based on the current trial objective score.
[0302] At block 945, in response to determining that the machine learning model has not reached convergence based on the current trial objective score, a new set of hyperparameters is determined for subsequent trials of the hyperparameter tuning process. The new set of hyperparameters is determined based on the current trial objective score, the trial objective scores from previous trials of the hyperparameter tuning process, the set of hyperparameter values, and all other sets of hyperparameter values from previous trials of the hyperparameter tuning process.
[0303] At block 950, in response to determining that the machine learning model has reached convergence, at least one of the multiple machine learning models is provided. In some instances, providing includes: deploying at least one of the multiple machine learning models in a digital assistant or chatbot system as described with respect to Figures 1 to 3 In some instances, providing includes: displaying and / or transmitting at least one of the multiple machine learning models to a user and / or other systems (e.g., external systems). In some instances, providing includes: deploying at least one of the multiple machine learning models in an inference phase to automatically analyze text and classify the text into intents. In some instances, providing includes: storing at least one of the multiple machine learning models in a system (e.g., an external system or storage device). In some instances, providing includes: deploying at least one of the multiple machine learning models in an inference phase to automatically analyze text to create a conversation with a user and / or respond to a query posed by the user.
[0304] Illustrative System
[0305] Figure 10Depicts a simplified diagram of a distributed system 1000. In the illustrated example, the distributed system 1000 includes one or more client computing devices 1002, 1004, 1006, and 1008 coupled to a server 1012 via one or more communication networks 1010. The client computing devices 1002, 1004, 1006, and 1008 may be configured to execute one or more applications.
[0306] In various examples, the server 1012 may be adapted to run one or more services or software applications that implement one or more embodiments described in the present disclosure. In certain examples, the server 1012 may also provide other services or software applications that may include non-virtual environments and virtual environments. In some examples, these services may be provided as web-based services or cloud services (such as under a software-as-a-service (SaaS) model) to users of the client computing devices 1002, 1004, 1006, and / or 1008. A user operating the client computing devices 1002, 1004, 1006, and / or 1008 may then utilize one or more client applications to interact with the server 1012 to utilize the services provided by these components.
[0307] In Figure 10 the depicted configuration, the server 1012 may include one or more components 1018, 1020, and 1022 that implement the functions performed by the server 1012. These components may include software components that may be executed by one or more processors, hardware components, or combinations thereof. It should be understood that various different system configurations may be possible that are different from the distributed system 1000. Thus, Figure 10 the example shown is one example of a distributed system for implementing an example system and is not intended to be limiting.
[0308] A user may use the client computing devices 1002, 1004, 1006, and / or 1008 to execute one or more applications, models, or chatbots that may generate one or more events or models that may then be implemented or served in accordance with the teachings of the present disclosure. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 10 only four client computing devices are depicted, any number of client computing devices may be supported.
[0309] Client devices can include various types of computing systems, such as portable handheld devices, general-purpose computers such as personal computers and laptop computers, workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computing devices can run various types and versions of software applications and operating systems (e.g., Microsoft Apple or UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome TM OS)), including various mobile operating systems (e.g., Microsoft Windows Windows Android TM 、 Palm ). Portable handheld devices can include cellular phones, smartphones (e.g., ), tablet computers (e.g., ), personal digital assistants (PDAs), etc. Wearable devices can include Google head-mounted displays and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices (e.g., Microsoft gaming consoles with or without gesture input devices, Sony systems, various gaming systems provided by and others). Client devices can be capable of executing various different applications, such as various Internet-related applications, communication applications (e.g., email applications, short message service (SMS) applications) and can use various communication protocols.
[0310] (Multiple) networks 1010 can be any type of network familiar to those skilled in the art that can support data communication using any of a variety of available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internetwork Packet Exchange), etc. By way of example only, (multiple) networks 610 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, public switched telephone network (PSTN), infrared network, wireless network (e.g., according to the Institute of Electrical and Electronics Engineers (IEEE) 1002.11 protocol suite, a network operating under any one of these protocols and / or any other wireless protocol) and / or any combination of these networks and / or other networks.
[0311] Server 1012 may consist of: one or more general-purpose computers, dedicated server computers (including, by way of example, PC (personal computer) servers, servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. Server 1012 may include one or more virtual machines running a virtual operating system or other computing architectures involving virtualization, such as one or more flexible pools of logical storage devices that can be virtualized to maintain the virtual storage device of the server. In various examples, server 1012 may be adapted to run one or more services or software applications that provide the functions described in the foregoing disclosure.
[0312] The computing system in server 1012 may run one or more operating systems, including any one of the operating systems discussed above and any commercially available server operating system. Server 1012 may also run any one of a variety of additional server applications and / or middleware applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, servers, database servers, etc. Exemplary database servers include, but are not limited to, those commercially available from (International Business Machines Corporation), etc.
[0313] In some embodiments, server 1012 may include one or more applications to analyze and combine data feeds and / or event updates received from users of client computing devices 1002, 1004, 1006, and 1008. By way of example, data feeds and / or event updates may include, but are not limited to feeds, updates or real-time updates received from one or more third-party information sources and continuous data streams, and the real-time updates may include real-time events related to sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc. Server 1012 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client computing devices 1002, 1004, 1006, and 1008.
[0314] The distributed system 1000 may also include one or more data repositories 1014, 1016. In some examples, these data repositories can be used to store data and other information. For example, one or more of the data repositories 1014, 1016 can be used to store information (such as information related to chatbot performance or generated models) for use by the chatbot when the server 1012 performs various functions according to various embodiments. The data repositories 1014, 1016 can reside in various locations. For example, the data repositories used by the server 1012 can be local to the server 1012 or can be remote from the server 1012 and communicate with the server 1012 via a network-based or dedicated connection. The data repositories 1014, 1016 can be of different types. In some examples, the data repositories used by the server 1012 can be databases, such as relational databases, such as those provided by Oracle and other vendors. One or more of these databases can be adapted to implement the storage, update, and retrieval of data to and from the database in response to commands in SQL format.
[0315] In some examples, one or more of the data repositories 1014, 1016 can also be used by an application to store application data. The data repositories used by the application can be of different types, such as, for example, key-value storage repositories, object storage repositories, or general storage repositories supported by a file system.
[0316] In some examples, the functions described in this disclosure can be provided as a service via a cloud environment. Figure 11 is a simplified block diagram of a cloud-based system environment in which various services according to some examples can be provided as cloud services. In Figure 11 the depicted example, the cloud infrastructure system 1102 can provide one or more cloud services that can be requested by a user using one or more client computing devices 1104, 1106, and 1108. The cloud infrastructure system 1102 can include one or more computers and / or servers, which can include those computers and / or servers described above for the server 1012. The computers in the cloud infrastructure system 1102 can be organized as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0317] (Multiple) networks 1110 can facilitate data communication and exchange between clients 1104, 1106, and 1108 and cloud infrastructure system 1102. (Multiple) networks 1110 can include one or more networks. The networks can be of the same or different types. (Multiple) networks 1110 can support one or more communication protocols (including wired and / or wireless protocols) to facilitate communication.
[0318] Figure 11 The depicted example is only one example of a cloud infrastructure system and is not intended to be restrictive. It should be understood that in some other examples, cloud infrastructure system 1102 can have more or fewer components than those depicted, can combine two or more components, or can have different component configurations or arrangements. For example, although Figure 11 three client computing devices are depicted, in alternative examples, any number of client computing devices can be supported. Figure 11 Depicted are three client computing devices, but in alternative examples, any number of client computing devices can be supported.
[0319] The term cloud service is generally used to refer to services provided on demand to a user via a communication network such as the Internet by a service provider's system (e.g., cloud infrastructure system 1102). Generally, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's on-premises servers and systems. The cloud service provider's system is managed by the cloud service provider. Thus, a customer can enable itself to utilize the cloud services provided by the cloud service provider without having to purchase separate licenses, support, or hardware and software resources for the services. For example, the cloud service provider's system can host an application, and a user can order and use the application on demand via the Internet without the user having to purchase the infrastructure resources for executing the application. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, Oracle of Redwood Shores, California offers several cloud services such as middleware services, database services, Java cloud services, and other services.
[0320] In certain examples, cloud infrastructure system 1102 can provide one or more cloud services using different models (such as under a software as a service (SaaS) model, a platform as a service (PaaS) model, an infrastructure as a service (IaaS) model, and other models (including hybrid service models)). Cloud infrastructure system 1102 can include a set of applications, middleware, databases, and other resources that implement the provisioning of various cloud services.
[0321] The SaaS model enables applications or software to be delivered as a service to customers over a communication network such as the Internet, without the customer having to purchase the hardware or software of the underlying application. For example, the SaaS model can be used to provide customers with access to on-demand applications hosted by the cloud infrastructure system 1102. Oracle Examples of SaaS services provided by Oracle include, but are not limited to, various services for human resources / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, social applications, and others.
[0322] The IaaS model is typically used to provide infrastructure resources (e.g., servers, storage, hardware, and networking resources) as cloud services to customers to provide elastic computing and storage capabilities. Provided by Oracle Provides various IaaS services.
[0323] The PaaS model is typically used to provide platform and environmental resources as a service that enables customers to develop, run, and manage applications and services without the customer having to procure, build, or maintain such resources. Provided by Oracle Examples of PaaS services provided by Oracle include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), Data Management Cloud Service, various application development solution services, and other services.
[0324] Cloud services are typically provided in an on-demand self-service basis, subscription-based, elastically scalable, reliable, highly available, and secure manner. For example, a customer can order one or more services provided by the cloud infrastructure system 1102 via a subscription order. Then, the cloud infrastructure system 1102 performs processing to provide the services requested in the customer's subscription order. For example, a user can use speech to request the cloud infrastructure system to take a certain action (e.g., an intent) as described above and / or to provide services for a chatbot system as described herein. The cloud infrastructure system 1102 can be configured to provide one or even multiple cloud services.
[0325] The cloud infrastructure system 1102 can provide cloud services via different deployment models. In the public cloud model, the cloud infrastructure system 1102 can be owned by a third-party cloud service provider, and cloud services are provided to any general public customers, where the customers can be individuals or enterprises. In some other examples, under the private cloud model, the cloud infrastructure system 1102 can operate within an organization (e.g., within an enterprise organization) and services are provided to customers within the organization. For example, the customers can be various departments of an enterprise such as the human resources department, payroll department, etc. or even individuals within the enterprise. In some other examples, under the community cloud model, the cloud infrastructure system 1102 and the provided services can be shared by several organizations in the relevant community. Various other models such as a hybrid of the models mentioned above can also be used.
[0326] The client computing devices 1104, 1106, and 1108 can be of different types (such as Figure 10 the client computing devices 1002, 1004, 1006, and 1008 depicted) and can be capable of operating one or more client applications. Users can use the client devices to interact with the cloud infrastructure system 1102, such as requesting services provided by the cloud infrastructure system 1102. For example, a user can use the client device to request information or actions from a chatbot as described in this disclosure.
[0327] In some examples, the processing for providing services performed by the cloud infrastructure system 1102 can involve model training and deployment. This analysis can involve using, analyzing, and manipulating data sets to train and deploy one or more models. The analysis can be performed by one or more processors, thus potentially processing data in parallel, performing simulations using the data, etc. For example, big data analysis can be performed by the cloud infrastructure system 1102 to generate and train one or more models for a chatbot system. The data used for this analysis can include structured data (e.g., data stored in a database or structured according to a structured model) and / or unstructured data (e.g., data chunks (binary large objects)).
[0328] As Figure 11 depicted in the example of, the cloud infrastructure system 1102 can include infrastructure resources 1130 that are used to facilitate the provision of various cloud services provided by the cloud infrastructure system 1102. The infrastructure resources 1130 can include, for example, processing resources, storage or memory resources, networking resources, etc. In some examples, a storage virtual machine available for serving storage requests from applications can be part of the cloud infrastructure system 1102. In other examples, the storage virtual machine can be part of a different system.
[0329] In some examples, to facilitate the efficient supply of these resources to support the various cloud services provided by the cloud infrastructure system 1102 for different customers, the resources can be bound into resource groups or resource modules (also referred to as "pods"). Each resource module or pod can include a pre-integrated and optimized combination of one or more types of resources. In some examples, different pods can be pre-provisioned for different types of cloud services. For example, a first set of pods can be provisioned for database services, a second set of pods can be provisioned for Java services (the second set of pods can include a different resource combination than the pods in the first set), and so on. For some services, the resources allocated for service provision can be shared among the services.
[0330] The cloud infrastructure system 1102 itself can internally use services 1132 that are shared by different components of the cloud infrastructure system 1102 and facilitate the service provision of the cloud infrastructure system 1102. These internal shared services can include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, services for enabling cloud support, email services, notification services, file transfer services, etc.
[0331] The cloud infrastructure system 1102 can include multiple subsystems. These subsystems can be implemented in software or hardware or a combination thereof. As Figure 11 depicted, the subsystems can include a user interface subsystem 1112 that enables users or customers of the cloud infrastructure system 1102 to interact with the cloud infrastructure system 1102. The user interface subsystem 1112 can include various different interfaces, such as a web interface 1114, an online store interface 1116 (wherein cloud services provided by the cloud infrastructure system 1102 are advertised and consumers can purchase them), and other interfaces 1118. For example, a customer can use a client device to request (service request 1134) one or more services provided by the cloud infrastructure system 1102 using one or more of the interfaces 1114, 1116, and 1118. For example, a customer can access an online store, browse the cloud services provided by the cloud infrastructure system 1102, and place a subscription order for one or more services provided by the cloud infrastructure system 1102 that the customer wishes to subscribe to. The service request can include information identifying the customer and the one or more services the customer expects to subscribe to. For example, a customer can place a subscription order for a service provided by the cloud infrastructure system 1102. As part of the order, the customer can provide information identifying the chatbot system for which the service is to be provided and optionally provide one or more credentials for the chatbot system.
[0332] In some examples (such as Figure 11In the depicted example, the cloud infrastructure system 1102 can include an Order Management Subsystem (OMS) 1120 configured to process new orders. As part of this processing, the OMS 1120 can be configured to: create an account for the customer (if not already created); receive from the customer billing and / or charging information to be used to bill the customer for the requested services; verify the customer information; after verification, book an order for the customer; and orchestrate various workflows to prepare the order for provisioning.
[0333] Once properly verified, the OMS 1120 can then call an Order Provisioning Subsystem (OPS) 1124 configured to provision resources for the order, including processing resources, memory resources, and networking resources. Provisioning can include allocating resources for the order and configuring the resources to facilitate the services requested by the customer order. The manner in which resources are provisioned for the order and the type of resources provisioned can depend on the type of cloud service the customer has ordered. For example, according to one workflow, the OPS 1124 can be configured to determine the specific cloud service being requested and identify the number of clusters that may have been pre-configured for that specific cloud service. The number of clusters allocated for the order can depend on the size / volume / tier / scope of the requested service. For example, the number of clusters to be allocated can be determined based on the number of users the service is to support, the duration of the service being requested, etc. Then, the allocated clusters can be customized for the specific requesting customer for providing the requested service.
[0334] In some examples, the setup phase processing as described above can be performed by the cloud infrastructure system 1102 as part of the provisioning process. The cloud infrastructure system 1102 can generate an application ID and select a storage virtual machine for the application from storage virtual machines provided by the cloud infrastructure system 1102 itself or from storage virtual machines provided by other systems other than the cloud infrastructure system 1102.
[0335] The cloud infrastructure system 1102 can send a response or notification 1144 to the requesting customer to indicate when the requested service is now ready for use. In some instances, information (e.g., a link) enabling the customer to start using and leveraging the benefits of the requested service can be sent to the customer. In some examples, for the customer requesting the service, the response can include a chatbot system ID generated by the cloud infrastructure system 1102 and information identifying the chatbot system selected by the cloud infrastructure system 1102 for the chatbot system corresponding to the chatbot system ID.
[0336] The cloud infrastructure system 1102 can provide services to multiple customers. For each customer, the cloud infrastructure system 1102 is responsible for managing information related to one or more subscription orders received from the customer, maintaining customer data related to the orders, and providing the requested services to the customer. The cloud infrastructure system 1102 can also collect usage statistics regarding the customer's use of the subscribed services. For example, statistics such as the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime can be collected. This usage information can be used to bill the customer. Billing can be completed, for example, on a monthly basis.
[0337] The cloud infrastructure system 1102 can provide services to multiple customers in parallel. The cloud infrastructure system 1102 can store information for these customers (which may include proprietary information). In some examples, the cloud infrastructure system 1102 includes an identity management subsystem (IMS) 1128 that is configured to manage customer information and provide separation of the managed information such that information related to one customer cannot be accessed by another customer. The IMS 1128 can be configured to provide various security-related services such as identity services, such as information access management, authentication and authorization services, services for managing customer identities and roles, and related functions.
[0338] Figure 12 An example of a computer system 1200 is illustrated. In some examples, the computer system 1200 can be used to implement any digital assistant or chatbot system within a distributed environment, as well as the various servers and computer systems described above. As Figure 12 shown, the computer system 1200 includes various subsystems, including a processing subsystem 1204 that communicates with a plurality of other subsystems via a bus subsystem 1202. These other subsystems can include a processing acceleration unit 1206, an I / O subsystem 1208, a storage subsystem 1218, and a communication subsystem 1224. The storage subsystem 1218 can include non-transitory computer-readable storage media, including a storage medium 1222 and a system memory 1210.
[0339] The bus subsystem 1202 provides the mechanism for allowing the various components and subsystems of the computer system 1200 to communicate with each other as expected. Although the bus subsystem 1202 is schematically shown as a single bus, alternative examples of the bus subsystem may utilize multiple buses. The bus subsystem 1202 can be any one of several types of bus structures including a memory bus or memory controller, a peripheral bus, a local bus using any one of various bus architectures, and the like. For example, such architectures can include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus (which can be implemented as a Mezzanine bus manufactured to the IEEE P1386.1 standard), among others.
[0340] The processing subsystem 1204 controls the operation of the computer system 1200 and can include one or more processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The processors can include single-core processors or multi-core processors. The processing resources of the computer system 1200 can be organized into one or more processing units 1232, 1234, etc. The processing units can include one or more processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some examples, the processing subsystem 1204 can include one or more dedicated coprocessors such as a graphics processor, a digital signal processor (DSP), etc. In some examples, some or all of the processing units in the processing subsystem 1204 can be implemented using custom circuits such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs).
[0341] In some examples, the processing units in the processing subsystem 1204 can execute instructions stored in the system memory 1210 or on the computer-readable storage medium 1222. In various examples, the processing units can execute various programs or code instructions and can maintain multiple programs or processes executing simultaneously. At any given time, some or all of the program code to be executed can reside in the system memory 1210 and / or on the computer-readable storage medium 1222 (potentially including residing on one or more storage devices). Through suitable programming, the processing subsystem 1204 can provide the various functions described above. In instances where the computer system 1200 is executing one or more virtual machines, one or more processing units can be allocated to each virtual machine.
[0342] In some examples, a processing acceleration unit 1206 may optionally be provided for performing custom processing or for offloading some of the processing performed by the processing subsystem 1204, thereby accelerating the overall processing performed by the computer system 1200.
[0343] The I / O subsystem 1208 may include devices and mechanisms for inputting information to and / or for outputting information from or via the computer system 1200. Generally, the term input device is intended to include all possible types of devices and mechanisms for inputting information to the computer system 1200. User interface input devices may include, for example, a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen incorporated into a display, a scroll wheel, a click wheel, a dial, a button, a switch, a keypad, an audio input device having a voice command recognition system, a microphone, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices such as Microsoft Motion Sensors, Microsoft 360 game controllers, devices that provide an interface for receiving input using gestures and voice commands. User interface input devices may also include eye gesture recognition devices such as Google ) that detect eye activity from a user (e.g., "blinking" when taking a picture and / or making a menu selection) and transform the eye gesture into an input to an input device such as Google Blink Detector. Additionally, user interface input devices may include voice recognition sensing devices that enable a user to interact with a voice recognition system (e.g., Navigator).
[0344] Other examples of user interface input devices include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, and graphics tablets, as well as audio / visual devices (such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, eye gaze tracking devices). Additionally, user interface input devices may include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound devices. User interface input devices may also include, for example, audio input devices such as MIDI keyboards, digital musical instruments, etc.
[0345] Generally, the term output device is intended to include all possible types of devices and mechanisms for outputting information from computer system 1200 to a user or other computers. User interface output devices can include a display subsystem, indicator lights, or non-visual displays such as audio output devices. The display subsystem can be a cathode ray tube (CRT), flat panel device (such as a flat panel device using a liquid crystal display (LCD) or a plasma display), a projection device, a touch screen, etc. For example, user interface output devices can include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, voice output devices, and modems.
[0346] Storage subsystem 1218 provides a repository or data storage for storing the information and data used by computer system 1200. Storage subsystem 1218 provides a tangible non-transitory computer-readable storage medium for storing the basic programming and data constructs that provide the functionality of some examples. Storage subsystem 1218 can store software (e.g., programs, code modules, instructions) that, when executed by processing subsystem 1204, provides the functionality described above. The software can be executed by one or more processing units of processing subsystem 1204. Storage subsystem 1218 can also provide authentication in accordance with the teachings of the present disclosure.
[0347] Storage subsystem 1218 can include one or more non-transitory memory devices, and the one or more non-transitory memory devices include volatile memory devices and non-volatile memory devices. As Figure 12 shown, storage subsystem 1218 includes system memory 1210 and computer-readable storage medium 1222. System memory 1210 can include multiple memories, including volatile main random access memory (RAM) for storing instructions and data during program execution and non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some embodiments, the basic input / output system (BIOS), which contains basic routines that help transfer information between elements within computer system 1200 during startup, can typically be stored in the ROM. The RAM generally contains the data and / or program modules that are currently being operated on and executed by processing subsystem 1204. In some embodiments, system memory 1210 can include various different types of memories such as static random access memory (SRAM), dynamic random access memory (DRAM), etc.
[0348] By way of example and not limitation, as Figure 12As depicted, the system memory 1210 can load the executing application programs 1212 (which can include various application programs such as web browsers, middleware application programs, relational database management systems (RDBMS), etc.), program data 1214, and the operating system 1216. By way of example, the operating system 1216 can include various versions of Microsoft Apple and / or Linux operating systems, various commercially available or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google OS, etc.) and / or mobile operating systems such as iOS, Phone, OS, OS, OS operating systems, and other operating systems.
[0349] The computer-readable storage medium 1222 can store programming and data constructs that provide some example functionality. The computer-readable medium 1222 can provide storage for the computer system 1200 for computer-readable instructions, data structures, program modules, and other data. Software (programs, code modules, instructions) that provides the functionality described above when executed by the processing subsystem 1204 can be stored in the storage subsystem 1218. By way of example, the computer-readable storage medium 1222 can include non-volatile memories such as hard disk drives, disk drives, optical disc drives (such as CD ROM, DVD, disc or other optical media). The computer-readable storage medium 1222 can include but is not limited to drives, flash memory cards, universal serial bus (USB) flash memory drives, secure digital (SD) cards, DVD discs, digital video tapes, etc. The computer-readable storage medium 1222 can also include non-volatile memory-based SSDs such as flash memory-based solid-state drives (SSDs), enterprise-class flash memory drives, solid-state ROMs, volatile memory-based SSDs such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs.
[0350] In some examples, the storage subsystem 1218 can further include a computer-readable storage medium reader 1220 that can be further connected to the computer-readable storage medium 1222. The reader 1220 can receive data from a memory device such as a disc, flash memory drive, etc. and is configured to read data from the memory device.
[0351] In some examples, computer system 1200 may support virtualization technologies, including but not limited to virtualization of processing and memory resources. For example, computer system 1200 may provide support for executing one or more virtual machines. In some examples, computer system 1200 may execute programs such as hypervisors that facilitate the configuration and management of virtual machines. Each virtual machine may be allocated memory, computing (e.g., processors, cores), I / O, and networking resources. Each virtual machine typically runs independently of other virtual machines. A virtual machine typically runs its own operating system, which may be the same as or different from the operating systems executed by other virtual machines of computer system 1200. Thus, multiple operating systems may potentially be run simultaneously by computer system 1200.
[0352] Communication subsystem 1224 provides an interface to other computer systems and networks. Communication subsystem 1224 serves as an interface for receiving data from other systems and transmitting data from computer system 1200 to other systems. For example, communication subsystem 1224 may enable computer system 1200 to establish a communication channel to one or more client devices via the Internet for receiving information from and sending information to the client devices. For example, when computer system 1200 is used to implement Figure 1 the depicted robotic system 120, the communication subsystem may be used to communicate with a chatbot system selected for the application.
[0353] Communication subsystem 1224 may support both wired communication protocols and / or wireless communication protocols. In some examples, communication subsystem 1224 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular phone technologies, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE 802.XX home standards, or other mobile communication technologies, or any combination thereof)), global positioning system (GPS) receiver components, and / or other components. In some examples, in addition to or instead of a wireless interface, communication subsystem 1224 may provide wired network connectivity (e.g., Ethernet).
[0354] Communication subsystem 1224 may receive and transmit various forms of data. In some examples, in addition to other forms, communication subsystem 1224 may also receive input communications in the form of structured and / or unstructured data feeds 1226, event streams 1228, event updates 1230, etc. For example, communication subsystem 1224 may be configured to receive (or send) data feeds 1226 from users of social media networks and / or other communication services in real time, such as feeds, Updates, web feeds (such as Rich Site Summary (RSS) feeds), and / or real-time updates from one or more third-party information sources.
[0355] In some examples, the communication subsystem 1224 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1228 of real-time events and / or event updates 1230 (which may be essentially continuous or unbounded without an explicit end). Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc.
[0356] The communication subsystem 1224 may also be configured to transfer data from the computer system 1200 to other computer systems or networks. The data may be transferred in various different forms such as structured and / or unstructured data feeds 1226, event streams 1228, event updates 1230, etc., to one or more databases that may communicate with one or more stream data source computers coupled to the computer system 1200.
[0357] The computer system 1200 may be one of various types, including handheld portable devices (e.g., cellular phones, computing tablet computers, PDAs), wearable devices (e.g., Google head-mounted displays), personal computers, workstations, mainframes, self-service terminals, server racks, or any other data processing system. Due to the ever-changing nature of computers and networks, the Figure 12 description of the depicted computer system 1200 is intended to be only a specific example. Many other configurations with more or fewer components than the Figure 12 depicted system are possible. Based on the present disclosure and the teachings provided herein, it should be understood that there are other ways and / or methods to implement the various examples.
[0358] Although specific examples have been described, various modifications, changes, alternative constructions, and equivalents are possible. The examples are not limited to operating in certain specific data processing environments, but are free to operate in multiple data processing environments. Additionally, although certain examples have been described using a specific series of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. While some flowcharts depict operations as sequential processes, many operations may be performed in parallel or simultaneously. Additionally, the order of operations may be rearranged. The process may have additional steps not included in the figures. The various features and aspects of the examples described above may be used alone or in combination.
[0359] Further, although certain examples have been described using a specific combination of hardware and software, it should be recognized that other combinations of hardware and software are possible. Certain examples may be implemented using only hardware, only software, or a combination thereof. The various processes described herein may be implemented in any combination on the same processor or different processors.
[0360] In cases where an apparatus, system, component, or module is described as being configured to perform certain operations or functions, such configuration may be accomplished, for example, by designing an electronic circuit to perform the operations, by programming a programmable electronic circuit (such as a microprocessor) to perform the operations (such as by executing computer instructions or code), or by a processor or core programmed to execute code or instructions stored on a non-transitory memory medium, or any combination thereof. Processes may communicate using a variety of techniques including, but not limited to, conventional techniques for interprocess communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0361] Specific details are given in this disclosure to provide a thorough understanding of the examples. However, the examples may be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques have been shown without unnecessary detail to avoid obscuring the examples. This description provides only exemplary examples and is not intended to limit the scope, applicability, or configuration of other examples. Rather, the previous description of the examples will provide those skilled in the art with an enabling description for implementing the various examples. Various changes may be made to the function and arrangement of the elements.
[0362] Accordingly, this specification and the drawings should be regarded in an illustrative rather than a restrictive sense. However, it will be apparent that additions, deletions, and other modifications and changes may be made thereto without departing from the broader spirit and scope set forth in the claims. Accordingly, although specific examples have been described, these examples are not intended to be restrictive. Various modifications and equivalents are within the scope of the following claims.
[0363] In the foregoing specification, aspects of the present disclosure have been described with reference to specific examples of the present disclosure, but those skilled in the art will recognize that the present disclosure is not limited thereto. The various features and aspects of the disclosure described above may be used alone or in combination. Further, the examples may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings should be regarded as illustrative rather than restrictive.
[0364] In the foregoing description, for purposes of illustration, the method has been described in a particular order. It should be appreciated that in alternative examples, the method may be performed in an order different from that described. It should also be appreciated that the methods described above may be performed by hardware components or may be embodied in a sequence of machine-executable instructions that may be used to cause a machine, such as a general or special purpose processor or logic circuitry programmed with the instructions, to perform the method. These machine-executable instructions may be stored on one or more machine-readable media, such as a CD-ROM or other type of optical disc, a floppy disc, a ROM, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, or other types of machine-readable media suitable for storing electronic instructions. Alternatively, the methods may be performed by a combination of hardware and software.
[0365] Where components are described as being configured to perform certain operations, such configuration may be accomplished, for example, by designing electronic circuitry or other hardware for performing the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuitry) for performing the operations, or any combination thereof.
[0366] Although illustrative examples of the present application have been described in detail herein, it should be understood that the inventive concepts may be otherwise variously embodied and employed, and the appended claims are intended to be construed to include such variations, except as limited by the prior art.
Claims
1. A computer-implemented method, comprising: initializing a machine learning algorithm using a set of hyperparameter values; accessing a hyperparameter objective function defined at least in part over multiple domains of a search space associated with the machine learning algorithm, wherein the search space includes a training data set and an evaluation data set, wherein each domain includes a subdivision of the search space, the subdivision having at least one training data set and at least one evaluation data set, and wherein the hyperparameter objective function includes a domain score for each domain, the domain score being calculated based on the number of instances correctly or incorrectly predicted by the machine learning algorithm within the at least one evaluation data set during a given trial; for each trial of a hyperparameter tuning process: training the machine learning algorithm for each domain using the at least one training data set associated with each domain and the set of hyperparameter values, wherein the training outputs a plurality of machine learning models, the plurality of machine learning models including a machine learning model for each domain; evaluating the machine learning model for each domain using the at least one evaluation data set associated with each domain and the set of hyperparameter values, wherein the evaluation includes generating the domain score for each domain; calculating a current trial objective score using the hyperparameter objective function based on the domain score for each domain and a domain weight associated with each domain; and determining whether the machine learning model has reached convergence based on the current trial objective score; and in response to determining that the machine learning model has reached convergence, providing at least one of the plurality of machine learning models.
2. The computer-implemented method of claim 1, wherein, the hyperparameter objective function is formulated to normalize instance-level improvements and instance-level regressions over the total number of instances to obtain an improvement score and a regression score, and wherein the domain score for each domain is calculated based on the improvement score and the regression score.
3. The computer-implemented method of claim 2, wherein, the improvement score is calculated based on: (i) the number of instances correctly predicted by the machine learning model within the at least one evaluation data set during the given trial, (ii) the number of instances incorrectly predicted by a baseline machine learning model within the at least one evaluation data set, and (iii) the total count of the number of instances within the at least one evaluation data set, and wherein the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model within the at least one evaluation data set during the given trial, (ii) the number of instances correctly predicted by a baseline machine learning model within the at least one evaluation data set, and (iii) the total count of the number of instances within the at least one evaluation data set.
4. The computer-implemented method of claim 1, wherein, the hyperparameter objective function is formulated to exclude unstable instances, the unstable instances being instances within the at least one evaluation data set that are determined to produce different prediction results from each other using the same machine learning model.
5. The computer-implemented method according to claim 4, wherein, excluding unstable instances includes: (i) subtracting the count of the excluded unstable instances from the number of instances correctly predicted by the machine learning model during the given trial within the at least one evaluation dataset, and (ii) subtracting the count of the excluded unstable instances from the number of instances incorrectly predicted by the machine learning model during the given trial within the at least one evaluation dataset.
6. The computer-implemented method according to claim 1, wherein, the hyperparameter objective function is formulated to include a parameter for an acceptable regression ratio m for one or more of the plurality of domains.
7. The computer-implemented method according to claim 6, wherein, the parameter is defined based on a regression score, and if the regression score is less than the acceptable regression ratio m, the regression score is set to zero, otherwise the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model during the given trial within the at least one evaluation dataset, (ii) the number of instances correctly predicted by a baseline machine learning model within the at least one evaluation dataset, and (iii) the total count of the number of instances within the at least one evaluation dataset.
8. A system, comprising: one or more processors; and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the system to perform operations including: initializing a machine learning algorithm with a set of hyperparameter values; accessing a hyperparameter objective function defined at least partially over a plurality of domains of a search space associated with the machine learning algorithm, wherein the search space includes a training dataset and an evaluation dataset, wherein each domain includes a subdivision of the search space having at least one training dataset and at least one evaluation dataset, and wherein the hyperparameter objective function includes a domain score for each domain, the domain score being calculated based on the number of instances correctly or incorrectly predicted by the machine learning algorithm during a given trial within the at least one evaluation dataset; for each trial of a hyperparameter tuning process: training the machine learning algorithm for each domain using the at least one training dataset associated with each domain and the set of hyperparameter values, wherein the training outputs a plurality of machine learning models, the plurality of machine learning models including a machine learning model for each domain; evaluating the machine learning models for each domain using the at least one evaluation dataset associated with each domain and the set of hyperparameter values, wherein the evaluation includes generating a domain score for each domain; calculating a current trial objective score using the hyperparameter objective function based on the domain score for each domain and a domain weight associated with each domain; and determining whether the machine learning model has reached convergence based on the current trial objective score; and in response to determining that the machine learning model has reached convergence, providing at least one of the plurality of machine learning models.
9. The system according to claim 8, Wherein, the hyperparameter objective function is formulated to normalize instance-level improvements and instance-level regressions over the total number of instances to obtain an improvement score and a regression score, and wherein the domain score for each domain is calculated based on the improvement score and the regression score.
10. The system of claim 9, wherein, the improvement score is calculated based on: (i) the number of instances correctly predicted by the machine learning model during the given trial in the at least one evaluation dataset, (ii) the number of instances incorrectly predicted by a baseline machine learning model in the at least one evaluation dataset, and (iii) the total count of the number of instances in the at least one evaluation dataset, and wherein the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model during the given trial in the at least one evaluation dataset, (ii) the number of instances correctly predicted by a baseline machine learning model in the at least one evaluation dataset, and (iii) the total count of the number of instances in the at least one evaluation dataset.
11. The system of claim 8, wherein, the hyperparameter objective function is formulated to exclude unstable instances, which are instances in the at least one evaluation dataset that are determined to produce different prediction results from each other using the same machine learning model.
12. The system of claim 11, wherein, excluding unstable instances includes: (i) subtracting the count of the excluded unstable instances from the number of instances correctly predicted by the machine learning model during the given trial in the at least one evaluation dataset, and (ii) subtracting the count of the excluded unstable instances from the number of instances incorrectly predicted by the machine learning model during the given trial in the at least one evaluation dataset.
13. The system of claim 8, wherein, the hyperparameter objective function is formulated to include a parameter for an acceptable regression ratio m for one or more of the plurality of domains.
14. The system of claim 13, wherein, the parameter is defined based on the regression score, and if the regression score is less than the acceptable regression ratio m, the regression score is set to zero, otherwise the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model during the given trial in the at least one evaluation dataset, (ii) the number of instances correctly predicted by a baseline machine learning model in the at least one evaluation dataset, and (iii) the total count of the number of instances in the at least one evaluation dataset.
15. One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to perform operations including the following: Initialize a machine learning algorithm using a set of hyperparameter values; Access a hyperparameter objective function defined at least partially over a plurality of domains of a search space associated with the machine learning algorithm, wherein, The search space includes a training data set and an evaluation data set, where each domain includes a subdivision of the search space, the subdivision having at least one training data set and at least one evaluation data set, and where the hyperparameter objective function includes a domain score for each domain, the domain score being calculated based on the number of instances correctly or incorrectly predicted by the machine learning algorithm within the at least one evaluation data set during a given trial; For each trial of the hyperparameter tuning process: The machine learning algorithm for each domain is trained using the at least one training data set associated with each domain and the set of hyperparameter values, where the training outputs a plurality of machine learning models, the plurality of machine learning models including a machine learning model for each domain; The machine learning models for each domain are evaluated using the at least one evaluation data set associated with each domain and the set of hyperparameter values, where the evaluation includes generating the domain score for each domain; The hyperparameter objective function is used to calculate the current trial objective score based on the domain score for each domain and the domain weights associated with each domain; and Based on the current trial objective score, it is determined whether the machine learning model has reached convergence; and in response to determining that the machine learning model has reached convergence, at least one of the plurality of machine learning models is provided.
16. One or more non-transitory computer-readable media as recited in claim 15, wherein, The hyperparameter objective function is formulated to normalize instance-level improvement and instance-level regression over the total number of instances to obtain an improvement score and a regression score, and where the domain score for each domain is calculated based on the improvement score and the regression score.
17. One or more non-transitory computer-readable media as recited in claim 16, wherein, The improvement score is calculated based on: (i) the number of instances correctly predicted by the machine learning model within the at least one evaluation data set during the given trial, (ii) the number of instances incorrectly predicted by a baseline machine learning model within the at least one evaluation data set, and (iii) the total count of the number of instances within the at least one evaluation data set, and where the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model within the at least one evaluation data set during the given trial, (ii) the number of instances correctly predicted by a baseline machine learning model within the at least one evaluation data set, and (iii) the total count of the number of instances within the at least one evaluation data set.
18. One or more non-transitory computer-readable media as recited in claim 15, wherein, The hyperparameter objective function is formulated to exclude unstable instances, the unstable instances being instances within the at least one evaluation data set that are determined to produce different prediction results from each other using the same machine learning model.
19. One or more non-transitory computer-readable media as recited in claim 18, wherein, Excluding unstable instances includes: (i) subtracting the count of the excluded unstable instances from the number of instances correctly predicted by the machine learning model during the given trial within the at least one evaluation dataset, and (ii) subtracting the count of the excluded unstable instances from the number of instances incorrectly predicted by the machine learning model during the given trial within the at least one evaluation dataset.
20. One or more non-transitory computer-readable media as recited in claim 15, wherein, the hyperparameter objective function is formulated to include a parameter for an acceptable regression ratio m for one or more of the plurality of domains, and wherein the parameter is defined based on a regression score, and if the regression score is less than the acceptable regression ratio m, the regression score is set to zero, otherwise the regression score is calculated based on: (i) the number of instances incorrectly predicted by the machine learning model during the given trial within the at least one evaluation dataset, (ii) the number of instances correctly predicted by a baseline machine learning model within the at least one evaluation dataset, and (iii) the total count of the number of instances within the at least one evaluation dataset.