Method and system for hyperparameter tuning based on constraints

By using a weighted cost function to optimize hyperparameters across multiple metrics and datasets, the method addresses the limitations of single-objective hyperparameter tuning in chatbot systems, achieving improved performance and effectiveness.

JP7692432B2Active Publication Date: 2025-06-13ORACLE INT CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022559647
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-29
Filing Date
2021-03-30
Publication Date
2025-06-13
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

Existing hyperparameter tuning algorithms for machine learning models in chatbot systems focus on a single objective, ignoring other important objectives and considering only a single dataset, which limits their effectiveness in optimizing performance across multiple metrics and datasets.

Method used

A method for tuning hyperparameters of machine learning models in chatbot systems that involves obtaining multiple datasets and selecting multiple metrics for evaluation. Each metric is assigned a weight to indicate its importance, and a cost function is created to measure performance based on these metrics and weights. The hyperparameters are then optimized to improve the model's performance across all selected metrics.

Benefits of technology

This approach enables multi-objective optimization of machine learning models, allowing them to perform well across multiple metrics and datasets, thereby improving the overall effectiveness and efficiency of chatbot systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692432000002
    Figure 0007692432000002
  • Figure 0007692432000003
    Figure 0007692432000003
  • Figure 0007692432000004
    Figure 0007692432000004
Patent Text Reader

Abstract

Techniques for tuning hyperparameters of a model are disclosed. A dataset is obtained to train the model, and a metric is selected to evaluate the model's performance. Each metric is assigned a weight that specifies its importance to the model's performance. A function is created to measure performance based on the weighted metrics. The hyperparameters are tuned to optimize model performance. Tuning the hyperparameters includes (i) training a model configured based on current values ​​of the hyperparameters, (ii) evaluating the model's performance using the function, (iii) determining whether the model is optimized for the metric, (iv) responsive to the model not being optimized, exploring new values ​​for the hyperparameters, reconfiguring the model with the new values, and repeating steps (i)-(iii) using the reconfigured model, and (v) responsive to the model being optimized for the metric, providing the trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority to (1) U.S. Provisional Application No. 63 / 002,159, filed Mar. 30, 2020; (2) U.S. Provisional Application No. 63 / 119,577, filed Nov. 30, 2020; (3) U.S. Non - Provisional Application No. 17 / 216,496, filed Mar. 29, 2021; and (4) U.S. Non - Provisional Application No. 17 / 216,498, filed Mar. 29, 2021. The applications referenced above are hereby incorporated by reference in their entireties for all purposes.

[0002] Field of the Invention The present disclosure generally relates to chatbot systems, and more particularly to techniques for tuning hyperparameters of machine - learning models used in chatbot systems.

Background Art

[0003] Background Many users around the world are on instant messaging or chat platforms to get instant responses. Organizations often engage in live conversations with customers (or end-users) using these instant messaging or chat platforms. However, hiring service staff to engage in live communication with customers or end-users can be very costly for an organization. Chatbots or bots have begun to be developed to simulate conversations with end-users, especially on the Internet. End-users can communicate with the bot via the messaging app that the end-user has already installed and used. Generally, intelligent bots driven by artificial intelligence (AI) can communicate more knowledgeably and contextually in a live conversation, and thus can enable a more natural conversation between the bot and the end-user for an improved conversation experience. Instead of the end-user learning a fixed set of keywords or commands that the bot knows how to respond to, intelligent bots can understand the end-user's intent based on the end-user's utterance in natural language and respond accordingly.

[0004] Typically, an individual bot is trained as a classifier and adopts a model that predicts or infers the class or category of the input from a set of classes or categories. When creating a machine learning model, the parameters of the model that define the architecture of the model must be determined. Such parameters are called hyperparameters of the model. The process of determining the ideal settings of the hyperparameters, i.e., the values to be assigned to each hyperparameter of the model, is called hyperparameter tuning.

[0005] Standard hyperparameter tuning algorithms perform hyperparameter tuning operations with a single objective (e.g., model accuracy) in mind. To achieve this objective, the hyperparameter tuning algorithm searches for the best hyperparameter settings that optimize the single objective. Thus, the model is tuned to optimize the single objective. For different objectives, different models have to be trained, and each of them is tuned to optimize the corresponding objective.

[0006] The embodiments described herein address these and other problems individually and collectively. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0007] Overview Techniques (e.g., methods, systems, non-transitory computer-readable media storing code or instructions executable by one or more processors) for tuning hyperparameters of a machine learning model used in a chatbot system are provided. Various embodiments are described herein, including methods, systems, non-transitory computer-readable storage media storing programs, code, or instructions executable by one or more processors, and the like.

[0008] According to one aspect of the present disclosure, a method for tuning a set of hyperparameters of a machine learning model used in a chatbot system is provided. The method includes obtaining one or more data sets for training the machine learning model and selecting a plurality of metrics (or targets) for evaluating the performance of the machine learning model with respect to the one or more data sets. A first weight is assigned to each metric of the plurality of metrics. The first weight specifies the importance of each metric with respect to the performance of the machine learning model. A cost function or loss function is created for measuring the performance of the machine learning model based on the plurality of metrics and the first weights assigned to each of the plurality of metrics. A set of hyperparameters associated with the machine learning model is tuned to optimize the machine learning model with respect to the plurality of metrics. The process of tuning the set of hyperparameters includes: (i) training a machine learning model configured based on the current set of values of the set of hyperparameters on the one or more data sets; (ii) using the cost function or loss function to evaluate the performance of the machine learning model with respect to the one or more data sets; (iii) based on the evaluation, determining whether the machine learning model is optimized with respect to the plurality of metrics; (iv) in response to the machine learning model not being optimized with respect to the plurality of metrics, the tuning process searches for a new set of values for the set of hyperparameters, reconstructs the machine learning model with the new set of values, and repeats steps (i) to (iii) using the reconstructed machine learning model; and (v) in response to the machine learning model being optimized with respect to the plurality of metrics, providing the machine learning model as a trained machine learning model.

[0009] According to one aspect of the present disclosure, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions. The instructions, when executed, cause the one or more data processors to perform some or all of the one or more methods described herein.

[0010] According to another aspect of the present disclosure, there is provided a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to execute all or part of one or more of the methods described herein.

[0011] The foregoing will be more apparent when considered in conjunction with the following specification, claims, and accompanying drawings, which include other features and embodiments.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Best Mode for Carrying Out the Invention

[0013] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of particular embodiments of the invention. It will be apparent, however, that the various embodiments may be practiced without these specific details. The figures and description are not intended to be limiting. The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.

[0014] Preface A digital assistant is an AI-driven interface that helps users accomplish various tasks in natural language conversations. For each digital assistant, a customer can assemble one or more skills. Skills (also described herein as chatbots, bots, or skillbots) are individual bots that focus on specific types of tasks such as inventory tracking, timecard submission, and expense report creation. When an end user engages with a digital assistant, the digital assistant evaluates the end user input, routes the conversation to the appropriate chatbot, and routes the conversation from the appropriate chat. The digital assistant can be made available to end users via various channels such as FACEBOOK (registered trademark) Messenger, SKYPE MOBILE (registered trademark) Messenger, or Short Message Service (SMS). The channel enables chat to flow to and from the end user to the digital assistant and its various chatbots on various messaging platforms. The channel may also support user agent incrementation, event-triggered conversations, and testing.

[0015] Based on the intent, the chatbot can understand what the user wants the chatbot to do. An intent is composed of a permutation of typical user requests and statements, also referred to as utterances (such as obtaining account balance, making a purchase, etc.). As used herein, an utterance or message can refer to a set of words (e.g., one or more sentences) exchanged during a conversation with the chatbot. An intent provides a name indicating some user action (e.g., ordering a pizza) and may be created by compiling a set of real-life user statements or utterances generally associated with triggering that action. Since the chatbot's understanding is derived from these intents, each intent is created from a dataset that is robust (from one to two dozen utterances) and may vary to enable the chatbot to interpret ambiguous user input. The rich set of utterances enables the chatbot to understand what the user desires when receiving messages that mean the same thing but are expressed differently, such as "Ignore this order" or "Cancel the delivery!". Collectively, the intents and the utterances belonging to them constitute a training corpus for chat. By using the corpus to train a machine learning model, the customer can essentially convert that machine learning model (hereinafter simply referred to as the "model") into a reference tool for resolving end-user input into a single intent. The customer can improve the sensitivity of chat understanding through the cycle of intent testing and intent training.

[0016] However, building a chatbot that can determine an end user's intent based on user utterances is, in part, a difficult task due to the nuances and ambiguities of natural language, as well as the dimensionality of the input space (e.g., possible user utterances) and the size of the output space (the number of intents). Therefore, it may be necessary to design, train, monitor, debug, and retrain the chatbot in order to improve the chatbot's performance and the user experience provided by the chatbot. In conventional systems, design and training systems are provided to design and improve the model architecture and to train and retrain models of digital assistants or chatbots in spoken language understanding (SLU) and natural language processing (NLP).

[0017] In some examples, the model is trained as a classifier and configured to predict or infer, for a given input (e.g., an utterance), a class or category from a set of target classes or categories for that input. Such a classifier is typically trained to generate a distribution of probabilities for the set of target classes, where the probabilities are generated by the classifier for each target class within the set, and the generated probabilities sum to 1 (or 100% when represented as a percentage). In a classifier such as a neural network, the output layer of the neural network may use the softmax function as its activation function to generate a distribution of probability scores for the set of classes. These probabilities are also referred to as confidence scores. The class with the highest associated confidence score may be output as the answer for the input.

[0018] Model training is performed using training data (which may also be called labeled data), where the inputs and the labels (ground truths) associated with those inputs are known. For example, the training data may include inputs x(i) and, for each input x(i), a target value or correct answer (also called the ground truth) y(i) for that input. A pair (x(i), y(i)) is called a training sample, and the training data may include many such training samples. For example, the training data used to train a model for a chatbot may include a set of utterances and, for each utterance in the set, a known (ground truth) class for that utterance. The space of all inputs x(i) in the training data may be represented by X, and the space of all corresponding targets y(i) may be represented by Y. The goal of training a neural network is to learn a hypothesis function "h()" that maps the training input space X to the target value space Y such that h(x) is a good predictor of the corresponding value of y. In some implementations, as part of deriving the hypothesis function, a cost function or loss function is defined that measures the difference between the ground truth value for an input and the value predicted for that input by the model. This cost function or loss function is minimized as part of training. Training techniques such as backpropagation training techniques used with neural networks may be used, which iteratively change / manipulate the weights associated with the inputs to a perceptron in a neural network with the goal of minimizing the loss function associated with the output provided by the neural network.

[0019] As described above, standard hyperparameter tuning algorithms perform hyperparameter tuning operations with a single objective (e.g., model accuracy) in mind. A drawback of such standard hyperparameter tuning algorithms is that they ignore other important objectives (e.g., regression error) during the optimization process. Further, standard hyperparameter tuning mechanisms consider only a single dataset for training / evaluating a machine learning model. Some hyperparameter tuning algorithms consider multiple datasets while evaluating the model, but each of the datasets is considered to have the same level of importance in training the machine learning model.

[0020] Accordingly, different approaches are needed to address these problems. The present disclosure provides hyperparameter tuning systems and techniques that optimize a function while considering multiple metrics simultaneously, i.e., perform multi-objective optimization. A weight indicating the level of importance of the metric with respect to the performance of the machine learning model is assigned to each of the metrics. The hyperparameter tuning systems and techniques also provide hyperparameter tuning while considering different datasets with different levels of importance. Specifically, a weight is assigned to each dataset to specify the importance of the dataset in training the machine learning model. Additionally, the hyperparameter tuning systems and techniques enable one or more constraints on the machine learning model being trained. The training infrastructure (also referred to herein as the hyperparameter tuning system) uses various automated techniques to automatically identify, set, and tune hyperparameters to train the model such that the trained model complies with and meets the constraints specified for that model. In this way, the hyperparameter tuning system of the present disclosure provides a single machine learning model that functions across different datasets and different metrics.

[0021] Bot system A bot (also referred to as a skill, chatbot, chatterbot, or talkbot) is a computer program that can execute conversations with end users. Bots can generally respond to natural language messages (such as questions or comments) through a messaging application that uses natural language messages. Companies can communicate with end users through one or more bot systems via a messaging application. The messaging application, sometimes called a channel, can be an end user's preferred messaging application that the end user has already installed and is familiar with. Thus, the end user does not need to download and install a new application to chat with the bot system. The messaging application can include, for example, an over-the-top (OTT) messaging channel (such as Facebook Messenger, Facebook WhatsApp, WeChat, Line, Kik, Telegram, Talk, Skype, Slack, or SMS), a virtual private assistant (such as Amazon Dot, Echo, or Show, Google® Home, Apple HomePod, etc.), a native or hybrid / responsive mobile app or web app with a chat feature or a mobile and web app extension that extends the chat feature, or a voice-based input (such as a device or app with an interface that uses Siri, Cortana, Google Voice, or other voice input for conversations).

[0022] FIG. 1 is a simplified block diagram of an environment 100 incorporating a chatbot system according to a particular embodiment. The environment 100 includes a Digital Assistant Builder Platform (DABP) 102 that enables users of the DABP 102 to create and deploy digital assistants or chatbot systems. The DABP 102 can be used to create one or more digital assistants (or DAs) or chatbot systems. For example, as shown in FIG. 1, a user 104 representing a particular enterprise can use the DABP 102 to create and deploy a digital assistant 106 for users of the particular enterprise. For example, a bank can use the DABP 102 to create one or more digital assistants for use by bank customers. Multiple enterprises can use the same DABP 102 platform to create digital assistants. As another example, the owner of a restaurant (e.g., a pizza shop) can use the DABP 102 to create and deploy a digital assistant that enables customers of the restaurant to order food (e.g., order pizza).

[0023] FIG. 1 is a simplified block diagram of an environment 100 incorporating a chatbot system according to a particular embodiment. The environment 100 includes a Digital Assistant Builder Platform (DABP) 102 that enables users of the DABP 102 to create and deploy digital assistants or chatbot systems. The DABP 102 can be used to create one or more digital assistants (or DAs) or chatbot systems. For example, as shown in FIG. 1, a user 104 representing a particular enterprise can use the DABP 102 to create and deploy a digital assistant 106 for users of the particular enterprise. For example, a bank can use the DABP 102 to create one or more digital assistants for use by bank customers. Multiple enterprises can use the same DABP 102 platform to create digital assistants. As another example, the owner of a restaurant (e.g., a pizza shop) can use the DABP 102 to create and deploy a digital assistant that enables restaurant customers to order food (e.g., order pizza).

[0024] Digital assistants such as the digital assistant 106 constructed using the DABP 102 can be used to perform various tasks via a natural language-based conversation between the digital assistant and its user 108. As part of the conversation, the user may provide one or more user inputs 110 to the digital assistant 106 and receive a response 112 from the digital assistant 106. The conversation can include one or more of the inputs 110 and responses 112. Through these conversations, the user can request that one or more tasks be performed by the digital assistant 106, and in response, the digital assistant 106 is configured to execute the user-requested task and respond to the user with an appropriate response.

[0025] The user input 110 is generally in natural language form and is referred to as an utterance. The user utterance 110 can be in text form, such as when the user types a sentence, question, text snippet, or even a single word and provides the text as input to the digital assistant 106. In some embodiments, the user utterance 110 can be in voice input or spoken form, such as when the user says or speaks something that is provided as input to the digital assistant 106. The utterance is typically the language spoken by the user 108. For example, the utterance can be in English or some other language. If the utterance is in voice form, the voice input is converted to text form of the utterance in that particular language, and then the text utterance is processed by the digital assistant 106. Various speech-to-text processing techniques can be used to convert the voice or auditory input to a text utterance, and the text utterance is then processed by the digital assistant 106. In some embodiments, the conversion from voice to text can be performed by the digital assistant 106 itself.

[0026] The utterance, which can be a text utterance or a voice utterance, can be a fragment, a sentence, multiple sentences, one or more words, one or more questions, a combination of the aforementioned types, etc. The digital assistant 106 is configured to apply natural language understanding (NLU) techniques to the utterance to understand the meaning of the user input. As part of the NLU processing for the utterance, the digital assistant 106 is configured to perform processing to understand the meaning of the utterance, which involves identifying one or more intents and one or more entities corresponding to the utterance. Once the meaning of the utterance is understood, the digital assistant 106 can perform one or more actions or operations in response to the understood meaning or intent. For the purposes of this disclosure, it is assumed that the utterance is either a text utterance directly provided by the user 108 of the digital assistant 106 or the result of the conversion of an input voice utterance to text form. However, this is not intended to be limiting or restrictive in any way.

[0027] For example, the user input 108 may request that a pizza be ordered by providing an utterance such as "I want to order a pizza". Upon receiving such an utterance, the digital assistant 106 is configured to understand the meaning of the utterance and take appropriate action. Appropriate actions may include, for example, asking questions that request user input regarding the type of pizza the user wants to order, the size of the pizza, any toppings for the pizza, etc., and responding to the user. The response provided by the digital assistant 106 may also be in natural language form and may typically be in the same language as the input utterance. As part of generating these responses, the digital assistant 106 may perform natural language generation (NLG). To order a pizza, through the conversation between the user and the digital assistant 106, the digital assistant may guide the user to provide all the necessary information for ordering a pizza, and then, at the end of the conversation, may cause the pizza to be ordered. The digital assistant 106 may end the conversation by outputting to the user information indicating that the pizza has been ordered.

[0028] At the concept level, digital assistant 106 executes various processes in response to an utterance received from a user. In some embodiments, this process includes, for example, understanding the meaning of the input utterance (using NLU), determining the actions to be executed in response to the utterance, causing the actions to be executed when appropriate, generating a response to be output to the user in response to the user utterance, outputting the response to the user, etc., and involves a series of processing steps or a pipeline of processing steps. The NLU process can include parsing the received input utterance to understand the structure and meaning of the utterance, and refining and restructuring the utterance to develop a more understandable form (e.g., logical form) or structure for the utterance. Generating a response may include using natural language generation (NLG technology). Thus, the natural language processing (NLP) performed by a digital assistant can include a combination of NLU and NLG processes. The NLU process performed by a digital assistant such as digital assistant 106 can include various NLU-related processes such as syntactic analysis (e.g., tokenization, rearrangement, identification of part-of-speech tags for sentences, identification of named entities in sentences, generation of dependency trees representing sentence structures, splitting of sentences into clauses, analysis of individual clauses, resolution of anaphoric forms, execution of chunking, etc.). In one embodiment, the NLU process or a part thereof is executed by digital assistant 106 itself. In some other embodiments, digital assistant 106 can use other resources to execute a part of the NLU process. For example, the syntax and structure of the input utterance sentence may be identified by processing the sentence using syntactic analysis, part-of-speech tagging, and / or named entity recognition. In one implementation example, in the case of English, syntactic analysis, part-of-speech tagging, and named entity recognition, such as those provided by the Stanford NLP Group, are used to analyze the sentence structure and syntax. These are provided as part of the Stanford CoreNLP toolkit.

[0029] Although the various examples provided in this disclosure show English utterances, this is meant as merely an example. In certain embodiments, digital assistant 106 can also process utterances in languages other than English. Digital assistant 106 may provide subsystems (e.g., components that implement NLU functionality) configured to perform processing for different languages. These subsystems may be implemented as pluggable units that can be invoked using service calls from the NLU core server. This makes the NLU processing flexible and extensible for each language, including enabling processing in different orders. A language pack may be provided for each individual language, and the language pack can register a list of subsystems that can be serviced from the NLU core server.

[0030] Digital assistants such as digital assistant 106 shown in FIG. 1 can be made available or accessible to its user 108 via various different channels, including but not limited to, via an application, via a social media platform, via various messaging services and applications (e.g., instant messaging applications), and via other applications or channels. Since a single digital assistant can constitute several channels for it, it can be run simultaneously on different services and accessed simultaneously by different services.

[0031] Digital assistants or chatbot systems generally include, or are associated with, one or more skills. In certain embodiments, these skills are individual chatbots (referred to as skillbots) configured to interact with a user and fulfill specific types of tasks such as inventory tracking, time card submission, expense report creation, food ordering, bank account verification, reservation creation, widget purchase, and the like. For example, in the embodiment shown in FIG. 1, digital assistant or chatbot system 106 includes skills 116-1, 116-2, etc. For the purposes of the present disclosure, the term "skill" is used synonymously with the term "skillbot".

[0032] Each skill associated with a digital assistant helps a user of the digital assistant complete a task through a conversation with the user, and the conversation can include a combination of text or auditory input provided by the user and responses provided by the skillbot. These responses may be in the form of text messages or auditory messages to the user and / or may be provided using simple user interface elements (e.g., selection lists) presented to the user to make selections.

[0033] There are various ways in which skills or skillbots can be associated with or added to a digital assistant. In one example, a skillbot is developed by an enterprise and then can be added to a digital assistant using DABP102, for example, via a user interface provided by DABP102 to register the skillbot with the digital assistant. In another example, a skillbot is developed and created using DABP102 and then can be added to a digital assistant created using DABP102. In yet another example, DABP102 provides an online digital store (referred to as a "skill store") that offers a plurality of skills directed to a wide range of tasks. Skills provided through the skill store may also expose various cloud services. To add a skill to a digital assistant generated using DABP102, a user of DABP102 can access the skill store via DABP102, select a desired skill, and indicate that the selected skill is to be added to the digital assistant created using DABP102. Skills from the skill store can be added to the digital assistant as is or in a modified form (e.g., a user of DABP102 can select and clone a particular skillbot provided by the skill store, customize or modify the selected skillbot, and then add the modified skillbot to a digital assistant created using DABP102).

[0034] To implement a digital assistant or chatbot system, various different architectures may be used. For example, in one embodiment, a digital assistant created and deployed using DABP102 may be implemented using a master bot / sub (or child) bot paradigm or architecture. According to this paradigm, the digital assistant is implemented as a master bot that interacts with one or more child bots that are skill bots. For example, in the embodiment shown in FIG. 1, digital assistant 106 includes master bot 114 and skill bots 116-1, 116-2, etc. that are child bots of master bot 114. In certain embodiments, digital assistant 106 itself is considered to operate as a master bot.

[0035] A digital assistant implemented according to a master - slave bot architecture enables a user of the digital assistant to interact with multiple skills via an integrated user interface, i.e., via the master bot. When the user engages with the digital assistant, the user input is received by the master bot. Then, the master bot performs a process to determine the meaning of the user input utterance. Next, the master bot determines whether the task requested by the user in the utterance can be processed by the master bot itself. If not, the master bot selects an appropriate skill bot to process the user request and routes the conversation to the selected skill bot. Thereby, the user can converse with the digital assistant via a common single interface and still be provided with the ability to use several skill bots configured to perform specific tasks. For example, in the case of a digital assistant developed for enterprise use, the master bot of the digital assistant can interface with skill bots having specific functions such as a CRM bot for performing functions related to customer relationship management (CRM), an ERP bot for performing functions related to enterprise resource planning (ERP), and an HCM bot for performing functions related to human capital management (HCM). Thus, the end - user or consumer of the digital assistant only needs to know how to access the digital assistant via the common master bot interface, and behind it, multiple skill bots are provided to process user requests.

[0036] In one embodiment, in a master bot / slave bot infrastructure, the master bot is configured to recognize a list of available skill bots. The master bot may have access to various available skill bots and, for each skill bot, metadata identifying the capabilities of each skill bot, including the tasks that can be performed by each skill bot. When receiving a user request in the form of speech, the master bot is configured to identify or predict a particular skill bot from among the plurality of available skill bots that is most likely to respond well to the user request or best process the user request. The master bot then routes that utterance (or a portion thereof) to that particular skill bot for further processing. Thus, control flows from the master bot to the skill bot. The master bot can support multiple input and output channels. In some embodiments, the routing may be performed with the assistance of processing performed by one or more available skill bots. For example, as discussed below, a skill bot can be trained to infer the intent of an utterance and determine whether the inferred intent matches the intent for which the skill bot is configured. Accordingly, the routing performed by the master bot can include the skill bot communicating to the master bot an indication of whether it is configured with an intent suitable for processing the utterance.

[0037] The embodiment of FIG. 1 shows a digital assistant 106 including a master bot 114 and skill bots 116-1, 116-2, and 116-3, but this is not intended to be limiting. The digital assistant can include various other components (e.g., other systems and subsystems) that provide the functionality of the digital assistant. These systems and subsystems may be implemented in examples using only software (e.g., code, instructions stored on a computer-readable medium and executable by one or more processors), only hardware, or a combination of software and hardware.

[0038] DABP102 provides an infrastructure as well as various services and features that enable a user of DABP102 to create a digital assistant that includes one or more skillbots associated with the digital assistant. In some cases, a skillbot can be created by cloning an existing skillbot, for example, by cloning a skillbot provided by a skill store. As described above, DABP102 may provide a skill store or skill catalog that provides multiple skillbots for performing various tasks. A user of DABP102 can clone a skillbot from the skill store. Optionally, the cloned skillbot may be modified or customized. In some other cases, a user of DABP102 creates a skillbot from scratch using the tools and services provided by DABP102.

[0039] In certain embodiments, at a high level, creating or customizing a skillbot includes the following steps: (1) Set settings for the new skillbot (2) Set one or more intents for the skillbot (3) Set one or more entities for one or more intents (4) Train the skillbot (5) Create a dialog flow for the skillbot (6) Optionally add custom components to the skillbot (7) Test and deploy the skillbot. Each step is briefly described below.

[0040] (1) Configure settings for the new skill bot - Various settings may be configured for the skill bot. For example, the skill bot designer can specify one or more call names for the created skill bot. These call names serve as identifiers for the skill bot and can then be used by the user of the digital assistant to explicitly call the skill bot. For example, the user can include the call name in the user's utterance to explicitly call the corresponding skill bot.

[0041] (2) Configure one or more intents and associated exemplary utterances for the skill bot - The skill bot designer specifies one or more intents (also called bot intents) for the created skill bot. The skill bot is then trained based on these specified intents. These intents represent categories or classes for which the skill bot is trained to make inferences about the input utterance. When receiving an utterance, the trained skill bot infers the intent of the utterance, and the inferred intent is selected from a predefined set of intents used to train the skill bot. The skill bot then takes an appropriate action in response to the utterance based on the inferred intent. In some cases, the intents for the skill bot represent tasks that the skill bot can perform for the user of the digital assistant. Each intent is given an intent identifier or intent name. For example, in the case of a skill bot trained for a bank, the intents specified for that skill bot may include "CheckBalance", "TransferMoney", "DepositCheck", etc.

[0042] For each intent defined for the skill bot, the skill bot designer may also provide one or more exemplary utterances that represent that intent. These exemplary utterances are meant to represent utterances that a user may enter into the skill bot for that intent. For example, for the intent of balance inquiry, exemplary utterances may include "What's my savings account balance?", "How much is in my checking account?", "How much money do I have in my account?", etc. Thus, various permutations of typical user utterances may be specified as utterance examples for the intent.

[0043] Intents and their associated exemplary utterances are used as training data for training the skill bot. Various different training techniques may be used. As a result of this training, a prediction model is generated, which is configured to take an utterance as input and output the intent inferred for the utterance by the prediction model. In some cases, the input utterance is provided to an intent analysis engine (e.g., a rule-based or machine learning-based or machine learning classifier executed by the skill bot) that is configured to predict or infer the intent for the input utterance using the trained model. The skill bot may then take one or more actions based on the inferred intent.

[0044] (3) Set one or more entities for one or more intents - In some examples, additional context may be required to enable the skill bot to respond appropriately to user utterances. For example, there may be situations where user input utterances resolve to the same intent in the skill bot. For example, in the above example, the utterances "What's my savings account balance?" and "How much is in my checking account?" both resolve to the same balance inquiry intent, but these utterances are different requests with different expectations. To clarify such requests, one or more entities may be added to the intent. Using the example of a banking skill bot, an entity called AccountType that defines values called "checking" and "saving" may enable the skill bot to analyze user requests and respond appropriately. In the above example, the utterances resolve to the same intent, but the values associated with the AccountType entity are different for the two utterances. This allows the skill bot to perform different actions for the two utterances in some cases, even though they resolve to the same intent. One or more entities may be specified for a particular intent set for the skill bot. Thus, entities are used to add context to the intent itself. Entities help to more fully describe the intent and enable the skill bot to complete user requests.

[0045] In one embodiment, there are two types of entities, namely, (a) the built-in entities provided by DABP102, and (2) the custom entities that can be specified by the skill bot designer. The built-in entities are general-purpose entities that can be used with a variety of bots. Examples of built-in entities include, but are not limited to, entities related to time, date, address, number, email address, duration, cycle period, currency, phone number, URL, and the like. Custom entities are used for more customized applications. For example, for banking skills, the AccountType entity may be defined by the skill bot designer to enable various banking transactions by checking user input for keywords such as current, regular, and credit card.

[0046] (4) Training the Skill Bot - The skill bot is configured to receive user input in the form of speech, analyze or otherwise process the received input, and identify or select an intent associated with the received user input. As described above, the skill bot must be trained for this purpose. In one embodiment, the skill bot is trained based on the intents set for the skill bot and exemplary utterances associated with the intents (collectively training data), whereby the skill bot can resolve a user input utterance to one of the set intents of the skill bot. In a particular embodiment, the skill bot is trained using training data and uses a prediction model that enables the skill bot to identify what the user is saying (or in some cases, what the user is trying to say). DABP102 provides various different training techniques that can be used by skill bot designers to train the skill bot, including various machine learning-based training techniques, rule-based training techniques, and / or combinations thereof. In one embodiment, a portion of the training data (e.g., 80%) is used to train the skill bot model and another portion (e.g., the remaining 20%) is used to test or validate the model. Once trained, the trained model (sometimes referred to as the trained skill bot) can then be used to process and respond to user utterances. In some cases, the user's utterance may be a question that requires only a single answer and no further conversation. To handle such situations, a Q&A (question and answer) intent may be defined for the skill bot. The Q&A intent is generated in the same manner as a normal intent. The dialog flow for the Q&A intent may differ from the dialog flow for a normal intent. For example, unlike a normal intent, the dialog flow for the Q&A intent may not require a prompt to solicit additional information (e.g., a value for a particular entity) from the user.

[0047] (5) Create a dialog flow for the skill bot - The dialog flow specified for the skill bot describes how the skill bot reacts when different intents for the skill bot are resolved in response to the received user input. The dialog flow defines the actions or operations that the skill bot takes, such as how the skill bot responds to user utterances, how the skill bot prompts the user for input, and how the skill bot returns data. The dialog flow is like a flowchart that the skill bot follows. The skill bot designer specifies the dialog flow using a language such as the Markdown language. In one embodiment, a version of YAML called OBotML can be used to specify the dialog flow for the skill bot. The dialog flow definition for the skill bot serves as a model of the conversation itself, enabling the skill bot designer to choreograph the conversation between the skill bot and the user to whom the skill bot corresponds.

[0048] In one embodiment, the dialog flow definition of the skill bot includes three sections: (a) Context section (b) Default transition section (c) State section.

[0049] Context section - The skill bot designer can define the variables used in the conversation flow in the context section. Other variables that can be named in the context section include, but are not limited to, variables for error handling, variables for built-in entities or custom entities, user variables that enable the skill bot to recognize and persist user preferences, and the like.

[0050] Default Transition Section - Transitions for the skill bot can be defined in the dialog flow state section or the default transition section. Transitions defined in the default transition section act as a fallback and are triggered when there are no applicable transitions defined within the state or when the conditions required to trigger a state transition cannot be met. The default transition section can be used to define routing that enables the skill bot to handle unexpected user actions seamlessly.

[0051] State Section - The dialog flow and its related operations are defined as a series of transient states that manage the logic within the dialog flow. Each state node within the dialog flow definition names the components that provide the functionality required at that point in the dialog. In this way, states are constructed around the components. States include component-specific characteristics and define transitions to other states that are triggered after the component has been executed.

[0052] Special case scenarios can be handled using the state section. For example, it may be desirable to give the user the option to temporarily step out of the first skill they are engaged in and do something in a second skill within the digital assistant. For instance, if the user is involved in a conversation with a shopping skill (e.g., the user has made some selection for a purchase), the user may want to jump to a banking skill (e.g., to check if they have sufficient funds for that purchase) and then return to the shopping skill to complete their order. To handle this, the state section in the dialog flow definition of the first skill can be configured to start a conversation with a second different skill within the same digital assistant and then return to the original dialog flow.

[0053] (6) Add custom components to the skill bot - As described above, the states specified in the dialog flow for the skill bot designate the components that provide the necessary functions corresponding to those states. The components enable the skill bot to execute functions. In one embodiment, DABP102 provides a set of preconfigured components for performing a wide range of functions. The skill bot designer can select one or more of these preconfigured components and associate them with states within the dialog flow for the skill bot. The skill bot designer can also create custom or new components using the tools provided by DABP102 and associate the custom components with one or more states within the dialog flow for the skill bot.

[0054] (7) Test and deploy the skill bot - DABP102 provides several features that enable the skill bot designer to test the skill bot under development. The skill bot can then be deployed in and included in the digital assistant.

[0055] The above description explains how to create a skill bot, but using a similar technique, it is also possible to create a digital assistant (or master bot). At the master bot or digital assistant level, built-in system intents can be set for the digital assistant. These built-in system intents are used to identify common tasks that the digital assistant itself (i.e., the master bot) can handle without invoking the skill bots associated with the digital assistant. Examples of system intents defined for the master bot include the following: (1) Exit: Applies when the user wants to end the current conversation or context in the digital assistant; (2) Help: Applies when the user requests help or directions; (3) Unresolved intent: Applies to user input that does not match well with the exit intent and the help intent. The digital assistant also stores information about one or more skill bots associated with the digital assistant. This information enables the master bot to select a specific skill bot to process the utterance.

[0056] At the master bot or digital assistant level, when the user inputs a phrase or utterance to the digital assistant, the digital assistant is configured to perform a process of determining how to route the utterance and the associated conversation. The digital assistant uses a routing model, which can be rule-based, AI-based, or a combination thereof, to make this determination. The digital assistant uses the routing model to determine whether the conversation corresponding to the user input utterance should be routed to a specific skill for processing, processed by the digital assistant or the master bot itself according to the built-in system intent, or processed as a different state in the current conversation flow.

[0057] In certain embodiments, as part of this process, the digital assistant determines whether the user input utterance explicitly identifies a skill bot by its call name. If the call name is present in the user input, it is treated as an explicit call to the skill bot corresponding to the call name. In such a scenario, the digital assistant can route the user input to the explicitly called skill bot for further processing. In the absence of a specific or explicit call, in some embodiments, the digital assistant evaluates the received user input utterance and calculates confidence scores for system intents and skill bots associated with the digital assistant. The scores calculated for a skill bot or system intent represent the likelihood that the user input represents a task that the skill bot is configured to perform or a system intent. System intents or skill bots with associated calculated confidence scores that exceed a threshold (e.g., the Confidence Threshold routing parameter) are selected as candidates for further evaluation. The digital assistant then selects a specific system intent or skill bot from the identified candidates for further processing of the user input utterance. In certain embodiments, after one or more skill bots are identified as candidates, the intents associated with those candidate skills are evaluated (using the trained models for each skill) and confidence scores are determined for each intent. Generally, intents with confidence scores exceeding a threshold (e.g., 70%) are treated as candidate intents. If a specific skill bot is selected, the user utterance is routed to that skill bot for further processing. If a system intent is selected, one or more actions are performed by the master bot itself according to the selected system intent.

[0058] Constraint and Target-Based Hyperparameter Tuning As described above, the chatbot may use one or more machine learning models to perform various functions. For example, the chatbot may use a machine learning model that is configured to take an utterance as input and infer or predict the intent of each utterance. The intent inferred for an utterance by the model may then be used by the chatbot to determine how to respond to that utterance. The implementation of a machine learning model (also referred to as a model) in a chatbot typically occurs in two stages, namely, (1) a training stage in which training data is run on one or more algorithms to create a trained model, and (2) an inference stage in which the trained model is used to make predictions based on new data. A training infrastructure is generally provided to implement the training stage for training the model. The training infrastructure may be provided by a tool or application, or software, used to perform the training. The training infrastructure is configured to run the training data on one or more algorithms to train or learn the algorithms and create a model. The training infrastructure generally provides hyperparameters that manage this training process. A set of values for these hyperparameters determines the network structure for the algorithm (e.g., the number of input layers, the number of hidden layers, activation functions, etc.) and how the algorithm is trained (e.g., learning rate, number of epochs, etc.).

[0059] FIG. 2 is a diagram showing exemplary types of hyperparameters according to various embodiments. Hyperparameters may include the number of layers in a model, the type of learning algorithm used to train the model, the learning rate, the number of training epochs, the number of hidden units in each layer, and the width, i.e., the number of units in each hidden layer. Typically, a user using a training infrastructure manually sets the values of the hyperparameters. However, this can be a very difficult task that requires a very deep knowledge of the training process. As will be described next with reference to FIG. 4, hyperparameter tuning systems according to various embodiments are provided, which are configured to perform multi-objective optimization, i.e., optimize functions of multiple metrics (e.g., loss functions), in an automated manner. As shown in FIG. 3, the metrics may include a stability metric, a regression error metric, a confidence score metric, a model size metric, a training time or execution time metric, an accuracy metric, or any combination thereof.

[0060] The stability metric ensures that the training process is stable, i.e., when there are minor changes to the training data (e.g., when one training example is added or removed), the predictions made by the model do not change fundamentally. The regression error metric minimizes the number of regressions of the model, i.e., the regression error corresponds to a model misclassifying a certain input while a previous version of that model correctly classified that input. The confidence score metric ensures that the model predicts a particular example with high confidence (i.e., the model not only makes a correct prediction but also makes it with high confidence). The model size metric corresponds to the size of the trained model that should be within a user-defined threshold, e.g., within 100 megabytes, and the training time metric corresponds to the amount of time utilized when training the model. The accuracy metric ensures that the trained model achieves a certain user-defined accuracy level, e.g., 95% accuracy for a certain validation dataset. In other words, when a machine learning model is trained on a particular training dataset, a particular target accuracy is achieved for a particular validation dataset. It is understood that the training and validation datasets can be selected to cover the entire range of the chatbot's use cases, i.e., datasets ranging from very small datasets to very large datasets across the scope of the application.

[0061] The hyperparameter tuning system of the present disclosure is configured to train a machine learning model with respect to a plurality of datasets (e.g., training datasets) and evaluate the performance of the machine learning model with respect to a plurality of metrics. According to some embodiments, each of the datasets used when training the machine learning model is assigned a weight indicating the importance of the dataset when training the machine learning model. In other words, the weight assigned to a dataset corresponds to the level of influence the dataset has on the training of the machine learning model.

[0062] Furthermore, the hyperparameter tuning system is configured to assign weights to each metric utilized in multi-objective optimization. Specifically, the weight assigned to a metric indicates the importance of the metric with respect to the performance of the machine learning model. As will be described in detail below, the assignment of metrics and weights to different datasets is performed according to one or more policies that govern the hyperparameter tuning system.

[0063] Furthermore, the hyperparameter tuning system enables the specification of one or more constraints when training a machine learning model. The constraints may be requirements imposed on the machine learning model, i.e., the quality or characteristics that the user desires to achieve in the trained model. The constraints may also be related to the training process itself. The constraints may be specified before the training of the model is initiated. Thus, for a given set of constraints, the training infrastructure automatically identifies hyperparameters, sets the values of the hyperparameters, and uses various automation techniques to tune the hyperparameter values in order to train a machine learning model such that the trained machine learning model complies with and meets the set of constraints.

[0064] In some embodiments, the hyperparameter tuning system provides for validating a machine learning model over a wide range of validation / test datasets. The hyperparameter tuning system associates specific target values with one or more metrics used to evaluate the performance of the machine learning model. A hyperparameter objective function (e.g., a loss function) is constructed based on the target values of the plurality of metrics. Additionally, as will be described in detail below with reference to FIG. 4, the hyperparameter tuning system employs an asymmetric loss mechanism (i.e., a larger penalty is imposed when the target is not met compared to when a reward is given for meeting or exceeding the target) to assign weights to different metrics when validating the machine learning model.

[0065] Referring to FIG. 4, a hyperparameter tuning system according to various embodiments is shown. The hyperparameter tuning system 400 includes a dataset weight assignment unit 410, a metric selection and weight assignment unit 420, a constraint establishment unit 430, and a hyperparameter tuner 450. The hyperparameter tuner 450 includes an optimizer 451 (also referred to herein as a tuning unit) and a set of hyperparameters 455.

[0066] The hyperparameter tuning system 400 is configured to train a machine learning model (e.g., a model associated with a chatbot) with respect to a plurality of datasets (i.e., training datasets) and evaluate the performance of the machine learning model based on a plurality of metrics. The dataset weight assignment unit 410 obtains one or more datasets, such as dataset 1 405A, dataset 2 405B, and dataset K 405C, and assigns weights to each dataset according to some policies 415. The weights assigned to the datasets correspond to the importance of the datasets in the training of the machine learning model. Assigning different weights to different datasets provides that the datasets have an appropriate influence (i.e., according to their respective weights) on the training of the machine learning model.

[0067] As an example, one of the policies 415 may require the machine learning model to achieve a low regression. Thus, the dataset weight assignment unit 410 assigns a higher weight to the regression dataset (e.g., dataset 1 405A) than to another type of dataset. As another example, dataset 1 405A may correspond to a dataset obtained from a first client of the hyperparameter tuning system 400, and dataset 2 405B may correspond to a dataset obtained from a second client different from the first client. The training data included in dataset 1 and dataset 2 may correspond to different types of user utterances with respect to the context (provided by their respective clients). Assuming that one of the policies 415 indicates that the first client is more important than the second client (e.g., the first client has a higher service level agreement (SLA) with the hyperparameter tuning system 400), the dataset corresponding to the first client, i.e., dataset 1 405A, is assigned a higher weight than dataset 2 405B. Further, it is understood that the system administrator of the hyperparameter tuning system 400 determines the policy 415 before training the machine learning model. The weighted dataset 405 is provided as a first input to the hyperparameter tuner 450.

[0068] The metric selection and weighting unit 420 selects a plurality of metrics 440, such as metric 1 440A, metric 2 440B, and metric M 440C, to evaluate the performance of a machine learning model with respect to one or more data sets 405. Note that the metrics 440 correspond to metrics such as stability, accuracy, model size, regression error, etc., as shown and described with respect to FIG. 3. In one embodiment, the metric selection and weighting unit 420 selects a plurality of metrics, such as metric 1 440A, metric 2 440B, etc., from the set 440 of available metrics based on a certain criterion. For example, a first client of the hyperparameter tuning system 400 may desire to emphasize the accuracy parameter for the machine learning model, and another client of the hyperparameter tuning system 400 may desire to emphasize another metric, such as the regression error metric, for the machine learning model. The requirements of different clients may be stored as one of the policies 415, based on which the metric selection and weighting unit 420 selects a plurality of metrics for evaluating the performance of the machine learning model.

[0069] The metric selection and weighting unit 420 is further configured to assign a weight to each of the selected metrics. The weight assigned to a particular metric indicates the importance of the metric with respect to the performance of the machine learning model. In one embodiment, the metric selection and weighting unit 420 assigns weights to the metrics based on the level of importance of the clients of the hyperparameter tuning system 400. For example, if a first client is more important than a second client (i.e., the first client has a higher SLA than the second client), a higher weight is assigned to the metrics requested by the first client than to the metrics requested by the second client. As another example, consider a machine learning model trained for domain detection. In such a case, there are two metrics, namely, the in-domain recall and the out-of-domain recall used to evaluate the performance of the machine learning model. In such a domain detection model, it is often desired that the model functions well, i.e., above a certain threshold level, with respect to in-domain detection as compared to out-of-domain detection. In this case, the weighting unit 420 assigns a higher weight to the in-domain recall metric as compared to the out-of-domain recall metric. The plurality of weighted metrics are provided as a second input to the hyperparameter tuner 450.

[0070] Each of the metrics 440 is associated with a corresponding set of specifications. The set of specifications includes a plurality of specification parameters that define or characterize the metric. As shown in FIG. 4, metric 1 440A is associated with specification set 1 442A, metric 2 440B is associated with specification set 2 442B, and metric M 440C is associated with specification set M 442C. It should be understood that the set of specifications associated with a particular metric can be set independently of the sets of specifications associated with other metrics. Referring to FIG. 5, the set of specifications for a metric includes (1) a training data set, (2) a validation data set, (3) a metric definition that defines a measure of how well the model meets the target goals for the data set, e.g., the training data set and the validation data set, (4) a target score for the metric (i.e., the score of the metric that the model is expected to meet), (5) a penalty coefficient for the metric, and (6) a bonus coefficient for the metric.

[0071] According to some embodiments, the metrics may be allowed to share a training data set and / or a validation data set. However, the machine learning model is expected to produce more robust results when there is diversity in the training data set and the validation data set among different metrics. Each of the specification parameters of specification sets 442A, 442B, and 442C can be set to a specific value to achieve the desired result. For example, for a metric of regression error, the corresponding set of specifications can be set as follows: 1. The training data set is modeled based on a specific set of customers, e.g., important customers. The validation data set includes examples that are expected to be correctly classified by a previous machine learning model and widely used by customers. 2. The metric definition of the regression error is set to the accuracy rate. 3. The target score is set to 95%, that is, it is expected to be achieved such that at least 95% of the examples correctly labeled by the old version of the machine learning model are also correctly labeled by the new version of the machine learning model. 4. The penalty coefficient is set to 100 and the bonus coefficient is set to 1. In doing so, each percentage point below 95% is penalized 100 times more than the percentage point above 95%.

[0072] In one embodiment, for the stability metric, the corresponding specification set can be set as follows: 1. Since a small dataset is expected to cause higher instability, the training dataset is set to be a small-sized dataset. The validation dataset may be set to be reasonably larger in size than the training dataset. 2. The stability metric definition is set to be the standard deviation in the accuracy score of the machine learning model on the validation dataset, for example, when the machine learning model is trained 10 times. 3. The target score is set to 5%, that is, it is desired that the variation in the accuracy score of the machine learning model is at most 5%. 4. The penalty coefficient is set to 10 and the bonus coefficient is set to 1. In doing so, each percentage point that does not meet the target score incurs a loss 10 times greater than the improvement in the percentage point exceeding the target score.

[0073] In one embodiment, for the confidence score metric, the corresponding specification set can be set as follows: 1. The training dataset can range from small to large in size and can vary depending on the domain. The validation dataset may include in-domain examples belonging to those intent class labels. 2. The metric definition of the confidence score metric is set to a fraction of the sentences for which the confidence threshold of the model is greater than 25% (i.e., the prediction is accurate and can be trusted at least 25% more than other predictions). 3. The target score is set to 85%, i.e., it is desired that at least 85% of the examples are confidently labeled. 4. The penalty coefficient is set to 10 and the bonus coefficient is set to 1.

[0074] It is understood that the settings of the specification set described above are exemplary and are intended to be non - limiting. The system administrator may set each of the specification sets in any other manner based on different requirements. Further, the use of different specification sets when validating a machine - learning model is described below with reference to the hyperparameter tuner 450.

[0075] According to some embodiments, the hyperparameter tuning system 400 enables a user (e.g., a system administrator) to specify one or more constraints for hyperparameter tuning of a machine - learning model. The one or more constraints are specified via the constraint establishment unit 430. The one or more constraints are provided as a third input to the hyperparameter tuner 450. Each constraint is a requirement imposed on the hyperparameter tuning of the machine - learning model, i.e., each constraint is a requirement to be satisfied by the trained machine - learning model. When one or more constraints are given, the hyperparameter tuner 450 identifies the set of hyperparameters that affect each constraint, specifies values for the identified hyperparameters, and is configured to iteratively tune the hyperparameters until the trained machine - learning model satisfies each of the one or more constraints as described below.

[0076] In one embodiment, for each of one or more constraints, the hyperparameter tuner 450 identifies one or more hyperparameters from the set 455 of hyperparameters that affect each constraint. The hyperparameter tuner 450 identifies one or more hyperparameters that affect a constraint by varying the values of the hyperparameters and determining whether the change in the values of the hyperparameters affects the values associated with the constraint. Further, it should be understood that the first set of hyperparameters that affect the first constraint may be different from the second set of hyperparameters that affect the second constraint.

[0077] Once it has identified one or more hyperparameters that affect each constraint, the hyperparameter tuner 450 specifies values for the set 455 of hyperparameters and iteratively tunes the hyperparameters until each of the constraints is satisfied. As an example, assume that the set 455 of hyperparameters includes five hyperparameters: H = [h1, h2, h3, h4, h5]. Further, for illustration purposes, assume that the user has specified two constraints C1 and C2, and the hyperparameter tuner 450 has identified that hyperparameters h1 and h3 affect constraint C1 and hyperparameters h2, h3, and h5 affect constraint C2.

[0078] The optimizer 451 of the hyperparameter tuner 450 (also referred to herein as the tuning unit) specifies values for a set of hyperparameters H, i.e., V(H)=[v(h1), v(h2), v(h3), v(h4), v(h5)], which is also referred to herein as the hyperparameter setting. Note that the optimizer 451 assigns the initial settings of the hyperparameters in a random manner. Further, with respect to the constraint C1, the optimizer iteratively changes the values of the hyperparameters h1 and / or h3 (while maintaining the values of the hyperparameters h2, h4, and h5) until the constraint C1 is satisfied. Note that a constraint is satisfied when the values of the hyperparameters that affect the constraint meet the requirements imposed by the constraint. With respect to the constraint C2, the optimizer 451 iteratively changes / modifies the values of the hyperparameters h2, h4, and / or h5 (while maintaining the values of the hyperparameters h1 and h3) until the constraint C2 is satisfied. Note that the optimizer performs the above iterations while training the machine learning model with respect to one or more datasets. Specifically, as will be described next, the optimizer 451 determines the optimal settings of the hyperparameters that satisfy each constraint while optimizing the objective function (e.g., cost or loss function) of the machine learning model for a plurality of metrics. In one embodiment, examples of user-specified constraints include the following constraints: - The inference latency for a given batch should be less than a specific threshold. For example, the latency should be less than 50 milliseconds for a batch size of 1. - The maximum size of the trained model needs to be below some specific threshold (e.g., 10 MB). - The training time of the model should be less than some user-specific time threshold. For example, the training time should be less than 5 minutes.

[0079] Furthermore, in some embodiments, a plurality of constraints specified by a user are prioritized. For example, each constraint is assigned an importance level that reflects the fulfillment of that constraint. In some embodiments, the constraints may be specified such that the trained machine learning model will necessarily have to meet some constraints, while the fulfillment of other constraints is desired but optional.

[0080] The optimizer 451 of the hyperparameter tuner 450 constructs / formulates the objective function to be optimized. The objective function is a loss function or cost function that serves as a performance metric for training / validating the machine learning model using one or more training / validation data sets. In one embodiment, the arguments of the objective function are a set of hyperparameters associated with the machine learning model that are optimized by the optimizer 451. The value of the objective function is a weighted combination of the difference between the actual value of each metric and the target value set for each metric. The weight of each metric in the weighted combination depends on whether the metric exceeds or does not exceed the target value. Specifically, an asymmetric loss technique is utilized where a higher weight associated with not being able to achieve the target value (rather than exceeding the target value) is assigned to the metric.

[0081] For example, according to one embodiment, the objective function can be formulated as follows: where v is a vector of hyperparameter values, m i (v) is shown as the value of the i-th metric on v, and t i is shown as the target value of the i-th metric, p i and b i are shown as the penalty coefficient and bonus coefficient of the i-th metric respectively, the objective function or loss function (L(v)) can be formulated as follows:

[0082]

Equation

[0083] Optimizer 451 adjusts a set 455 of hyperparameters associated with a machine learning model to optimize an objective function (e.g., obtain a minimum value of a loss function) across a plurality of metrics. The optimizer may tune the set of hyperparameters using one of a grid-based method, a gradient search method, and a Bayesian method. Next, with reference to FIGS. 5 and 6, details regarding hyperparameter tuning will be described. In this way, hyperparameter tuner 450 trains / validates the machine learning model by tuning the set of hyperparameters, achieving optimal performance for different weighted metrics while ensuring that each of one or more constraints is satisfied. In other words, hyperparameter tuning system 400 performs optimization of a plurality of weighted metrics by training a machine learning model on different weighted datasets while supporting one or more user-specified constraints. When optimizing a machine learning model for a plurality of metrics, hyperparameter tuning system 400 outputs the trained / validated ML model along with a setting of hyperparameters that not only satisfies one or more constraints but also optimizes the objective function. In the above-described embodiment, optimizer 451 constructs and optimizes the objective function. Note that the configuration of hyperparameter tuner 450 described above does not limit the scope of the present disclosure. For example, hyperparameter tuner 450 may include an objective function formulation unit (not shown) that formulates the objective function to be optimized by optimizer 451.

[0084] FIG. 6 shows a simplified flowchart 600 illustrating a training process executed by a hyperparameter tuning system according to some embodiments. The processes shown in FIG. 6 may be implemented by software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of each system. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The methods presented in FIG. 6 and described below are intended to be exemplary and non-limiting. FIG. 6 shows various processing steps that occur in a particular sequence or order, which is not intended to be limiting. In certain alternative embodiments, those steps may be executed in any different order or some steps may be executed in parallel.

[0085] In step 610, a dataset for training a machine learning model is obtained. In step 620, the hyperparameter tuning system assigns weights to each of the obtained datasets according to a policy. Note that assigning different weights to different datasets enables the datasets to have an appropriate impact (i.e., according to their respective weights) on the training of the machine learning model.

[0086] In step 630, a plurality of metrics are selected to evaluate the performance of the machine learning model with respect to the obtained dataset. For example, one or more metrics as shown in FIG. 3 are selected by a user (e.g., a system administrator) to evaluate the performance of the machine learning model. For each selected metric, a weight is assigned to the metric according to another policy to indicate the importance of the metric with respect to the performance of the machine learning model (step 640). In step 650, the user establishes one or more constraints. Each constraint is a quality or characteristic that the user desires to achieve in the trained machine learning model. In other words, each constraint is a requirement imposed on the machine learning model.

[0087] In step 660, the process formulates / constructs a function (i.e., an objective function) based on the set 455 of input weighted metrics and hyperparameters. In one embodiment, the objective function is a loss function or a cost function that serves as a performance metric for training a machine learning model using one or more datasets.

[0088] In step 670, the process iteratively tunes the set of hyperparameters associated with the machine learning model to optimize the machine learning model (e.g., obtain the optimal value of the objective function) for a plurality of metrics. For example, when training a machine learning model on a weighted dataset, one or more hyperparameters that affect one or more constraints and / or functions are identified by varying the values of the hyperparameters and determining whether the variation in the values of the hyperparameters affects the values associated with the constraints and / or functions.

[0089] In one embodiment, the process of tuning the hyperparameters evaluates the value of the function for the current setting of the hyperparameters (i.e., the value of the hyperparameters) and determines whether the current setting satisfies each of one or more constraints. If at least one of the constraints is violated and / or the value of the function is not optimal, the tuning process modifies the values of one or more hyperparameters to obtain a new setting of the hyperparameters and continues to train the machine learning model based on the new setting. Note that the determination of whether the current setting violates a particular constraint is performed by determining whether the values of one or more hyperparameters that affect the constraint satisfy the requirements imposed by the constraint. Further, the determination of whether the value of the function (indicating the performance of the machine learning model for the current setting) is optimal is made by comparing the value of the function with new values of the function obtained through different settings of the hyperparameters.

[0090] In this way, the tuning process iterates through the space of hyperparameter values until a setting is achieved that yields the optimal value of the function (and does not violate any constraints). Further, it is understood that the process of tuning hyperparameters can start with an initial setting of the hyperparameters that is assigned in a random manner. Further, the tuning process can implement one of a random search method, a Bayesian search method, a branch-and-bound method, a grid search method, a genetic algorithm, etc. when exploring the space of hyperparameter values to obtain a new hyperparameter setting. When the machine learning model is optimized, the machine learning model is output to the user as a trained machine learning model (along with the values of the hyperparameters that achieve the optimized machine learning model).

[0091] FIG. 7 shows a simplified flowchart 700 illustrating a validation process executed by a hyperparameter tuning system, according to an embodiment. The processes shown in FIG. 7 may be implemented by software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of each system. The software may be stored on a non-transitory storage medium (e.g., on a memory device). The methods presented in FIG. 7 and described below are intended to be exemplary and non-limiting. FIG. 7 shows various processing steps that occur in a particular sequence or order, which is not intended to be limiting. In certain alternative embodiments, those steps may be executed in any different order or some steps may be executed in parallel.

[0092] In step 710, a plurality of metrics are selected to evaluate the performance of the machine learning model and regarding which hyperparameters of the machine learning model should be tuned. In step 720, a set of specifications associated with each selected metric is set according to a certain criterion. Specifically, based on the criterion, values are assigned to the specification parameters included in the set of specifications. For example, if it is desirable to reduce the regression error, the set of specifications associated with the regression error metric is set as follows: a low value is set for the target score, and a high value is set for the penalty coefficient corresponding to the regression error metric. As an example, the target score may be set to 90%, that is, it is expected that at least 90% of the training examples that were previously correctly labeled will also be correctly labeled by the current version of the machine learning model. Further, setting a high penalty coefficient, for example, a penalty coefficient of 100, means that each percentage point below 90% (i.e., the performance of the machine learning model) will be penalized 100 times more than the percentage points above 90%. When setting the set of specifications for the regression error metric in this way, the regression error dominates the objective function (i.e., the loss function), and the hyperparameter tuner in FIG. 4 performs tuning of the hyperparameters that minimize the regression error.

[0093] In step 730, a metric score is calculated for each metric. Specifically, the selected metric is evaluated on the validation dataset to generate a metric score for each metric. The metric scores calculated for each metric are compared in step 740 with the corresponding target score for that metric. In one embodiment, the difference between the metric score and the target score of the metric is calculated. If the metric score is higher than the target score, the difference is multiplied by a bonus coefficient associated with that metric. However, if the metric score is lower than the target score, the difference is multiplied by a penalty coefficient. In step 750, an objective function, such as a loss function, is formulated based on the process performed in step 740. Specifically, the loss function is determined to be the sum of the differences (between the metric score and the target score) multiplied by either the corresponding bonus coefficient or penalty coefficient. Thereafter, the process proceeds to step 760 to optimize the formulated objective function.

[0094] In step 760, the hyperparameter tuner iteratively adjusts the hyperparameters associated with the machine learning model to optimize the weighted loss function formulated in step 750 (e.g., to obtain the minimum value of the loss function). In one embodiment, when tuning the hyperparameters, the hyperparameter tuner evaluates the value of the loss function for the current settings (i.e., the values of the hyperparameters). The hyperparameter tuner further determines whether the value of the loss function is optimal by comparing the value of the loss function with new values of the loss function obtained through different settings of the hyperparameters.

[0095] In this way, the tuning process iterates through the space of hyperparameter values until a setting is achieved that results in an optimal value of the loss function. Further, the process of tuning hyperparameters can start with an initial setting of the hyperparameters that is assigned in a random manner. Additionally, the tuning process can implement a search algorithm for exploring the hyperparameter space to obtain new hyperparameter settings. The hyperparameter tuner can utilize any one of a random search method, a Bayesian search method, a branch-and-bound method, etc. in exploring the hyperparameter space. When the loss function is optimized (i.e., the minimum value of the loss function is obtained), the hyperparameter tuner provides, as output, the verified machine learning model to the user, along with the values of the hyperparameters that achieve the optimized loss function.

[0096] Exemplary System FIG. 8 shows a schematic diagram of a distributed system 800. In the illustrated example, the distributed system 800 includes one or more client computing devices 802, 804, 806, and 808 coupled to a server 812 via one or more communication networks 810. The client computing devices 802, 804, 806, and 808 can be configured to execute one or more applications.

[0097] In various examples, server 812 may be adapted to execute one or more services or software applications that enable one or more embodiments described in this disclosure. In one example, server 812 may also provide other services or software applications that may include non-virtual and virtual environments. In some examples, these services may be provided to users of client computing devices 802, 804, 806, and / or 808 as web-based services or cloud services, such as under a Software as a Service (SaaS) model. Users operating client computing devices 802, 804, 806, and / or 808 may utilize one or more client applications to interact with server 812 to utilize the services provided by these components.

[0098] In the configuration shown in FIG. 8, server 812 may include one or more components 818, 820, and 822 that implement the functions executed by server 812. These components may include software components that may be executed by one or more processors, hardware components, or combinations thereof. It should be recognized that a wide variety of system configurations may be possible that may differ from distributed system 800. Accordingly, the example shown in FIG. 8 is an example of a distributed system for implementing an example system and is not intended to be limiting.

[0099] The user may execute one or more applications, models, or chatbots using client computing devices 802, 804, 806, and / or 808, which may generate one or more events or models, which may then be realized or processed in accordance with the teachings of the present disclosure. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although FIG. 8 shows only four client computing devices, any number of client computing devices may be supported.

[0100] The client device can include various types of computing systems such as portable handheld devices, general-purpose computers such as personal computers and laptops, workstation computers, wearable devices, game systems, thin clients, various messaging devices, sensors or other sensing devices. These computing devices can include various types and versions of software applications and operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-based operating systems, Linux® or Linux-based operating systems, e.g., various mobile operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android®, BlackBerry®, Palm OS®), Google Chrome® OS). Portable handheld devices can include cellular phones, smartphones (e.g., iPhone®), tablets (e.g., iPad®), personal digital assistants (PDAs), etc. Wearable devices can include Google Glass® head-mounted displays and other devices. Game systems can include various handheld game devices, Internet-connected game devices (e.g., Microsoft Xbox® game consoles with / without Kinect® gesture input devices, Sony PlayStation® systems, various game systems provided by Nintendo®, etc.). The client device may be capable of executing a variety of applications such as various Internet-related applications, communication applications (e.g., e-mail applications, short message service (SMS) applications), and may use various communication protocols.

[0101] Network 810 can be any type of network well-known to those skilled in the art that can support data communication using any of a variety of available protocols, and the above protocols include, but are not limited to, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Exchange), AppleTalk®, etc. By way of example only, network 810 can include a Local Area Network (LAN), an Ethernet®-based network, Token Ring, Wide Area Network (WAN), the Internet, virtual network, Virtual Private Network (VPN), intranet, extranet, Public Switched Telephone Network (PSTN), infrared network, wireless network (e.g., a wireless network operating under any of the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suites, Bluetooth® and / or any other wireless protocol), and / or any combination of these and / or other networks.

[0102] Server 812 may be configured by one or more general-purpose computers, dedicated server computers (including, by way of example, PC (Personal Computer) servers, UNIX® servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or other suitable configurations and / or combinations. Server 812 can include one or more virtual machines that execute a virtual operating system, or other computing architectures involving virtualization. This can be, for example, one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server. In various examples, server 812 may be adapted to execute one or more services or software applications that provide the functions described in the above disclosure.

[0103] The computing system within server 812 can include one or more operating systems, including any of the above operating systems, and can execute a commercially available server operating system. Further, server 812 can execute any of a variety of other server applications and / or middleware applications, including an HTTP (Hypertext Transfer Protocol) server, an FTP (File Transfer Protocol) server, a CGI (Common Gateway Interface) server, a JAVA (registered trademark) server, a database server, etc. Exemplary database servers include, but are not limited to, those commercially available from Oracle (registered trademark), Microsoft (registered trademark), Sybase (registered trademark), IBM (registered trademark) (International Business Machines), etc.

[0104] In some implementations, server 812 can include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of client computing devices 802, 804, 806, and 808. As an example, the data feeds and / or event updates can include real-time events related to sensor data applications, financial stock market dashboards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc., and can include, but are not limited to, Twitter (registered trademark) feeds, Facebook (registered trademark) updates, or real-time updates received from one or more third-party information sources and continuous data streams. Server 812 can also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 802, 804, 806, and 808.

[0105] The distributed system 800 may also include one or more data repositories 814, 816. In certain examples, these data repositories can be used to store data and other information. For example, one or more of the data repositories 814, 816 can be used to store information related to the generated models used by the chatbot performance or the chatbot used by the server 812 when performing various functions according to various embodiments. The data repositories 814, 816 can be located in various places. For example, the data repository used by the server 812 may be at a local location of the server 812 or at a remote location from the server 812 and communicate with the server 812 via a network-based connection or a dedicated connection. The data repositories 814, 816 may be of different types. In certain examples, the data repository used by the server 812 may be a database, such as a relational database provided by Oracle Corporation (registered trademark) and other manufacturers. One or more of these databases may be adapted to enable storage, update, and retrieval between the data and the database in response to SQL format commands.

[0106] In certain examples, one or more of the data repositories 814, 816 may be used by an application to store application data. The data repository used by the application may be of various types, such as, for example, a key-value store repository, an object store repository, or a general-purpose storage repository supported by a file system.

[0107] In certain examples, the functions described in this disclosure may be provided as services via a cloud environment. FIG. 9 is a simplified block diagram of a cloud-based system environment that may provide various services as cloud services according to a particular example. In the example shown in FIG. 9, the cloud infrastructure system 902 may provide one or more cloud services that a user may request using one or more client computing devices 904, 906, and 908. The cloud infrastructure system 902 may include one or more computers and / or servers that may include what was previously described with respect to server 812. The computers within the cloud infrastructure system 902 may be configured as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.

[0108] Network 910 may facilitate the communication and exchange of data between clients 904, 906, and 908 and the cloud infrastructure system 902. Network 910 may include one or more networks. The networks may be of the same type or different types. Network 910 may support one or more communication protocols, including wired and / or wireless protocols, to facilitate communication.

[0109] The example shown in FIG. 9 is merely an example of a cloud infrastructure system and is not intended to be limiting. It should be understood that in some other examples, the cloud infrastructure system 902 may have more or fewer components than shown in FIG. 9, may combine two or more components, or may have components with different configurations or arrangements. For example, although FIG. 9 shows three client computing devices, in alternative examples, any number of client computing devices may be supported.

[0110] The term "cloud service" is generally used to refer to services that are made available on demand to users via a communication network such as the Internet by a service provider's system (e.g., cloud infrastructure system 902). Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premises servers and systems. The cloud service provider's system is managed by the cloud service provider. Thus, customers can utilize cloud services provided by the cloud service provider without having to separately purchase licenses, support, or hardware and software resources for the services. For example, the cloud service provider's system can host applications, and users can order and use applications on demand via the Internet without having to purchase infrastructure resources to run the applications. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, several cloud services such as middleware services, database services, Java (registered trademark) cloud services, etc. are offered by Oracle Corporation (registered trademark) of Redwood Shores, California.

[0111] In certain examples, cloud infrastructure system 902 can provide one or more cloud services using various models such as the software as a service (SaaS) model, the platform as a service (PaaS) model, the infrastructure as a service (IaaS) model, including a hybrid service model. Cloud infrastructure system 902 can include a suite of applications, middleware, databases, and other resources that enable the provisioning of various cloud services.

[0112] The SaaS model enables applications or software to be delivered as a service to customers over a communication network such as the Internet without the customer having to purchase the underlying hardware or software for the application. For example, by using the SaaS model, customers can be given access to on-demand applications hosted by a cloud infrastructure system 902. Examples of SaaS services provided by Oracle Corporation (registered trademark) include, but are not limited to, various services for human resources / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, social applications, and the like.

[0113] The IaaS model is generally used to provide flexible computing and storage capabilities by providing infrastructure resources (such as servers, storage, hardware, and networking resources) to customers as cloud services. Various IaaS services are provided by Oracle Corporation (registered trademark).

[0114] The PaaS model is generally used to provide as a service platform and environmental resources that enable customers to develop, run, and manage applications and services without having to procure, build, or manage the environmental resources. Examples of PaaS services provided by Oracle Corporation (registered trademark) include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), data management cloud services, various application development solution services, and the like.

[0115] Cloud services are generally provided in an on-demand self-service, subscription-based, flexibly scalable, reliable, highly available, and secure manner. For example, a customer may order one or more services provided by a cloud infrastructure system 902 via a subscription order. The cloud infrastructure system 902 then provides the services requested in the customer's subscription order by performing processing. For example, a user can use speech to cause the cloud infrastructure system to take a specific action (such as an intent) as described above and / or request that a service for a chatbot system be provided as described herein. The cloud infrastructure system 902 may be configured to provide one cloud service or multiple cloud services.

[0116] The cloud infrastructure system 902 may provide cloud services via various deployment models. In the public cloud model, the cloud infrastructure system 902 may be owned by a third-party cloud service provider, and the cloud services are provided to general public customers. This customer may be an individual or an enterprise. In another example, under the private cloud model, the cloud infrastructure system 902 may function within an organization (such as within a corporate organization), and the services are provided to customers within this organization. For example, this customer may be various departments of an enterprise such as the human resources department or the payroll department, or an individual within the enterprise. In another example, under the community cloud model, the cloud infrastructure system 902 and the provided services may be shared by various organizations within the relevant community. Other various models such as a hybrid model of the above models may also be used.

[0117] Client computing devices 904, 906, and 908 may be of different types (e.g., client computing devices 802, 804, 806, and 808 shown in FIG. 8), and may be operable with one or more client applications. A user may interact with the cloud infrastructure system 902, such as by using the client device to request services provided by the cloud infrastructure system 902. For example, a user may use the client device to request information or actions from a chatbot, as described in this disclosure.

[0118] In some examples, the processes executed by the cloud infrastructure system 902 to provide services may include model training and deployment. This analysis may include training and deploying one or more models by using, analyzing, and processing a dataset. This analysis may be performed by one or more processors, optionally in parallel, and may include executing simulations using the data. For example, big data analysis may be performed by the cloud infrastructure system 902 to generate and train one or more models for a chatbot system. The data used for this analysis may include structured data (e.g., data stored in a database or data structured according to a structured model) and / or unstructured data (e.g., a data blob (binary large object)).

[0119] As shown in the example of FIG. 9, the cloud infrastructure system 902 may include infrastructure resources 930 that are utilized to facilitate the provisioning of various cloud services provided by the cloud infrastructure system 902. The infrastructure resources 930 may include, for example, processing resources, storage or memory resources, networking resources, and the like. In a particular example, a storage virtual machine that is available to process storage requested from an application may be part of the cloud infrastructure system 902. In other examples, the storage virtual machine may be part of a different system.

[0120] In a particular example, to facilitate the efficient provisioning of these resources to support the various cloud services provided by the cloud infrastructure system 902 to different customers, the resources may be grouped into sets of resources or resource modules (also referred to as "pods"). Each resource module or pod may include a pre-integrated and optimized combination of one or more types of resources. In a particular example, different pods may be pre-provisioned for different types of cloud services. For example, a first set of pods may be provisioned for database services, and a second set of pods, which may include a different combination of resources than the pods within the first set of pods, may be provisioned for Java services or the like. For some services, the resources allocated to provision these services may be shared among the services.

[0121] The cloud infrastructure system 902 itself may internally use a service 932 that is shared by different components of the cloud infrastructure system 902 and facilitates the provisioning of services by the cloud infrastructure system 902. These internal shared services may include, but are not limited to, security identity services, integration services, enterprise repository services, enterprise manager services, virus scan whitelist services, high availability, backup recovery services, services enabling cloud support, email services, notification services, file transfer services, etc.

[0122] The cloud infrastructure system 902 may include a plurality of subsystems. These subsystems may be implemented in software, or in hardware, or a combination thereof. As shown in FIG. 9, the subsystem may include a user interface subsystem 912 that enables a user or customer of the cloud infrastructure system 902 to interact with the cloud infrastructure system 902. The user interface subsystem 912 may include various different interfaces such as a web interface 914, an online store interface 916 where cloud services provided by the cloud infrastructure system 902 are advertised and available for purchase by consumers, and other interfaces 918. For example, a customer may use a client device to request (service request 934) one or more services provided by the cloud infrastructure system 902 using one or more of the interfaces 914, 916, and 918. For example, a customer may access an online store, browse the cloud services provided by the cloud infrastructure system 902, and place a subscription order for one or more services provided by the cloud infrastructure system 902 and desired by the customer to subscribe to. This service request may include information identifying the customer and one or more services the customer desires to subscribe to. For example, a customer may place an order for a service provided by the cloud infrastructure system 902. As part of the order, the customer may provide information identifying the chatbot system through which the service will be provided and, optionally, one or more qualification information for the chatbot system.

[0123] In a specific example, such as the example shown in FIG. 9, the cloud infrastructure system 902 may include an order management subsystem (OMS) 920 configured to process new orders. As part of this process, the OMS 920 creates a customer account if one has not already been created, receives billing and / or account information from the customer that is used to bill the customer for providing the requested service to the customer, verifies the customer information, and after verification, reserves the order for the customer and may be configured to prepare the order for provisioning by coordinating various workflows.

[0124] Once properly authenticated, the OMS 920 may call an order provisioning subsystem (OPS) 924 configured to provision resources for this order, including processing, memory, and networking resources. Provisioning may include allocating resources for the order and configuring the resources to facilitate the service requested by the customer order. The way resources are provisioned for an order and the type of resources provisioned may depend on the type of cloud service the customer ordered. For example, according to a certain workflow, the OPS 924 may be configured to determine the specific cloud service requested and identify the number of pods that would be preconfigured for this specific cloud service. The number of pods allocated for an order may depend on the size / volume / level / range of the service requested. For example, the number of pods to allocate may be determined based on the number of users the service is to support, the period for which the service is requested, etc. Next, the allocated pods may be customized for the specific customer making the request to provide the requested service.

[0125] In certain examples, the setup stage processing can be performed by the cloud infrastructure system 902 as part of the provisioning process as described above. The cloud infrastructure system 902 can generate an application ID and select a storage virtual machine for the application from among the storage virtual machines provided by the cloud infrastructure system 902 itself or from storage virtual machines provided by other systems outside the cloud infrastructure system 902.

[0126] The cloud infrastructure system 902 may send a response or notification 944 to the requesting customer to indicate when the requested service will be available. In some examples, information (e.g., a link) enabling the customer to initiate use and utilization of the benefits of the requested service may be sent to the customer. In certain examples, for a customer requesting a service, the response may include a chatbot system ID generated by the cloud infrastructure system 902 and information identifying the chatbot system selected by the cloud infrastructure system 902 for the chatbot system corresponding to the chatbot system ID.

[0127] The cloud infrastructure system 902 can provide services to multiple customers. For each customer, the cloud infrastructure system 902 manages information related to one or more subscription orders received from the customer, maintains customer data related to the orders, and serves to provide the requested services to the customer. Also, the cloud infrastructure system 902 may collect usage statistics regarding use of the subscribed services by the customers. For example, the statistics may be collected regarding the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime. This usage information may be used to bill the customers. Billing may be performed, for example, monthly.

[0128] The cloud infrastructure system 902 may provide services to multiple customers in parallel. The cloud infrastructure system 902 may store information about these customers, which may include copyright information in some cases. In a particular example, the cloud infrastructure system 902 includes an Identity Management Subsystem (IMS) 928 configured to manage customer information and separate the information being managed so that information about one customer is not accessible from information about another customer. The IMS 928 may be configured to provide various security-related services such as identity services like information access management, authentication and authorization services, services for managing customer identities and roles and related capabilities, etc.

[0129] FIG. 10 shows an example of a computer system 1000. In some examples, the computer system 1000 may be either any digital assistant or chatbot system within a distributed environment, and may be used to implement the various servers and computer systems described above. As shown in FIG. 10, the computer system 1000 includes various subsystems including a processing subsystem 1004 that communicates with several other subsystems via a bus subsystem 1002. These other subsystems may include a processing acceleration unit 1006, an I / O subsystem 1008, a storage subsystem 1018, and a communication subsystem 1024. The storage subsystem 1018 may include a non-transitory computer-readable storage medium including a storage medium 1022 and a system memory 1010.

[0130] The bus subsystem 1002 provides a mechanism for enabling the various components and subsystems of the computer system 1000 to communicate with each other as intended. Although the bus subsystem 1002 is schematically shown as a single bus, alternative examples of bus subsystems may utilize multiple buses. The bus subsystem 1002 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a local bus, etc., using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, which may be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.

[0131] The processing subsystem 1004 controls the operation of the computer system 1000 and may include one or more processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The processor may include a single-core or multi-core processor. The processing resources of the computer system 1000 can be organized into one or more processing units 1032, 1034, etc. The processing unit may include one or more processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some examples, the processing subsystem 1004 may include one or more dedicated coprocessors such as a graphics processor, a digital signal processor (DSP), etc. In some examples, some or all of the processing units of the processing subsystem 1004 may use customized circuits such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).

[0132] In some examples, the processing units within the processing subsystem 1004 may execute instructions stored in the system memory 1010 or the computer-readable storage medium 1022. In various examples, the processing units may execute various programs or code instructions and may maintain multiple programs or processes to be executed simultaneously. At any given point in time, some or all of the program code to be executed may reside in the system memory 1010 and / or the computer-readable storage medium 1022 that potentially includes one or more storage devices. Through appropriate programming, the processing subsystem 1004 may provide the various functions described above. In an example where the computer system 1000 is executing one or more virtual machines, one or more processing units may be assigned to each virtual machine.

[0133] In certain examples, a processing acceleration unit 1006 can optionally be provided to execute customized processing to accelerate the overall processing performed by the computer system 1000, or to offload a portion of the processing performed by the processing subsystem 1004.

[0134] The I / O subsystem 1008 can include devices and mechanisms for inputting information into the computer system 1000 and / or outputting information from, or via, the computer system 1000. In general, the use of the term "input device" is intended to include all conceivable types of devices and mechanisms for inputting information into the computer system 1000. User interface input devices can include, for example, a keyboard, a mouse or other pointing device such as a trackball, a touchpad or touch screen incorporated into a display, a scroll wheel, a click wheel, a dial, a button, a switch, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices can also include motion sensing and / or gesture recognition devices such as a Microsoft Kinect (registered trademark) motion sensor that enables a user to control and interact with the input device, a Microsoft Xbox (registered trademark) 360 game controller, a device that provides an interface for receiving input using gestures and voice commands. User interface input devices can also include gesture recognition devices such as a Google Glass (registered trademark) blink detector that detects a user's eye movement (e.g., a "blink" while taking a photo and / or making a menu selection) and converts the eye gesture into an input to the input device (e.g., Google Glass (registered trademark)). Also, user interface input devices can include a voice recognition sensing device that enables a user to interact with a voice recognition system (e.g., a Siri (registered trademark) navigator) via voice commands.

[0135] Other examples of user interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, game pads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser range finders, and eye-tracking devices. Also, the user interface input device may include medical imaging input devices such as, for example, computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound examination devices. The user interface input device may also include audio input devices such as, for example, MIDI keyboards, digital musical instruments, etc.

[0136] In general, the use of the term output device is intended to include all possible types of devices and mechanisms for outputting information from the computer system 1000 to the user or another computer. The user interface output device may include, for example, a display subsystem, indicator lights, or non-visual displays such as audio output devices. The display subsystem may be a flat panel device using a cathode ray tube (CRT), liquid crystal display (LCD), or plasma display, a projected device, a touch screen, etc. For example, the user interface output device may include, but is not limited to, various display devices for visually conveying text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, audio output devices, and modems.

[0137] Storage subsystem 1018 provides a repository or data store for storing information and data used by computer system 1000. Storage subsystem 1018 provides a tangible non-transitory computer-readable storage medium for storing basic programming and data configurations that provide some example functionality. Software (e.g., programs, code modules, instructions) that provides the functionality described above when executed by processing subsystem 1004 may be stored in storage subsystem 1018. The software may be executed by one or more processing units of processing subsystem 1004. Storage subsystem 1018 may also provide authentication in accordance with the teachings of the present disclosure.

[0138] Storage subsystem 1018 may include one or more non-transitory memory devices including volatile and non-volatile memory devices. As shown in FIG. 10, storage subsystem 1018 includes system memory 1010 and computer-readable storage medium 1022. System memory 1010 may include several memories including volatile main random access memory (RAM) for storing instructions and data during program execution and non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some implementations, a basic input / output system (BIOS) including basic routines that assist in the transfer of information between elements within computer system 1000, such as during startup, may typically be stored in ROM. Typically, RAM includes data and / or program modules that are currently being operated on and executed by processing subsystem 1004. In some implementations, system memory 1010 may include multiple different types of memories such as static random access memory (SRAM), dynamic random access memory (DRAM), and the like.

[0139] As an example, without limitation, as shown in FIG. 10, the system memory 1010 may load an application program 1012, program data 1014, and an operating system 1016 that may include various applications such as a web browser, a middle layer application, a relational database management system (RDBMS), etc. As an example, the operating system 1016 may include Microsoft Windows (registered trademark), Apple Macintosh (registered trademark) and / or Linux operating systems, various commercially available UNIX (registered trademark) or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google Chrome (registered trademark) OS, etc.), and / or various versions of mobile operating systems such as iOS (registered trademark), Windows Phone, Android (registered trademark) OS, BlackBerry (registered trademark) OS, Palm (registered trademark) OS operating systems, etc.

[0140] The computer-readable storage medium 1022 can store programming and data configurations that provide the functionality of several examples. The computer-readable storage medium 1022 can provide storage of computer-readable instructions, data structures, program modules, and other data for the computer system 1000. Software (programs, code modules, instructions) that provides the above functionality when executed by the processing subsystem 1004 may be stored in the storage subsystem 1018. As an example, the computer-readable storage medium 1022 may include non-volatile memory such as a hard disk drive, a magnetic disk drive, an optical disk drive such as a CD ROM, a DVD, a Blu-Ray (registered trademark) disk, or other optical media. The computer-readable storage medium 1022 may include, but is not limited to, a Zip (registered trademark) drive, a flash memory card, a universal serial bus (USB) flash drive, a secure digital (SD) card, a DVD disk, a digital video tape, etc. The computer-readable storage medium 1022 may also include solid state drives (SSDs) based on non-volatile memory such as flash memory-based SSDs, enterprise flash drives, solid state ROMs, SSDs based on volatile memory such as solid state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs.

[0141] In certain examples, the storage subsystem 1018 may further include a computer-readable storage medium reader 1020 that is connectable to the computer-readable storage medium 1022. The reader 1020 may be configured to receive and read data from a memory device such as a disk, a flash drive, etc.

[0142] In certain examples, computer system 1000 may support virtualization techniques including, but not limited to, virtualization of processing and memory resources. For example, computer system 1000 may provide support for running one or more virtual machines. In certain examples, computer system 1000 may execute a program such as a hypervisor that facilitates the configuration and management of virtual machines. Memory, computing (e.g., processors, cores), I / O, and networking resources may be allocated to each virtual machine. Each virtual machine typically runs independently of other virtual machines. A virtual machine may run its own operating system, which may be the same as or different from the operating systems run by other virtual machines typically executed by computer system 1000. Thus, potentially multiple operating systems may be executed simultaneously by computer system 1000.

[0143] Communication subsystem 1024 provides an interface to other computer systems and networks. Communication subsystem 1024 functions as an interface for sending and receiving data between other systems and computer system 1000. For example, communication subsystem 1024 may enable computer system 1000 to establish a communication channel to one or more client devices via the Internet for sending and receiving information between the computer system 1000 and the one or more client devices. For instance, if computer system 1000 is used to implement the bot system 120 shown in FIG. 1, the communication subsystem may be used to communicate with the chatbot system selected for the application.

[0144] The communication subsystem 1024 may support both wired and / or wireless communication protocols. In one example, the communication subsystem 1024 may include, for example, radio frequency (RF) transceiver components for accessing a wireless voice and / or data network using cellular phone technology, advanced data network technologies such as 3G, 4G or EDGE (Enhanced Data rates for Global Evolution), WiFi (IEEE802.XX family of standards, or other mobile communication technologies, or any combination thereof), a Global Positioning System (GPS) receiver component, and / or other components. In some examples, in addition to or instead of a wireless interface, the communication subsystem 1024 may provide a wired network connection (e.g., Ethernet (registered trademark)).

[0145] The communication subsystem 1024 may receive and transmit data in various formats. In some examples, in addition to other formats, the communication subsystem 1024 may receive input communications in the form of structured data feeds and / or unstructured data feeds 1026, event streams 1028, event updates 1030, etc. For example, the communication subsystem 1024 may be configured to receive (or transmit) data feeds 1026 in real-time from users of other communication services such as social media networks and / or Twitter (registered trademark) feeds, Facebook (registered trademark) updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.

[0146] In certain examples, the communication subsystem 1024 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1028 and / or event updates 1030 of real-time events that are inherently continuous or infinite and have no explicit end. Examples of applications that generate continuous data include, for example, sensor data applications, financial stock market dashboards, network performance measurement tools (such as network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and the like.

[0147] The communication subsystem 1024 may be configured to transmit data from the computer system 1000 to other computer systems or networks. This data may be communicated to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 1000 in various different forms such as structured and / or unstructured data feeds 1026, event streams 1028, event updates 1030, and the like.

[0148] The computer system 1000 may be any one of various types, including a handheld portable device (e.g., an iPhone (registered trademark) cellular phone, an iPad (registered trademark) computing tablet, a PDA), a wearable device (e.g., a Google Glass (registered trademark) head-mounted display), a personal computer, a workstation, a mainframe, a kiosk, a server rack, or other data processing systems. Since the nature of computers and networks is constantly changing, the description of the computer system 1000 shown in FIG. 10 is only intended as a specific example. Many other configurations are possible with more or fewer components than the system shown in FIG. 10. It should be recognized that there are other aspects and / or methods for implementing various examples based on the disclosure and teachings herein.

[0149] Although specific examples have been described, various modifications, changes, alternative configurations, and equivalents are possible. The examples are not limited to operating within a specific data processing environment and can operate freely within multiple data processing environments. Further, although the examples have been described using a specific set of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. Some of the flowcharts describe the operations as sequential processes, but many of these operations may be performed in parallel or simultaneously. Additionally, the order of the operations may be re-specified. The process may have additional steps not included in the figures. The various features and aspects of the above examples may be used individually or together.

[0150] Furthermore, while specific examples have been described using specific combinations of hardware and software, it should be understood that other combinations of hardware and software are possible. The specific examples may be implemented using only hardware, only software, or a combination thereof. The various processes described herein may be implemented on the same processor or different processors in any combination.

[0151] When a device, system, component, or module is described as being configured to perform a particular operation or function, such a configuration can be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, by programming computer instructions or code that execute code or instructions stored in a non-transitory memory medium or any combination thereof, or by executing a processor or core, etc. The processes can communicate using a variety of techniques including, but not limited to, conventional techniques for inter-process communication, different pairs of processes may use different techniques, and the same pair of processes may use different techniques at different times.

[0152] In the present disclosure, examples are made to be well understood by showing specific details. However, the examples can be implemented without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary details so as not to obscure the examples. This specification provides only exemplary examples and is not intended to limit the scope, applicability, or configuration of other examples. Rather, the above description of the examples provides those skilled in the art with an explanation that enables the implementation of various examples. Various changes are possible within the scope of the functions and configurations of the elements.

[0153] Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. However, it will be apparent that additions, deletions, omissions, as well as other modifications and alterations can be made to them without departing from the broader spirit and scope as set forth in the claims. Thus, specific examples have been described, but these are not intended to be limiting. Various variations and equivalents are within the scope of the appended claims.

[0154] In the foregoing specification, specific examples of aspects of the present disclosure have been described with reference thereto, but one of ordinary skill in the art will recognize that the present disclosure is not limited thereto. The various features and aspects of the foregoing disclosure may be used individually or in combination. Further, the examples can be utilized in a variety of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

[0155] In the foregoing description, for purposes of illustration, methods were described in a particular order. In alternative examples, it should be understood that the methods may be performed in an order different from that described. Also, the foregoing methods may be performed by hardware components or may be embodied in a sequence of machine-executable instructions that, when used, cause a machine such as a general or special purpose processor or logic circuitry programmed with such instructions to perform the methods. These machine-executable instructions can be stored on one or more machine-readable media such as a CD-ROM or other type of optical disk, a floppy disk, ROM, RAM, EPROM, EEPROM, magnetic or optical card, flash memory, or other type of machine-readable media suitable for storing electronic instructions. Alternatively, these methods may be performed by a combination of hardware and software.

[0156] If an element is described as being configured to perform a particular operation, such a configuration may be achieved, for example, by designing an electronic circuit or other hardware to perform the particular operation, programming a programmable electronic circuit (such as a microprocessor or other suitable electronic circuit) to perform the particular operation, or any combination thereof.

[0157] Examples for the description of the present application have been described in detail herein, but it should be understood that the concept of the present invention can be embodied and adopted in various other aspects, and the claims are intended to be construed to include such variations, except when limited by the prior art.

Claims

1. A method comprising: obtaining one or more data sets for training a machine learning model; selecting a plurality of metrics for evaluating the performance of the machine learning model with respect to the one or more data sets; each metric of the plurality of metrics is associated with specification parameters including a target score of the metric, a bonus coefficient of the metric, and a penalty coefficient of the metric, and the method further comprises: evaluating the performance of the machine learning model using the one or more data sets and obtaining a metric score for each metric of the plurality of metrics; assigning to each metric of the plurality of metrics a first weight based on the penalty coefficient or the bonus coefficient of the metric and the difference between the metric score of the metric and the target score of the metric, the first weight specifying the importance of each metric with respect to the performance of the machine learning model; creating a cost function or a loss function for measuring the performance of the machine learning model based on the plurality of metrics and the first weights assigned to each of the plurality of metrics; tuning a set of hyperparameters associated with the machine learning model to optimize the machine learning model with respect to the plurality of metrics, the tuning comprising: (i) training the machine learning model with the one or more data sets, the machine learning model being configured based on a current set of values of the set of hyperparameters, and the tuning further comprising: (ii) using the cost function or the loss function to evaluate the performance of the machine learning model with respect to the one or more data sets; (iii) determining, based on the evaluation, whether the machine learning model is optimized with respect to the plurality of metrics; (iv) in response to the machine learning model not being optimized with respect to the plurality of metrics, exploring a new set of values for the set of hyperparameters, reconfiguring the machine learning model with the new set of values, and repeating steps (i) to (iii) using the reconfigured machine learning model. A method comprising: (v) providing the machine learning model as a trained machine learning model in response to the machine learning model being optimized for the plurality of metrics. **Claim 2** The plurality of metrics at least include the size of the machine learning model, the training time of the machine learning model, the accuracy rate of the machine learning model, the stability of the machine learning model, the regression error of the machine learning model, and the confidence score of the machine learning model, the method according to claim 1. **Claim 3** The method according to claim 1 or 2, further comprising assigning a second weight to each of the one or more data sets, the second weight specifying the importance of each data set when training the machine learning model. **Claim 4** The method according to any one of claims 1 to 3, further comprising establishing one or more constraints based on one or more of the hyperparameters in the set of hyperparameters, and the optimized machine learning model satisfies each of the one or more constraints. **Claim 5** The first of the one or more constraints corresponds to requiring that the model size of the machine learning model is less than a threshold model size, or requiring that the training time of the machine learning model is less than a threshold time limit, the method according to claim 4. **Claim 6** The machine learning model is a neural network model, and the set of hyperparameters at least includes the number of layers of the machine learning model, the learning rate of the machine learning model, the number of hidden units in each layer of the machine learning model, and the learning algorithm used to train the machine learning model, the method according to any one of claims 1 to 5. **Claim 7** A computing device, comprising a processor and a memory containing instructions that, when executed by the processor, cause the computing device to, at least, acquire one or more data sets for training a machine learning model, select a plurality of metrics for evaluating the performance of the machine learning model with respect to the one or more data sets, Each of the plurality of metrics is associated with specification parameters including a target score of the metric, a bonus coefficient of the metric, and a penalty coefficient of the metric. When the instruction is further executed by the processor, the computing device is caused to, at least, evaluate the performance of the machine learning model using the one or more data sets and obtain a metric score for each of the plurality of metrics, assign a first weight to each of the plurality of metrics based on the penalty coefficient or the bonus coefficient of the metric and the difference between the metric score of the metric and the target score of the metric, and the first weight specifies the importance of each metric with respect to the performance of the machine learning model, create a cost function or a loss function for measuring the performance of the machine learning model based on the plurality of metrics and the first weights assigned to each of the plurality of metrics, tune a set of hyperparameters associated with the machine learning model to optimize the machine learning model for the plurality of metrics, and tuning the set of hyperparameters (i) includes training the machine learning model with the one or more data sets, the machine learning model being configured based on a current set of values of the set of hyperparameters, and tuning the set of hyperparameters further (ii) includes evaluating the performance of the machine learning model with respect to the one or more data sets using the cost function or the loss function, (iii) includes determining, based on the evaluation, whether the machine learning model is optimized for the plurality of metrics, (iv) includes, in response to the machine learning model not being optimized for the plurality of metrics, searching for a new set of values for the set of hyperparameters, reconfiguring the machine learning model with the new set of values, and repeating steps (i) to (iii) using the reconfigured machine learning model. A computing device comprising: (v) providing the machine learning model as a trained machine learning model in response to the machine learning model being optimized for the plurality of metrics.

8. The computing device according to claim 7, wherein the plurality of metrics includes at least the size of the machine learning model, the training time of the machine learning model, the accuracy rate of the machine learning model, the stability of the machine learning model, the regression error of the machine learning model, and the confidence score of the machine learning model.

9. The computing device according to claim 7 or 8, wherein the processor is further configured to assign a second weight to each of the one or more data sets, and the second weight specifies the importance of each data set when training the machine learning model.

10. The computing device according to any one of claims 7 to 9, wherein the processor is further configured to establish one or more constraints based on one or more of the set of hyperparameters, and the optimized machine learning model satisfies each of the one or more constraints.

11. The computing device according to claim 10, wherein a first constraint among the one or more constraints requires that the model size of the machine learning model is less than a threshold model size, or requires that the training time of the machine learning model is less than a threshold time limit.

12. The computing device according to any one of claims 7 to 11, wherein the machine learning model is a neural network model, and the set of hyperparameters includes at least the number of layers of the machine learning model, the learning rate of the machine learning model, the number of hidden units in each layer of the machine learning model, and the learning algorithm used to train the machine learning model.

13. A program for causing a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • weighted profit evaluator on training data

    JP2017500637A

  • Information processing method and information processing system

    JP2019096285A

  • Machine learning hyperparameter tuning tool

    US20190236487A1