Irrelevant utterance detection in chatbot system

The master bot system with a classifier model efficiently routes user inputs to relevant chatbots, addressing inefficiencies in handling irrelevant queries and enhancing user experience by optimizing resource utilization and response time.

JP2025098075APending Publication Date: 2025-07-01ORACLE INT CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025040153
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-09-10
Filing Date
2025-03-13
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing chatbot systems face inefficiencies in handling irrelevant user inputs, leading to wasted computational resources and suboptimal user experiences due to misidentification of user queries by general-purpose chatbots.

Method used

A master bot system equipped with a classifier model that utilizes training feature vectors and set representations to determine whether an input utterance is relevant to any skillbot, routing irrelevant inputs directly or prompting users for clarification, thereby optimizing resource utilization and improving user interaction.

Benefits of technology

The system effectively filters out irrelevant inputs, conserves computational and network resources, and enhances user experience by quickly routing queries to appropriate chatbots, reducing processing time and improving overall system responsiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098075000001_ABST
    Figure 2025098075000001_ABST
Patent Text Reader

Abstract

To provide a system and method for determining whether or not an input utterance is irrelevant to a set of skill bots associated with a master bot.SOLUTION: A system comprises a training system 350 and a master bot 114. The training system 350 trains a classifier of the master bot 114, in which the training includes accessing a training utterance associated with the skill bot and generating a training feature vector from the training utterance. The master bot 114 executes generating an input feature vector together with accessing the input utterance and comparing the input feature vector with a plurality of sets of expressions using the classifier, to thereby determine whether or not the input feature is out of a range and hence it is impossible to cope therewith by the skill bot.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cross - reference to Related Applications This application claims the benefit and priority of U.S. Provisional Application No. 62 / 899,700, filed on September 12, 2019, entitled "Detecting Unrelated Utterances in a Chatbot System", under 35 U.S.C. § 119(e), the content of which is incorporated herein by reference for all purposes.

[0002] Background A chatbot is an artificial - intelligence - based software application or device that provides an interface for conversations with human users. A chatbot can be programmed to perform various tasks in response to user input provided during the conversation. The user input can be supplied in various forms, including, for example, voice input and text input. For this reason, natural language understanding, speech - to - text conversion, and other language - processing technologies can be employed as part of the processing performed by the chatbot. In some computing environments, multiple chatbots are available for interacting with a user, and each chatbot addresses a different set of tasks.

Summary of the Invention

[0003] Summary The technology described herein is for determining that an input utterance from a user is not related to any of the skillbots, also referred to as chatbots, within a set of one or more skillbots available to a master bot. In some embodiments, the master bot can evaluate the input utterance and determine whether the input utterance is unrelated to the skillbots or route the input utterance to an appropriate skillbot.

Means for Solving the Problems

[0004] In some embodiments, the systems described herein include a training system and a master bot. The training system is configured to train a classifier model. Training the classifier model includes accessing training utterances associated with the skill bots, where the training utterances include respective training utterances associated with each of the skill bots. Each skill bot is configured to provide an interaction with a user. Training further includes generating training feature vectors from the training utterances. The training feature vectors include respective training feature vectors associated with each of the skill bots. Training further includes generating a plurality of set representations of the training feature vectors. Each set representation of the plurality of set representations corresponds to a subset of the training feature vectors. Training further includes configuring the classifier model to compare an input feature vector with the plurality of set representations. The master bot is configured to access an input utterance as user input and generate an input feature vector from the input utterance. The master bot further is configured to compare the input feature vector with the plurality of set representations of the training feature vectors using the classifier model and output an indication that the skill bots are unable to handle the user input based on the input feature vector being outside the range of the plurality of set representations.

[0005] In additional or alternative embodiments, the methods described herein include the step of accessing, by a computer system, training utterances associated with skill bots including P, the training utterance includes respective subsets of training utterances for each skill bot of the skill bots. Each skill bot is configured to provide an interaction with a user. The method further includes generating a training feature vector from the training utterance, the training feature vector including respective training feature vectors for each training utterance. The method further includes determining a centroid position for a cluster in a feature space and assigning each training feature vector to a respective cluster having the centroid position to which the training feature vector is closest among the clusters. The method further includes repeatedly modifying the cluster until a stop condition is met. The step of modifying the cluster includes increasing the count of the cluster to an updated count, determining a new centroid position for the cluster by an amount equal to the updated count, and reassigning the training feature vector to the cluster based on proximity to the new centroid position. The method further includes determining a boundary of the cluster, the boundary including respective boundaries for each cluster of the cluster. The method further includes accessing an input utterance, converting the input utterance into an input feature vector, and determining that the input feature vector is outside the range of the boundary of the cluster by comparing the input feature vector with the boundary of the cluster. Additionally, the method includes outputting an indication that the skill bot cannot handle the input utterance based on the input feature vector being outside the range of the cluster in the feature space.

[0006] In further additional or alternative embodiments, the methods described herein include the step of a computer system accessing training utterances associated with a skillbot. The training utterances include respective subsets of training utterances for each skillbot of the skillbots. Each skillbot is configured to provide an interaction with a user. The method further includes the step of generating a training feature vector from the training utterances, the training feature vector including respective training feature vectors for each training utterance of the training utterances, the method further includes the step of dividing the training utterances into conversation categories. The method further includes the step of generating a synthetic feature vector corresponding to the conversation categories. The step of generating the synthetic feature vector includes, for each conversation category of the conversation categories, generating a respective synthetic feature vector as a set of the respective training feature vectors of the training utterances within the conversation category. The method further includes the steps of accessing an input utterance, converting the input utterance into an input feature vector, and determining that the input feature vector is not sufficiently similar to the synthetic feature vector by comparing the input feature vector with the synthetic feature vector. Additionally, the method includes the step of outputting an indication that the skillbot is unable to handle the input utterance based on the input feature vector not being sufficiently similar to the synthetic feature vector.

[0007] In further additional or alternative embodiments, the system described herein includes a master bot. The master bot is configured to perform operations including accessing an input utterance as user input, generating an input feature vector from the input utterance, comparing the input feature vector with multiple set representations of training feature vectors using a classifier model, and outputting an instruction indicating that the skill bot cannot handle the user input based on the input feature vector being outside the range of the multiple set representations. Each skill bot of the skill bots is configured to provide interaction with the user.

[0008] In further additional or alternative embodiments, the method described herein includes a step of accessing an input utterance The method further includes a step of converting the input utterance into an input feature vector. The method further includes a step of determining that the input feature vector is outside the range of the boundary of the cluster in the feature space by comparing the input feature vector with the boundary of the cluster in the feature space. The method further includes a step of outputting an instruction indicating that the skill bot cannot handle the input utterance based on the input feature vector being outside the range of the cluster in the feature space. Each skill bot of the skill bots is configured to provide interaction with the user.

[0009] In further additional or alternative embodiments, the method described herein includes the step of accessing an input utterance. The method further includes the step of converting the input utterance into an input feature vector. The method further includes the step of determining that the input feature vector is not sufficiently similar to the composite feature vector by comparing the input feature vector with the composite feature vector. The composite feature vector corresponds to a conversation category. The method further includes the step of outputting an indication that the input utterance cannot be handled by the skill bot based on the input feature vector not being sufficiently similar to the composite feature vector. Each skill bot of the skill bot is configured to provide a conversation with the user.

[0010] In further additional or alternative embodiments, the method described herein includes the step of accessing an input utterance as user input. The method further includes the step of generating an input feature vector from the input utterance. The method further includes the step of comparing the input feature vector with a plurality of set representations of the training feature vector using a classifier model. The method further includes the step of outputting an indication that the user input cannot be handled by the skill bot based on the input feature vector being outside the range of the plurality of set representations. Each skill bot of the skill bot is configured to provide a conversation with the user.

[0011] In further additional or alternative embodiments, the method described herein is used to train a classifier model. The method includes accessing training utterances associated with a skill bot, where the training utterances include respective training utterances associated with each skill bot of the skill bot. Each skill bot of the skill bot is configured to provide an interaction with a user. The method further includes generating training feature vectors from the training utterances, where the training feature vectors include respective training feature vectors associated with each skill bot of the skill bot. The method further includes generating a plurality of set representations of the training feature vectors. Each set representation of the plurality of set representations corresponds to a subset of the training feature vectors. The method further includes configuring the classifier model to compare an input feature vector with the plurality of set representations of the training feature vectors.

[0012] In further additional or alternative embodiments, the method described herein is used to generate a cluster that can be used to determine whether a skill bot can handle an input utterance. The method includes accessing, by a computer system, training utterances associated with a skill bot, where the training utterances include respective subsets of training utterances for each skill bot of the skill bot. Each skill bot of the skill bot is configured to provide an interaction with a user. The method further includes generating training feature vectors from the training utterances, where the training feature vectors include respective training feature vectors for each training utterance of the training utterances. The method further includes determining a centroid position for clusters in the feature space. The method The method further includes, for each of the clusters, assigning each training feature vector of the training feature vectors to the cluster having the centroid position closest to the training feature vector among the clusters, and repeatedly modifying the clusters until a stop condition is satisfied. The step of modifying the clusters includes increasing the count of the clusters to an updated count, determining a new centroid position for the clusters by an amount equal to the updated count, and reassigning the training feature vectors to the clusters based on proximity to the new centroid position. The method further includes determining the boundaries of the clusters, where the boundaries include respective boundaries for each of the clusters of the clusters.

[0013] In further additional or alternative embodiments, the method described herein can be used to generate a synthetic feature vector that can be used to determine whether a skill bot can handle an input utterance. The method further includes accessing, by a computer system, training utterances associated with the skill bot, where the training utterances include respective subsets of training utterances for each of the skill bots of the skill bots. Each of the skill bots of the skill bots is configured to provide a dialogue with a user. The method further includes generating training feature vectors from the training utterances, where the training feature vectors include respective training feature vectors for each of the training utterances of the training utterances. The method further includes splitting the training utterances into conversation categories. The method further includes generating synthetic feature vectors corresponding to the conversation categories. The step of generating the synthetic feature vectors includes generating, for each of the conversation categories of the conversation categories, each synthetic feature vector as a set of the respective training feature vectors of the training utterances within the conversation category.

[0014] In further additional or alternative embodiments, the system described herein includes means for performing any of the methods described above.

[0015] The foregoing will become more apparent when considered in conjunction with other features and embodiments, with reference to the following specification, appended claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Modes for Carrying Out the Invention

[0017] Detailed Description In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and the description are not intended to be limiting. The term "exemplary" as used herein is employed to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments or designs. This specification describes various embodiments including methods, systems, programs, codes, or instructions executable by one or more processors, and non-transitory computer-readable storage media storing the same.

[0018]

[0019] As described above, a particular environment includes a plurality of chatbots, also referred to herein as skillbots, and thus each chatbot is specialized to handle a respective set of tasks or skills. In that case, it would be advantageous to automatically direct a user input from a user to the chatbot most suitable for handling that user input, and further to quickly identify when the user input is irrelevant to the chatbots and thus cannot be handled by any of the available chatbots. Some embodiments described herein enable a computing system, such as a master bot, to perform preliminary processing on a user input to determine whether the user input is irrelevant to the chatbots (i.e., cannot be handled by the chatbots). As a result, computing resources are not wasted by the chatbots attempting to process irrelevant user inputs. Thus, the particular embodiments described herein save the computing resources of the chatbots and detect irrelevant user inputs early to prevent chatbots in an environment from processing such user inputs. Further, the particular embodiments save network resources by preventing the master bot from sending such user inputs to one or more chatbots that cannot handle such user inputs.

[0020] In an environment that includes multiple chatbots, a master chatbot may include a classifier that determines which chatbot should process an input utterance. For example, the classifier can be implemented as a neural network or some other ML model that outputs a set of probabilities. Each probability is associated with a chatbot and indicates a level of confidence that the chatbot can handle the input utterance. This type of master chatbot selects the chatbot with the highest confidence level and forwards the input utterance to that chatbot. However, the master chatbot may also accidentally provide a high confidence score for an irrelevant input utterance. This is because the confidence is substantially split among the available chatbots without any consideration of whether the chatbot is configured to handle the input utterance. This can result in the chatbot processing an input utterance that it is not defined to handle. The chatbot may ultimately provide an indication that it cannot handle the input utterance or a request for clarification from the user. However, this can occur after the chatbot has consumed resources to process the input utterance.

[0021] Alternatively, the developer may train the master chatbot to help it recognize utterances that cannot be handled by any of the available chatbots. For example, the master chatbot can be trained with labeled training data that includes relevant utterances (i.e., those that can be handled by a chatbot) and irrelevant utterances (i.e., those that cannot be handled by a chatbot) to educate the master chatbot to recognize irrelevant utterances. However, the training data The range of data is likely to be insufficient. Input utterances that do not resemble any of the irrelevant utterances used during training may ultimately be transferred to the chatbot for processing. As a result, the computational resources for training the master bot will increase, and the master bot will still fail to recognize a wide range of irrelevant utterances.

[0022] The specific embodiments described herein address the drawbacks of the above-described technologies and can be used instead of, or in conjunction with, such technologies. In a particular embodiment, a classifier model, as referred to herein, a trained machine-learning (ML) model, is used to determine whether an input utterance provided as user input (i.e., a natural language phrase that may be in text form) is related to any of the available skillbots within a set of skillbots. This determination can be made, for example, by a master bot that utilizes the classifier model. The master bot generates an input feature vector from the input utterance. In one example, the classifier model of the master bot compares the feature vector with a set of clusters of training feature vectors, which are the feature vectors of the training data, to determine whether the input feature vector falls within the range of any of the clusters. If the input feature vector is outside the range of all the clusters, the master bot determines that the input utterance is unrelated to the skillbots. In another example, the classifier model of the master bot compares the input feature vector with a set of synthetic feature vectors. Each synthetic feature vector represents one or more training feature vectors belonging to each category. If the input feature vector is not sufficiently similar to any of the synthetic feature vectors, the master bot determines that the input utterance is unrelated to the skillbots. If the input utterance is determined to be unrelated to any of the bots, the input utterance may be considered to be in the "none class" and not routed to any of the handling bots. Instead, the processing of the input utterance may end, or the master bot may prompt the user to clarify what the user intended.

[0023] Such improved routing can prevent computational resources from being diverted to queries that may ultimately not be effectively addressed by the system. This can improve the overall responsiveness of the system. Additionally, the user can be guided to rephrase or clarify the query so that it can be more easily addressed by a properly trained specialized chatbot or set of chatbots, thus improving the user experience. For this reason, in addition to saving processing and network resources, the average time to handle user queries can be shortened by such improved routing through the bot network.

[0024] If the input utterance is determined to be related to at least one bot, the input utterance may be subject to intent classification. In intent classification, the intent that most closely matches the input utterance is determined in order to initiate the conversation flow associated with that intent. For example, each intent of a skill bot may be associated with a state machine that defines various conversation states regarding the conversation with the user. Intent classification can be performed at the individual bot level. For example, each skill bot registered with the master bot may have its own classifier (e.g., an ML-based classifier) trained with respect to a predetermined utterance associated with that particular bot. The input utterance can be input to the intent classifier of the bot that is most closely related to the input utterance in order to determine which of the bot's intents best matches the utterance.

[0025] As a result, the user experience can be improved. Because the user's query can be better equipped to handle the user's query than a non-specialized general bot or can handle the user's query more quickly than a non-specialized general bot. This is because it can be quickly routed to the selected skill bot or set of selected skill bots. With improved routing, additionally or alternatively, as a result, fewer computing resources may be used. This is because the processing resources consumed by the selected skill bot or selected set of skill bots when handling a user query can be less than those of a non-specialized general bot.

[0026] Overview of an exemplary chatbot system FIG. 1 is a block diagram of an environment including a master bot 114 that communicates with various skill bots 116, also referred to as chatbots, in accordance with some embodiments described herein. This environment includes a digital assistant builder platform (DABP) 102 that enables developers to program and deploy a digital assistant (DA) 106 or a chatbot system. The DA 106 includes or accesses a master bot 114 that includes or accesses one or more skill bots 116. Each skill bot 116 is configured to provide one or more skills or tasks to a user. In some embodiments, the master bot 114 and the skill bots 116 are executed on top of the DA 106 itself. However, alternatively, only the master bot 114 is executed on the DA 106 and communicates with skill bots 116 executed elsewhere (e.g., on another computing device).

[0027] In some embodiments, DABP102 can be used to program one or more DAs106. For example, as shown in FIG. 1, a developer can use DABP102 to create and deploy a digital assistant 106 for users to access. For example, DABP102 can be used by a bank to create one or more digital assistants for use by bank customers. The same DABP102 platform can be used by multiple companies to create digital assistants 106. As another example, the owner of a restaurant (e.g., a pizza parlor) may use DABP102 to create and deploy a digital assistant that enables restaurant customers to order food (e.g., order pizza). Additionally or alternatively, for example, DABP102 can be used by a developer to deploy one or more skill bots 116 such that one or more skill bots 116 can access the master bot 114 of an existing digital assistant 106.

[0028] Additionally or alternatively, in some embodiments, as further described below, DABP102 is configured to enable the master bot 114 of a digital assistant to recognize when an input utterance is unrelated to any of the available skill bots 116 by training the master bot 114.

[0029] For the purposes of the present disclosure, a "digital assistant" is an entity that assists a user of the digital assistant in accomplishing various tasks through a natural language conversation. The digital assistant can be implemented using only software (e.g., the digital assistant is a digital entity realized using a program, code, or instructions executable by one or more processors), using hardware, or using a combination of hardware and software. The digital assistant can be embodied or realized in a variety of physical systems or devices that may include general-purpose or dedicated hardware such as a computer, a mobile phone, a wristwatch, an appliance, a vehicle, etc. The digital assistant may also be referred to as a chatbot system. Thus, for the purposes of the present disclosure, the terms digital assistant and chatbot system may be synonymous.

[0030] In some embodiments, the digital assistant 106 can be used to perform various tasks through a natural language-based conversation between the digital assistant and its user 108. As part of the conversation, the user may provide a response 112 to the user input 110 and optionally (e.g., if the user input includes instructions for performing a task) provide the user input 110 to the digital assistant 106 that can perform one or more tasks related to the user input 110. The conversation or dialogue may include one or more user inputs 110 and responses 112. Through the conversation, the user can request one or more tasks to be performed by the digital assistant 106, and the digital assistant 106 is configured to, in response, perform the tasks requested by the user and reply to the user with an appropriate response.

[0031] User input 110 can be in the form of natural language called speech. The speech can be in text form when the user types a sentence, question, text fragment, or even a single word such as a phrase and provides the text as an input to digital assistant 106. In some embodiments, the user speech can be in voice input or speech form when the user speaks something provided as an input to digital assistant 106. Digital assistant 106 can include or access a microphone for capturing such speech. The speech is typically the language spoken by the user. When the speech is in the form of voice input, the voice input may be converted to text speech in the same language, and digital assistant 106 may process the text speech as user input. Various speech-to-text processing techniques can be used to convert voice input to text speech. In some embodiments, the speech-to-text conversion is performed by digital assistant 106 itself, but various implementations are within the scope of the present disclosure. For the purposes of the present disclosure, the input speech (i.e., the speech provided as user input) is assumed to be either text speech directly provided by user 108 of digital assistant 106 or the result of converting input voice speech to text form. However, this is not intended to limit or restrict in any way.

[0032] An utterance can be a fragment, a sentence, multiple sentences, one or more words, one or more questions, or a combination of the above types. In some embodiments, a digital assistant 106 including corresponding master bot 114 and skill bot 116 is configured to apply natural language understanding (NLU) techniques to an utterance to understand the meaning of the user input. As part of the NLU processing for the utterance, the digital assistant 106 may perform processing to understand the meaning of the utterance, which involves identifying one or more intents and one or more entities corresponding to the utterance. Once the meaning of the utterance is understood, the digital assistant 106 including the master bot 114 and skill bot 116 may perform one or more actions or operations according to the understood meaning or intent.

[0033] For example, the user input may request an order for pizza by providing an input utterance such as "I want to order pizza". Upon receiving such an utterance, the digital assistant 106 determines the meaning of the utterance and takes an appropriate action. Appropriate actions may include, for example, responding to the user with a question that requests user input regarding the type of pizza the user wants to order, the size of the pizza, or any toppings for the pizza. The response 112 provided by the digital assistant 106 may also be in natural language form and typically in the same language as the input utterance. As part of generating these responses 112, the digital assistant 106 may perform natural language generation (NLG). For the user to order pizza, through the conversation between the user and the digital assistant 106, the digital assistant may guide the user to provide all the information necessary to order pizza, and further, at the end of the conversation, may cause the pizza to be ordered. The digital assistant 106 may end the conversation by outputting to the user information indicating that the pizza has been ordered.

[0034] At the conceptual level, digital assistant 106, along with its master bot 114 and associated skill bots 116, performs various processes in response to utterances received from the user. In some embodiments, this process includes, for example, understanding the meaning of the input utterance (using NLU), determining the actions to be performed in response to the utterance, causing the actions to be executed at the appropriate time, generating a response to be output to the user in response to the utterance, and outputting the response to the user, including a series or sequence of processing steps. NLU processing may include parsing the input utterance received to understand the structure and meaning of the utterance, and refining and improving the utterance to develop a form (e.g., logical form) that is more easily parsed and understood. Generating a response may include using NLG technology. Thus, the natural language processing performed by digital assistant 106 may include a combination of NLU processing and NLG processing.

[0035] The NLU processing performed by digital assistant 106 may include various NLU processes such as parsing (e.g., tokenization, classification by headword, identification of part-of-speech tags, identification of named entities, generation of a dependency tree to represent the sentence structure, splitting of the utterance into clauses of the sentence, analysis of individual clauses, resolution of anaphora, execution of chunking, etc.). In certain embodiments, the NLU processing or a part thereof is performed by digital assistant 106 itself. Additionally or alternatively, digital assistant 106 may use other resources to perform some parts of the NLU processing. For example, the syntax and structure of the input utterance may be identified by processing the input utterance using a parser, a part-of-speech tagging tool, or a named entity recognizer separate from digital assistant 106.

[0036] Although the various examples provided in this disclosure show utterances in English, this is intended as an example only. In certain embodiments, digital assistant 106 can, additionally or alternatively, handle utterances in languages other than English. Digital assistant 106 can provide subsystems (e.g., components that implement NLU functionality) configured to perform processing for various languages. These subsystems may be implemented as pluggable units that can be invoked using service calls from an NLU core server that can run on digital assistant 106. This makes the NLU processing flexible and extensible for each language, including enabling processing in various orders. Language packs may be provided for individual languages. The language packs can register a list of subsystems that can be provided from the NLU core server.

[0037] In some embodiments, digital assistant 106 can be made available or accessible to its user 108 via various different channels, such as via a particular application, via a social media platform, via various messaging services and applications (e.g., an instant messaging application), or via other applications or channels. A single digital assistant can have several channels through which it can be run and simultaneously accessed by different services. Additionally or alternatively, digital assistant 106 may be implemented on a device local to the user and thus may be a personal digital assistant used by the user or other nearby users.

[0038] Digital assistant 106 may be associated with one or more skills. In certain embodiments, these skills are passed through individual chatbots referred to as skillbots 116. are implemented thereby, each of which is configured to interact with a user to perform a specific type of task such as tracking inventory, submitting a time card, creating an expense report, ordering food, checking a bank account, making a reservation, purchasing a widget, or other tasks. For example, in the embodiment shown in FIG. 1, digital assistant 106 includes or accesses three skill bots 116, each of which implements a specific skill or set of skills. However, various amounts of skill bots 116 may be supported by digital assistant 106. Each skill bot 116 may be implemented as hardware, software, or a combination of both.

[0039] Each skill associated with and thereby implemented as a skill bot 116 for digital assistant 106 is configured to assist the user of digital assistant 106 in completing a task through a conversation with the user. The conversation may include a combination of user input 110, which may be text or voice provided by the user, and a response 112 provided by skill bot 116 via digital assistant 106. These responses 112 may be in the form of a text message or a voice message to the user, or may be provided using simple user interface elements (e.g., a selection list) presented to the user for the user to make a selection.

[0040] There are various ways in which the skill bot 116 can be associated with or added to the digital assistant. In some cases, the skill bot 116 can be developed by an enterprise and further added to the digital assistant 106 using the DABP 102 via, for example, a user interface provided by the DABP 102 for registering the skill bot 116 with the digital assistant 106. In other cases, the skill bot 116 can be developed and created using the DABP 102 and further added to the digital assistant 106 using the DABP 102. In still other cases, the DABP 102 provides an online digital store (referred to as a "skill store") that offers a plurality of skills directed to a wide range of tasks. Skills provided via the skill store can also expose various cloud services. To add a skill to the digital assistant 106 using the DABP 102, a developer can access the skill store via the DABP 102, select a desired skill, and indicate that the selected skill should be added to the digital assistant 106. Skills from the skill store can be added to the digital assistant as is or in a modified form. The DABP 102, which is added to the digital assistant 106 in the form of the skill bot 116 to add a skill, can configure the master bot 114 of the digital assistant 106 to communicate with the skill bot 116. In addition, the DABP 102 can configure the master bot 114 with skill bot data that enables the master bot 114 to recognize utterances that the skill bot 116 can handle and thereby update the data used by the master bot 114 to determine whether an input utterance is irrelevant to any of the skill bots 116 of the digital assistant 106. The operations for configuring the master bot 114 in this way will be described in more detail below.

[0041] The skill bots 116 available for use with the digital assistant 106 and the skills thereby can be very different. For example, in the case of the digital assistant 106 developed for an enterprise, the master bot 114 of the digital assistant can interface with skill bots 116 having specific functions, such as a CRM bot for executing functions related to customer relationship management (CRM), an ERP bot for executing functions related to enterprise resource planning (ERP), an HCM bot for executing functions related to human capital management (HCM), and the like. Various other skills can also be available for the digital assistant and can be and can be influenced by the intended use of the digital assistant 106.

[0042] To implement the digital assistant 106, various architectures can be used. In certain embodiments, the digital assistant 106 may be implemented using a master - child paradigm or architecture. According to this paradigm, the digital assistant 106 functions as the master bot 114 by including or accessing the master bot 114 and interacts with one or more child bots that are the skill bots 116. The skill bots 116 may or may not be executed directly on the digital assistant 106. In the example shown in FIG. 1, the digital assistant 106 includes (i.e., accesses and uses) the master bot 114 and three skill bots 116. However, an excessive amount of skill bots 116 can also be used, and that amount can change over time depending on whether skills are added to or removed from the digital assistant 106.

[0043] The digital assistant 106 implemented according to the master-child architecture enables the user of the digital assistant 106 to interact with multiple skill bots 116, thereby enabling the use of multiple skills that can be implemented separately through an integrated user interface, i.e., via the master bot 114. In some embodiments, when the user engages with the digital assistant 106, the user input is received by the master bot 114. The master bot 114 then performs preprocessing to determine the meaning of the input utterance acting as the user input. In this case, the input utterance can be, for example, the user input itself or a text version of the user input. The master bot 114 determines whether the input utterance is irrelevant to the skill bots 116 available for that input utterance. This can apply, for example, when the input utterance requires a skill other than the skills of the skill bots 116. If the input utterance is irrelevant to the skill bots 116, the master bot 114 may return to the user an indication that the input utterance is irrelevant to the skill bots 116. For example, the digital assistant 106 can ask the user to clarify or report that the input utterance was not understood. However, if the master bot 114 identifies an appropriate skill bot 116, the master bot 114 may route the input utterance and the ongoing conversation thereby to that skill bot 116. This enables the user to interact with the digital assistant 106 having multiple skill bots 116 via a common interface.

[0044] The embodiment of FIG. 1 shows a digital assistant 106 that includes a master bot 114 and a skill bot 116, but this is not intended to be limiting. In some embodiments, the digital assistant 106 may include various other components such as other systems or subsystems that provide the functionality of the digital assistant 106. These systems and subsystems may be implemented in software only (e.g., as code stored on a computer-readable medium and executable by one or more processors), in hardware, or in a combination of software and hardware in implementation examples.

[0045] In certain embodiments, the master bot 114 is configured to recognize available skill bots 116. For example, the master bot 114 may access metadata that identifies the various available skill bots 116 and the capabilities of the skill bots 116, including the tasks that each skill bot 116 can perform. When receiving a user request in the form of an input utterance, the master bot 114 may identify or predict a particular skill bot 116 from among the plurality of available skill bots 116 that may best function or respond to the user request, or alternatively, in an alternative example, it may be configured to determine that the input utterance is unrelated to any of the skill bots 116. If the skill bot 116 is determined to be able to handle the input utterance, the master bot 114 may route the input utterance or at least a portion of the input utterance to that skill bot 116 for further handling. Thus, control continues from the master bot 114 to the skill bot 116.

[0046] In some embodiments, DABP102 provides an infrastructure that enables developer users of DABP102 to create a digital assistant 106 that includes one or more skillbots 116, as well as various services and features. In some cases, skillbots 116 can be created by cloning existing skillbots 116, for example, by cloning skillbots 116 provided in a skill store. As described above, DABP102 can provide a skill store that provides multiple skillbots 116 for performing various tasks. Users of DABP102 can clone skillbots 116 from the skill store and, if necessary, the cloned skillbots 116 may be modified or customized. In some other cases, developer users of DABP102 create skillbots 116 from a zero state, such as by using tools and services provided by DABP102.

[0047] In certain embodiments, creating or customizing a skillbot 116 at a high level involves the following operations.

[0048] (1) Configure settings for the new skillbot. (2) Configure one or more intents for the skillbot.

[0049] (3) Configure one or more entities for one or more intents. (4) Train the skillbot.

[0050] (5) Create a dialogue flow for the skillbot. (6) Optionally, add custom components to the skillbot.

[0051] (7) Test and deploy the skillbot. Each of the above operations will be briefly described below.

[0052] (1) Configure settings for the new skill bot 116. Various settings can be configured for the skill bot 116. For example, a skill bot developer can specify one or more call names for the skill bot 116 being created. These call names, which function as identifiers for the skill bot 116, can further be used by the user of the digital assistant 106 to explicitly call the skill bot 116. For example, the user can include the call name in the user's input utterance to explicitly call the corresponding skill bot 116.

[0053] (2) Configure one or more intents for the skill bot 116 and related exemplary utterances. The designer of the skill bot 116 specifies one or more intents, also referred to as chatbot intents, for the skill bot 116 being created. The skill bot 116 is then trained based on these specified intents. These intents represent the categories or classes for which the skill bot 116 is trained to infer about input utterances. When receiving an utterance, the trained skill bot 116 infers (i.e., determines) the intent of the utterance. In this case, the inferred intent is selected from a set of predefined intents used to train the skill bot 116. The skill bot 116 then takes an appropriate action in response to the utterance based on the intent inferred for the utterance. In some cases, the intents for the skill bot 116 represent tasks that the skill bot 116 can perform for the user of the digital assistant. Each intent is given an intent identifier or intent name. For example, in the case of a skill bot 116 trained for a bank, the intents specified for the skill bot 116 can include "CheckBalance", "TransferMoney", "DepositCheck", etc.

[0054] For each intent defined for the Skillbot 116, the designer of the Skillbot 116 may also provide one or more exemplary utterances that represent and illustrate the intent. These exemplary utterances are intended to represent the utterances that a user may input to the Skillbot 116 regarding that intent. For example, in the case of the CheckBalance intent, exemplary utterances may include "What's my savings account balance?", "How much is in my checking account? ?", "How much money do I have in my account?", etc. Thus, various permutations of typical user utterances can be specified as utterance examples for the intent.

[0055] The intents and their associated exemplary utterances, also referred to as training utterances, are used as training data for training the Skillbot 116. Various different training techniques may be used. As a result of this training, a prediction model is generated that is configured to receive an utterance as input and output the intent inferred for that utterance. In some cases, the input utterance provided as user input is input into an intent analysis engine (e.g., a rule-based or ML-based classifier executed by the Skillbot 116) that is configured to predict or infer the intent regarding the input utterance using the trained model. The Skillbot 116 may then take one or more actions based on the inferred intent.

[0056] (3) Configure entities for one or more intents of the Skillbot 116. In some cases, additional context may be required to enable the Skillbot 116 to respond appropriately to the input utterance. For example, there may be situations where multiple input utterances result in the same intent in the Skillbot 116. For example, utterances such as "What is the balance of my regular savings account?" and "How much is in my current account?" both result in the same CheckBalance intent, but these are different requests for different things. To clarify such requests, one or more entities can be added to the intent. An entity called Account_Type that defines values such as "Current Account" and "Savings" can, using the example of the banking Skillbot 116, enable the Skillbot 116 to analyze user requests and respond appropriately. In the above example, multiple utterances result in the same intent, but in the case of these two utterances, the values associated with the Account_Type entity are different. This enables the Skillbot 116 to perform different actions for these two utterances in some cases, even though they result in the same intent. One or more entities can be specified for a particular intent configured for the Skillbot 116. Thus, entities are used to add context to the intent itself. Entities help to more fully describe the intent and enable the Skillbot 116 to fulfill user requests.

[0057] In certain embodiments, (a) an embedded entity that can be provided by the DABP 102 There are two types of entities, namely, (1) embedded entities, and (2) custom entities that can be specified by developers. Embedded entities are general-purpose entities that can be used with a variety of skill bots 116. Examples of embedded entities include entities related to time, date, address, number, email address, duration, cycle period, currency, phone number, uniform resource locator (URL), and the like. Custom entities are used for more customized applications. For example, in the case of banking skills, an Account_Type entity can be defined by a developer to support various banking transactions by checking user input for keywords such as current account, savings account, and credit card.

[0058] (4) Train the skill bot 116. The skill bot 116 is configured to receive user input in the form of speech, analyze or process the received user input, and identify or select an intent associated with the received user input. In some embodiments, as described above, the skill bot 116 must be trained for this purpose. In certain embodiments, the skill bot 116 is trained based on intents configured for the skill bot 116 and exemplary utterances (i.e., training utterances) associated with the intents, whereby the skill bot 116 can break down an input utterance into one of its configured intents. In certain embodiments, the skill bot uses a prediction model trained using training data to enable the skill bot to identify what the user is saying (or, in some cases, what the user is trying to say). The DABP 102 can provide various different training techniques that developers can use to train the skill bot 116, including various ML-based training techniques, rule-based training techniques, or combinations thereof. In certain embodiments, a portion of the training data (e.g., 80%) is used to train the skill bot model, and another portion (e.g., the remaining 20%) is used to test or validate the model. Once trained, the trained model, also referred to as the trained skill bot 116, can be used to handle and respond to input utterances. In certain cases, the input utterance can be a question that requires only a single answer and no further conversation. To handle such situations, a question-and-answer (Q&A) intent may be defined for the skill bot 116. In some embodiments, Q&A An intent is created in a similar scenario as other intents, but the dialogue flow for a Q&A intent may be different from that for a normal intent. For example, the dialogue flow for a Q&A intent may not need to prompt for additional information from the user (e.g., values for specific entities) unlike in the case of other intents.

[0059] (5) Create a dialogue flow for the skill bot 116. The dialogue flow specified for the skill bot 116 describes how the skill bot 116 reacts when various intents regarding the skill bot 116 are resolved in response to the received user input 110. The dialogue flow defines the actions or operations that the skill bot 116 will take, such as how the skill bot 116 responds to user utterances, how the skill bot 116 prompts the user for input, and how the skill bot 116 returns data. The dialogue flow is similar to a flowchart that the skill bot 116 follows. The designer of the skill bot 116 specifies the dialogue flow using a language such as the Markdown language. In a particular embodiment, a version of YAML called OBotML can be used to specify the dialogue flow for the skill bot 116. The dialogue flow definition for the skill bot 116 functions as a model for the conversation itself, causing the designer of the skill bot 116 to choreograph the dialogue between the skill bot 116 and the user to whom the skill bot 116 corresponds.

[0060] In a particular embodiment, the dialogue flow definition for the skill bot 116 includes three sections described below.

[0061] (a) Context section (b) Default transition section (c) State section Context Section: The developer of Skillbot 116 can define variables used in the conversation flow in the context section. Other variables that can be named in the context section include, for example, variables for error handling, variables for built-in entities or custom entities, and user variables that enable Skillbot 116 to recognize and maintain the user's preferences.

[0062] Default Transition Section: Transitions for Skillbot 116 can be defined in the dialogue flow state section or the default transition section. Transitions defined in the default transition section act as a fallback and are triggered when there are no applicable transitions defined within a state or when the conditions required to trigger a state transition are not met. The default transition section can be used to define routing that enables Skillbot 116 to handle unexpected user actions gracefully.

[0063] State Section: The dialogue flow and its related actions are defined as a series of temporary states that manage the logic within the dialogue flow. Each state node within the dialogue flow definition names the components that provide the functionality required at that point in the conversation. In this way, states are constructed around the components. A state includes component-specific characteristics and defines transitions to other states that are triggered after the component has been executed.

[0064] Special case scenarios can be addressed using the state section. For example, it may be desirable to give the user the option to temporarily leave a first skill the user is engaged in and do something with a second skill within digital assistant 106. For example, if the user is engaged in a conversation with the shopping skill (e.g., the user has made some selection for purchase), the user may want to jump to the banking skill (e.g., to check if the user has sufficient funds for the purchase) and then return to the shopping skill to complete the user's order. To address this, the state section within the conversation flow definition of the first skill can be configured to start a conversation with another second skill in the same digital assistant and then return to the original conversation flow.

[0065] (6) Add custom components to skillbot 116: As described above, the states specified in the conversation flow for skillbot 116 name the components that provide the functions required for that state. The components enable skillbot 116 to perform the functions. In certain embodiments, DABP 102 provides a set of preconfigured components for performing a wide range of functions. The developer can select one or more of these preconfigured components and associate them with the states within the conversation flow for skillbot 116. The developer can also create custom or new components using the tools provided by DABP 102 and associate the custom components with one or more states within the conversation flow for skillbot 116.

[0066] (7) Test and deploy skillbot 116: DABP 102 may provide some features that enable the developer to test the skillbot 116 under development Furthermore, skillbot 116 can be deployed and included in digital assistant 106.

[0067] Although the method of creating the Skillbot 116 has been described above, the same technique may be used to create the Digital Assistant 106 or the Masterbot 114. At the Masterbot level or the Digital Assistant level, built-in system intents may be configured for the Digital Assistant 106. In some embodiments, these built-in system intents are used to identify common tasks that the Masterbot 114 can handle without invoking the Skillbot 116. Examples of system intents defined for the Masterbot 114 include the following. (1) Exit: Applies when the user indicates that they desire to end the current conversation or context in the Digital Assistant. (2) Help: Applies when the user seeks assistance or direction; (3) UnresolvedIntent: Applies to user input that does not successfully match the Exit intent and the Help intent. The Masterbot 114 may store information about one or more Skillbots 116 associated with the Digital Assistant 106. This information enables the Masterbot 114 to select a particular Skillbot 116 to handle the utterance or, alternatively, to determine that the utterance is not related to any of the Skillbots 116 of the Digital Assistant 106.

[0068] When a user inputs a phrase or utterance to the digital assistant 106, the master bot 114 is configured to execute a process to determine how to route the utterance and the related conversation. The master bot 114 determines this using a routing model, which can be rule-based, ML-based, or a combination thereof. Using the routing model, the master bot 114 determines whether the conversation corresponding to the utterance should be routed to a specific skill bot 116 for handling, should be handled by the digital assistant 106 or the master bot 114 itself for each built-in system intent, should be handled as a different state in the current conversation flow, or is unrelated to any of the skill bots 116 associated with the digital assistant 106.

[0069] In certain embodiments, as part of this process, master bot 114 determines whether the input utterance clearly identifies skill bot 116 using its call name. If the call name is present in the input utterance, the call name is treated as a clear call to the skill bot 116 corresponding to the call name. In such a scenario, master bot 114 may further route the input utterance to the clearly called skill bot 116 for further processing. In the absence of a specific call or clear call, in certain embodiments, master bot 114 evaluates the input utterance and calculates (e.g., using a logistic regression model) a confidence score for the system intents associated with digital assistant 106 and for skill bot 116. The score calculated for skill bot 116 or the system intent represents the degree to which the input utterance is likely to represent a task configured for skill bot 116 to perform or to represent a system intent. Any system intent or skill bot 116 for which the associated calculated confidence score exceeds a threshold may be selected as a candidate for further evaluation. Next, master bot 114 selects a specific system intent or skill bot 116 from the identified candidates to further process the input utterance. In certain embodiments, after one or more skill bots 116 are identified as candidates, the intents associated with those candidate skill bots 116 are evaluated, for each skill bot 116, such as by using a trained model, and further, a confidence score is determined for each intent. Generally, any intent having a confidence score exceeding a threshold (e.g., 70%) is treated as a candidate intent. If a specific skill bot 116 is selected, the input utterance is routed to that skill bot 116 for further processing. When a system intent is selected, one or more actions are performed by master bot 114 itself according to the selected system intent.

[0070] In some embodiments of the master bot 114 described herein, when applicable, not only direct the input utterance to the appropriate skill bot 116, but also determine whether to prompt some indication that the input utterance is irrelevant when it becomes irrelevant to the available skill bot 116. FIG. 2 is a flowchart showing a method 200 of configuring and using a master bot to direct an input utterance to a skill bot 116 and, when applicable, determine that a particular input utterance is irrelevant to the available skill bot 116, according to some embodiments described herein. The method 200 of FIG. 2 is an overall overview of various specific cases described in detail below.

[0071] The method 200 shown in FIG. 2 and other methods described herein may be implemented in software (e.g., as code, instructions, or programs) by one or more processing units (e.g., a processor or processor core), in hardware, or a combination thereof. The software may be stored in a non-transitory storage medium such as a memory device. This method 200 is intended to be exemplary and non-limiting. FIG. 2 shows various operations that occur in a particular order, which is not intended to be limiting. In certain embodiments, for example, these operations may be performed in a different order, or one or more operations of method 200 may be performed in parallel. In certain embodiments, method 200 may be performed by a training system and the master bot 114.

[0072] As shown in FIG. 2, in block 205 of method 200, the classifier model of the master bot 114 is initialized. In some embodiments, the classifier model is configured to determine whether the input utterance is irrelevant to any of the available skill bots 116. Additionally, in some embodiments, the classifier model may also be configured to route the input utterance to the appropriate skill bot 116 suitable for processing if the input utterance is considered relevant to the skill bot 116.

[0073] In some embodiments, initializing the classifier model may include training the classifier model on how to recognize irrelevant input utterances. For this purpose, a training system, which may be part of the DABP 102, may have access to training utterances (i.e., exemplary utterances) for each skill bot 116. The training system may generate respective training feature vectors that describe and represent each training utterance of the various skill bots 116 that are or will be available to the master bot 114. Each training feature vector is the feature vector of the corresponding training utterance. The training system may then generate various set representations from the training feature vectors. Each set representation may represent a set of training feature vectors. For example, such a set may be a cluster or a composite feature vector, as will be described in detail below. To initialize the classifier model, the training system may configure the classifier model to compare the input vector with the various set representations that represent a set of training feature vectors.

[0074] In block 210, the master bot 114 receives the user input 110 in the form of an input utterance. For example, the user may have typed the input utterance as the user input 110, or the user may have spoken the user input 110 that the digital assistant 106 converted into the input utterance. The input utterance may include a user request to be handled by the digital assistant 106 and thus by the master bot 114.

[0075] ​In block 215, master bot 114 determines, using a classifier model, whether any of the skill bots 116 associated with master bot 114 can handle the input utterance. In some embodiments, the classifier model of master bot 114 generates an input feature vector that describes and represents the input utterance. The classifier model can compare the input feature vector with a set representation of training feature vectors, and based on this comparison, the classifier model determines whether the input utterance is irrelevant to any of the skill bots 116 available to master bot 114 (i.e., the skill bots 116 in which the training utterances are represented in the set representation).

[0076] Examples of classification at the digital assistant or skill bot level FIG. 3 is a block diagram showing a master bot 114, also referred to as a master bot (MB) system, according to a particular embodiment described herein. Master bot 114 can be implemented in software only, hardware only, or a combination of hardware and software. In some embodiments, master bot 114 includes a preprocessing subsystem 310 and a routing subsystem 320. The master bot 114 shown in FIG. 3 is merely an example of the arrangement of components within master bot 114. Those skilled in the art will recognize many possible variations, alternatives, and modifications. For example, in some implementations, master bot 114 may have more or fewer systems or components than shown in FIG. 3, may combine two or more subsystems, or may have a different configuration or arrangement of subsystems.

[0077] In some embodiments, the language processing subsystem 310 is configured to process user input 110 provided by a user. Such processing may include, for example, converting user input 110 to text input utterance 303 using automatic speech recognition or some other tool if the user input is in a form other than voice or text, such as some other form.

[0078] In some embodiments, the routing subsystem 320 is configured to determine (a) whether the input utterance 303 is unrelated to any of the available skillbots 116, and (b) which skillbot 116 is most suitable for handling the input utterance 303 if the input utterance 303 is related to at least one skillbot 116. In particular, the classifier model 324 of the routing subsystem 320 may be rule-based or ML-based, or a combination thereof, and may be configured to determine whether the input utterance 303 is unrelated to any of the available skillbots 116, represents a particular skillbot 116, or represents a particular intent configured for a particular skillbot 116. For example, as described above, the skillbot 116 may be composed of one or more chatbot intents. Each chatbot intent may have its own dialogue flow and may be associated with one or more tasks that the skillbot 116 can perform. When it is determined that the input utterance 303 represents a particular skillbot 116 or represents an intent configured for a particular skillbot 116, the routing subsystem 320 can call the particular skillbot 116 and transmit the input utterance 303 as an input 335 for the particular skillbot 116. However, if the input utterance 303 is considered to be unrelated to any of the available skillbots 116, the input utterance 303 is considered to belong to the non-class 316, which is a class of utterances that cannot be handled by the available skillbots. In that case, the digital assistant 106 may indicate to the user that it cannot handle the input utterance 303.

[0079] In some embodiments, classifier model 324 can be implemented using a rule-based model, an ML-based model, or both. For example, in some embodiments, classifier model 324 can include a rule-based model trained on training data 354 that includes exemplary utterances. Thus, using the rule-based model, it is determined whether an input utterance is irrelevant to any of the available skillbots 116 (i.e., any skillbot 116 associated with digital assistant 106). Additionally or alternatively, in some embodiments, classifier model 324 can include a neural network trained on training data 354. Training data 354 can include a corresponding set of exemplary utterances for each skillbot 116 (e.g., two or more exemplary utterances for each intent that skillbot 116 is configured to handle). For example, in the example of FIG. 3, the available skillbots 116 include a first skillbot 116a, a second skillbot 116b, and a third skillbot 116c, and training data 354 includes first skillbot data 358a corresponding to the first skillbot 116a, second skillbot data 358b corresponding to the second skillbot 116b, and third skillbot data 358c corresponding to the third skillbot 116c, including respective skillbot data for each such skillbot 116. Skillbot data for a skillbot 116 can include training utterances (i.e., exemplary utterances) that represent utterances that can be handled by that particular skillbot 116. Thus, training data 354 includes training utterances for each of the various skillbots 116. Training system 350 may be incorporated into DABP 102 but need not necessarily be incorporated, and may use training data 354 to train classifier model 324 of master bot 114 to perform these tasks.

[0080] In some embodiments, the training system 350 trains the classifier model 324 using the training data 354 to determine whether an input utterance is unrelated to any of the available skillbots 116. Generally, the training system 350 may generate training feature vectors to describe and represent the training utterances within the training data 354. As will be described in detail below, the training system 350 may generate set representations. Each set representation represents a set of training feature vectors. After training and during operation, the classifier model 324 may compare the input utterance 303 to the set representations to determine whether the input utterance 303 is unrelated to any of the available skillbots 116.

[0081] In some embodiments, classifier model 324 includes one or more submodels, and each submodel is configured to perform various tasks related to classifying input utterance 303. In addition to or instead of the above, for example, classifier model 324 may include an ML model or other type of model configured to determine which skill bot 116 is most suitable for an input utterance 303 when the input utterance 303 is considered to be related to (i.e., not unrelated to) at least one available skill bot 116. In some embodiments, classifier model 324 may determine the most appropriate skill bot 116 while determining that the input utterance is related to at least one skill bot 116. In this case, there is no need to make yet another determination to identify the most suitable skill bot 116. However, alternatively, classifier model 324 may determine, for each skill bot 116, a relevant confidence score (e.g., using a logistic regression model) indicating the likelihood that skill bot 116 is most suitable for processing input utterance 303 (i.e., most suitable for the input utterance). For this purpose, the neural network included in classifier model 324 may be trained by a training system to determine the likelihood that input utterance 303 represents each skill bot 116 or one of the intents configured for skill bot 116. For example, the neural network may determine and output respective confidence scores associated with each skill bot 116. In this case, the confidence score for a skill bot indicates the likelihood that skill bot 116 can handle input utterance 303 or that it is the best available skill bot for handling the input utterance. Considering the confidence scores, routing subsystem 320 may select the skill bot 116 associated with the highest confidence score and route input utterance 303 to that skill bot 116 for processing.

[0082] FIG. 4 is a diagram showing a skill bot 116, also referred to as a skill bot system, according to a particular embodiment described herein. An instance of the skill bot 116 shown in FIG. 4 can be used as the skill bot 116 of FIG. 1 and can be implemented in software only, hardware only, or a combination of hardware and software. As shown in FIG. 4, the skill bot 116 may include a bot classifier model 424 that determines the intent of the input utterance 303 and a conversation manager 430 that generates a response 435 based on the intent.

[0083] The bot classifier model 424 can be implemented using a rule-based model or an ML-based model, or both, and can receive the input utterance 303 routed to the skill bot 116 by the master bot 114 as input. The bot classifier model 424 may access the rules 452 and intent data 454 within the data store 450 on the skill bot 116, or may be accessible to the skill bot 116. For example, the intent data 454 may include exemplary utterances or other data for each intent, and the rules 452 may describe how to use the intent data to determine the intent for the input utterance 303. The bot classifier model 424 can apply these rules to the input utterance 303 and the intent data 454 to determine the intent of the input utterance 303.

[0084] More specifically, in some embodiments, the bot classifier model 424 may operate in a manner similar to how the classifier model 324 of FIG. 3 determines the skill bot 116 to handle the input utterance 303. For example, the bot classifier model 424 may determine each confidence level (e.g., using a logistic regression model), and in a manner similar to how the classifier model 324 may assign a confidence level to the skill bot 116, each such confidence level may be assigned to each chatbot intent that the skill bot 116 is configured for. Thus, the confidence levels indicate the likelihood that each chatbot intent is most likely applicable to the input utterance 303. The bot classifier model 424 may then select the intent to which the highest confidence level is assigned. In additional or alternative embodiments, the bot classifier model 424 determines the intent for the input utterance 303 by comparing the input feature vector that describes the input utterance 303 to either or both of (a) a composite feature vector that includes one composite feature vector to represent the training feature vectors for each intent, or (b) a cluster of training feature vectors. In this case, each cluster represents one or more intents. The use of such feature vectors is described in more detail below.

[0085] Once the intent that the utterance 202 best represents is identified, the bot classifier model 424 may communicate an intent indication 422 (i.e., an indication of the identified intent) to the conversation manager 430. In the embodiment of FIG. 4, the conversation manager 430 is shown as local to the skill bot 116. However, the conversation manager 430 may be shared across the master bot 114 and / or multiple skill bots 116. Thus, in some embodiments, the conversation manager 430 is local to the digital assistant or the master bot 114.

[0086] In response to receiving the intent instruction 422, the conversation manager 430 may determine an appropriate response 435 to the utterance 202. For example, the response 435 may be an action or message specified in the dialogue flow definition 455 configured for the system 400 of the skill bot 116, and the response 435 may be used as the DA response 112 in the embodiment of FIG. 1. For example, the data store 450 may include various dialogue flow definitions including respective dialogue flow definitions for each intent, and the conversation manager 430 may access the dialogue flow definition 455 based on the identification of the intent corresponding to the dialogue flow definition 455. The conversation manager 430 may determine, based on the dialogue flow definition 455, according to the dialogue flow definition 455, a certain dialogue flow state as the next state to transition to. The conversation manager 430 may determine the response 435 based on yet another process of the input utterance 303. For example, when the input utterance 303 is "Check balance in savings", the conversation manager 430 may transition to a dialogue flow state in which a dialogue related to the user's savings account is presented to the user. The conversation manager 430 may transition to this state based on the intent instruction 422 indicating that the identified intent is the "CheckBalance" intent configured for the skill bot 116, and further based on the recognition that the value of "saving" has been extracted for the "Account_Type" entity.

[0087] Use of feature vectors for describing utterances As described above, the classifier model 324 may use the feature vector when determining whether the input utterance is irrelevant or relevant to any of the available skill bots 116. For the purposes of the present disclosure, the feature vector is a set of vectors or coordinates that describe the features of the utterance and thereby describe the utterance itself. The feature vector describing the utterance may be used to represent the utterance in a particular situation, as described herein.

[0088] The concept of feature vectors is based on the concept of word embeddings. Generally, word embeddings are a type of language modeling where words are mapped to corresponding vectors. A particular word embedding can map semantically similar words to similar regions of the vector space. As a result, similar words will be close to each other within the vector space, and dissimilar words will be far apart. In a simple example of word embeddings, "one hot" encoding is used. In this case, each word in the dictionary is mapped to a vector having an amount of dimensions equal to the size of the dictionary, and thus, the vector will have a value of 1 in the dimension corresponding to the word itself and a value of zero in all other dimensions. For example, the first two words of the sentence "Go intelligent bot service artificial intelligence, Oracle" can be represented using the following "one hot" encoding. has a vector, and thus, the vector will have a value of 1 in the dimension corresponding to the word itself and a value of zero in all other dimensions. For example, the first two words of the sentence "Go intelligent bot service artificial intelligence, Oracle" can be represented using the following "one hot" encoding.

[0089] [Table 1]

[0090] Feature vectors can be used to represent words, sentences, or various types of phrases. Considering the above simple example of word embeddings, the corresponding feature vector can map a sequence of words, such as an utterance, to a feature vector that is a set of the word embeddings of that sequence of words. The set can be, for example, a sum, an average, or a weighted average. The set can be, for example, a sum, an average, or a weighted average.

[0091] In the case of different utterances, such utterances contain words with significantly different semantic meanings, so their respective feature vectors can also be different. However, utterances that are semantically similar and thus contain the same or semantically similar words between utterances can have similar feature vectors (i.e., feature vectors that are located close to each other within the vector space). Each feature vector corresponds to a single point within a vector space, also referred to as a feature space, and that point is the result of adding the feature vector to the origin of the feature space. In some embodiments, the points for utterances that are semantically similar to each other are located close to each other. Throughout the present disclosure, the feature vector and its corresponding point are referred to interchangeably because they provide different visual appearances for the same information.

[0092] Exemplary method for initializing the classifier model of the master bot In some embodiments, the classifier model 324 of the master bot 114 utilizes the feature vectors of the training utterances as a criterion for determining whether an input utterance is unrelated or related to any of the available skill bots 116. FIG. 5 is a diagram showing a method 500 for initializing the classifier model 324 of the master bot 114 to perform this task according to some embodiments described herein. For example, this method 500 or a similar method can be executed in block 205 of method 200 for configuring and using the master bot 114 to route the input utterance 303. FIG. 5 is a schematic method 500, and a more detailed example thereof is illustrated and described with reference to FIGS. 14 and 18.

[0093] Method 500 shown in FIG. 5 and other methods described herein can be implemented in software (e.g., as code, instructions, or programs) executed by one or more processing units (e.g., a processor or a processor core), in hardware, or in a combination thereof. The software can be stored in a non-transitory storage medium such as a memory device. This method 500 is intended to be illustrative and non-limiting. FIG. 5 shows various operations that occur in a particular order, which is not intended to be limiting. In certain embodiments, for example, these operations may be executed in a different order or one or more operations of method 500 may be executed in parallel. In certain embodiments, method 500 may be executed by a training system 350 that may be part of DABP 102.

[0094] In block 505, the training system 350 accesses training utterances (also referred to as exemplary utterances) for various skill bots 116 associated with the master bot 114. For example, the training utterances may be stored as skill bot data in training data 354 accessible to the training system 350. In some embodiments, each skill bot 116 available to the master bot 114 may be associated with a subset of the training utterances, and each such subset of training utterances may include training utterances for each intent that the skill bot 116 is configured for. Thus, the set of training utterances may include training utterances for each intent of each skill bot 116 such that all intents of all skill bots 116 are represented.

[0095] In block 510, the training system 350 generates training feature vectors from the training utterances accessed in block 505. As described above, the training utterances may include subsets associated with each skill bot 116 and may further include training utterances for each intent of each skill bot 116. Thus, in some embodiments, the training feature vectors may each include a respective subset for each skill bot 116, and further may each include a respective feature vector for each intent of each skill bot.

[0096] FIG. 6 shows the generation of a training feature vector 620 from a training utterance 615 according to some embodiments described herein. Specifically, FIG. 6 pertains to the training utterance 615 of the skill bot data 358 within the training data 354, and the skill bot data 358 is associated with a particular skill bot 116. The skill bot 116 is configured to handle input utterances 303 of a plurality of intents including intent A and intent B. Thus, the training utterance 615 includes a training utterance 615 representing intent A and a training utterance 615 representing intent B.

[0097] As described above, the training system 350 may generate a training feature vector 620 to describe and represent each training utterance 615. Various techniques are known for converting a sequence of words such as the training utterance 615 into a feature vector such as the training feature vector 620, and one or more of such techniques may be used. For example, the training system 350 may encode each training utterance 615 as a corresponding training feature vector 620 using one-hot encoding or some other encoding, although encoding is not necessarily required.

[0098] Also, as described above, the feature vector can be represented as a point. In some embodiments, each training feature vector 620 can be represented as a point 640 within a feature space 630. In this case, the feature space has a number of dimensions equal to the number of features (i.e., the number of dimensions) within the training feature vector 620. In the example of FIG. 6, two training feature vectors 620 that represent the same intent, specifically intent A, are plotted as points 640 that are close to each other because these two training feature vectors 620 are semantically similar. However, not all training feature vectors 620 or all feature vectors related to a particular intent need to be represented as points that are close to each other.

[0099] Returning to FIG. 5, at block 515, the training system 350 generates a plurality of set representations of the training feature vectors 620 generated at block 510. Each set representation represents a set of training feature vectors 620. As will be described in detail below, the set representation can be, for example, a cluster of training feature vectors 620 or a composite feature vector that is a set of multiple training feature vectors 620. In essence, the set representation can be a way of representing a plurality of training feature vectors 620 grouped together. Each set of training feature vectors represented by a corresponding set representation can share a common intent, a group of common intents, a common skill bot 116, or a common region of the feature space 630, or alternatively, the training feature vectors represented by a single set representation do not need to have any commonality other than being based on the training data 354.

[0100] In block 520, the training system 350 may configure the classifier model 324 to compare the input utterance 303 provided as the user input 110 with various set representations. For example, the set representations may be stored in a storage device accessible to the classifier model 324 of the master bot 114. The classifier model 324 may be programmed with rules on how to determine whether the input utterance matches or does not match various set representations. The definition of a match may depend on the specific set representation being used, as will be described in more detail below.

[0101] FIG. 7 is a diagram showing a method 700 for determining whether an input utterance 303 provided as a user input 110 is unrelated to any available skill bot 116 associated with the master bot 114 using the classifier model 324 of the master bot 114 according to some embodiments described herein. This method 700 or a similar method may be further executed for each received input utterance 303 after initialization of the classifier model 324. For example, this method 700 or a similar method may be executed in block 215 of method 200 for configuring and using the master bot 114 to route the input utterance 303. FIG. 7 is a schematic method 700, and a more detailed example thereof is illustrated and described with reference to FIGS. 16 and 21.

[0102] Method 700 shown in FIG. 7 and other methods described herein may be implemented in software (e.g., as code, instructions, or programs), in hardware, or in a combination thereof, by one or more processing units (e.g., a processor or a processor core). The software may be stored in a non-transitory storage medium such as a memory device. This method 700 is intended to be exemplary and non-limiting. FIG. 7 shows various operations that occur in a particular order, but this is not intended to be limiting. In certain embodiments, for example, these operations may be performed in a different order, or one or more operations of method 700 may be performed in parallel. In certain embodiments, method 700 may be performed by master bot 114.

[0103] In block 705 of method 700, master bot 114 accesses input utterance 303 provided as user input 110. For example, in some embodiments, the user may provide user input 110 in the form of voice input, and digital assistant 106 may convert that user input 110 into text input utterance 303 for use by master bot 114.

[0104] In block 710, master bot 114 may generate an input feature vector from input utterance 303 accessed in block 705. The input feature vector may describe and represent input utterance 303. Various techniques for converting a sequence of words, such as an input utterance, into a feature vector are known, and one or more of such techniques may be used. For example, training system 350 may encode the input utterance as a corresponding input feature vector using one-hot encoding or some other encoding, but encoding is not necessarily required. However, one embodiment of master bot 114 uses the same technique as was used to generate training feature vector 620 from training utterances when training classifier model 324.

[0105] In determination block 715, master bot 114 may cause classifier model 324 to compare the input feature vector generated in block 710 with the set representation determined in block 515 of method 500 for initializing classifier model 324. Thus, master bot 114 may determine whether the input feature vector matches any of the skill bots 116 available to master bot 114. The specific techniques for comparison and matching may depend on the nature of the set representation. For example, as described below, if the set representation is a cluster of training feature vectors 620, classifier model 324 may compare the input feature vector (i.e., the point representing the input feature vector) with the cluster to determine whether the input feature vector is within the range of any of the clusters and thus may match at least one skill bot 116. Or, if the set representation is a composite vector of training feature vectors, classifier model 324 may compare the input feature vector with the composite feature vector to determine whether the input feature vector is sufficiently similar to any such composite feature vector and thus may match at least one skill bot 116. Various implementations are possible and within the scope of the present disclosure.

[0106] In determination block 715, if the input feature vector is considered to match at least one skill bot 116, in block 720, master bot 114 may route the input utterance to the skill bot 116 to which the input feature vector is considered to match. However, in decision block 715, if the input feature vector is considered not to match any of the skill bots 116, in block 725, master bot 114 may indicate that the utterance cannot be processed by any of the skill bots 116. This indication may be passed to the digital assistant, and the digital assistant may provide an output to the user indicating that the user input 110 cannot be processed or addressed by the digital assistant.

[0107] FIG. 8 shows another example of a method 800 for determining whether an input utterance 303 provided as a user input 110 is unrelated to any of the available skill bots 116 associated with the master bot 114, using the classifier model 324 of the master bot 114 according to some embodiments described herein. This method 800 or a similar method can be further executed for each received input utterance 303 after the initialization of the classifier model 324. For example, this method 800 or a similar method can be executed in block 215 of method 200 for configuring and using the master bot 114 to route the input utterance 303. Similar to the method 700 of FIG. 7, the method 800 of FIG. 8 is a schematic method 800, and more detailed examples of some of its method blocks are illustrated and described with reference to FIGS. 16 and 21. However, in contrast to the method 700 of FIG. 7, this method 800 exemplifies preliminary filtering operations in decision block 810 and block 815, which can be used to ensure that certain input utterances 303 similar to the training utterance 615 are not classified as belonging to the non-class 316. In other words, this method 800 includes a filter that filters out any input utterance 303 that is considered to be sufficiently similar to the training utterance 615 from the non-class considerations, thereby ensuring that such input utterances 303 are not routed to the skill bot 116.

[0108] Method 800 shown in FIG. 8 and other methods described herein can be implemented in software (e.g., as code, instructions, or programs) executed by one or more processing units (e.g., a processor or a processor core), in hardware, or in a combination thereof. The software can be stored in a non-transitory storage medium such as a memory device. This method 800 is intended to be exemplary and non-limiting. FIG. 8 shows various operations that occur in a particular order, which is not intended to be limiting. In certain embodiments, for example, these operations may be performed in a different order, or one or more operations of method 800 may be performed in parallel. In certain embodiments, method 800 may be executed by master bot 114.

[0109] In block 805 of method 800, master bot 114 accesses input utterance 303 provided as user input 110. For example, in some embodiments, the user may provide user input 110 in the form of voice input, and digital assistant 106 may convert that user input 110 into text input utterance 303 for use by master bot 114.

[0110] In decision block 810, master bot 114 causes classifier model 324 to determine whether all words, or a predetermined percentage of the words, of input utterance 303 accessed in block 805 are found within training utterance 615. For example, in some embodiments, classifier model 324 utilizes a Bloom filter based on the words of training utterance 615 and applies this Bloom filter to input utterance 303. obtain. For example, in some embodiments, classifier model 324 utilizes a Bloom filter based on the words of training utterance 615 and applies this Bloom filter to input utterance 303. apply to input utterance 303.

[0111] If the input utterance 303 contains only words found in the training utterance 615 in the determination block 810, the master bot 114 may determine that the input utterance is related to at least one skill bot 116 and thus does not belong to the non-class 316. In that case, at block 815, the master bot 114 may route the input utterance 303 to the skill bot 116 that most closely matches the input utterance 303. For example, as described above, the classifier model 324 may be configured to assign a confidence score to each skill bot 116 for the input utterance (e.g., using a logistic regression model). The master bot 114 may select the skill bot 116 having the highest confidence score and route the input utterance 303 to that skill bot 116.

[0112] However, if in the determination block 810 the input utterance 303 contains any words not in the training utterance 615 or contains a percentage of words greater than a threshold, the method 800 proceeds to block 820. At block 820, the master bot 114 may generate an input feature vector from the input utterance 303 accessed at block 805. The input feature vector may describe and represent the input utterance 303. Various techniques for converting a sequence of words such as an input utterance into a feature vector are known, and one or more of such techniques may be used. For example, the training system 350 may encode each training utterance 615 as a corresponding training feature vector 620 using one-hot encoding or some other encoding, but it is not necessarily required to encode. However, one embodiment of the master bot 114 uses the same technique as used to generate the training feature vector 620 from the training utterance.

[0113] In determination block 825, master robot 114 may cause classifier model 324 to compare the input feature vector generated in block 820 with the set representation determined in block 515 of method 500 for initializing classifier model 324, whereby master robot 114 can determine whether the input feature vector matches any of the skill robots 116 available to master robot 114. Specific techniques for comparison and matching may depend on the nature of the set representation. For example, as described below, if the set representation is a cluster of training feature vectors 620, classifier model 324 can compare the input feature vector (i.e., the point representing the input feature vector) with the cluster to determine whether the input feature vector is within the range of any of the clusters and thus whether it matches at least one skill robot 116, or if the set representation is a composite vector of training feature vectors, classifier model 324 can compare the input feature vector with the composite feature vector to determine whether the input feature vector is sufficiently similar to any such composite feature vector and thus whether it matches at least one skill robot 116. Various implementations are possible and within the scope of the present disclosure.

[0114] In determination block 825, if the input feature vector is considered to match at least one skill robot 116, in block 830, master robot 114 may route the input utterance to the skill robot 116 to which the input feature vector is considered to match. However, in determination block 825, if the input feature vector is considered not to match any of the skill robots 116, in block 835, master robot 114 may indicate that the utterance cannot be processed by any of the skill robots 116. This indication may be passed to the digital assistant, which may provide the user with an indication that user input 110 cannot be processed or addressed by the digital assistant. Strength may be provided to the user.

[0115] Examples of types of clusters available for use by the classifier model As described above, an exemplary type of set representation that can be used by the classifier model 324 is a cluster of training feature vectors 620. Generally, it can be handled by the available skill bot 116, and thus, it can be assumed that the input utterances related to the available skill bot 116 have some semantic similarity to the training utterances 615 regarding those skill bots 116. Therefore, the input utterance 303 related to the available skill bot 116 is likely to be represented as an input feature vector close to one or more training utterances 615 within the feature space 630.

[0116] Considering the proximity of the feature vectors of semantically similar utterances, a boundary can be defined to separate the feature vectors plotted at points with a common intent, or to separate the feature vectors with different intents. In a two-dimensional space, the boundary can be a line or a circle, and thus, the points on one side of the line, such as a line, belong to the first intent class (i.e., corresponding to the utterances with the first intent), and the points on the other side of the line belong to the second intent class. In three dimensions, the boundary can be represented as a plane or a sphere. More generally, in different dimensions, the boundary can be a hyperplane, a hypersphere, or a hypervolume. The boundary can have various shapes and does not need to be completely spherical or symmetric.

[0117] FIG. 9 shows an example of the feature space 630 including points representing the feature vectors of exemplary utterances according to some embodiments described herein. In this example, some of the exemplary utterances belong to the balance class and represent a first intent related to a request for account balance information, and the rest of the exemplary utterances belong to the transaction class and represent a second intent related to a request for information about transactions. In FIG. 9, the exemplary utterances within the balance class are labeled with b, and the exemplary utterances within the transaction class are labeled with t. These intent classes can be defined, for example, with respect to the financial-related skill bot 116.

[0118] For example, the example utterances in the table below may belong to the balance class in the first column and the transaction class in the second column.

[0119] [Table 2]

[0120] FIG. 10 illustrates an example of the feature space 630 of FIG. 9 with a class boundary 1010 between the intent classes of the feature vectors of the example utterances, according to some embodiments described herein. Specifically, as shown in FIG. 10, a line may be drawn as a class boundary 1010 to separate points in the balance class (i.e., feature vectors of the example utterances) from points in the transaction class. This class boundary 1010 is a rough approximation of the division between the two intent classes. To create a more precise boundary 1010 between the intent classes, it may be necessary to create a circle or other geometric volume that defines the clusters of feature vectors, such that each cluster only contains feature vectors in a single corresponding intent class (i.e., has an intent associated with that intent class).

[0121] FIG. 11 illustrates a common intent-related example according to some embodiments described herein. 11 shows an example of feature space 630 of FIG. 9 with class boundaries 1010 that separate, specifically, isolate, the feature vectors associated with an intent class into respective clusters. In this example, as in some embodiments, not all feature vectors of an intent class are in a single cluster, and no cluster includes feature vectors from more than one intent class. Specifically, in the illustrated example, a first cluster 1110a defined by a first class boundary 1010 includes only feature vectors in a balance class, and a second cluster 1110b defined by a second class boundary 1010 includes only feature vectors in a transaction class. As described in more detail below, some embodiments described herein can form class boundaries 1010 to create clusters such as those shown in FIG. 11.

[0122] FIG. 12 shows another example of the feature space 630 of FIG. 9 having class boundaries 1010 that separate, specifically isolate, the feature vectors associated with a common intent into respective clusters, according to some embodiments described herein. In this example, as in some embodiments, not all of the feature vectors of an intent class are within a single cluster, and there are no clusters that contain feature vectors from two or more intent classes. Specifically, in the illustrated example, the first cluster 1110c defined by the first class boundary 1010 contains only feature vectors within the balance class, and the second cluster 1110d defined by the second class boundary 1010 contains only feature vectors within the transaction class. However, in contrast to the example of FIG. 11, the class boundaries 1010, and thus the clusters, overlap. Some embodiments described herein support overlapping clusters, as shown in FIG. 12. As will be described in detail below, some embodiments described herein can form class boundaries 1010 to create clusters such as those shown in FIG. 12.

[0123] FIG. 13 shows another example of the feature space 630 of FIG. 9 having a class boundary 1010 that separates feature vectors into clusters, according to some embodiments described herein. In this example, as in some embodiments, not all feature vectors of a given intent class are within a single cluster, and furthermore, a given cluster can represent different intents by including feature vectors with different intents. Specifically, in the illustrated example, the first cluster 1110e defined by the first class boundary 1010 includes only feature vectors within the balance class, and the second cluster 1110f defined by the second class boundary 1010 includes some feature vectors within the balance class (i.e., having balance-related intents) and some feature vectors within the transaction class (i.e., having transaction-related intents). Some embodiments described herein support clusters with different intents, as shown in FIG. 13. As will be described in detail below, some embodiments described herein can form class boundaries 1010 to create clusters such as those shown in FIG. 13.

[0124] Clustering for identifying irrelevant input utterances FIG. 14 is a diagram showing a method 1400 for determining whether an input utterance is unrelated or related to an available skill bot 116 by initializing a classifier model 324 of a master bot 114 and utilizing clusters of training feature vectors, according to some embodiments described herein. This method 1400 or a similar method can be used in block 205 of the above-described method 200 for configuring and using the master bot 114 to direct an input utterance 303 to the skill bot 116. Further, the method 1400 of FIG. 14 is a more specific variation of the method 500 of FIG. 5. In some embodiments of this method, as described below, k-means clustering for configuring the classifier model 324 is performed. K-means clustering may enable more accurate formation of clusters compared to other clustering techniques such as k-nearest neighbors. However, various clustering techniques may be used instead of or in addition to k-means clustering.

[0125] The method 1400 shown in FIG. 14 and other methods described herein may be implemented in software (e.g., as code, instructions, or a program) by one or more processing units (e.g., a processor or a processor core), in hardware, or in a combination thereof. The software may be stored in a non-transitory storage medium such as a memory device. This method 1400 is intended to be illustrative and non-limiting. FIG. 14 shows various operations that occur in a particular order, which is not intended to be limiting. In certain embodiments, for example, these operations may be performed in a different order, or one or more operations of the method 1400 may be performed in parallel. In certain embodiments, the method 1400 may be performed by a training system 350 that may be part of the DABP 102.

[0126] In block 1405 of method 1400, training system 350 accesses training utterances 615, also referred to as exemplary utterances, for the various skill bots 116 associated with master bot 114. For example, training utterances 615 may be stored as skill bot data in training data 354 accessible by training system 350. In some embodiments, each skill bot 116 available to master bot 114 may be associated with a subset of training utterances 615, and each such subset of training utterances 615 may include training utterances 615 for each intent that skill bot 116 is configured to handle. Thus, the set of training utterances 615 may include training utterances 615 for each intent of each skill bot 116 such that all intents of all skill bots 116 are represented.

[0127] FIG. 15 shows an example of the execution of aspects of method 1400 of FIG. 14, according to some embodiments described herein. As shown in FIG. 15, training system 350 may access training data 354. Training data 354 may include skill bot data related to skill bots 116 available to master bot 114 to which classifier model 324 during training belongs. In this example, the training data 354 accessed by training system 350 may include training utterances 615 from first skill bot data 358d, second skill bot data 358e, and third skill bot data 358f, each of which may include training utterances 615 representing a respective skill bot 116 available to master bot 114. Further, for a given set of skill bot data, as shown with respect to first skill bot data 358d of FIG. 15, each training utterance 615 may be associated with an intent that the associated skill bot 116 is configured to evaluate and handle.

[0128] In block 1410 of FIG. 14, the training system 350 generates a training feature vector 620 from the training utterance 615 accessed at block 1405. As described above, the training utterance 615 may include a subset associated with each skill bot 116, and may further include the training utterance 615 for each intent of each skill bot 116. Thus, in some embodiments, the training feature vector 620 may include a respective subset for each skill bot 116, and may further include a respective feature vector for each intent of each skill bot. As described above, the training system 350 may generate a training feature vector 620 to describe and represent each training utterance 615. Various techniques are known for converting a series of words, such as the training utterance 615, into a feature vector, such as the training feature vector 620, and one or more of such techniques may be used. For example, the training system 350 may encode each training utterance 615 as a corresponding training feature vector 620 using one-hot encoding or some other encoding, but it is not necessarily required to encode.

[0129] As shown in FIG. 15, for example, each training utterance 615 from the various skill bot data 358 may be converted into a respective training feature vector 620. Thus, in some embodiments, the resulting training feature vector 620 may represent all of the skill bots 116 for which the training utterance 615 was provided, and all of the intents for which the training utterance 615 was provided.

[0130] In block 1415 of FIG. 14, the training system 350 may set (i.e., initialize) a count that is the amount of clusters to be generated. In some embodiments, for example, the count may first be set to an amount n, where n is the total number of intents across the various skill bots 116 available to the master bot 114. However, other various amounts may be used as the initial value of the count.

[0131] In some embodiments, the training system 350 executes the step of repeatedly determining a set of clusters of the training feature vector 620 one, two, or more times. This method 1400 utilizes the step of determining the clusters two times, but in some embodiments, one or more times may be used. In block 1420, the training system 350 begins the step of determining the first set of clusters. The first time utilizes an iterative loop. In this iterative loop, the training system 350 generates clusters of the training feature vector 620 and then determines whether such clusters are sufficient. As described below, if the clusters are considered sufficient, the training system 350 may end the first iteration.

[0132] In some embodiments, block 1425 is the start of the first iteration loop for determining clusters. Specifically, in block 1425, the training system 3509 can determine the centroid positions (i.e., the position for each centroid) for each of the various clusters to be generated in this iteration. The number of centroid positions is equal to the count determined in block 1415. In some embodiments, in the first iteration of this loop, the training system 350 can select a set of randomly selected centroid positions having a number equal to the count within the feature space 630 in which the training utterance 615 fits. More specifically, for example, the training system 350 can determine a bounding box such as a minimum bounding box with respect to the points corresponding to the training feature vectors 620. Then, the training system 350 can randomly select a set of centroid positions within that bounding box. In this case, the number of selected positions is equal to the count determined for the cluster. During iterations other than the first, the count has increased from the previous iteration, and thus, in some embodiments, only the centroid positions of the newly added centroids are randomly determined. The centroids carried over from the previous iteration can retain their positions. Each centroid for the corresponding cluster can be positioned for each centroid position.

[0133] In block 1430, the training system 350 can determine the clusters by assigning each training feature vector 620 to the nearest centroid among the various centroids whose positions were determined in block 1425. For example, for each training feature vector 620, the training system 350 can calculate the distance to each centroid position and assign that training feature vector 620 to the centroid having such a minimum distance. The set of training feature vectors 620 assigned to a common centroid can together form the cluster associated with that centroid. Thus, there can be a number of clusters equal to the number of centroids, which is the value of the count.

[0134] In the example of FIG. 15, only a portion of the feature space 630 is shown, and within that portion, the first centroid 1510a and the second centroid 1510b are visible. The training system 350 determined that, of the three training feature vectors 620 shown, two are closest to the first centroid 1510a and one is closest to the second centroid 1510b. In this example, the two training feature vectors 620 closest to the first centroid 1510a form the first cluster 1110g, and the one training feature vector 620 closest to the second centroid 1510b forms the second cluster 1110h.

[0135] In block 1435 of FIG. 14, for the clusters determined in block 1430, the training system 350 recalculates the positions of the centroids of each cluster. For example, in some embodiments, the centroid of each cluster is calculated to be the average (e.g., arithmetic mean) of the training feature vectors 620 assigned to that centroid and thus assigned to that cluster.

[0136] In decision block 1440, the training system 350 can determine whether a stop condition is met. In some embodiments, the training system 350 repeatedly increases the number of centroids (i.e., increases the count) using this method 1400 or a similar method until convergence occurs such that there is no longer a possibility of a significant improvement in the clustering, and the training feature vectors 620 can be assigned to their closest associated centroids. In some embodiments, the stop condition defines a sufficient level of convergence.

[0137] In some embodiments, the stopping condition may be satisfied when one or both of the following conditions are true. That is, (1) the average cluster cost meets a first threshold. For example, the average cluster cost is less than 1.5 or some other predetermined value. Or, (2) the outlier ratio meets a second threshold. For example, the outlier ratio is 0.25 or less or some other predetermined value or less. The cluster cost for a particular cluster may be defined as (a) the sum of the squared distances between the centroid of the cluster recalculated in block 1435 and each training feature vector 620 assigned to that centroid, divided by the amount of the training feature vectors 620 assigned to that centroid (b). Thus, the average cluster cost may be the average of the various cluster costs of the various centroids. The outlier ratio may be defined as the total number of outliers between clusters divided by the number of clusters (i.e., count). There are various techniques for defining outliers, and one or more of such techniques may be used by the training system 350. In some embodiments, the stopping condition is satisfied if and only if (1) the average cluster cost meets the first threshold (e.g., less than 1.5), and (2) the outlier ratio meets the second threshold (e.g., 0.25 or less).

[0138] Generally, the average cluster cost tends to decrease as the value of k (i.e., count) increases, whereas the outlier ratio tends to increase as k increases. In some embodiments, when the above-described stopping condition that takes both factors into account is applied, the maximum k value that meets the respective thresholds for both the cluster count and the outlier ratio is the final k value for the classifier model 324.

[0139] If the stop condition is not met in decision block 1440, method 1400 may proceed to block 1445. In block 1445, the training system increments a count for the next clustering. In some embodiments, the count may be incremented incrementally by an amount that increases the likelihood that the stop condition will be met after a certain number of loop iterations. For example, in some embodiments, the count can be incremented by a value of step size equal to

[0140] [Number]

[0141] Here, n is the total number of intents across all available skillbots 116, and u is the total number of training utterances 615 across all available skillbots 116. This step size will ensure that the stop condition is eventually met. Specifically, for this step size, starting from a count equal to n, the 20th iteration will have a count that is u or more in terms of the amount of utterances. If the count is u or more, each utterance may have its own cluster, or otherwise, there is still a possibility that the average cluster cost is less than 1.5 and the outlier ratio is 0.25 or less, which meets an exemplary stop condition. More generally, the step size can be selected to ensure that the above iterations do not waste computing resources by loops that take an unreasonable amount of time. After the count is updated in block 1445, method 1400 returns to block 1425 to perform another loop iteration.

[0142] However, if the stop condition is met in decision block 1440, method 1400 may end the current iteration loop and proceed to block 1450. In block 1450, training system 350 starts the second clustering. In this case, the clusters may be further defined based on the work done in the first pass.

[0143] In some embodiments, block 1455 is the start of the second iteration loop for determining clusters. Specifically, in block 1455, the training system 3509 can determine the respective centroid positions (i.e., the respective positions for each centroid) for the various clusters to be generated in this iteration. The number of centroid positions is equal to the current value of the count. In some embodiments, in the first iteration of this loop, the training system 350 uses the centroid positions as recalculated at block 1435 before the first end. In iterations other than the first, the count has increased from the previous iteration. In that case, the centroids from the previous iteration can hold their centroid positions, and in some embodiments, the centroid positions of the newly added centroids due to the increase in the count can be determined randomly. Each centroid for the corresponding cluster can be positioned for each centroid position.

[0144] In block 1460, the training system 350 can determine the clusters by assigning each training feature vector 620 to the nearest centroid among the various centroids whose positions were determined in block 1455. For example, for each training feature vector 620, the training system 350 can calculate the distance to each centroid position and assign the training feature vector 620 to the centroid having such a minimum distance. The set of training feature vectors 620 assigned to a common centroid can together form the cluster associated with that centroid. Thus, there can be a number of clusters equal to the number of centroids, which is the value of the count.

[0145] In block 1465, for the clusters determined in block 1460, the training system 350 recalculates the position of the centroid of each cluster. For example, in some embodiments, the centroid of each cluster is calculated to be the average (e.g., arithmetic mean) of the training feature vectors 620 assigned to that centroid and thus assigned to that cluster.

[0146] In determination block 1470, the training system 350 can determine whether a stop condition is satisfied. In some embodiments, the training system 350 repeatedly increases the number of centroids (i.e., increments the count) using this method 1400 or a similar method until convergence occurs such that there is no possibility of a significant improvement in clustering, and can assign the training feature vectors 620 to their nearest associated centroids. In some embodiments, the stop condition defines a sufficient level of convergence.

[0147] In some embodiments, the stop condition can be satisfied if one or both of the following conditions are true. That is, (1) the average cluster cost meets a first threshold. For example, the average cluster cost is less than 1.5 or some other predetermined value. Or, (2) the outlier ratio meets a second threshold. For example, the outlier ratio is 0.25 or less or some other predetermined value or less. The cluster cost for a particular cluster may be defined as (a) the sum of the squared distances between the centroid of the cluster recalculated in block 1465 and each training feature vector 620 assigned to that centroid, divided by the amount of training feature vectors 620 assigned to that centroid (b). Thus, the average cluster cost can be the average of the various cluster costs of the various centroids. The outlier ratio may be defined as the total number of outliers between clusters divided by the number of clusters (i.e., the count). There are various techniques for defining outliers, and one or more of such techniques can be used by the training system 350. In some embodiments, the stop condition is satisfied if and only if (1) the average cluster cost meets a first threshold (e.g., less than 1.5) and (2) the outlier ratio meets a second threshold (e.g., 0.25 or less).

[0148] If the stop condition is not met in the determination block 1470, the method 1400 may proceed to block 1475. In block 1475, the training system increases the count for the next clustering. In some embodiments, the count may be incremented gradually by an amount corresponding to an amount that increases the likelihood that the stop condition will be met after a certain number of loop iterations. For example, in some embodiments, the count can be incremented by a value of the step size equal to

[0149]

Number

[0150] This step size ensures that the stop condition will ultimately be met. Specifically, in the case of this step size, the stop condition is likely to be met by the end of the fifth iteration. More generally, the step size can be selected to ensure that the above iterations do not waste computing resources by loops that take an unreasonable amount of time. After the count is updated in block 1475, the method 1400 returns to block 1455 and performs another loop iteration.

[0151] However, if the stop condition is met in the determination block 1470, the method 1400 may end the current iteration loop and proceed to block 1480. At block 1480, the training system 350 determines respective boundaries 1010 for each cluster determined above. In some embodiments, the boundary 1010 for a cluster is defined to be centered on the centroid of the cluster and to include all training feature vectors 620 assigned to the cluster. In some embodiments, for example, the boundary 1010 of a cluster is a hypersphere (e.g., a circle or a sphere) centered on the centroid. In some embodiments, the radius of the boundary 1010 is, with respect to the margin value (i.e., the padding amount), either (1) in the cluster that is farthest from the centroid, the larger of the maximum distances from the center to the training feature vectors 620 is added, or (2) the average of the respective distances from the training feature vectors 620 within the cluster to the centroid is added, and further, a value obtained by adding three times the standard deviation of such distances may be added. That is, the radius can be set to radius = margin + max(max(distances), mean(distances) + 3σ(distances)). In this case, the distances are a set of the respective distances from the training feature vectors of the cluster to the centroid of the cluster, max(distances) is the maximum value of the set, mean(distances) is the average of the set and σ(distances) is the standard deviation of the set. Further, the margin value (margin) may be a margin of error and may have a value of zero or greater.

[0152] In some embodiments, the margin is used to define a boundary 1010 that includes a wider effective range than otherwise, so as to reduce the possibility that the associated input utterance 303 is outside the range of all clusters and thus may be labeled as a member of the non-class 316. In other words, the margin can fill the boundary 1010. For example, the margin

[0153]

Number

[0154] may have a value of. Here, u is the total number of training utterances 615 used. The margin can take into account the following situation. Specifically, potentially the amount of training utterances is too low (for example, 20 - 30), so that the training feature vector 620 does not cover a significant part of the feature space 630, and thus the situation where the cluster is too small to capture the relevant input utterance 303.

[0155] Returning to the example of FIG. 15, the training system 350 determines respective boundaries for each cluster determined based on the assignment of the training feature vector 620. Specifically, in this example, the first boundary 1010a is determined for the first cluster 1110g that includes two training feature vectors 620, and the second boundary 1010b is determined for the second cluster 1110h that includes one training feature vector 620. The amount of training feature vectors may be as small as in this example, or may be large, such as the number of training feature vectors 620 per cluster. It is understood that this simplified example is non - limiting and is provided for illustrative purposes only.

[0156] As shown in FIG. 14, in block 1485, the training system 350 may configure the classifier model 324 of the master bot 114 to utilize the boundary 1010, also referred to as the cluster boundary, determined in block 1480. For example, the training system 350 may store an indication of the cluster boundary in a storage device accessible to the classifier model (for example, on the digital assistant 106). The classifier model 324 may be configured to compare an input utterance with such a cluster boundary 1010 that functions as a set representation of the training feature vector 620.

[0157] Various modifications can be made to the above-described method 1400, and these modifications are within the scope of the present disclosure. For example, some examples of the training system 350 perform the improvement of the cluster only once. In that case, the operations from block 1445 to block 1475 can be skipped so that the method 1400 proceeds to block 1480 when the stop condition is satisfied at the determination block 1440. Some other examples of the training system 350 perform the improvement of the cluster more than twice. In that case, the operations from block 1445 to block 1475 can be repeated each time the number of times is increased after the second execution. These and other implementations are within the scope of the present disclosure.

[0158] In some embodiments, the k value (i.e., the value of the count and thereby the number of clusters) determined in the above-described method 1400 depends on the total number of training utterances 615 and the distribution of the training feature vectors 620 across the entire feature space 630. Accordingly, the k value can vary for each master bot 114 depending on the available skill bots 116 and the training utterances 615 available to represent the intents of those skill bots 116. The optimal k value is a balanced value such that each cluster is large enough that the input utterances 303 related to each cluster fall within the range of the cluster boundary 1010 while the unrelated utterances fall outside the range of the boundary 1010. If the cluster is too large, the risk of false matches increases. If the cluster is too small (e.g., consisting of a single training utterance 615), the usefulness of such a cluster is limited because the classifier model 324 can be overfitted.

[0159] In the above example of method 1400, the training feature vectors 620 are not split or grouped based on intent, and thus, a cluster may include training feature vectors 620 representing various skillbots 116 or various intents of one or more skillbots 116. Additionally or alternatively, one embodiment of the training system 350 may ensure that each cluster includes only training feature vectors 620 representing a single skillbot 116, a single intent, or a single subbot. In this case, the subbot is associated with a subset of training utterances representing a single skillbot 116. Various techniques may be used to limit the clusters in this way. For example, the training utterances may be separated into groups based on intent, subbot, or skillbot 116, and each instance of method 1400 may be executed for each group. In this way, the clusters determined in one instance of method 1400 for a corresponding group (e.g., training utterances 615 representing a particular skillbot 116) may include only training feature vectors 620 from that corresponding group. This may be limited to training utterances of a single intent, subbot, or skillbot 116. Various other implementations are possible and are within the scope of the present disclosure.

[0160] FIG. 16 illustrates a method 1600 for determining whether an input utterance 303 provided as a user input 110 is unrelated to any available skill bot 116 associated with a master bot 114, using a classifier model 324 of the master bot 114 according to some embodiments described herein. This method 1600 or a similar method can be used in block 215 of the above-described method 200 for configuring and using the master bot 114 to direct the input utterance 303 to a skill bot 116. The method 1600 of FIG. 16 is a more specific variation of the method 700 of FIG. 7, and like the method 700 of FIG. 7, the method 1600 of FIG. 16 can be used with the preliminary filtering operation described with respect to the method 800 of FIG. 8. More specifically, as will be described below, the classifier model 324 of the master bot 114 can utilize clusters of training feature vectors 620 to determine whether the input utterance 303 belongs to a non-class 316 (i.e., is unrelated to the available skill bots 116).

[0161] The method 1600 shown in FIG. 16 and other methods described herein can be implemented in software (e.g., as code, instructions, or programs) by one or more processing units (e.g., a processor or a processor core), in hardware, or in a combination thereof. The software can be stored in a non-transitory storage medium such as a memory device. This method 1600 is intended to be illustrative and non-limiting. FIG. 16 shows various operations that occur in a particular order, but this is not intended to be limiting. In certain embodiments, for example, these operations may be performed in a different order, or one or more operations of the method 1600 may be performed in parallel. In certain embodiments, the method 1600 may be performed by a master bot 114 associated with a set of available skill bots 116.

[0162] In block 1605 of method 1600, master bot 114 accesses input utterance 303 provided as user input 110. For example, in some embodiments, the user may provide user input 110 in the form of voice input, and digital assistant 106 may convert that user input 110 into text input utterance 303 for use by master bot 114.

[0163] In block 1610, master bot 114 may generate an input feature vector from input utterance 303 accessed in block 1605. More specifically, in some embodiments, master bot 114 causes classifier model 324 of master bot 114 to generate an input feature vector from input utterance 303. The input feature vector may describe and represent input utterance 303. Various techniques for converting a sequence of words, such as an input utterance, into a feature vector are known, and one or more of such techniques may be used. For example, training system 350 may encode input utterance 303 as a corresponding input feature vector using one-hot encoding or some other encoding, although encoding is not necessarily required. However, one embodiment of master bot 114 uses the same technique used to generate training feature vector 620 from training utterance 615 when training classifier model 324.

[0164] In determination block 1615, master robot 114 compares the input feature vector generated in block 1610 with the clusters. More specifically, in some embodiments, master robot 114 causes classifier model 324 to compare the input feature vector generated in block 1610 with the clusters determined during training, specifically, the boundaries of the clusters. As described above, each cluster may include a set of training feature vectors 620 and may include a boundary 1010 based on the training feature vectors 620 of that cluster. For example, boundary 1010 includes all the training feature vectors 620 assigned to the cluster and, in some cases, some additional space outside those training feature vectors 620. Specifically, in some embodiments, classifier model 324 determines whether the input feature vector (i.e., the point corresponding to the input feature vector) is within any boundary 1010 of any cluster of training feature vectors 620. Various techniques for determining whether a point is within a boundary exist in the art, and one or more such techniques can be used to determine whether the input feature vector is within the range of any of the cluster boundaries 1010.

[0165] In determination block 1620, classifier model 324 makes a determination based on comparing the input feature vector with the cluster boundaries. If the input feature vector is not within any cluster boundary and thus outside the range of all cluster boundaries 1010, method 1600 proceeds to block 1625.

[0166] Figure 17 shows an example of executing this method 1600 when the input feature vector 1710 is outside the range of all cluster boundaries. In some embodiments, the master bot 114 provides the input utterance 303 to the classifier model 324, which causes the classifier model 324 to convert the input utterance 303 into the input feature vector 1710 and compare the input feature vector with the cluster boundary 1010. In the example of Figure 17, five clusters 1110 are shown in the feature space 630. However, a greater or lesser number of clusters 1110 may be used. In this example, the input feature vector 1710 is outside the range of all cluster boundaries 1010, and thus the classifier model 324 outputs an indication to the master bot 114 that the input utterance 303 belongs to the non-class 316.

[0167] Returning to Figure 16, at block 1625, based on the classifier model 324 indicating that the input feature vector is outside the range of all cluster boundaries 1010, the master bot 114 indicates that it cannot process (i.e., cannot further process) the input utterance 303. For example, the digital assistant 106 may respond to the user by requesting clarification or reporting that the user input 110 is not related to the skills of the digital assistant.

[0168] However, if the input feature vector is within the range of one or more cluster boundaries 1010, method 1600 skips to block 1630. At block 1630, the master bot 114 determines a skill bot 116 for handling (i.e., further processing the input utterance 303 and determining a response 435 to the input utterance 303) the input utterance 303 by selecting one of the available skill bots 116. The determination of the skill bot 116 can be performed in various ways, and the techniques used may depend on the configuration of the one or more cluster boundaries 1010 within which the input feature vector falls.

[0169] To select the skill bot 116 for the input utterance 303, the master bot 114 may consider various training utterances (i.e., training feature vectors that are members of the cluster 1110 within the boundary 1010 in which the input feature vector 1710 lies) represented by the training feature vectors 620 that share a cluster with the input feature vector 1710. For example, if the training feature vector 620 falls within two or more overlapping clusters 1110, the training utterances 615 having the corresponding training feature vector 1710 in any of those two or more clusters 1110 may be considered. Similarly, if the training feature vector 620 falls within only a single cluster 1110, the training utterance 615 having the corresponding training feature vector 620 within that cluster 1110 is considered. If all of the training utterances 615 being considered represent a single skill bot 116, then in some cases, for example, if the input feature vector 1710 falls within a cluster 1110 that includes a training feature vector 620 associated with only a single skill bot 116, the master bot 114 may select that skill bot 116 to handle the input utterance 303.

[0170] In some embodiments, classifier model 324 may be able to identify a particular intent of a particular skill bot 116 to handle an input utterance. For example, if one or more clusters 1110 that the training feature vectors 620 fall into contain only the training feature vectors 620 of training utterances 615 that represent a single intent of a single skill bot 116, classifier model 324 may identify that the particular intent is applicable to input utterance 303. If classifier model 324 can identify a particular intent of a particular skill bot 116, master bot 114 may route input utterance 303 to that skill bot 116 and may indicate that intent to skill bot 116. As a result, skill bot 116 may omit performing its own classification of input utterance 303 to infer the intent, but may instead infer the intent indicated by master bot 114.

[0171] However, if the training utterances 615 under consideration represent multiple skill bots 116, in some cases, for example, if input feature vector 1710 falls within a cluster 1110 composed of training feature vectors 620 that represent multiple skill bots 116, or if input feature vector 1710 falls within multiple overlapping clusters 1110, master bot 114 may need to further classify input utterance 303 to select a skill bot 116. Various techniques may be used to further It may be classified. In some embodiments, the classifier model 324 implements a machine learning mode to calculate a confidence score for the input utterance 303 (e.g., using a logistic regression model) for each skill bot 116 having a training feature vector 620 associated within one or more clusters 1110 in which the input feature vector 1710 falls. For example, the machine learning model may be trained using synthetic feature vectors. The master bot 114 may then select the skill bot 116 having the highest confidence score for handling the input utterance 303. In contrast to the conventional use of confidence scores to identify relevant skill bots 116, in some embodiments, it has already been determined that the input utterance 303 is related to the skill bot 116. Thus, the risk of routing the input utterance 303 to an irrelevant skill bot 116 is reduced or eliminated.

[0172] In additional or alternative embodiments, the classifier model 324 may utilize a k-nearest neighbor technique to select a skill bot 116 from among two or more skill bots 116 that share the input feature vector 1710 and the related training feature vector 620 belonging to one or more clusters 1110. For example, the classifier model 324 may select a value for k and identify the k-nearest neighbor training feature vectors 620 for the input feature vector 1710 from among the training feature vectors 620 that fall within one or more clusters 1110 of the input feature vector 1710. The classifier model 324 may identify the skill bot 116 having the greatest number of related training feature vectors 620 within that set of k-nearest neighbor training feature vectors 620. The master bot 114 may select that skill bot 116 for handling the input utterance. Various other implementations for selecting the skill bot 116 are possible and within the scope of the present disclosure.

[0173] In block 1635, master bot 114 may transfer input utterance 303 to the skill bot selected in block 1630. Next, skill bot 116 may process the input utterance 303 and respond to user input 110.

[0174] Use of composite vectors for identifying irrelevant input utterances As described above, generally, some embodiments described herein utilize a set representation of training feature vectors 620 to determine whether input utterance 303 is related to a set of available skill bots 116. Also, as described above, the set representation may be cluster 1110. However, additionally or alternatively, the set representation may be a higher-level feature vector, referred to herein as a composite feature vector.

[0175] FIG. 18 is a diagram showing a method 1800 for initializing classifier model 324 of master bot 114 to determine whether input utterance 303 is unrelated or related to available skill bots 116 using a composite feature vector, according to some embodiments described herein. This method 1800 or a similar method can be used in block 205 of method 200 above for configuring and using master bot 114 to direct input utterance 303 to skill bot 116. Further, method 1800 of FIG. 18 is a more specific variation of method 500 of FIG. 5.

[0176] The method 1800 shown in FIG. 18 and other methods described herein may be implemented in software (e.g., as code, instructions, or a program) by one or more processing units (e.g., a processor or a processor core), in hardware, or in a combination thereof. The software may be stored in a non-transitory storage medium such as a memory device. This method 1800 is intended to be exemplary and non-limiting. FIG. 18 shows various operations that occur in a particular order, but this is not intended to be limiting. In a particular embodiment, for example, these operations are performed in a different order It may be done, or one or more operations of method 1800 may be executed in parallel. In certain embodiments, method 1800 may be executed by a training system 350 that may be part of DABP 102.

[0177] In block 1805 of method 1800, the training system 350 accesses training utterances 615 (also referred to as exemplary utterances) for various skill bots 116 associated with the master bot 114. For example, the training utterances 615 may be stored as skill bot data in training data 354 accessible to the training system 350. In some embodiments, each skill bot 116 available to the master bot 114 may be associated with a subset of the training utterances 615, and each such subset of the training utterances 615 may include the training utterances 615 for each intent that the skill bot 116 is configured for. Thus, the set of training utterances 615 may include the training utterances 615 for each intent of each skill bot 116 such that all intents of all skill bots 116 are represented.

[0178] In block 1810, the training system 350 generates a training feature vector 620 from the training utterance 615 accessed in block 1805. As described above, the training utterance 615 may include a subset associated with each skill bot 116, and may further include the training utterance 615 for each intent of each skill bot 116. Thus, in some embodiments, the training feature vector 620 may include respective subsets for each skill bot 116, and may further include respective feature vectors for each intent of each skill bot. As described above, the training system 350 may generate a training feature vector 620 to describe and represent each training utterance 615. Various techniques are known for converting a sequence of words, such as the training utterance 615, into a feature vector, such as the training feature vector 620, and one or more of such techniques may be used. For example, the training system 350 may encode each training utterance 615 as a corresponding training feature vector 620 using one-hot encoding or some other encoding, but it is not necessarily required to encode.

[0179] In block 1815, the training system 350 divides the training feature vector 620, and thus the corresponding training utterance 615, into conversation categories. The conversation categories can be defined, for example, based on intents, sub-bots, or skill bots 116. In some embodiments, when the conversation categories are based on intents, each conversation category includes a training feature vector 620 that represents a single corresponding intent for which the skill bot 116 is configured. In that case, the number of conversation categories may be equal to the number of intents across the various skill bots 116 available to the master bot 114. In some embodiments, when the conversation categories are based on skill bots 116, each conversation category includes a training feature vector 620 that represents a single corresponding skill bot 116. In that case, the number of conversation categories may be equal to the number of skill bots 116 available to the master bot 114. In some embodiments, when the conversation categories are based on sub-bots, each conversation category includes a subset of the training feature vectors 620 that represent a single corresponding skill bot 116. In that case, the number of conversation categories may be greater than or equal to the number of skill bots 116 available to the master bot 114.

[0180] In block 1820, the training system 350 generates a composite feature vector from the training feature vectors 620, whereby each training feature vector 620 within a conversation category represents and pertains to that conversation category They are aggregated into corresponding synthetic feature vectors. In other words, for each conversation category to which the training feature vector 620 is assigned in block 1815, the training system 350 can combine each training feature vector 620 into a synthetic feature vector for the conversation category. Various techniques can be used for aggregation. For example, the synthetic feature vector can be the average (e.g., arithmetic mean) of the training feature vectors 620 in each category. Based on the conversation category, the synthetic feature vector can be an intent vector, a sub-bot vector, or a bot vector depending on whether the conversation category is defined based on an intent, a sub-bot, or a skill bot 116.

[0181] The synthetic feature vector can be generated as the arithmetic mean of the training feature vectors 620, but other mathematical functions including other types of linear combinations can be additionally or alternatively used to generate the synthetic feature vector. For example, in some embodiments, the synthetic feature vector can be a weighted average of the training feature vectors 620 within the conversation category. The weighting can be based on various factors such as the priority of specific keywords in the corresponding training utterance 615, and thus, a greater weight will be given to the training utterance 615 having specific keywords. When generating a sub-bot vector or a bot vector, the training feature vectors 620 can be aggregated as a weighted average such that the training feature vectors 620 corresponding to a specific intent are given a greater weight than other training feature vectors 620. For example, the training feature vectors 620 corresponding to an intent with a larger number of representative training utterances 615 can be given a greater weight compared to the training feature vectors 620 corresponding to an intent with a smaller number of representative training utterances 615 in a collective state. Various implementations are possible and are within the scope of the present disclosure.

[0182] Figures 19 and 20 illustrate the concept of a synthetic feature vector. Specifically, FIG. 19 shows the generation of a synthetic feature vector using intent-based conversation categories according to some embodiments described herein. As a result, the synthetic feature vector shown in FIG. 19 is at the intent level and is thus an intent vector. In the example of FIG. 19, at least two skillbots 116 are available to the master bot 114. The first skillbot 116 is associated with first skillbot data 358g that includes a training utterance 615 representing intent A, which is the first intent, and another training utterance 615 representing intent B, which is the second intent. Thus, the first skillbot 116 is configured to handle input utterances associated with intent A or intent B. The second skillbot 116 is associated with second skillbot data 358h that includes a training utterance 615 representing intent C, which is the third intent. Thus, the second skillbot 116 is configured to handle input utterances associated with intent C.

[0183] In the example of FIG. 19, the training system 350 converts all training utterances 615 regarding the skillbot 116 into respective feature vectors 620. The training system 350 groups the training feature vectors into intent-based conversation categories. Thus, the training feature vectors 620 corresponding to the training utterances 615 for intent A form a first conversation category, the training feature vectors 620 corresponding to the training utterances 615 for intent B form a second conversation category, and the training feature vectors 620 corresponding to the training utterances 615 for intent C form a third conversation category. In this example, as described above, the training feature vectors 620 of a given intent-based conversation category are aggregated (e.g., averaged) into a synthetic feature vector for the conversation category and thus for the associated intent. Specifically, the training system 350 for intent A The training feature vectors 620 corresponding to the training utterances 615 are aggregated into the first composite feature vector 1910a, and the training system 350 aggregates the training feature vectors 620 corresponding to the training utterances 615 for intent B into the second composite feature vector 1910b, and the training system 350 aggregates the training feature vectors 620 corresponding to the training utterances 615 for intent C into the third composite feature vector 1910c. Thus, in this example, each composite feature vector represents a respective intent of the skill bot 116.

[0184] FIG. 20 shows the generation of composite feature vectors using a skill bot-based conversation category, also referred to as a bot-based conversation category, according to some embodiments described herein. As a result, the composite feature vectors shown in FIG. 20 are at the skill bot level and thus are bot vectors. In the example of FIG. 20, at least two skill bots 116 are available to the master bot 114. The first skill bot 116 is associated with first skill bot data 358g that includes training utterances 615 representing intent A, which is the first intent, and other training utterances 615 representing intent B, which is the second intent. Thus, the first skill bot 116 is configured to handle input utterances associated with intent A or intent B. The second skill bot 116 is associated with second skill bot data 358h that includes training utterances 615 representing intent C, which is the third intent. Thus, the second skill bot 116 is configured to handle input utterances associated with intent C.

[0185] In the example of FIG. 20, the training system 350 converts all training utterances 615 regarding the skillbot 116 into respective feature vectors 620. The training system 350 groups the training feature vectors into skillbot-based conversation categories. Thus, the training feature vectors 620 corresponding to the training utterances 615 for intent A form a first conversation category together with the training feature vectors 620 corresponding to the training utterances 615 for intent B, and the training feature vectors 620 corresponding to the training utterances 615 for intent C form a second conversation category. In this example, as described above, the training feature vectors 620 of a given skillbot-based conversation category are aggregated (e.g., averaged) into a composite feature vector for the conversation category and thus for the associated skillbot 116. Specifically, the training system 350 aggregates the training feature vectors 620 corresponding to the training utterances 615 for intent A and the training feature vectors 620 corresponding to the training utterances 615 for intent B into a first composite feature vector 1910d, and the training system 350 aggregates the training feature vectors 620 corresponding to the training utterances 615 for intent C into a second composite feature vector 1910e. Thus, in this example, each composite feature vector represents a respective skillbot 116.

[0186] In some embodiments, the training system 350 is not limited to generating only one type of composite feature vector. For example, the training system 350 may generate intent vectors, sub-bot vectors, and skillbot vectors, or the training system 350 may generate some other combination of these or other types of composite feature vectors.

[0187] Returning to FIG. 18, in block 1825, the training system 350 may configure the classifier model 324 of the master bot 114 to utilize the synthetic feature vector determined in block 1820. For example, the training system 350 may store an indication of the synthetic feature vector in a memory device accessible by the classifier model 324 (e.g., on the digital assistant 106). The classifier model 324 may be configured to compare an input utterance to such a synthetic feature vector that functions as a set representation of the training feature vectors 620.

[0188] FIG. 21 is a diagram illustrating a method 2100 for determining whether an input utterance 303 provided as a user input 110 is unrelated to any available skill bot 116 associated with the master bot 114, according to some embodiments described herein. This method 2100 or a similar method can be used in block 215 of the above-described method 200 for configuring and using the master bot 114 to direct the input utterance 303 to the skill bot 116. The method 2100 of FIG. 21 is a more specific variation of the method 700 of FIG. 7, and similar to the method 700 of FIG. 7, the method 2100 of FIG. 21 can be used in conjunction with the preliminary filtering operation described with respect to the method 800 of FIG. 8. More specifically, as described below, the classifier model 324 of the master bot 114 may utilize the synthetic feature vector to determine whether the input utterance 303 belongs to the non-class 316 (i.e., is unrelated to the available skill bot 116).

[0189] Method 2100 shown in FIG. 21 and other methods described herein can be implemented in software (e.g., as code, instructions, or programs) executed by one or more processing units (e.g., a processor or processor core), in hardware, or in a combination thereof. The software can be stored in a non-transitory storage medium such as a memory device. This method 2100 is intended to be exemplary and non-limiting. FIG. 21 shows various operations that occur in a particular order, which is not intended to be limiting. In certain embodiments, for example, these operations may be performed in a different order or one or more operations of method 2100 may be performed in parallel. In certain embodiments, method 200 may be performed by master bot 114 associated with a set of available skill bots 116.

[0190] In block 2105 of method 2100, master bot 114 accesses input utterance 303 provided as user input 110. For example, in some embodiments, the user may provide user input 110 in the form of voice input, and digital assistant 106 may convert that user input 110 into text input utterance 303 for use by master bot 114.

[0191] In block 2110, the Masterbot 114 may generate an input feature vector 1710 from the input utterance 303 accessed in block 2105. More specifically, in some embodiments, the Masterbot 114 causes the classifier model 324 of the Masterbot 114 to generate the input feature vector 1710 from the input utterance 303. The input feature vector 1710 may describe and represent the input utterance 303. Various techniques are known for converting a sequence of words, such as the input utterance 303, into a feature vector, and one or more of such techniques may be used. For example, the training system 350 may, but need not, encode the input utterance 303 as the corresponding input feature vector 1710 using one-hot encoding or some other encoding. However, one embodiment of the Masterbot 114 uses the same technique used to generate the training feature vector 620 from the training utterance 615 in training the classifier model 324.

[0192] In block 2115, the Masterbot 114 compares the input feature vector 1710 generated in block 2110 with the composite feature vector. More specifically, in some embodiments, the Masterbot 114 compares the input feature vector 1710 generated in block 2110 with the composite feature vector. The determined input feature vector 1710 is compared to a composite feature vector, which may be determined as described above with respect to Figure 18. In comparing the input feature vector 1710 to the composite feature vectors, the classifier model 324 may determine a similarity or distance between the input feature vector 1710 and each composite feature vector previously constructed for the classifier model 324.

[0193] The similarity or distance between the input feature vector 1710 and another feature vector, specifically the synthetic feature vector, can be calculated in various ways. For example, to determine the similarity, the classifier model 324 may calculate the absolute value of the arithmetic difference (i.e., the Euclidean distance) between the synthetic feature vector and the input feature vector 1710 by subtracting one from the other and obtaining the absolute value. In another example, the classifier model 324 may multiply the input feature vector 1710 by the synthetic feature vector. For instance, if one-hot encoding is used for the input feature vector 1710 and the synthetic feature vector, each vector entry (i.e., each dimension) will have a value of 1 or 0. In this case, a value of 1 may represent the presence of a particular feature. If the input feature vector 1710 and the synthetic feature vector have mostly the same features, the result of the vector-vector multiplication will be a vector having approximately the same number of 1 values as either of the two vectors being multiplied. Otherwise, the resulting vector will mostly be 0. As another example, cosine similarity can be used. For example, the cosine similarity between the input feature vector 1710 and the synthetic feature vector can be calculated as the dot product of the two vectors divided by the product of the Euclidean norms of both vectors. Various other techniques for measuring similarity are feasible and within the scope of the present disclosure.

[0194] In the determination block 2120, the classifier model 324 of the master robot 114 determines whether the input feature vector 1710 is sufficiently similar to any of the synthetic feature vectors. The classifier model 324 may use a predetermined threshold such that the similarity is considered sufficient when the similarity satisfies the threshold. Similar to the case of the distance metric, when two vectors are similar, if the similarity metric used provides a small value, the threshold can be, for example, an upper threshold such that the input feature vector 1710 is sufficiently close to the synthetic feature vector when the similarity is below the threshold. However, when two vectors are similar and the similarity metric used provides a large value, the threshold can be, for example, a lower threshold such that the input feature vector 1710 is sufficiently close to the synthetic feature vector when the similarity is above the threshold.

[0195] In some embodiments, the determination of whether there is sufficient similarity can be made using a hierarchy of composite feature vectors. For example, the input feature vector 1710 is compared with the bot vector until sufficient similarity is found or until it is determined that the input feature vector 1710 is not sufficiently similar to any such composite feature vector, and then, if necessary, with the sub-bot vector, and then, if necessary, with the intent vector. Since the bot vector is less than the intent vector and the intent vector is less than the sub-bot vector, the computing resources required to determine similarity to the bot vector may be less than the computing resources required to determine similarity to the intent vector, and may be less than the computing resources required to determine similarity to the sub-bot vector. Thus, by comparing the input feature vector 1710 with the composite feature vectors at various levels starting from the highest available level (e.g., the skill bot level), the classifier model 324 can use less computationally intensive processing while using more computationally intensive processing only if necessary, and determine which skill bot 116, if any, is most suitable for handling the input utterance 303.

[0196] In some embodiments, the input utterance 303 may be considered irrelevant to any of the available skill bots 116 if the input feature vector 1710 is sufficiently far from all of the bot vectors (e.g., beyond a predetermined distance that is considered excessive). However, in additional or alternative embodiments, the input utterance 303 is not necessarily considered irrelevant based solely on the difference from the bot vectors. For example, if the input feature vector 1710 is not similar to all of the bot vectors (i.e., not sufficiently similar), this is not necessarily a clear indication that the input utterance 303 belongs to the non-class 316. Since the corresponding intent is not similar to the other intents for which the skill bot 116 is configured, there may be a situation where the input feature vector 1710 is far from (i.e., not similar to) any of the bot vectors but still close to a certain intent vector. In some embodiments, comparing the input feature vector 1710 with both the bot vectors and the intent vectors is typically sufficient to identify irrelevant input utterances 303. Thus, the input utterance 303 may be considered irrelevant to the available skill bots 116 if the input feature vector 1710 is not similar to all of the bot vectors and not similar to all of the intent vectors. In some embodiments, the input feature vector 1710 may alternatively or additionally be compared with sub-bot vectors or other composite feature vectors as part of determining whether the input utterance 303 is a member of a non-class. Various implementations are possible and within the scope of the present disclosure.

[0197] In determination block 2120, regardless of what comparison is performed, if based on that, classifier model 324 determines that input feature vector 1710 is not sufficiently similar to any of the synthetic feature vectors, method 2100 proceeds to block 2125. In block 2125, master bot 114 determines that the input utterance is unrelated to any of the available skill bots 116, and for this reason, master bot 114 may indicate that input utterance 303 cannot be processed. As a result, for example, the digital assistant may request the user to clarify user input 110.

[0198] However, in determination block 2120, if classifier model 324 determines that input feature vector 1710 is sufficiently similar to one or more synthetic feature vectors, method 2100 proceeds to block 2130. In block 2130, based on the comparison with the synthetic feature vectors, master bot 114 determines (i.e., selects) a skill bot 116 for handling input utterance 303. For example, master bot 114 may select the skill bot 116 associated with the synthetic feature vector that is considered to be the most similar to input feature vector 1710. For example, if the most similar synthetic feature vector is an intent vector, master bot 114 may select the skill bot 116 configured to handle the intent corresponding to the intent vector. Further, if the most similar synthetic feature vector is a sub-bot vector or a bot vector, master bot 114 may select the skill bot 116 corresponding to the sub-bot vector or the bot vector. In another example, master bot 114 may use the k-nearest neighbor technique to select a skill bot 116, as further described below.

[0199] In block 2135, master bot 114 may route an input utterance to the skill bot 116 selected in block 2130. In some embodiments, if master bot 114 identifies a particular intent for the input utterance (e.g., if input feature vector 1710 is deemed to be close enough to a single intent vector), master bot 114 may indicate that intent to skill bot 116, thereby enabling skill bot 116 to skip the process of inferring the intent. Skill bot 116 may then process input utterance 303 and respond to user input 110.

[0200] FIG. 22 is a diagram showing a method 2200 for selecting a skill bot 116 to handle an input utterance, according to some embodiments described herein. This method 2200 or a similar method may be used in block 2130 of method 2100 described above after it is determined that input utterance 303 is related to at least one available skill bot 116. This method 2200 provides a k-nearest neighbor technique for selecting a skill bot 116, although other techniques may be used in addition to or instead of this technique. Specifically, this method 2200 utilizes k-nearest neighbor training feature vectors 620 for input feature vector 1710 to determine which skill bot 116 should be selected.

[0201] The method 2200 shown in FIG. 22 and the other methods described herein can be implemented in software (e.g., as code, instructions, or programs) executed by one or more processing units (e.g., a processor or a processor core), in hardware, or in a combination thereof. The software can be stored in a non-transitory storage medium such as a memory device. This method 2200 is intended to be illustrative and non-limiting. FIG. 22 shows various operations that occur in a particular order, which is not intended to be limiting. In certain embodiments, for example, these operations may be executed in a different order or one or more operations of method 2200 may be executed in parallel. In certain embodiments, method 2200 may be executed by master bot 114.

[0202] As shown in FIG. 22, at block 2205, master bot 114 determines a value for k, where k is the number of neighbors that will be considered. In some embodiments, the value of k can be a factor of the total number of training feature vectors 620 (i.e., the total number of training utterances 615).

[0203] At block 2210, using the value of k determined at block 2205, master bot 114 determines a set of the k training feature vectors 620 that are closest (i.e., most similar) to input feature vector 1710. Master bot 114 may use the same or a different similarity metric as used above when determining whether input feature vector 1710 is sufficiently similar to any of the composite feature vectors. For example, the similarity metric used can be Euclidean distance, vector multiplication, or cosine similarity.

[0204] In block 2215, master bot 114 selects skill bot 116 that has the most training feature vectors 620 within the set determined in block 2210. In some embodiments, it is not necessary for the training feature vectors 620 of skill bot 116 to constitute a majority of the set; rather, it is only necessary that there is no other skill bot 116 having a greater number of training feature vectors 620 within the set. As described above, after this selection of skill bot 116, master bot 114 may route input utterance 303 to the selected skill bot 116 for processing.

[0205] In addition to, or instead of, considering the training feature vector 620 closest to input feature vector 1710, some embodiments of master bot 114 may take into account the closest composite feature vector, such as the closest intent vector. FIG. 23 is a diagram illustrating another exemplary method 2300 for selecting skill bot 116 to handle an input utterance, according to some embodiments described herein. This method 2300, or a similar method, may be used in block 2130 of method 2100 described above after it is determined that input utterance 303 is related to at least one available skill bot 116. This method 2300 provides a k-nearest neighbor approach for selecting skill bot 116, although other techniques may be used in addition to or instead of this approach. Specifically, FIG. 2 In contrast to method 2200 of FIG. 2, this method 2300 utilizes the k-nearest neighbor intent vector for input feature vector 1710 to determine which skill bot 116 should be selected.

[0206] The method 2300 shown in FIG. 23 and the other methods described herein can be implemented in software (e.g., as code, instructions, or programs) executed by one or more processing units (e.g., a processor or a processor core), in hardware, or in a combination thereof. The software can be stored in a non-transitory storage medium such as a memory device. The method 2300 is intended to be exemplary and non-limiting. FIG. 23 shows various operations that occur in a particular order, but this is not intended to be limiting. In certain embodiments, for example, these operations may be performed in a different order or one or more operations of the method 2300 may be performed in parallel. In certain embodiments, the method 2300 may be executed by the master bot 114.

[0207] As shown in FIG. 23, in block 2305, the master bot 114 determines a value for k. Here, k is the number of neighbors that will be considered. In some embodiments, the value of k can be a factor of the total number of intent vectors. The total number of intent vectors may be equal to the total number of intents represented in the training utterance 615, may be equal to the total number of intents that the available skill bot 116 is configured to handle, and the value of k may be selected as a factor of that amount.

[0208] In block 2310, the master bot 114 uses the value of k determined in block 2305 to determine a set of the k intent vectors that are closest (i.e., most similar) to the input feature vector 1710. The master bot 114 may use the same or a different similarity metric as that used above when determining whether the input feature vector 1710 is sufficiently similar to any of the composite feature vectors in the method 2100 above. For example, the similarity metric used can be Euclidean distance, vector multiplication, or cosine similarity.

[0209] In block 2315, master bot 114 selects the skill bot 116 that has the most intent vectors within the set determined in block 2310. In some embodiments, the intent vectors of the skill bot 116 do not have to constitute a majority of the set; rather, it is only necessary that there is no other skill bot 116 within the set that has more intent vectors. As described above, after this selection of the skill bot 116, the master bot 114 may route the input utterance 303 to the selected skill bot 116 for processing.

[0210] Implementation example FIG. 24 shows a schematic diagram of a distributed system 2400 for implementing one embodiment. In the illustrated embodiment, the distributed system 2400 includes one or more client computing devices 2402, 2404, 2406, and 2408 coupled to a server 2412 via one or more communication networks 2410. The client computing devices 2402, 2404, 2406, and 2408 may be configured to execute one or more applications.

[0211] In various examples, the server 2412 may be adapted to execute one or more services or software applications that enable the processes described in this disclosure.

[0212] In certain embodiments, the server 2412 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments these services may be web-based services or apps to the users of the client computing devices 2402, 2404, 2406 and / or 2408, such as under a software as a service (SaaS) model of the service. It can be provided as a loud service. A user operating client computing devices 2402, 2404, 2406, and / or 2408 can utilize the services provided by these components by interacting with server 2412 using one or more client applications.

[0213] In the configuration shown in FIG. 24, server 2412 may include one or more components 2418, 2420, and 2422 that implement the functions executed by server 2412. These components may include software components that can be executed by one or more processors, hardware components, or combinations thereof. It should be recognized that a wide variety of system configurations different from distributed system 2400 may be implemented. Accordingly, the embodiment shown in FIG. 24 is an example of a distributed system for implementing the system of the embodiment and is not intended to be limiting.

[0214] The user may interact with server 2412 in accordance with the teachings of the present disclosure using client computing devices 2402, 2404, 2406, and / or 2408. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although FIG. 24 shows only four client computing devices, any number of client computing devices may be supported.

[0215] The client device can include various types of computing systems such as portable handheld devices, general-purpose computers such as personal computers and laptops, workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors or other sensing devices. These computing devices can run various types and versions of software applications and operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-based operating systems, Linux® or Linux-based operating systems, such as various mobile operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android®, BlackBerry®, Palm OS®), Google Chrome® OS). Portable handheld devices can include cellular phones, smartphones (e.g., iPhone®), tablets (e.g., iPad®), personal digital assistants (PDAs), etc. Wearable devices can include Google Glass® head-mounted displays and other devices. Gaming systems can include various handheld gaming devices, Internet-connected gaming devices (e.g., Microsoft Xbox® game consoles with / without Kinect® gesture input devices, Sony PlayStation® systems, various gaming systems provided by Nintendo®, etc.). The client device may be capable of running a wide variety of applications such as various Internet-related applications, communication applications (e.g., email applications, short message service (SMS) applications), and may use various communication protocols.

[0216] Network 2410 may be any type of network well-known to those skilled in the art that can support data communication using any of a variety of available protocols, and the above protocols include, but are not limited to, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Exchange), AppleTalk®, etc. By way of example only, network 2410 may include a Local Area Network (LAN), an Ethernet®-based network, Token Ring, Wide Area Network (WAN), Internet, virtual network, Virtual Private Network (VPN), intranet, extranet, Public Switched Telephone Network (PSTN), infrared network, wireless network (e.g., a wireless network operating under any of the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suites, Bluetooth® and / or any other wireless protocol), and / or any combination of these and / or other networks.

[0217] Server 2412 may be configured by one or more general-purpose computers, dedicated server computers (including, by way of example, PC (Personal Computer) servers, UNIX® servers, midrange servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or other suitable configurations and / or combinations. Server 2412 may include one or more virtual machines that execute a virtual operating system, or other computing architectures with virtualization. This may be, for example, one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server. In various embodiments, server 2412 may be adapted to execute one or more services or software applications that provide the functions described in the above disclosure.

[0218] The computing system within server 2412 can execute one or more operating systems including any of the above, as well as commercially available server operating systems. Further, server 2412 can execute any of a variety of additional server applications and / or middle-tier applications including, for example, an HTTP (Hypertext Transfer Protocol) server, an FTP (File Transfer Protocol) server, a CGI (Common Gateway Interface) server, a JAVA (registered trademark) server, a database server, etc. Exemplary database servers include, but are not limited to, those commercially available from Oracle (registered trademark), Microsoft (registered trademark), Sybase (registered trademark), IBM (registered trademark) (International Business Machines), etc. It is not limited thereto.

[0219] In some implementations, server 2412 can include one or more applications for analyzing and collating data feeds and / or event updates received from users of client computing devices 2402, 2404, 2406, and 2408. As an example, the data feeds and / or event updates can include real-time events related to sensor data applications, financial stock market dashboards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc., and can include, but are not limited to, Twitter (registered trademark) feeds, Facebook (registered trademark) updates, or real-time updates received from one or more third-party information sources and continuous data streams. Server 2412 can also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 2402, 2404, 2406, and 2408.

[0220] The distributed system 2400 may also include one or more data repositories 2414, 2416. In certain embodiments, these data repositories may be used to store data and other information. For example, one or more of the data repositories 2414, 2416 may be used to store data or information generated by the processes described herein and / or data or information used for the processes described herein. The data repositories 2414, 2416 may be located in various locations. For example, a data repository used by the server 2412 may be at a local location of the server 2412 or at a remote location from the server 2412 and communicate with the server 2412 via a network-based connection or a dedicated connection. The data repositories 2414, 2416 may be of different types. In certain embodiments, a data repository used by the server 2412 may be a database, such as a relational database provided by Oracle Corporation (registered trademark) and other manufacturers. One or more of these databases may be adapted to enable storage, update, and retrieval of data in response to SQL-format commands.

[0221] In certain embodiments, one or more of the data repositories 2414, 2416 may be used by an application to store application data. A data repository used by an application may be of various types, such as, for example, a key-value store repository, an object store repository, or a general-purpose storage repository supported by a file system.

[0222] In certain embodiments, the functions described in this disclosure may be provided as services via a cloud environment. FIG. 25 is a simplified block diagram of a cloud-based system environment that may provide the functions described herein as cloud services according to certain embodiments. In the embodiment shown in FIG. 25, the cloud infrastructure system 2502 may provide one or more cloud services that a user may request using one or more client computing devices 2504, 2506, and 2508. The cloud infrastructure system 2502 may include one or more computers and / or servers that may include those previously described with respect to server 2412. The computers within the cloud infrastructure system 2502 may be configured as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.

[0223] Network 2510 may facilitate the communication and exchange of data between clients 2504, 2506, and 2508 and the cloud infrastructure system 2502. Network 2510 may include one or more networks. The networks may be of the same type or different types. Network 2510 may support one or more communication protocols, including wired and / or wireless protocols, to facilitate communication.

[0224] The embodiment shown in FIG. 25 is merely an example of a cloud infrastructure system and is not intended to be limiting. In some other embodiments, it should be understood that the cloud infrastructure system 2502 may have more or fewer components than those shown in FIG. 25, may combine two or more components, or may have components with different configurations or arrangements. For example, although FIG. 25 shows three client computing devices, in alternative embodiments, any number of client computing devices may be supported.

[0225] The term "cloud service" is generally used to refer to services that are made available on demand to users via a communication network such as the Internet by a service provider's system (e.g., cloud infrastructure system 2502). Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premises servers and systems. The cloud service provider's system is managed by the cloud service provider. Thus, customers can utilize cloud services provided by the cloud service provider without having to separately purchase licenses, support, or hardware and software resources for the services. For example, the cloud service provider's system can host an application, and users can order and use the application on demand via the Internet without having to purchase infrastructure resources to run the application. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, several cloud services such as middleware services, database services, Java (registered trademark) cloud services, etc. are provided by Oracle Corporation (registered trademark) of Redwood Shores, California.

[0226] In certain embodiments, cloud infrastructure system 2502 can provide one or more cloud services using various models such as a software as a service (SaaS) model, a platform as a service (PaaS) model, an infrastructure as a service (IaaS) model, including a hybrid service model. Cloud infrastructure system 2502 can include a set of applications, middleware, databases, and other resources that enable the provisioning of various cloud services.

[0227] The SaaS model enables applications or software to be delivered to customers as a service over a communication network such as the Internet without the customer having to purchase the underlying hardware or software for the application. For example, by using the SaaS model, customers can be given access to on-demand applications hosted by the cloud infrastructure system 2502. Examples of SaaS services provided by Oracle Corporation (registered trademark) include, but are not limited to, various services for human resources / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, social applications, and the like.

[0228] The IaaS model is generally used to provide flexible computing and storage capabilities by offering infrastructure resources (such as servers, storage, hardware, and networking resources) to customers as cloud services. Various IaaS services are provided by Oracle Corporation (registered trademark).

[0229] The PaaS model is generally used to provide as a service a platform and environmental resources that enable customers to develop, run, and manage applications and services without having to procure, build, or manage the environmental resources. Examples of PaaS services provided by Oracle Corporation (registered trademark) include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), data management cloud services, various application development solution services, and the like.

[0230] Cloud services are generally on an on-demand self-service basis, with a subscription It is provided in a Yon base, flexibly scalable, highly reliable, highly available, and secure manner. For example, a customer may order one or more services provided by the cloud infrastructure system 2502 via a subscription order. Then, the cloud infrastructure system 2502 provides the services requested in the customer's subscription order by executing the processing. For example, in certain embodiments, the chatbot-related functions described herein may be provided as cloud services provided by a user / subscriber. The cloud infrastructure system 2502 may be configured to be provided by one cloud service or a plurality of cloud services.

[0231] The cloud infrastructure system 2502 may provide cloud services via various deployment models. In the public cloud model, the cloud infrastructure system 2502 may be owned by a third-party cloud service provider, and the cloud services are provided to general public customers. This customer may be an individual or an enterprise. In other specific embodiments, under the private cloud model, the cloud infrastructure system 2502 may function within an organization (e.g., within a corporate organization), and the services are provided to customers within this organization. For example, this customer may be various departments of an enterprise such as the personnel department and the salary department, or an individual within the enterprise. In yet another embodiment, under the community cloud model, the cloud infrastructure system 2502 and the provided services may be shared by various organizations within the relevant community. Other various models such as a hybrid model of the above models may also be used.

[0232] Client computing devices 2504, 2506, and 2508 may be of different types (e.g., devices 2402, 2404, 2406, and 2408 shown in FIG. 24), and may be operable with one or more client applications. A user may interact with the cloud infrastructure system 2502, such as by using a client device to request services provided by the cloud infrastructure system 2502. For example, a user may use a client device to request chatbot-related services described in this disclosure.

[0233] In some embodiments, the processing performed by the cloud infrastructure system 2502 may include big data analytics. This analytics may include using, analyzing, and processing large data sets to detect and visualize various trends, behaviors, relationships, etc. within this data. This analytics may be performed by one or more processors, optionally, by processing the data in parallel and performing simulations using the data. The data used in this analytics may include structured data (e.g., data stored in a database or data structured according to a structured model) and / or unstructured data (e.g., a data blob (binary large object)). ject).

[0234] As shown in the embodiment of FIG. 25, the cloud infrastructure system 2502 may include infrastructure resources 2530 that are utilized to facilitate the provision of various cloud services provided by the cloud infrastructure system 2502. The infrastructure resources 2530 may include, for example, processing resources, storage or memory resources, networking resources, and the like.

[0235] In certain embodiments, these resources for supporting the various cloud services provided by the cloud infrastructure system 2502 to different customers may be efficiently provisioned by grouping the resources into sets of resources or resource modules (also referred to as "pods"). Each resource module or pod may include a pre-integrated and optimized combination of one or more types of resources. In certain embodiments, different pods may be pre-provisioned for different types of cloud services. For example, a first set of pods may be provisioned for database services, and a second set of pods, which may include a different combination of resources than the pods within the first set of pods, may be provisioned for Java services or the like. For some services, the resources allocated for provisioning these services may be shared among the services.

[0236] The cloud infrastructure system 2502 itself may internally use a service 2532 that is shared by various components of the cloud infrastructure system 2502 and that facilitates the provisioning of services by the cloud infrastructure system 2502. These internal shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scan / whitelist services, high availability, backup and recovery services, services enabling cloud support, email services, notification services, file transfer services, and the like.

[0237] The cloud infrastructure system 2502 may include a plurality of subsystems. These subsystems may be implemented in software, or in hardware, or a combination thereof. As shown in FIG. 25, the subsystems may include a user interface subsystem 2512 that enables a user or customer of the cloud infrastructure system 2502 to interact with the cloud infrastructure system 2502. The user interface subsystem 2512 may include various different interfaces such as a web interface 2514, an online store interface 2516 where cloud services provided by the cloud infrastructure system 2502 are advertised and can be purchased by consumers, and other interfaces 2518. For example, a customer may use a client device to request (service request 2534) one or more services provided by the cloud infrastructure system 2502 using one or more of the interfaces 2514, 2516, and 2518. For example, a customer may access an online store, browse the cloud services provided by the cloud infrastructure system 2502, and place a subscription order for one or more services provided by the cloud infrastructure system 2502 and desired by the customer to subscribe to. This service request may include information identifying the customer and one or more services the customer desires to subscribe to.

[0238] In certain embodiments, such as the embodiment shown in FIG. 25, the cloud infrastructure system 2502 may include an order management subsystem (OMS) 2520 configured to process new orders. As part of this process, the OMS 2520 may create a customer account if it does not already exist, receive billing and / or account information from the customer for use in billing the customer for the requested services provided to the customer, verify the customer information, and after verification, reserve the order for the customer and configure various workflows to prepare the order for provisioning.

[0239] Upon proper validation, the OMS 2520 may call an order provisioning subsystem (OPS) 2524 configured to provision resources for this order, including processing, memory, and networking resources. Provisioning may include allocating resources for the order and configuring the resources to facilitate the services requested by the customer order. The way resources are provisioned for an order and the type of resources provisioned may depend on the type of cloud service the customer ordered. For example, according to a certain workflow, the OPS 2524 may be configured to determine the specific cloud service requested and identify the number of pods that would be pre-configured for this specific cloud service. The number of pods allocated for an order may depend on the size / volume / level / scope of the service requested. For example, the number of pods to allocate may be determined based on the number of users the service is to support, the period for which the service is requested, etc. Next, the allocated pods may be customized for the specific customer making the request to provide the requested services.

[0240] The cloud infrastructure system 2502 may send a response or notification 1044 to the requesting customer to indicate when the requested service will be available. In some examples, the customer may be sent information (e.g., a link) that enables the customer to begin using and utilizing the benefits of the requested service.

[0241] The cloud infrastructure system 2502 may provide services to multiple customers. For each customer, the cloud infrastructure system 2502 manages information related to one or more subscription orders received from the customer, maintains customer data related to the orders, and serves to provide the requested services to the customer. Also, the cloud infrastructure system 2502 may collect usage statistics regarding the use of the subscribed services by the customer. For example, the statistics may be collected regarding the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime. This usage information may be used to bill the customer. Billing may be done, for example, monthly.

[0242] The cloud infrastructure system 2502 may provide services to multiple customers in parallel. The cloud infrastructure system 2502 may store information about these customers, which may include copyright information in some cases. In certain embodiments, the cloud infrastructure system 2502 includes an identity management subsystem (IMS) 2528 configured to manage customer information and separate the information being managed so that information about one customer is not accessible to another customer. The IMS 2528 may be configured to provide various security-related services such as identity services such as information access management, authentication and authorization services, and services for managing the customer's identity and role and related capabilities.

[0243] FIG. 26 shows an exemplary computer system 2600 that may be used to implement certain embodiments. For example, in some embodiments, computer system 2600 may be any of the systems and subsystems of a chatbot system and may be used to implement the various servers and computer systems described above. As shown in FIG. 26, computer system 2600 includes various subsystems including a processing subsystem 2604 that communicates with several other subsystems via a bus subsystem 2602. These other subsystems may include a processing acceleration unit 2606, an I / O subsystem 2608, a storage subsystem 2618, and a communication subsystem 2624. Storage subsystem 2618 may include a non-transitory computer-readable storage medium including a storage medium 2622 and a system memory 2610.

[0244] Bus subsystem 2602 provides a mechanism for enabling the various components and subsystems of computer system 2600 to communicate with each other as intended. Bus subsystem 2602 is shown schematically as a single bus, although alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 2602 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, a local bus, using any of a variety of bus architectures. For example, such architectures may include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and IEEE P1386. ​It may include a Peripheral Component Interconnect (PCI) bus, which can be implemented as a mezzanine bus manufactured according to specifications, and the like.

[0245] The processing subsystem 2604 controls the operation of the computer system 2600 and may include one or more processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The processor may include a single-core or multi-core processor. The processing resources of the computer system 2600 can be organized into one or more processing units 2632, 2634, etc. The processing unit may include one or more processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some embodiments, the processing subsystem 2604 may include one or more dedicated coprocessors such as a graphics processor, a digital signal processor (DSP), etc. In some embodiments, some or all of the processing units of the processing subsystem 2604 may be implemented using customized circuits such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).

[0246] In some embodiments, the processing units within the processing subsystem 2604 may execute instructions stored in the system memory 2610 or the computer-readable storage medium 2622. In various embodiments, the processing units may execute various programs or code instructions and may maintain multiple programs or processes to be executed simultaneously. At any given time, some or all of the program code to be executed may reside in the system memory 2610 and / or the computer-readable storage medium 2622, which may potentially include one or more storage devices. Through appropriate programming, the processing subsystem 2604 may provide the various functions described above. In an example where the computer system 2600 is running one or more virtual machines, one or more processing units may be assigned to each virtual machine.

[0247] In certain embodiments, a processing acceleration unit 2606 may optionally be provided to execute customized processing to accelerate the overall processing executed by computer system 2600, or to offload a portion of the processing executed by processing subsystem 2604.

[0248] I / O subsystem 2608 may include devices and mechanisms for inputting information into computer system 2600 and / or for outputting information from, or via, computer system 2600. In general, the use of the term "input device" is intended to include all conceivable types of devices and mechanisms for inputting information into computer system 2600. User interface input devices include, for example, a keyboard, a mouse or other pointing device such as a trackball, a touchpad or touch screen incorporated in a display, a scroll wheel, a click wheel, a dial, a button, a switch, a key The input device may include a headphone, a voice input device with a voice command recognition system, a microphone, and other types of input devices. The user interface input device may also include a motion sensing and / or gesture recognition device, such as a Microsoft Kinect (registered trademark) motion sensor, a Microsoft Xbox (registered trademark) 360 game controller, a device with an interface for receiving inputs using gestures and voice commands, etc., which enables the user to control and interact with the input device. The user interface input device may also include an eye gesture recognition device, such as a Google Glass (registered trademark) blink detector, which detects the user's eye movements (e.g., "blinks" while taking a photo and / or while making a menu selection) and converts the eye gesture into an input to the input device (e.g., Google Glass (registered trademark)). In addition, the user interface input device may include a voice recognition sensing device that enables the user to interact with a voice recognition system (e.g., Siri (registered trademark) navigator) via voice commands.

[0249] Other examples of user interface input devices include, but are not limited to, auditory / visual devices such as a three-dimensional (3D) mouse, a joystick or a pointing stick, a game pad and a graphic tablet, and a speaker, a digital camera, a digital camcorder, a portable media player, a webcam, an image scanner, a fingerprint scanner, a barcode reader, a 3D scanner, a 3D printer, a laser range finder, and an eye tracking device. In addition, the user interface input device may include a medical imaging input device such as, for example, a computed tomography, a magnetic resonance imaging, a positron emission tomography, and a medical ultrasonic examination device. The user interface input device may also include a voice input device such as, for example, a MIDI keyboard, a digital musical instrument, etc.

[0250] In general, the use of the term output device is intended to include all possible types of devices and mechanisms for outputting information from computer system 2600 to a user or another computer. User interface output devices may include non-visual displays such as a display subsystem, indicator lights, or audio output devices. The display subsystem may be a flat panel device such as one using a cathode ray tube (CRT), liquid crystal display (LCD), or plasma display, a projection device, a touch screen, etc. For example, user interface output devices may include, but are not limited to, various display devices such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, audio output devices, and modems that visually convey text, graphics, and audio / video information.

[0251] Storage subsystem 2618 provides a repository or data store for storing information and data used by computer system 2600. Storage subsystem 2618 provides a tangible non-transitory computer-readable storage medium for storing the basic programming and data configurations that provide the functionality of some embodiments. Software (e.g., programs, code modules, instructions) that provides the above-described functionality when executed by processing subsystem 2604 may be stored in storage subsystem 2618. The software may be executed by one or more processing units of processing subsystem 2604. Storage subsystem 2618 may also provide a repository for storing data used in accordance with the teachings of the present disclosure.

[0252] Storage subsystem 2618 may include one or more non-transitory memory devices including volatile and non-volatile memory devices. As shown in FIG. 26, storage subsystem 2618 includes system memory 2610 and computer-readable storage medium 2622 . The system memory 2610 may include several memories, including a volatile main random access memory (RAM) for storing instructions and data during program execution, and a non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some embodiments, a basic input / output system (BIOS) including basic routines that assist in transferring information between elements within the computer system 2600, such as during startup, may typically be stored in the ROM. . Typically, the RAM includes data and / or program modules that are currently being operated on and executed by the processing subsystem 2604. In some embodiments, the system memory 2610 may include multiple different types of memories, such as static random access memory (SRAM), dynamic random access memory (DRAM), and the like.

[0253] As an example, without limitation, as shown in FIG. 26, the system memory 2610 may load an application program 2612, program data 2614, and an operating system 2616 that are in execution and may include various applications such as a web browser, a middle-tier application, a relational database management system (RDBMS), and the like. As an example, the operating system 2616 may include Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or various versions of mobile operating systems such as iOS®, Windows Phone, Android® OS, BlackBerry® OS, Palm® OS operating systems.

[0254] The computer-readable storage medium 2622 may store programming and data configurations that provide the functionality of some embodiments. The computer-readable storage medium 2622 may provide storage of computer-readable instructions, data structures, program modules, and other data for the computer system 2600. Software (programs, code modules, instructions) that provides the above functionality when executed by the processing subsystem 2604 may be stored in the storage subsystem 2618. As an example, the computer-readable storage medium 2622 may include non-volatile memory such as a hard disk drive, a magnetic disk drive, a CD ROM, a DVD, an optical disk drive such as a Blu-Ray (registered trademark) disk, or other optical media. The computer-readable storage medium 2622 may include, but is not limited to, a Zip (registered trademark) drive, a flash memory card, a Universal Serial Bus (USB) flash drive, a Secure Digital (SD) card, a DVD disk, a digital video tape, etc. The computer-readable storage medium 2622 may also include solid-state drives (SSDs) based on non-volatile memory such as flash memory-based SSDs, enterprise flash drives, solid-state ROMs, etc., SSDs based on volatile memory such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs.

[0255] In certain embodiments, the storage subsystem 2618 may further include a computer-readable storage medium reader 2620 that is connectable to the computer-readable storage medium 2622. The reader 2620 may be configured to receive and read data from a memory device such as a disk, a flash drive, etc.

[0256] In certain embodiments, computer system 2600 may support virtualization techniques including, but not limited to, virtualization of processing and memory resources. For example, computer system 2600 may provide support for running one or more virtual machines. In certain embodiments, computer system 2600 may execute a program such as a hypervisor that facilitates the configuration and management of virtual machines. Memory, computing (e.g., processors, cores), I / O, and networking resources may be allocated to each virtual machine. Each virtual machine typically runs independently of other virtual machines. A virtual machine may execute its own operating system, which may be the same as or different from the operating systems executed by other virtual machines typically run by computer system 2600. Thus, potentially multiple operating systems may be executed simultaneously by computer system 2600.

[0257] Communication subsystem 2624 provides an interface to other computer systems and networks. Communication subsystem 2624 functions as an interface for the transfer of data between other systems and computer system 2600. For example, communication subsystem 2624 may enable computer system 2600 to establish a communication channel to one or more client devices via the Internet for the transfer of information between computer system 2600 and the one or more client devices.

[0258] The communication subsystem 2624 may support both wired and / or wireless communication protocols. For example, in certain embodiments, the communication subsystem 2624 may include, for example, radio frequency (RF) transceiver components for accessing a wireless voice and / or data network using (e.g., cellular phone technology, advanced data network technologies such as 3G, 4G or EDGE (Enhanced Data rates for Global Evolution), WiFi (IEEE802.XX family of standards, or other mobile communication technologies, or any combination thereof), a global positioning system (GPS) receiver component, and / or other components. In some embodiments, the communication subsystem 2624 may provide a wired network connection (e.g., Ethernet (registered trademark)) in addition to or instead of a wireless interface.

[0259] The communication subsystem 2624 may receive and transmit data in various formats. For example, in some embodiments, the communication subsystem 2624 may receive input communications in the form of, in addition to other formats, structured data feeds and / or unstructured data feeds 2626, event streams 2628, event updates 2630, etc. For example, the communication subsystem 2624 may be configured to receive (or transmit) data feeds 2626 in real time from users of other communication services such as social media networks and / or Twitter (registered trademark) feeds, Facebook (registered trademark) updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.

[0260] In certain embodiments, the communication subsystem 2624 may be configured to receive data in the form of a continuous data stream, which may include an event stream 2628 and / or event updates 2630 of real-time events that are inherently continuous or infinite and have no explicit end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial stock market dashboards, network performance measurement tools (such as network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and the like.

[0261] The communication subsystem 2624 may be configured to communicate data from the computer system 2600 to other computer systems or networks. This data may be communicated to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 2600 in various different forms such as structured and / or unstructured data feeds 2626, event streams 2628, event updates 2630, and the like.

[0262] The computer system 2600 may be of any one of various types, including a handheld portable device (e.g., an iPhone (registered trademark) cellular phone, an iPad (registered trademark) computing tablet, a PDA), a wearable device (e.g., a Google Glass (registered trademark) head-mounted display), a personal computer, a workstation, a mainframe, a kiosk, a server rack, or other data processing systems. Since the nature of computers and networks is constantly changing, the description of the computer system 2600 shown in FIG. 26 is intended as a specific example only. Many other configurations are possible with more or fewer components than the system shown in FIG. 26. Those skilled in the art will recognize other aspects and / or methods for implementing various embodiments based on the disclosure and teachings herein.

[0263] Although specific embodiments have been described, various modifications, variations, alternative configurations, and equivalents are possible. Embodiments are not limited to operating within a particular data processing environment and can operate freely within multiple data processing environments. Additionally, although specific embodiments have been described using a particular set of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. Some of the flowcharts describe operations as sequential processes, but many of these operations can be performed in parallel or simultaneously. Additionally, the order of operations may be re-specified. The process may have additional steps not included in the figures. The various features and aspects of the above embodiments may be used individually or together.

[0264] Furthermore, although specific embodiments have been described using specific combinations of hardware and software, it should be understood that other combinations of hardware and software are possible. The specific embodiments may be implemented using only hardware, only software, or a combination thereof. The various processes described herein may be implemented on the same processor or on separate processors in any combination.

[0265] When a device, system, component, or module is described as being configured to perform a particular operation or function, such a configuration can be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (such as a microprocessor) to perform the operation, by programming computer instructions or code that execute code or instructions stored in a non-transitory memory medium or any combination thereof, or by executing a processor or core, etc. Processes can communicate using a variety of techniques including, but not limited to, conventional techniques for inter-process communication, different pairs of processes may use different techniques, and the same pair of processes may use different techniques at different times.

[0266] In the present disclosure, embodiments are made to be fully understood by showing specific details. However, the embodiments can be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary detail in order not to obscure the embodiments. This specification provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of other embodiments. Rather, the above description of the embodiments provides those skilled in the art with an explanation that enables the implementation of various embodiments. Various changes are possible within the scope of the functions and configurations of the elements.

[0267] ​Accordingly, the specification and the accompanying drawings are to be regarded as illustrative rather than restrictive. However, it will be apparent that additions, subtractions, deletions, as well as other modifications and changes can be made to these without departing from the broader spirit and scope described in the claims. Thus, while specific embodiments have been described, these are not intended to be limiting. Various variations and equivalents as well as any combination of the disclosed features are within the scope of the appended patent claims.

Claims

1. 1. A system comprising: a training system configured to train a classifier model, training the classifier model comprising: accessing training utterances associated with a skill bot, the training utterances including respective training utterances associated with each of the skill bots, each of the skill bots being configured to provide a dialogue with a user; and training the classifier model further comprising: generating training feature vectors from the training utterances, the training feature vectors including a respective training feature vector associated with each of the skill bots; and training the classifier model further comprises: generating a plurality of set representations of the training feature vector, each set representation of the plurality of set representations corresponding to a subset of the training feature vector, and training the classifier model further comprises: configuring the classifier model to compare an input feature vector to the plurality of set representations of the training feature vector, the system further comprising: A master bot, the master bot comprising: accessing an input utterance as user input; generating an input feature vector from the input utterance; comparing the input feature vector with the plurality of sets of representations of the training feature vectors using the classifier model; and outputting an indication that the user input cannot be addressed by the skill bot based on the input feature vector being outside of the plurality of set representations.

2. generating the plurality of set representations of the training feature vectors includes generating clusters to which the training feature vectors are assigned; The system of claim 1 , wherein comparing the input feature vector to the plurality of set representations comprises determining that the input feature vector does not fall within a boundary of the cluster.

3. generating the clusters determining respective centroid positions for the initial clusters in the feature space; and assigning each training feature vector to an initial cluster having a centroid position to which the training feature vector is closest among the initial clusters, wherein generating the clusters further includes: determining boundaries of the initial clusters, each of the initial clusters including a corresponding assigned training feature vector, and generating the clusters further includes: updating the initial clusters in response to determining that a stopping condition is not satisfied, wherein updating the initial clusters includes: determining an incremental count of a cluster such that the incremental count of the cluster is greater than an initial count of one of the initial clusters; determining respective centroid positions for the clusters in the feature space; assigning each training feature vector to one of the clusters having a respective centroid position to which the training feature vector is closest; determining the boundaries of the clusters, each of the clusters corresponding to The system of claim 2 , further comprising assigned training feature vectors for each of the plurality of training feature vectors.

4. The action of the master bot further comprises: accessing a second input utterance as a second user input; generating a second input feature vector from the second input utterance; determining that the second input feature vector is within a boundary of a cluster of the clusters; and forwarding the second input utterance to a skill bot associated with the certain cluster for processing based on the second input feature vector being within the boundary of the certain cluster.

5. The system of claim 4 , further comprising the skillbot, the skillbot configured to process the input utterance to perform an action in response to the user input.

6. Generating the plurality of set representations of the training feature vectors comprises: Segmenting the training utterances into conversational categories; and generating a composite feature vector corresponding to the conversation categories, wherein generating the composite feature vector comprises: generating a respective composite feature vector for each conversation category of the conversation categories as a collection of training feature vectors for each of the training utterances in the conversation category.

7. 7. The system of claim 6, wherein for each conversation category, generating the composite feature vector as a collection of the respective training feature vectors of the training utterances in the conversation category comprises averaging the respective training feature vectors of the training utterances in the conversation category.

8. 7. The system of claim 6, wherein the conversation categories are defined based on intents for which the skillbot is configured, such that each conversation category corresponds to a respective skillbot intent and includes training utterances that represent the respective skillbot intent.

9. The system of claim 6 , wherein each of the conversation categories corresponds to a respective one of the skill bots and includes training utterances representative of the respective skill bot.

10. 7. The system of claim 6, wherein comparing the input feature vector to the multiple set representations of the training feature vector comprises determining that the input feature vector is not sufficiently similar to any of the synthetic feature vectors.

11. The action of the master bot further comprises: accessing a second input utterance as a second user input; generating a second input feature vector from the second input utterance; determining that the second input feature vector is sufficiently similar to a composite feature vector of the composite feature vectors; and forwarding the second input utterance to a skill bot associated with the synthetic feature vector for processing based on the second input feature vector being sufficiently similar to the synthetic feature vector.

12. The skill bot further includes a skill bot that is responsive to the user input to The system of claim 11 , configured to process the input utterance to perform a speech recognition action.

13. 1. A method comprising: accessing, by a computer system, training utterances associated with skill bots, the training utterances including a respective subset of training utterances for each of the skill bots, each of the skill bots being configured to provide an interaction with a user, the method further comprising: generating training feature vectors from the training utterances, the training feature vectors including a respective training feature vector for each training utterance of the training utterances, the method further comprising: determining centroid positions for the clusters in the feature space; assigning each training feature vector to a respective one of the clusters having a respective centroid position to which the training feature vector is closest; and iteratively modifying the clusters until a stopping condition is met, the step of modifying the clusters comprising: incrementing the count of the clusters to an updated count; determining a new centroid position for the cluster by an amount equal to the updated count; and reassigning the training feature vectors to the clusters based on their proximity to the new centroid locations, the method further comprising: determining boundaries of the clusters, the boundaries including a respective boundary for each of the clusters; accessing an input utterance; converting the input utterance into an input feature vector; determining that the input feature vector is outside the boundary of the cluster by comparing the input feature vector to the boundary of the cluster; and outputting an indication that the input utterance cannot be addressed by the skill bot based on the input feature vector being outside of the cluster in the feature space.

14. accessing a second input utterance as a second user input; generating a second input feature vector from the second input utterance; determining that the second input feature vector is within one or more of the clusters; and forwarding the second input utterance to a skill bot associated with the one or more clusters for processing based on the second input feature vector being within the range of the one or more clusters.

15. The one or more clusters include respective training utterances of the skill bot and respective training utterances of a second skill bot, and the method further comprises: The method of claim 14 , comprising selecting the skill bot from between the skill bot and the second skill bot to process the input utterance based on respective confidence scores calculated for the skill bot and the second skill bot.

16. The one or more clusters include respective training utterances of the skill bot and respective training utterances of a second skill bot, and the method further comprises: The respective training utterances of the skill bot and the second skill bot 15. The method of claim 14, comprising selecting the skill bot from among the skill bot and the second skill bot to process the input utterance based on applying a k-nearest neighbor technique to the respective training utterances.

17. 1. A method comprising: accessing, by a computer system, training utterances associated with skill bots, the training utterances including a respective subset of training utterances for each skill bot of the skill bots, each skill bot of the skill bots configured to provide an interaction with a user, the method further comprising: generating training feature vectors from the training utterances, the training feature vectors including a respective training feature vector for each training utterance of the training utterances, the method further comprising: Segmenting the training utterances into speech categories; and generating composite feature vectors corresponding to the conversation categories, wherein generating the composite feature vectors comprises, for each conversation category of the conversation categories, generating a respective composite feature vector as a collection of training feature vectors for each of the training utterances in the conversation category, the method further comprising: accessing an input utterance; converting the input utterance into an input feature vector; determining that the input feature vector is not sufficiently similar to the composite feature vector by comparing the input feature vector to the composite feature vector; and outputting an indication that the input utterance cannot be addressed by the skill bot based on the input feature vector being not sufficiently similar to the synthetic feature vector.

18. accessing a second input utterance; converting the second input utterance into a second input feature vector; determining that the second input feature vector is sufficiently similar to a composite feature vector of the composite feature vectors by comparing the second input feature vector to the composite feature vectors; and forwarding the second input utterance to a skill bot associated with the synthetic feature vector for processing based on the second input feature vector being sufficiently similar to the synthetic feature vector.

19. the conversation categories are defined based on the intents that the skill bot is configured to address, such that each conversation category corresponds to a respective one or more skill bot intents and includes training utterances representative of the respective one or more skill bot intents; The method of claim 18 , wherein the composite feature vector corresponds to a skillbot intent that the skillbot is configured to address.

20. determining that the second input feature vector is sufficiently similar to the composite feature vector by comparing the second input feature vector to the composite feature vector, determining that the second input feature vector is sufficiently similar to one or more additional composite feature vectors corresponding to one or more additional skillbot intents of the skillbot; and performing a k-nearest neighbor analysis, the performing the k-nearest neighbor analysis comprising: identifying neighboring composite feature vectors in a predefined amount, the neighboring composite feature vectors being closest to the input feature vector, and performing the k-nearest neighbor analysis; The step of further comprising: determining that a majority of the composite feature vectors of the neighborhood correspond to a skillbot intent that the skillbot is configured to address; and selecting the skillbot based on the majority of the composite feature vectors of the neighborhood that correspond to a skillbot intent that the skillbot is configured to address.

21. A computer configured to carry out the method of any one of claims 13 to 20.

22. 1. A method comprising: accessing an input utterance as user input; generating an input feature vector from the input utterance; comparing the input feature vector to a plurality of sets of representations of the training feature vectors using a classifier model; and outputting an indication that the user input cannot be addressed by a skill bot based on the input feature vector being outside the range of the plurality of set representations, wherein each skill bot is configured to provide interaction with a user.

23. Training the classifier model may further include training the classifier model, the training of the classifier model comprising: accessing training utterances associated with the skill bots, the training utterances including respective training utterances associated with each of the skill bots; and training the classifier model further comprising: generating training feature vectors from the training utterances, the training feature vectors including a respective training feature vector associated with each of the skill bots; and training the classifier model further comprising: generating a plurality of set representations of the training feature vectors, each set representation of the plurality of set representations corresponding to a subset of the training feature vectors, and training the classifier model further comprises:

23. The method of claim 22, comprising configuring the classifier model to compare an input feature vector to the multiple set representations of the training feature vector.

24. generating the plurality of set representations of the training feature vectors includes generating clusters to which the training feature vectors are assigned; 24. The method of claim 23, wherein comparing the input feature vector to the plurality of set representations comprises determining that the input feature vector does not fall within a boundary of the cluster.

25. The step of generating clusters includes: determining respective centroid positions for the initial clusters in the feature space; assigning each training feature vector to one of the initial clusters having a respective centroid position to which the training feature vector is closest; determining boundaries of the initial clusters, each initial cluster including a corresponding assigned training feature vector, and generating the clusters further comprises: updating the initial clusters in response to determining that a stopping condition has not yet been satisfied, the updating of the initial clusters comprising: determining an incremental count of a cluster such that the incremental count of the cluster is greater than an initial count of one of the initial clusters; determining respective centroid positions for the clusters in the feature space; assigning each training feature vector to one of the clusters having a respective centroid position to which the training feature vector is closest; and determining the boundaries of the clusters, each of the clusters including a corresponding assigned training feature vector.

26. accessing a second input utterance as a second user input; generating a second input feature vector from the second input utterance; determining that the second input feature vector is within a boundary of one of the clusters; 26. The method of claim 25, further comprising: forwarding the second input utterance to a skill bot associated with the certain cluster for processing based on the second input feature vector being within the boundary of the certain cluster.

27. 27. The method of claim 26, further comprising configuring the skill bot to process the input utterance to perform an action in response to the user input.

28. The step of generating the plurality of sets of representations of the training feature vectors comprises: Segmenting the training utterances into speech categories; and generating a composite feature vector corresponding to the conversation categories, wherein generating the composite feature vector comprises: for each conversation category of the conversation categories, generating a respective composite feature vector as a collection of training feature vectors for each of the training utterances in the conversation category.

29. 30. The method of claim 28, wherein for each conversation category, generating a composite feature vector as a collection of the respective training feature vectors of the training utterances in the conversation category comprises averaging the respective training feature vectors of the training utterances in the conversation category.

30. 30. The method of claim 28, wherein the conversation categories are defined based on intents for which the skillbot is configured, such that each conversation category corresponds to a respective skillbot intent and includes training utterances representative of the respective skillbot intent.

31. 30. The method of claim 28, wherein each of the conversation categories corresponds to a respective one of the skill bots and includes training utterances representative of the respective skill bot.

32. 30. The method of claim 28, wherein comparing the input feature vector to the multiple set representations of the training feature vector comprises determining that the input feature vector is not sufficiently similar to any of the synthetic feature vectors.

33. The action of the master bot further comprises: accessing a second input utterance as a second user input; generating a second input feature vector from the second input utterance; The second input feature vector corresponds to a composite feature vector among the composite feature vectors. determining that the images are sufficiently similar; and forwarding the second input utterance to a skill bot associated with the synthetic feature vector for processing based on the second input feature vector being sufficiently similar to the synthetic feature vector.

34. 34. The method of claim 33, further comprising configuring the skill bot to process the input utterance to perform an action in response to the user input.

35. Dialogue system comprising means for carrying out the method according to any one of claims 22 to 34.

Citation Information

Patent Citations

  • Dialogue system and domain determining method

    JP2019070957A

  • Voice inquiry system, voice inquiry processing method, smart speaker operation server device, chatbot portal server device, and program.

    JP6555838B1

  • Method and system for facilitating a user-machine conversation

    US20180181558A1