Pervasive Deep Networks for Language Detection Using Hash Embedding
Patent Information
- Application Number
- JP2024526927
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-04
- Filing Date
- 2022-11-07
- Publication Date
- 2025-06-11
AI Technical Summary
Existing language detection techniques are inefficient and costly for organizations using instant messaging platforms, particularly in identifying the language of text units generated by speech-to-text conversion, which is crucial for effective multilingual bot interactions.
A computer-implemented method using a deep network with an attention mechanism and hash embeddings to process n-grams, combining convolutional neural networks with an embedding layer and classifier for accurate language prediction.
The method achieves performance comparable to or better than existing language detection APIs, reducing costs and improving the accuracy of language identification in diverse text units.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications This application claims priority to U.S. Provisional Application No. 63 / 263,728, filed November 8, 2021, entitled "WIDE AND DEEP NETWORK FOR LANGUAGE DETECTION USING HASHED EMBEDDINGS," and to U.S. Nonprovisional Application No. 18 / 052,694, filed November 4, 2022, entitled "WIDE AND DEEP NETWORK FOR LANGUAGE DETECTION USING HASH EMBEDDINGS," which are hereby incorporated by reference in their entireties for all purposes.
[0002] FIELD OF THEINVENTION The present disclosure relates generally to natural language processing, and more specifically to techniques for language detection. [Background technology]
[0003] background Many users around the world rely on instant messaging and chat platforms to get instant responses. Organizations often use these instant messaging and chat platforms to have live conversations with customers (or end users). However, it can be very costly for organizations to employ service representatives to engage in live communication with customers or end users. Chatbots (or "bots") have begun to be developed to simulate conversations with end users, especially over the Internet. End users can communicate with the bots through the messaging apps they already have installed and use. Intelligent bots generally leverage artificial intelligence (AI) and can communicate more intelligently and contextually in live conversations, allowing for a more natural conversation between the bot and the end user, improving the conversation experience. Instead of end users learning a fixed set of keywords or commands that the bot knows how to respond to, intelligent bots may be able to understand the end user's intent based on the user's utterances in natural language and respond accordingly.
[0004] Language detection is the task of identifying the language of a text unit. Examples of text units include sentences, emails, posts, text messages, product reviews, paragraphs, or documents. Text units may be generated by a speech-to-text module in response to an utterance. Language detection is one of the first steps in many text processing tasks, such as machine translation, text classification, etc. For example, accurate language detection can be important for the successful deployment of a multilingual bot. Summary of the Invention [Means for solving the problem]
[0005] overview The techniques disclosed herein generally relate to language detection (e.g., natural language processing). Examples of machine learning (ML) models that can be used to perform language detection include wide networks. For example, machine learning approaches to language detection can include presenting input text to a wide network as strings or as sequences of n-grams or subwords. The techniques disclosed herein can provide text-based language detection.
[0006] In various embodiments, a computer-implemented method for language detection includes obtaining a sequence of n-grams of a text unit, obtaining a plurality of ordered embedding vectors of the sequence of n-grams using an embedding layer, obtaining an encoding vector based on the plurality of ordered embedding vectors using a deep network, and obtaining a language prediction for the text unit based on the encoding vector using a classifier. The embedding layer includes a trained model having a plurality of component vectors, and the deep network includes a trained convolutional neural network with an attention mechanism (e.g., one or more attention layers). In this method, obtaining the plurality of ordered embedding vectors using the embedding layer includes, for each n-gram in the sequence of n-grams, obtaining a first hash value of the n-gram and a second hash value of the n-gram, selecting the first component vector from among the plurality of component vectors based on the first hash value, selecting the second component vector from among the plurality of component vectors based on the second hash value, and obtaining an embedding vector for the n-gram by concatenating the first component vector and the second component vector. In some embodiments, the deep network includes a convolutional neural network that is trained with an attention mechanism.
[0007] In some embodiments, the sequence of n-grams includes multiple character level n-grams and multiple word level n-grams, in which the value of n for the multiple character level n-grams is different from the value of n for the multiple word level n-grams.
[0008] In some embodiments, for each n-gram in the sequence of n-grams, obtaining a first hash value of the n-gram includes applying a hash function having a first random seed value to the n-gram, and obtaining a second hash value of the n-gram includes applying a hash function having a second random seed value to the n-gram, where the second seed value is different from the first seed value.
[0009] In some embodiments, obtaining the multiple embedded vectors ordered using the embedding layer includes, for each n-gram in the sequence of n-grams, applying a modulo function to the first hash value to obtain a first index and applying the modulo function to the second hash value to obtain a second index, wherein a selection of the first component vector is based on the first index and a selection of the second component vector is based on the second index.
[0010] In some embodiments, for each n-gram in the sequence of n-grams, obtaining an embedding vector for the n-gram includes concatenating the first component vector and the second component vector.
[0011] In some embodiments, a deep network, including a convolutional neural network trained with an attention mechanism, is used on the sequence of n-gram embedding vectors to generate a final encoded vector representing the text unit, taking into account the order of the n-grams that appear in the text unit.
[0012] In some embodiments, the classifier comprises a feed-forward neural network. In some embodiments, for the text unit coding vectors, using the classifier comprises applying a softmax function to the output of a final layer of the feed-forward neural network.
[0013] In various embodiments, an apparatus is provided that includes a processing circuit for performing some or all of one or more methods disclosed herein and a memory, coupled to the processing circuit, for storing a sequence of n-grams.
[0014] In various embodiments, a system is provided that includes one or more data processors and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.
[0015] In various embodiments, a computer program product tangibly embodied in one or more non-transitory machine-readable media includes instructions configured to cause one or more data processors to perform some or all of one or more methods disclosed herein.
[0016] The techniques described above and below can be implemented in many ways and in many contexts. As described in more detail below, some example implementations and contexts are provided with reference to the following figures. However, the following implementations and contexts are only a few of the many possible implementations and contexts. [Brief description of the drawings]
[0017] [Figure 1] FIG. 1 is a simplified block diagram of a distributed environment incorporating an illustrative embodiment. [Diagram 2]FIG. 2 is a simplified block diagram of a computing system implementing a Masterbot according to one embodiment. [Diagram 3] FIG. 1 is a simplified block diagram of a computing system implementing a skillbot according to one embodiment. [Figure 4] FIG. 2 illustrates an example of a model architecture according to various embodiments. [Diagram 5] FIG. 2 illustrates another example of a model architecture according to various embodiments. [Figure 6] 6 illustrates an example in which the model architecture of FIG. 5 is modified in accordance with various embodiments. [Figure 7] FIG. 2 illustrates an example of a request to an API according to various embodiments. [Figure 8] 4 illustrates an example response from an API according to various embodiments. [Figure 9] FIG. 2 shows a table describing the OPUS source dataset. [Figure 10] 1 illustrates language detection test results according to various embodiments. [Figure 11] FIG. 1 illustrates a block diagram of an apparatus according to various embodiments. [Figure 12] FIG. 1 illustrates an example of a deep network with attention that may be included in an apparatus according to various embodiments. [Figure 13] 5A-5C illustrate examples of operations that may be performed by an embedding layer in accordance with various embodiments. [Figure 14] FIG. 2 illustrates a process flow for language detection according to various embodiments. [Figure 15] FIG. 2 illustrates a process flow for language detection according to various embodiments. [Figure 16] FIG. 1 illustrates a simplified diagram of a distributed system for implementing various embodiments. [Figure 17]A simplified block diagram of one or more components of a system environment in which services provided by one or more components of an embodiment system may be provided as cloud services, according to various embodiments. [Figure 18] FIG. 1 illustrates an example computer system that can be used to implement various embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] Detailed Description In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and descriptions are not restrictive. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments or designs.
[0019] As used herein, when an action is "based on" something, this means that the action is based at least in part on at least a portion of that something. The use of "based on" is meant to be open and inclusive, and a process, step, calculation, or other action that is "based on" one or more described conditions, items, or values may in fact be based on additional conditions, items, or values other than those described. The terms "substantially," "approximately," and "about" as used herein are defined as being largely, but not necessarily completely, specified (including completely what is specified), as understood by those of skill in the art. In any of the disclosed embodiments, the terms "substantially," "approximately," or "about" may be replaced with "within [a percentage]" of what is specified, where the percentage includes 0.1, 1, 5, and 10 percent.
[0020] introduction Previous research has formulated the language detection task as a text classification task. One such approach utilizes traditional feature-based machine learning (e.g., Naive Bayes with n-gram features) to handle the task. Another such approach utilizes deep learning (e.g., Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) networks) to handle the task.
[0021] The techniques described herein include utilizing an attention CNN with n-gram features (i.e., a CNN with attention mechanism) to handle the task of language detection. For example, we describe an approach that uses deep learning to build a language detection application programming interface (API) for 135 languages. Experiments on publicly available datasets show that such models perform as well as or better than language detection APIs from fastText, Google®, and Microsoft.
[0022] Natural language processing has many applications. For example, digital assistants are artificial intelligence-driven interfaces that use natural language conversation to help users perform various tasks. For each digital assistant, customers can assemble one or more skills. A skill (also referred to herein as a chatbot, bot, or skillbot) is an individual computer program that focuses on a specific type of task, such as tracking inventory, submitting a time card, ordering a pizza, getting banking information, creating an expense report, etc. To perform a task, a bot can have a conversation with an end user. A bot can respond to natural language messages (e.g., questions or comments) through a messaging application that typically uses natural language messages. An enterprise may use one or more bot systems to communicate with end users through messaging applications. The messaging application, which may be referred to as a channel, may be an end user's preferred messaging application that the end user already has installed and is familiar with. Thus, an end user does not need to download and install a new application to chat with a bot system. Messaging applications may include, for example, over-the-top (OTT) messaging channels (e.g., Facebook Messenger, Facebook WhatsApp, WeChat, Line, Kik, Telegram, Talk, Skype, Slack, or SMS), virtual private assistants (e.g., Amazon Dot, Echo, or Show, Google Home, Apple HomePod), mobile app or web app extensions that extend native or hybrid / responsive mobile apps or web applications with chat capabilities, or voice-based input (e.g., devices or apps with an interface that uses Siri, Cortana, Google Voice, or other voice input for interaction).
[0023] In some examples, the bot system may be associated with a Uniform Resource Identifier (URI). The URI may use a string to identify the bot system. The URI may be used as a webhook for one or more messaging application systems. The URI may include, for example, a Uniform Resource Locator (URL) or a Uniform Resource Name (URN). The bot system may be designed to receive a message (e.g., a HyperText Transfer Protocol (HTTP) post-call message) from the messaging application system. An HTTP post-call message may be sent from the messaging application system to the URI. In some embodiments, the message may differ from an HTTP post-call message. For example, the bot system may receive a message from a Short Message Service (SMS). Although the description herein may refer to a communication received by the bot system as a message, it should be understood that the message may be an HTTP post-call message, an SMS message, or any other type of communication between the two systems.
[0024] End users can interact with bot systems through conversational interactions (sometimes called conversational user interfaces (UIs)), just like human-to-human interactions. In some cases, the interaction may involve the end user saying "hello" to the bot, which responds with "hello" and asks the end user how it can help. In some cases, the interaction may be a transactional interaction with a banking bot, e.g., transferring money from one account to another, an informational interaction with a human resources bot, e.g., checking vacation balances, or an interaction with a retail bot, e.g., discussing the return of a purchased item or asking for technical support.
[0025] In some embodiments, the bot system can intelligently handle end user interactions without interaction with an administrator or developer of the bot system. For example, an end user may send one or more messages to the bot system to achieve a desired goal. The messages may include certain content, such as text, emojis, voice, images, videos, or other message delivery methods. In some embodiments, the bot system can convert the content into a standardized format (e.g., a Representational State Transfer (REST) call to an enterprise service with appropriate parameters) and generate a natural language response. The bot system can also prompt the end user for additional input parameters or request other additional information. In some embodiments, the bot system can also initiate communication with the end user rather than passively responding to an utterance of the end user. Various techniques are described herein for identifying explicit invocations of the bot system and determining inputs to the bot system that are invoked. In some embodiments, the explicit invocation analysis is performed by the master bot based on detection of an invocation name in the utterance. In response to detection of an invocation name, the utterance can be refined for input to a skill bot associated with the invocation name.
[0026] A conversation with a bot can follow a particular conversation flow that includes multiple states. The flow can define what happens next based on the input. In some embodiments, a bot system can be implemented using a state machine that includes user-defined states (e.g., the end user's intent) and actions to take in the states or from state to state. A conversation may take different paths based on the end user's input, which can affect the flow decisions made by the bot. For example, at each state, based on the end user's input or utterance, the bot can determine the end user's intent and determine the appropriate action to take next. As used herein and in the context of utterances, the term "intent" refers to the intent of the user who provided the utterance. For example, a user may intend to converse with a bot to order a pizza, and thus the user's intent may be expressed through the utterance "order a pizza". The user's intent can be directed to a particular task that the user wants the chatbot to perform on their behalf. Thus, an utterance can be an expression of a question, command, request, etc. that reflects the user's intent. An intent can include a goal that the end user wants to achieve.
[0027] In the context of configuring a chatbot, the term "intent" is used herein to refer to configuration information for mapping a user's utterance to a specific task / action or category of tasks / actions that the chatbot can perform. To distinguish between an utterance intent (i.e., a user's intent) and a chatbot's intent, the latter may be referred to herein as a "bot's intent." A bot's intent may include a set of one or more utterances associated with the intent. For example, an intent to order a pizza may include various permutations of utterances expressing a desire to order a pizza. These related utterances may be used to train the chatbot's intent classifier, which may then determine whether an input utterance from a user matches the intent of ordering a pizza. A bot's intent may be associated with one or more dialog flows for initiating a conversation with a user at a state. For example, the first message of an intent to order a pizza may be the question, "What kind of pizza do you like?" In addition to the related utterances, a bot's intent may further include named entities associated with the intent. For example, an intent to order a pizza may include variables or parameters that are used to perform the task of ordering a pizza, e.g., topping 1, topping 2, type of pizza, size of pizza, amount of pizza, etc. The values of the entities are typically obtained through conversation with the user.
[0028] FIG. 1 is a simplified block diagram of an environment 100 incorporating a chatbot system according to an embodiment. The environment 100 comprises a digital assistant builder platform (DABP) 102 that allows users of the DABP 102 to create and deploy digital assistants or chatbot systems. The DABP 102 can be used to create one or more digital assistants (or DAs) or chatbot systems. For example, as shown in FIG. 1, a user 104 representing a particular business can use the DABP 102 to create and deploy a digital assistant 106 for users of the particular business. For example, a bank can use the DABP 102 to create one or more digital assistants for use by customers of the bank. The same DABP 102 platform can be used by multiple businesses to create digital assistants. As another example, a restaurant (e.g., a pizza place) owner can use the DABP 102 to create and deploy a digital assistant that enables customers of the restaurant to order food (e.g., order pizza).
[0029] For purposes of this disclosure, a "digital assistant" is an entity that assists a user of the digital assistant in accomplishing various tasks through natural language conversation. A digital assistant may be implemented using only software (e.g., a digital assistant is a digital entity implemented using programs, codes, or instructions executable by one or more processors), using hardware, or using a combination of hardware and software. A digital assistant may be embodied or implemented in a variety of physical systems or devices, such as a computer, a mobile phone, a watch, an appliance, a vehicle, etc. A digital assistant may also be referred to as a chatbot system. Thus, for purposes of this disclosure, the terms digital assistant and chatbot system are interchangeable.
[0030] A digital assistant, such as a digital assistant 106 built using DABP 102, can be used to perform a variety of tasks via natural language-based conversations between the digital assistant and its user 108. As part of the conversation, the user can provide one or more user inputs 110 to the digital assistant 106 and obtain responses 112 from the digital assistant 106. A conversation can include one or more of the inputs 110 and responses 112. Through these conversations, the user can request one or more tasks to be performed by the digital assistant, and in response, the digital assistant is configured to perform the tasks requested by the user and respond to the user with an appropriate response.
[0031] User input 110 is generally in a natural language format and is referred to as an utterance. User utterance 110 can be in text format, such as when a user inputs a sentence, a question, a fragment of text, or a single word and provides it as input to the digital assistant 106. In some embodiments, user utterance 110 can be in voice input or audio format, such as when a user says or speaks something that is provided as input to the digital assistant 106. The utterance is typically in a language that the user 108 speaks. For example, the utterance may be in English or another language. If the utterance is in audio format, the voice input is converted into a text format utterance in that particular language, and the text utterance is processed by the digital assistant 106. Various speech-to-text processing techniques can be used to convert the voice or voice input into a text utterance, which is then processed by the digital assistant 106. In some embodiments, the speech-to-text conversion may be performed by the digital assistant 106 itself.
[0032] The utterance may be a text or audio utterance and may be a fragment, a sentence, multiple sentences, one or more words, one or more questions, a combination of the aforementioned types, etc. The digital assistant 106 is configured to apply natural language understanding (NLU) techniques to the utterance to understand the meaning of the user input. As part of the NLU processing on the utterance, the digital assistant 106 is configured to perform processing to understand the meaning of the utterance, including identifying one or more intents and one or more entities corresponding to the utterance. Upon understanding the meaning of the utterance, the digital assistant 106 may perform one or more actions or behaviors depending on the understood meaning or intent. For purposes of this disclosure, it is assumed that the utterance is a text utterance provided directly by a user 108 of the digital assistant 106 or is the result of converting an input audio utterance into text format. However, this is not intended to be limiting or restrictive in any way.
[0033] For example, a user 108 input may request a pizza order by providing an utterance such as "I want to order a pizza." Upon receiving such an utterance, the digital assistant 106 is configured to understand the meaning of the utterance and perform an appropriate action. The appropriate action may include, for example, responding to the user with a question requesting user input about the type of pizza the user wants to order, the size of the pizza, the pizza toppings, etc. The responses provided by the digital assistant 106 may also be in natural language form, typically in the same language as the input utterance. As part of generating these responses, the digital assistant 106 may perform natural language generation (NLG). In the case of a user ordering a pizza, through a conversation between the user and the digital assistant 106, the digital assistant may guide the user to provide all the information required to order the pizza, and then have the pizza ordered at the end of the conversation. The digital assistant 106 may end the conversation by outputting information to the user indicating that a pizza is ordered.
[0034] At a conceptual level, the digital assistant 106 performs various processing in response to utterances received from a user. In some embodiments, this processing includes a series or pipeline of processing steps, such as understanding the meaning of the input utterance (also called natural language understanding (NLU)), determining an action to be performed in response to the utterance, appropriately triggering the execution of the action therein, generating a response to be output to the user in response to the user's utterance, outputting the response to the user, etc. NLU processing can include analyzing the received input utterance to understand the structure and meaning of the utterance, refining and reshaping the utterance to develop a more understandable form (e.g., logical form) or structure of the utterance. Generating the response can include the use of NLG techniques.
[0035] NLU processing performed by a digital assistant such as digital assistant 106 may include various NLP-related processing such as sentence parsing (e.g., tokenization, lemmatization, identifying part-of-speech tags for a sentence, identifying named entities within a sentence, generating a dependency tree representing the sentence structure, splitting the sentence into clauses, analyzing the individual clauses, resolving anaphora, performing chunking, etc.). In some embodiments, the NLU processing or parts thereof are performed by the digital assistant 106 itself. In some other embodiments, the digital assistant 106 may use other resources to perform parts of the NLU processing. For example, the syntax and structure of an input spoken sentence may be identified by processing the sentence using a parser, a part-of-speech tagger, and / or a named entity recognizer. In one implementation, for English, a parser, a part-of-speech tagger, and a named entity recognizer such as those provided by the Stanford Natural Language Processing (NLP) Group are used to analyze the structure and syntax of the sentence. These are provided as part of the Stanford CoreNLP toolkit.
[0036] Although various examples provided in this disclosure show speech in English, this is meant as an example only. In some embodiments, the digital assistant 106 can also process speech in languages other than English. The digital assistant 106 can provide subsystems (e.g., components that implement NLU functionality) configured to perform processing for different languages. These subsystems can be implemented as pluggable devices that can be invoked using service calls from the NLU core server. This makes NLU processing flexible and extensible for each language, including allowing different processing orders. Language packs can be provided for individual languages, and the language packs can register a list of subsystems that can be provided by the NLU core server.
[0037] A digital assistant, such as the digital assistant 106 shown in Figure 1, may be available or accessible to its user 108 through a variety of different channels, including, but not limited to, through an application, through social media platforms, through various messaging services and applications, and through other applications and channels. A single digital assistant may be configured with multiple channels to run on different services and be accessed simultaneously.
[0038] A digital assistant or chatbot system typically includes or is associated with one or more skills. In some embodiments, these skills are individual chatbots (called skillbots) that are configured to interact with a user and perform specific types of tasks, such as tracking inventory, submitting a timecard, creating an expense report, ordering food, checking a bank account, making a reservation, purchasing a widget, etc. For example, for the embodiment shown in FIG. 1, the digital assistant or chatbot system 106 includes skills 116-1, 116-2, etc. For purposes of this disclosure, the terms "skill" and "skills" are used synonymously with the terms "skillbot" and "skillbots," respectively.
[0039] Each skill associated with a digital assistant helps a user of the digital assistant complete a task through a conversation with the user, which may include a combination of text or voice input provided by the user and responses provided by the skill bot. These responses may take the form of text or voice messages to the user and / or use simple user interface elements (e.g., selection lists) presented to the user for the user to select from.
[0040] There are various ways in which skills or skillbots can be associated with or added to a digital assistant. In some cases, skillbots can be developed by companies and then added to a digital assistant using DABP 102. In other examples, skillbots can be developed and created using DABP 102 and then added to a digital assistant created using DABP 102. In yet other examples, DABP 102 provides an online digital store (referred to as a "skill store") that offers multiple skills directed to a wide range of tasks. Skills offered through the skill store can also expose various cloud services. To add a skill to a digital assistant being created using DABP 102, a user of DABP 102 can access the skill store via DABP 102, select the desired skill, and indicate that the selected skill is to be added to the digital assistant being created using DABP 102. Skills from the skill store can be added to a digital assistant as is or in a modified form (e.g., a user of DABP102 can select and clone a particular skill bot provided by the skill store, customize or modify the selected skill bot, and then add the modified skill bot to a digital assistant created using DABP102).
[0041] A variety of different architectures can be used to implement a digital assistant or chatbot system. For example, in one embodiment, a digital assistant created and deployed using DABP 102 may be implemented using a masterbot / child (or sub)bot paradigm or architecture. According to this paradigm, a digital assistant is implemented as a masterbot that interacts with one or more child bots, which are skillbots. For example, in the embodiment shown in FIG. 1, the digital assistant 106 comprises a masterbot 114 and skillbots 116-1, 116-2, etc., that are child bots of the masterbot 114. In one embodiment, the digital assistant 106 itself is considered to function as a masterbot.
[0042] A digital assistant implemented according to the master-child bot architecture allows a user of the digital assistant to interact with multiple skills through a unified user interface, i.e., a master bot. When a user operates the digital assistant, the user input is received by the master bot. The master bot then performs processing to determine the meaning of the utterance of the user input. The master bot then determines whether the task requested by the user in the utterance can be handled by the master bot itself. If not, the master bot selects an appropriate skill bot to handle the user's request and routes the conversation to the selected skill bot. This allows a user to converse with the digital assistant through a common single interface and also provides the ability to use multiple skill bots configured to perform specific tasks. For example, in the case of a digital assistant developed for an enterprise, the master bot of the digital assistant can be coordinated with skill bots with specific functions, such as a CRM bot to perform functions related to customer relationship management (CRM), an ERP bot to perform functions related to enterprise resource planning (ERP), an HCM bot to perform functions related to human capital management (HCM), etc. In this way, the end user or consumer of the digital assistant only needs to know how to access the digital assistant through a common master bot interface, and multiple skill bots are provided behind the scenes to handle the user's requests.
[0043] In one embodiment, in a masterbot / childbot infrastructure, the masterbot is configured to know a list of available skillbots. The masterbot has access to metadata identifying the various available skillbots and, for each skillbot, the capabilities of the skillbot, including tasks that can be performed by the skillbot. Upon receiving a user request in the form of an utterance, the masterbot is configured to identify or predict, among multiple available skillbots, a specific skillbot that can best serve or process the user request. The masterbot then routes the utterance (or a portion of the utterance) to that specific skillbot for further processing. Thus, control flows from the masterbot to the skillbot. The masterbot can support multiple input and output channels. In one embodiment, the routing can be performed utilizing processing performed by one or more available skillbots. For example, as described below, a skillbot can be trained to infer the intent of the utterance and determine whether the inferred intent matches an intent configured for the skillbot. Thus, the routing performed by the masterbot can include communicating to the masterbot an indication of whether the skillbot is configured with an intent suitable for processing the utterance.
[0044] 1 illustrates a digital assistant 106 with a masterbot 114 and skillbots 116-1, 116-2, and 116-3, but is not limited thereto. The digital assistant may include various other components (e.g., other systems and subsystems) that provide the functionality of the digital assistant. These systems and subsystems may be implemented solely in software (e.g., code, instructions stored on a computer-readable medium and executable by one or more processors), solely in hardware, or in an implementation using a combination of software and hardware.
[0045] DABP 102 provides infrastructure and various services and features that enable a user of DABP 102 to create a digital assistant including one or more skill bots that are associated with the digital assistant. In some cases, a skill bot can be created by duplicating an existing skill bot, for example, by duplicating a skill bot provided by a skill store. As previously indicated, DABP 102 provides a skill store or skill catalog that provides multiple skill bots for performing various tasks. A user of DABP 102 can clone a skill bot from the skill store. If necessary, modifications and customizations can be made to the cloned skill bot. In other examples, a user of DABP 102 creates a skill bot from scratch using tools and services provided by DABP 102. As previously indicated, a skill store or skill catalog provided by DABP 102 can provide multiple skill bots for performing various tasks.
[0046] In one embodiment, at a high level, creating or customizing a skillbot includes the following steps:
[0047] (1) Configure the settings for a new skill bot (2) Configure one or more intents for the skill bot (3) the composition of one or more entities for one or more intents (4) Skill Bot Training (5) Creating a dialogue flow for the skill bot (6) Add custom components to your skill bot as needed (7) Testing and Deploying Skill Bots Each of the above steps is briefly described below.
[0048] (1) Configuring Settings for a New Skill Bot - Various settings can be configured for a skill bot. For example, a skill bot designer can specify one or more invocation names for the skill bot they create. These invocation names can then be used by users of the digital assistant to explicitly invoke the skill bot. For example, a user can enter an invocation name in a user utterance to explicitly invoke the corresponding skill bot.
[0049] (2) Configuring one or more intents and associated example utterances for a skill bot - A skill bot designer specifies one or more intents (also called bot intents) for the skill bot to be created. The skill bot is then trained based on these specified intents. These intents represent categories or classes for which the skill bot is trained to infer input utterances. Upon receiving an utterance, the skill bot being trained infers the intent of the utterance. The inferred intent is selected from a set of predefined intents used to train the skill bot. The skill bot then performs an appropriate action in response to the utterance based on the inferred intent for the utterance. In some cases, the intents of the skill bot represent tasks that the skill bot can perform for a user of the digital assistant. Each intent is given an intent identifier or intent name. For example, for a skill bot being trained for banking, the intents specified for the skill bot may include "CheckBalance", "TransferMoney", "DepositCheck", etc.
[0050] For each intent defined for a skillbot, the skillbot designer can also provide one or more example utterances that represent and explain that intent. These example utterances are intended to represent utterances that a user can input to the skillbot for that intent. For example, for the CheckBalance intent, example utterances might include "What's the balance in my savings account?", "How much is in my checking account?", "How much money is in my account?", etc. Thus, various permutations of typical user utterances can be specified as example utterances for an intent.
[0051] The intents and associated example utterances are used as training data to train the skill bot. A variety of different training techniques may be used. This training results in a predictive model configured to receive an utterance as input and output an intent that is inferred for the utterance by the predictive model. In some cases, the input utterance is provided to an intent analysis engine, which is configured to predict or infer an intent for the input utterance using the trained model. The skill bot can perform one or more actions based on the inferred intent.
[0052] (3) Configuring an entity for one or more intents of a skill bot - In some cases, additional context may be required to enable the skill bot to respond appropriately to a user utterance. For example, there may be situations where user input utterances resolve to the same intent in a skill bot. For example, in the above example, the utterances "What is the balance in my savings account?" and "How much do I have in my checking account?" both resolve to the same CheckBalance intent, but these utterances are different requests that ask for different things. To disambiguate such requests, one or more entities are added to the intent. Using the banking skill bot example, an entity called AccountType that defines values "checking" and "saving" allows the skill bot to parse the user request and respond appropriately. In the above example, the utterances resolve to the same intent, but the values associated with the AccountType entity are different for the two utterances. This may enable the skill bot to take different actions for the two utterances even though they are resolving to the same intent. When configured for a skill bot, one or more entities can be specified for an intent. Thus, entities are used to add context to the intent itself. Entities help to more completely describe the intent, enabling the skill bot to complete the user request.
[0053] In one embodiment, there are two types of entities: (a) built-in entities provided by DABP 102, and (2) custom entities that can be specified by the skill bot designer. Built-in entities are generic entities that can be used by various bots. Examples of built-in entities include, but are not limited to, entities related to time, date, address, number, email address, duration, periodic period, currency, phone number, URL, etc. Custom entities are used for more customized applications. For example, for a banking skill, an AccountType entity may be defined by the skill bot designer that allows various banking transactions by checking user input against keywords such as check, savings, credit card, etc.
[0054] (4) Training the Skillbot - The skillbot is configured to receive user input in the form of utterances, parse or process the received input, and identify or select an intent associated with the received user input. As indicated above, the skillbot needs to be trained for this. In one embodiment, the skillbot is trained based on intents configured for the skillbot and example utterances associated with those intents (collectively, training data), so that the skillbot can resolve user input utterances to one of its configured intents. In one embodiment, the skillbot uses a predictive model that is trained using the training data to enable the skillbot to identify what the user is saying (or, in some cases, trying to say). DABP 102 provides various training techniques that the skillbot designer can use to train the skillbot, such as various machine learning based training techniques, rule-based training techniques, and / or combinations thereof. In one embodiment, a portion of the training data (e.g., 80%) is used to train the skillbot model, and another portion (e.g., the remaining 20%) is used to test or validate the model. Once training is complete, the trained model (also referred to as the trained skillbot) can be used to process and respond to user utterances. In some cases, a user utterance may be a question that requires only one answer and no further conversation. To handle such situations, you can define a Q&A (Question and Answer) intent for your skill bot. This allows your skill bot to output a response to a user request without updating the dialog definition. A Q&A intent is created in a similar way to a regular intent. The dialog flow for a Q&A intent may differ from that of a regular intent.
[0055] (5) Creating a Dialog Flow for a Skill Bot -- The dialog flow specified for a skill bot describes how the skill bot reacts as different intents of the skill bot are resolved depending on the user input received. The dialog flow defines the behavior or actions that the skill bot performs, such as how the skill bot responds to user utterances, how the skill bot prompts the user for input, and how the skill bot returns data. The dialog flow is like a flowchart that the skill bot follows. Skill bot designers specify the dialog flow using a language such as Markdown language. In one embodiment, a version of YAML called OBotML can be used to specify the dialog flow for a skill bot. The dialog flow definition for a skill bot serves as a model of the conversation itself, allowing the skill bot designer to choreograph the interaction between the skill bot and the user that the skill bot serves.
[0056] In one embodiment, a skill bot's dialog flow definition contains three sections:
[0057] (a) Context Section (b) Default transition section (c) Status Section Context Section - Skill bot designers can define variables that will be used in the conversation flow in the context section. Other variables that can be named in the context section include but are not limited to variables for error handling, variables for built-in or custom entities, user variables that allow the skill bot to know and preserve user preferences, etc.
[0058] Default Transitions Section - Transitions for a skill bot can be defined in the dialog flow states section or in the default transitions section. Transitions defined in the default transitions section act as fallbacks and are triggered when there is no corresponding transition defined within a state or when the conditions required to trigger a state transition are not met. The default transitions section allows you to define routing that enables your skill bot to gracefully handle unexpected user actions.
[0059] State Section - A dialog flow and its associated behavior are defined as a set of temporary states that govern the logic within the dialog flow. Each state node in a dialog flow definition names a component that provides the functionality required at that point in the dialog. Thus, states are built around components. States contain component-specific properties and define transitions to other states that are triggered after the component is executed.
[0060] Special case scenarios can be handled using the states section. For example, you may want to offer a user the option to temporarily leave a first skill they are using to do something in a second skill within the digital assistant. For example, if a user is engaged in a conversation with a shopping skill (e.g., the user is making some choices about a purchase), the user may want to jump to a banking skill (e.g., the user may want to ensure they have enough money for the purchase), and then return to the shopping skill to complete the user's order. To address this, you can configure an action in the first skill to initiate an interaction with a second, different skill within the same digital assistant, and then return to the original flow.
[0061] (6) Adding Custom Components to a Skillbot - As described above, a state specified in a skillbot's dialog flow names a component that corresponds to that state and provides the required functionality. A component enables a skillbot to perform a function. In an embodiment, DABP 102 provides a set of pre-configured components to perform a wide range of functions. A skillbot designer can select one or more of these pre-configured components and associate them with a state in the skillbot's dialog flow. A skillbot designer can also create custom or new components using tools provided by DABP 102 and associate the custom components with one or more states in the skillbot's dialog flow.
[0062] (7) Testing and Deploying Skillbots - DABP102 provides several features that enable skillbot designers to test the skillbots they are developing, after which they can be deployed and included in a digital assistant.
[0063] While the above discussion describes how to create a skill bot, similar techniques can also be used to create a digital assistant (or master bot). At the master bot or digital assistant level, built-in system intents can be configured for the digital assistant. These built-in system intents are used to identify common tasks that the digital assistant itself (i.e., the master bot) can handle without invoking a skill bot associated with the digital assistant. Examples of system intents defined for a master bot are: (1) Exit: applies when a user signals that they want to end the digital assistant's current conversation or context. (2) Help: applies when a user asks for help or direction. (3) UnresolvedIntent: applies to user input that does not match well with the exit and help intents. The digital assistant also stores information about one or more skill bots associated with the digital assistant. This information allows the master bot to select a specific skill bot to process an utterance.
[0064] At the MasterBot or Digital Assistant level, when a user inputs a phrase or utterance into the digital assistant, the digital assistant is configured to perform processing to determine how to route the utterance and associated conversation. The digital assistant determines this using a routing model, which can be rule-based, AI-based, or a combination thereof. The digital assistant uses the routing model to determine whether the conversation corresponding to the user input utterance should be routed to a specific skill for processing, should be handled by the digital assistant or MasterBot itself according to a built-in system intent, or should be treated as a separate state in the current conversation flow.
[0065] In one embodiment, as part of this processing, the digital assistant determines whether the user input utterance explicitly identifies a skill bot using its invocation name. If an invocation name is present in the user input, it is treated as an explicit invocation of the skill bot corresponding to the invocation name. In such a scenario, the digital assistant may route the user input to the explicitly invoked skill bot for further processing. In the absence of a specific or explicit invocation, in one embodiment, the digital assistant evaluates the received user input utterance and calculates a confidence score for the system intent and the skill bot associated with the digital assistant. The calculated score for the skill bot or system intent represents the likelihood that the user input represents the task that the skill bot is configured to perform, or represents the system intent. The system intent or skill bot whose associated calculated confidence score exceeds a threshold (e.g., a confidence threshold routing parameter) is selected as a candidate for further evaluation. The digital assistant then selects a particular system intent or skill bot from the identified candidates for further processing the user input utterance. In one embodiment, after one or more skill bots are identified as candidates, the intents associated with those candidate skills are evaluated (according to the intent model of each skill), and a confidence score is determined for each intent. Generally, intents with confidence scores above a threshold (e.g., 70%) are treated as candidate intents. If a particular skill bot is selected, the user's utterance is routed to that skill bot for further processing. If a system intent is selected, one or more actions are performed by the master bot itself according to the selected system intent.
[0066] FIG. 2 is a simplified block diagram of a Masterbot (MB) system 201 according to an embodiment. The MB system 201 can be implemented in software only, hardware only, or a combination of hardware and software. The MB system 201 includes a pre-processing subsystem 210, a multi-intent subsystem (MIS) 220, an explicit invocation subsystem (EIS) 230, a skillbot invoker 240, and a data store 250. The MB system 201 shown in FIG. 2 is only one example of an arrangement of components in a Masterbot. Those skilled in the art will recognize many possible variations, alternatives, and modifications. For example, in some implementations, the MB system 201 may have more or fewer systems or components than those shown in FIG. 2, may combine two or more subsystems, or may have a different configuration or arrangement of the subsystems.
[0067] The pre-processing subsystem 210 receives an utterance "A" 202 from a user and processes the utterance via a language detector 212 and a language parser 214. As indicated above, the utterance can be provided in a variety of ways, such as audio or text. The utterance 202 can be a sentence fragment, a complete sentence, multiple sentences, etc. The utterance 202 can include punctuation. For example, if the utterance 202 is provided as audio, the pre-processing subsystem 210 can convert the speech to text using a speech-to-text converter (not shown) that inserts punctuation marks (e.g., commas, semicolons, periods, etc.) into the resulting text.
[0068] The language detector 212 detects the language of the utterance 202 based on the text of the utterance 202. The way in which the utterance 202 is processed is language dependent, since each language has its own grammar and semantics. Differences between languages are taken into account when analyzing the syntax and structure of the utterance.
[0069] The language parser 214 parses the utterance 202 to extract part-of-speech (POS) tags for individual linguistic units (e.g., words) in the utterance 202. POS tags include, for example, noun (NN), pronoun (PN), verb (VB), etc. The language parser 214 may also tokenize the linguistic units of the utterance 202 (e.g., to convert each word into a separate token) and lemmatize the words. Lemmats are the main form of a set of words represented in a dictionary (e.g., "run" is a lemma for run, runs, ran, running, etc.). Other types of preprocessing that the language parser 214 may perform include chunking of compound expressions, for example, combining "credit" and "card" into a single expression "credit_card". The language parser 214 may also identify relationships between words in the utterance 202. For example, in some embodiments, the language parser 214 generates a dependency tree that indicates which parts of the utterance (e.g., particular nouns) are direct objects, which parts of the utterance are prepositions, etc. The results of the processing performed by the language parser 214 form the extracted information 205 and are provided as input to the MIS 220 along with the utterance 202 itself.
[0070] As indicated above, an utterance 202 may contain multiple sentences. For purposes of detecting multiple intents and explicit invocations, an utterance 202 may be treated as a single unit even if it contains multiple sentences. However, in an embodiment, preprocessing may be performed, for example by preprocessing subsystem 210, to identify a single sentence among multiple sentences for multiple intent and explicit invocation analysis. In general, the results generated by MIS 220 and EIS 230 are substantially the same regardless of whether utterance 202 is processed at the level of individual sentences or as a single unit containing multiple sentences.
[0071] The MIS 220 determines whether the utterance 202 expresses multiple intents. Although the MIS 220 can detect the presence of multiple intents in the utterance 202, the process performed by the MIS 220 does not include determining whether the intent of the utterance 202 matches any intent configured for the bot. Instead, the process for determining whether the intent of the utterance 202 matches the intent of the bot may be performed by the intent classifier 242 of the MB system 201 or by the intent classifier of the skill bot (e.g., as shown in the embodiment of FIG. 3). The process performed by the MIS 220 assumes that there is a bot (e.g., a specific skill bot or the master bot itself) that can process the utterance 202. Thus, the process performed by the MIS 220 does not require knowledge of the bots in the chatbot system (e.g., the IDs of the skill bots registered with the master bot) or knowledge of the intents configured for a specific bot.
[0072] To determine that an utterance 202 includes multiple intents, the MIS 220 applies one or more rules from a rule set 252 in the data store 250. The rules applied to the utterance 202 depend on the language of the utterance 202 and may include a sentence pattern that indicates the presence of multiple intents. For example, the sentence pattern may include a coordinating conjunction that joins two parts of a sentence (e.g., a conjunction), both parts corresponding to separate intents. If the utterance 202 matches the sentence pattern, it can be inferred that the utterance 202 represents multiple intents. Note that an utterance that includes multiple intents does not necessarily have different intents (e.g., intents directed to different bots or intents directed to different intents within the same bot). Instead, the utterance may include separate instances of the same intent. For example, "order pizza using payment account X, then order pizza using payment account Y."
[0073] As part of determining that the utterance 202 represents multiple intents, the MIS 220 also determines which portions of the utterance 202 are associated with each intent. For each intent expressed in the utterance that includes multiple intents, the MIS 220 constructs a new utterance for separate processing in place of the original utterance, e.g., utterance “B” 206 and utterance “C” 208, as shown in FIG. 2. Thus, the original utterance 202 may be split into two or more separate utterances that are processed one at a time. The MIS 220 determines which of the two or more utterances should be processed first using the extracted information 205 and / or from an analysis of the utterance 202 itself. For example, the MIS 220 may determine that the utterance 202 includes an indicator word that indicates that a particular intent should be processed first. The newly formed utterance corresponding to this particular intent (e.g., one of the utterances 206 or utterance 208) will be sent first for further processing by the EIS 230. After the conversation caused by the first utterance has ended (or has been temporarily interrupted), the next highest priority utterance (e.g., the other of utterance 206 or utterance 208) can be sent to EIS 230 for processing.
[0074] The EIS 230 determines whether the received utterance (e.g., utterance 206 or utterance 208) includes an invocation name of the skillbot. In an embodiment, each skillbot in the chatbot system is assigned a unique invocation name that distinguishes the skillbot from other skillbots in the chatbot system. A list of invocation names can be maintained as part of the skillbot information 254 in the data store 250. If the utterance contains words that match the invocation name, the utterance is considered to be an explicit invocation. If the bot is not explicitly invoked, the utterance received by the EIS 230 is considered an implicit invoking utterance 234 and is input to the masterbot's intent classifier (e.g., intent classifier 242) to determine which bot to use to process the utterance. In some cases, the intent classifier 242 will determine that the masterbot should process an implicit invoking utterance. In other examples, the intent classifier 242 will determine which skillbot to route the utterance to for processing.
[0075] The explicit call functionality provided by the EIS 230 has several advantages. It can reduce the amount of processing that the masterbot needs to perform. For example, when there is an explicit call, the masterbot may not need to perform an intent classification analysis (e.g., using the intent classifier 242) or may need to perform a reduced intent classification analysis to select a skillbot. Thus, the explicit call analysis may enable the selection of a particular skillbot without relying on an intent classification analysis.
[0076] There may also be situations where functionality overlaps between multiple skillbots. This can occur, for example, when the intents handled by two skillbots overlap or are very close to each other. In such situations, it may be difficult for the masterbot to identify which of multiple skillbots to select based on intent classification analysis alone. In such scenarios, an explicit invocation would make clear the specific skillbot to be used.
[0077] In addition to determining that the utterance is an explicit call, EIS 230 is responsible for determining whether any portion of the utterance should be used as input to the explicitly invoked skillbot. In particular, EIS 230 may determine whether any portion of the utterance is not relevant to the call. EIS 230 may perform this determination through analysis of the utterance and / or analysis of extracted information 205. Instead of sending the entire utterance received by EIS 230, EIS 230 may send the portion of the utterance that is not associated with the call to the invoked skillbot. In some cases, the input to the invoked skillbot is formed by simply removing the portion of the utterance that is associated with the call. For example, "I would like to order a pizza using PizzaBot" can be shortened to "I would like to order a pizza" because "using PizzaBot" is relevant to the call of PizzaBot but not the processing performed by PizzaBot. In some cases, EIS 230 may reformat the portion sent to the invoked bot to form, for example, a complete sentence. Thus, EIS 230 not only determines that there is an explicit call, but also determines what to send to the skillbot if there is an explicit call. In some cases, there may be no text to input to the bot being invoked. For example, if the utterance was "pizzabot," the EIS 230 may determine that the pizzabot is being invoked, but there is no text to be processed by the pizzabot. In such a scenario, the EIS 230 may indicate to the skillbot invoker 240 that there is nothing to send.
[0078] The skillbot invoker 240 invokes a skillbot in a variety of ways. For example, the skillbot invoker 240 can invoke the bot in response to receiving an indication 235 that a particular skillbot is selected as a result of an explicit invocation. The indication 235 can be sent by the EIS 230 along with the input of the explicitly invoked skillbot. In this scenario, the skillbot invoker 240 hands over control of the conversation to the explicitly invoked skillbot. The explicitly invoked skillbot determines an appropriate response to the input by treating the input from the EIS 230 as a standalone utterance. For example, the response can be to perform a particular action or to start a new conversation in a particular state, where the initial state of the new conversation depends on the input sent from the EIS 230.
[0079] Another way that the skillbot invoker 240 can invoke a skillbot is through an implicit invocation using the intent classifier 242. The intent classifier 242 can be trained using machine learning and / or rule-based training techniques to determine the likelihood that an utterance represents a task that a particular skillbot is configured to perform. The intent classifier 242 is trained with different classes, one class for each skillbot. For example, each time a new skillbot is registered with the masterbot, the intent classifier 242 can be trained using a list of example utterances associated with the new skillbot to determine the likelihood that a particular utterance represents a task that the new skillbot can perform. The parameters (e.g., a set of values for the parameters of the machine learning model) generated as a result of this training can be stored as part of the skillbot information 254.
[0080] In one embodiment, the intent classifier 242 is implemented using a machine learning model, as described in more detail herein. Training the machine learning model may include inputting at least a subset of utterances from example utterances associated with various skill bots, and generating, as an output of the machine learning model, an inference about which bot is the correct bot to process a particular training utterance. For each training utterance, an indication of the correct bot to use for the training utterance may be provided as ground truth information. The operation of the machine learning model may then be adapted (e.g., through backpropagation) to minimize the discrepancy between the generated inference and the ground truth information.
[0081] In one embodiment, the intent classifier 242 determines a confidence score for each skill bot registered with the master bot, indicating the likelihood that the skill bot can process the utterance (e.g., the implicit invocation utterance 234 received from the EIS 230). The intent classifier 242 can also determine a confidence score for each configured system-level intent (e.g., help, quit). If a particular confidence score meets one or more conditions, the skill bot invoker 240 invokes the bot associated with the particular confidence score. For example, a confidence score threshold may need to be met. Thus, the output 245 of the intent classifier 242 is either an identification of the system intent or an identification of a particular skill bot. In some embodiments, in addition to meeting the confidence score threshold, the confidence score must exceed the next highest confidence score by a certain win rate. Imposing such a condition allows routing to a particular skill bot when the confidence scores of multiple skill bots each exceed the confidence score threshold.
[0082] After identifying the bot based on the evaluation of the confidence score, the skillbot invoker 240 hands over the processing to the identified bot. In case of system intent, the identified bot becomes the master bot. Otherwise, the identified bot is the skillbot. Furthermore, the skillbot invoker 240 decides what to provide as input 247 to the identified bot. As indicated above, in case of an explicit invocation, the input 247 may be based on a part of the utterance that is not associated with the invocation, or there may be no input 247 (e.g., an empty string). In case of an implicit invocation, the input 247 may be the entire utterance.
[0083] The data store 250 comprises one or more computing devices that store data used by various subsystems of the masterbot system 201. As described above, the data store 250 includes rules 252 and skillbot information 254. The rules 252 include, for example, rules for determining by the MIS 220 when an utterance expresses multiple intents and how to split an utterance expressing multiple intents. The rules 252 further include rules for determining by the EIS 230 which part of an utterance that explicitly invokes a skillbot is sent to the skillbot. The skillbot information 254 includes the invocation names of the skillbots in the chatbot system, for example, a list of the invocation names of all skillbots registered to a particular masterbot. The skillbot information 254 can also include information used by the intent classifier 242 to determine a confidence score for each skillbot in the chatbot system, for example, parameters of a machine learning model.
[0084] 3 is a simplified block diagram of a Skillbot system 300 according to an embodiment. The Skillbot system 300 is a computing system that can be implemented in software only, hardware only, or a combination of hardware and software. In an embodiment, such as the embodiment shown in FIG. 1, the Skillbot system 300 can be used to implement one or more Skillbots within a digital assistant.
[0085] The Skillbot system 300 includes an MIS 310, an intent classifier 320, and a conversation manager 330. The MIS 310 is similar to the MIS 220 of FIG. 2 and provides similar functionality, including being operable to determine using rules 352 in a data store 350: (1) whether an utterance represents multiple intents, and if so, (2) how to split the utterance into separate utterances for each of the multiple intents. In one embodiment, the rules applied by the MIS 310 to detect multiple intents and split the utterance are the same as the rules applied by the MIS 220. The MIS 310 receives the utterance 302 and the information to be extracted 304. The information to be extracted 304 is similar to the information to be extracted 205 of FIG. 1 and can be generated using the language parser 214 or a language parser local to the Skillbot system 300.
[0086] The intent classifier 320 can be trained in a manner similar to the intent classifier 242 described above in connection with the embodiment of FIG. 2 and in further detail herein. For example, in one embodiment, the intent classifier 320 is implemented using a machine learning model. The machine learning model of the intent classifier 320 is trained for a particular skill bot using at least a subset of the example utterances associated with the particular skill bot as training utterances. The ground truth for each training utterance becomes the intent of the particular bot associated with the training utterance.
[0087] The utterance 302 can be received directly from a user or provided via a masterbot. For example, if the utterance 302 is provided via a masterbot as a result of processing via the MIS 220 and the EIS 230 in the embodiment shown in FIG. 2, the MIS 310 can be bypassed to avoid repeating the processing already performed by the MIS 220. However, if the utterance 302 is received directly from a user, for example during a conversation that occurs after routing to a skillbot, the MIS 310 can process the utterance 302 to determine whether the utterance 302 represents multiple intents. If so, the MIS 310 applies one or more rules to split the utterance 302 into separate utterances for each intent, for example, utterance "D" 306 and utterance "E" 308. If the utterance 302 does not represent multiple intents, the MIS 310 forwards the utterance 302 to the intent classifier 320 for intent classification without splitting the utterance 302.
[0088] The intent classifier 320 is configured to match the received utterance (e.g., utterance 306 or 308) with an intent associated with the skillbot system 300. As explained above, a skillbot can be configured with one or more intents, with each intent including at least one example utterance associated with the intent and used to train the classifier. In the embodiment of FIG. 2, the intent classifier 242 of the masterbot system 201 is trained to determine a confidence score for each individual skillbot and a confidence score for the system intent. Similarly, the intent classifier 320 can be trained to determine a confidence score for each intent associated with the skillbot system 300. The classification performed by the intent classifier 242 is at the bot level, whereas the classification performed by the intent classifier 320 is at the intent level and is therefore more granular. The intent classifier 320 has access to intent information 354. For each intent associated with the skillbot system 300, the intent information 354 represents and illustrates the meaning of the intent and typically includes a list of utterances associated with tasks that can be performed by the intent. The intent information 354 may further include parameters generated as a result of training on this utterance list.
[0089] The conversation manager 330 receives as an output of the intent classifier 320 an indication 322 of a particular intent identified by the intent classifier 320 as the best match to the utterance input to the intent classifier 320. In some cases, the intent classifier 320 may not be able to determine a match. For example, if the utterance is directed to a system intent or to an intent of a different skill bot, the confidence score calculated by the intent classifier 320 may fall below a confidence score threshold. When this occurs, the skill bot system 300 may refer the utterance to a master bot for processing, e.g., routing to another skill bot. However, if the intent classifier 320 is successful in identifying the intent within the skill bot, the conversation manager 330 will initiate a conversation with the user.
[0090] A conversation initiated by the conversation manager 330 is specific to the intent identified by the intent classifier 320. For example, the conversation manager 330 may be implemented using a state machine configured to execute a dialog flow for the identified intent. The state machine may include a default starting state (e.g., when the intent is invoked without additional input) and one or more additional states, each state having associated therewith an action to be performed by the skill bot (e.g., performing a purchase transaction) and / or a dialog to be presented to the user (e.g., questions, responses). Thus, the conversation manager 330 may determine an action / dialog 335 upon receiving an instruction 322 identifying the intent, and may determine additional actions or dialogs depending on subsequent utterances received during the conversation.
[0091] The data store 350 comprises one or more computing devices that store data used by various subsystems of the Skillbot system 300. As shown in Figure 3, the data store 350 includes rules 352 and intent information 354. In some embodiments, the data store 350 can be integrated into a masterbot or digital assistant data store, such as data store 250 of Figure 2.
[0092] Model Architecture for Language Detection 4 illustrates an example model architecture 400 according to various embodiments. In this example, input text units are split into n-grams at the word level, and each word in the text unit is also split into character-based n-grams to generate a sequence of n-grams for the text unit (e.g., according to the order in which the n-grams appear in the input text unit). The operations of splitting the text units into n-grams at the word level and splitting the text units into n-grams at the character level can be performed serially (e.g., word-level split followed by character-level split) or in parallel.
[0093] The value of n for the word-level n-grams may be the same as or different from the value of n for the character-level n-grams. In the example 400 of FIG. 4, the value of n for the word-level n-grams is 1 (word-based unigrams) and the value of n for the character-level n-grams is 2 (character-based bigrams). Using this scheme, the text unit "hello there" is converted (by the parser 410) into the sequence of n-grams [hello, _h, he, el, ll, lo, o_, there, _t, th, he, er, re, e_] (or the sequence of n-grams [_h, he, el, ll, lo, o_, hello, _t, th, he, er, re, e_, there]). Note that in this example, a special character (e.g., the underscore character "_") is added to the beginning and end of each word in the text unit to indicate word boundaries before the word is split into character-level n-grams.
[0094] Each n-gram in a sequence of n-grams is fed into an embedding layer 420 to generate a representation (e.g., a feature vector or "embedding vector") corresponding to the n-gram. The embedding layer 420 contains an embedding model (e.g., an embedding matrix that associates each n-gram with its corresponding embedding vector) that is trained. A CNN with attention mechanism 430 is used to capture the relationships between n-gram features (which may be indicated by aspects such as the relative order and / or relative weights of the n-grams in the sequence) to generate an encoding vector for the text unit. The encoding vector is classified using a feed-forward network (FFN) 440 with a softmax activation function to generate an output prediction (e.g., identification of the predicted dominant language of the input text unit).
[0095] In the example 400 described above with reference to FIG. 4, the vocabulary size is 13, with two word-level unigrams ("hello", "there") and eleven character-level bigrams ("_h", "he", "el", "ll", "lo", "o_", "_t", "th", "er", "re", "e_"). In practice, training datasets for language detection tasks are large, and more than 100 available languages may be represented, so the vocabulary is typically huge (e.g., more than 30 million for an internal dataset of 105 million sentences). For example, the vocabulary may include a complete n-gram set for Japanese or Chinese, another n-gram set for Vietnamese, another n-gram set for English, etc. Even if the dimensionality of the embedding vector space is relatively small (e.g., tens or hundreds of dimensions), the number of parameters in the embedding model may become very large, making the search very slow.
[0096] Hash embedding can be used to reduce the size of the vocabulary and thereby reduce the number of parameters in the embedding model. FIG. 5 shows an embodiment 500 of the model architecture of FIG. 4, where the embedding layer 520 employs hash embedding by converting each n-gram from the parser 410 (in a hash operation 515) into a corresponding hash identifier (ID), which is then input into the embedding model 525 to obtain the corresponding n-gram representation (e.g., embedding vector). Since the hash IDs in this example range from 0 to 9, the size of the vocabulary is reduced from 13 in FIG. 4 to a fixed number of 10 in this case. As in the example 400 shown in FIG. 4, a CNN with attention model 530 can be used to capture the relationships between n-gram features to generate text-wise encoding vectors, which can be classified using a FFN 540 and a softmax activation function to generate output (language) predictions.
[0097] In one example, the size of the output range of the hash function is equal to the desired vocabulary size, and each hash ID is obtained directly from the corresponding n-gram by applying the hash function to the n-gram. That is, the hash ID of an n-gram is a hash value generated by applying the hash function to the n-gram. In another example, each of the hash IDs is obtained by applying the hash function to the n-gram and then applying a modulo B function to the resulting hash value, where B is the desired size of the vocabulary. For example, in FIG. 5, the hash operation 515 may obtain each hash ID by applying a version of the MurmurHash algorithm (e.g., MurmurHash1, MurmurHash2, or MurmurHash3) to the n-gram to obtain a corresponding hash value (e.g., a 32-bit hash value), and then applying a modulo 10 function to the hash value to obtain the corresponding hash ID.
[0098] Hash embedding creates collisions because the vocabulary size is smaller than the number of unique n-grams. For example, as shown in Figure 5, the n-grams "hello" and "ll" have the same hash ID of 1. Using the Bloom embedding algorithm, the occurrence of collisions can be significantly reduced. Specifically, instead of mapping an n-gram to a single hash ID, each n-gram can be mapped to two (or more) hash IDs. The probability that both (or all) of any two n-grams will have the same hash ID is much lower than the probability that two n-grams will map to the same hash ID.
[0099] FIG. 6 shows an example 600 in which an implementation 615 of the hash operation 515 of the model architecture of FIG. 5 performs Bloom embedding to generate two hash IDs for each n-gram. This implementation 620 of the embedding layer 520 also includes an implementation 625 of the embedding model 525 that generates an embedding vector for each of the two hash IDs. The embedding vectors of the hash IDs are combined to obtain an embedding vector for the n-gram, which is input to a deep learning encoder 630. As in FIG. 5, the vocabulary size is set to the desired size B of the vocabulary (10 in this example), in which case no collisions occur in the hash buckets. As in the example 400 shown in FIG. 4, a CNN with attention model 630 can be used to capture the relationships between n-gram features to generate text-wise encoding vectors, which can be classified using a FFN 640 and a softmax activation function to generate output (language) predictions.
[0100] The n-gram based wide model described above with reference to Figures 4-6 has been found to perform better than character based wide models (which also require a much larger CNN). The n-gram model can be implemented to include a lookup layer (e.g., running in log(n) time) at the input to the deep network 430 (530, 630) and then a small CNN layer on top of it. This results in a much faster run time than character based models with a very large CNN layer on top of the deep network 430 (530, 630).
[0101] The model architecture shown in FIG. 5 or FIG. 6 may include some adjustable hyperparameters. For example, it may be desirable to set the value of n for word-based n-grams to 1 (unigrams) and select three values of n for character-based n-grams: 2, 3, and 4. For hash embedding, it may be desirable to set the number of buckets (B) to three million (3M) and the number of hashes to 2 to handle collision issues. A grid search can be performed to determine the values of the CNN window size and dropout probability. It may be found that the difference between the hyperparameter settings is not significant. It may be desirable to set the maximum number of n-grams in each sentence (e.g., 512) to speed up training.
[0102] During training, script information of the input characters (e.g., Latin, Devanagari) can be applied to limit prediction candidates. For example, if the coding of an input text unit is only CJK (Chinese, Japanese, Korean) script (e.g., as indicated by the Unicode encoding of the text unit), all predictions of Latin-based languages for that text unit may be blocked. Additionally or alternatively, since words may be used in many languages (e.g., the word "estas" may be used in Spanish and Esperanto), it may be desirable to integrate the relative popularity of languages into the model predictions. For example, a higher weight may be applied to predictions for more popular languages.
[0103] For comparison, the model architectures described above with reference to Figures 5 and 6 are designated "ODA Single API" and "ODA API", respectively. To demonstrate and evaluate these model architectures, we built RESTful service-providing application programming interfaces (APIs) (i.e., APIs that conform to Representational State Transfer (REST) constraints) using the FastAPI web framework. Figure 7 shows an example request to the API, and Figure 8 shows an example of the corresponding response from the API.
[0104] Training and Assessment Training data for the model architecture described above was exported from the Open Parallel Corpus Project (OPUS), Common Crawl data, and Wikipedia. Figure 9 shows a table describing the OPUS source dataset. The dataset obtained from the Common Crawl data to be cleaned contains text in 176 languages, including over 1,000 (1K+) tokens from each of 165 languages, over 1 million (1M+) tokens from each of 127 languages, and over 1 billion (1B+) tokens from each of 40 languages.
[0105] If a model works well on short text units, it is likely to also work well on longer text units. For training, short text units containing up to 10 million short sentences (<15 words) and up to 1 million long sentences (>=15 words, <30 words) were extracted from the OPUS dataset. Due to the large size of the Common Crawl dataset, we first extracted all page titles that did not contain numbers or special characters, and then extracted sentences from the body. There is a limit of 1.5 million sentences per language. The resulting training dataset contained text in 135 languages. 35 of the 135 languages were eventually removed due to lack of training data, so a total of 100 languages were supported in this example.
[0106] The following systems were selected for comparison: 1) FastText supports over 170 languages, is freely accessible, and has been shown to outperform other free language detection toolkits (e.g., langdetect, langid, Google's Compact Language Detector 2 (cld2), Google's Compact Language Detector 3 (cld3)). 2) Google Language Detection API (version supporting 109 languages). 3) Microsoft Language Detection API (version supporting 92 languages). 4) Amazon Language Detection API (version supporting 104 languages).
[0107] Two variants of the model architecture shown in Figure 6 (designated "ODA API") were used as baselines. The first variant (designated "CNN API") uses only CNN without attention mechanisms. In the second variant (designated "AVG API"), the CNN layer is omitted and an average pooling layer is used instead (e.g., for fastText). A character-based CNN model ("Char-CNN") was also used as a pure deep learning baseline (e.g., to determine whether a combination of deep neural networks (DNNs) with a wide range of features outperforms pure DNNs). A CNN with multi-kernel window sizes was used to mimic a wide range of features. A set of 69 overlapping languages supported by all comparison systems (including fastText, Google API, and Microsoft API) was selected.
[0108] Figure 10 shows the results of language detection tests using the ODA dataset (containing 335,051 (335K) utterances) for validation and early stopping. This experiment is considered an ablation study (e.g., a study in which a component of the system is removed). The CNN and ODA API were found to achieve better performance than the AGV API, which may indicate that the CNN layer is critical to the success of the model. The ODA API (with attention layer) was found to perform better than the CNN API on the validation set. The ODA API (with bloom embedding shown in Figure 6) was found to perform better on the validation set than the ODA single API (shown in Figure 5). Note that there is no significant difference in terms of the number of parameters between the models.
[0109] Results on the Chatterbot dataset show that the performance of the API is comparable to that of the Google, Microsoft, and Amazon APIs. Results are also obtained on the EuroParl dataset (a dataset of short texts extracted from the minutes of the European Parliament and containing sentence-aligned texts in 21 European languages), the wiLI-2018 dataset (a dataset of short texts extracted from Wikipedia, containing 235,000 paragraphs in 235 languages), the LanideNN dataset, the first simple English test (ODA-10K examples), and the second simple English test (336 examples). We find that fastText performs best among free language detection tools, but underperforms all commercial APIs on public datasets. Among our APIs (ODA, CNN, AVG), ODA, which combines CNN with an attention layer, performs best on all datasets except the simple English test. The performance produced by the ODA API on the Chatterbot dataset and our in-house simple English test is also comparable to that produced by the Google, Microsoft, and Amazon APIs. For other datasets (e.g., LanideNN, EuroParl), we find that our ODA API outperforms the Google and Microsoft APIs. We also find that combining a wide range of features with DNNs can outperform pure DNNs.
[0110] Language detection technology FIG. 11 illustrates a block diagram of an apparatus 1100 according to various embodiments. The elements illustrated in FIG. 11 are implemented in software (e.g., code, instructions, modules, programs) executed by processing circuitry (e.g., one or more processing units such as processors or cores) of the respective system, hardware, or combination thereof, and coupled to a memory (e.g., for storing text units, sequences of n-grams, and / or parameters of a network to be trained). The apparatus 1100 includes a parser 1110 that receives a text unit as input and generates a corresponding sequence of n-grams for the text unit (e.g., as described above with reference to splitting 410). The corresponding sequence of n-grams may include word-level n-grams and / or character-level n-grams, and the n values of the multiple character-level n-grams may be the same or different from the n values of the multiple word-level n-grams. In one example, the corresponding sequence of n-grams includes word-level unigrams and character-level bigrams.
[0111] The apparatus 1100 also includes an embedding layer 1120 that receives the sequence of n-grams and generates ordered embedding vectors corresponding to the sequence of n-grams. The ordered embedding vectors may be based on component vectors (e.g., an embedding model such as a trained embedding matrix). The order of the ordered embedding vectors may indicate or correspond to an order of occurrence of the corresponding n-grams within the text unit.
[0112] The apparatus 1100 also includes a deep network 1130 that receives the ordered embedding vectors and generates an encoding vector for the text unit. The deep network may include at least one hidden layer between the input layer and the output layer and may include a CNN that is trained. The deep network may include an attention mechanism (e.g., one or more attention layers) that generates attention weights that indicate, for example, which n-grams should be paid more attention to (e.g., which n-grams should be weighted more) when performing prediction.
[0113] 12 illustrates another example 1210 of a deep network 1130 that includes an attention mechanism. The mechanism includes an attention layer 1230 that is configured to assign attention weights (depicted by thick dashed lines) to the outputs of a CNN layer 1220. The final encoding vector of the input text is a weighted sum of the CNN outputs (e.g., weighted using the attention weights).
[0114] The apparatus 1100 also includes a classifier 1140 that receives the encoded vectors and generates a linguistic prediction for the text unit. The classifier may include a feedforward neural network. In such a case, the classifier may be configured to apply a softmax function to the output of the final layer of the feedforward neural network (e.g., weighted using an attention layer, as described with reference to FIG. 12).
[0115] As mentioned above, the input text unit may be parsed into a sequence of n-grams, which may include word-level n-grams and / or character-level n-grams, where the value of n in each case may be a tunable parameter. FIG. 13 illustrates an example of operations that may be performed by an implementation 1325 of the embedding layer 1120 to obtain, for each n-gram in the sequence of n-grams, a corresponding embedding vector in a plurality of ordered embedding vectors. In this example, a first hash is performed on the n-gram to obtain a first hash value, and a modulo B operation is applied to the first hash value to obtain a first index, where B is the number of component vectors in the plurality of component vectors to be trained (e.g., the embedding model to be trained). Similarly, a second hash is performed on the n-gram to obtain a second hash value, and a modulo B operation is applied to the second hash value to obtain a second index. The component vectors indicated by the first index and the second index are combined (e.g., concatenated, weighted, and / or added) to obtain an embedding vector for the n-gram. The configuration of the combining operation may include one or more adjustable parameters (e.g., whether the component vectors are concatenated / weighted / added, how the weights are determined, etc.).
[0116] FIG. 14 is a flow chart illustrating a process 1400 of language detection according to an embodiment. The process illustrated in FIG. 14 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores), hardware, or a combination thereof of the respective systems. The software may be stored in a non-transitory storage medium (e.g., memory device). The method illustrated in FIG. 14 and described below is for illustrative purposes and not limiting. FIG. 14 illustrates various processing steps performed in a particular sequence or order, but is not intended to be limiting. In an alternative embodiment, the steps may be performed in some different order, or some steps may be performed in parallel. In an embodiment, such as the embodiment illustrated in FIGS. 1-3, the process illustrated in FIG. 14 may be performed by a pre-processing subsystem (e.g., language detector 212) to generate extracted information for use by one or more other subsystems (e.g., multiple intent subsystems 220 or 310 and / or explicit invocation subsystem 110 or intent classifier 320).
[0117] At block 1404, a sequence of n-grams for the text units is obtained by a data processing system (e.g., chatbot systems 106, 201, and / or 300 described with respect to Figures 1-3, respectively). Obtaining the sequence of n-grams may include receiving the text units as input and parsing the text units to generate a sequence of n-grams. The sequence of n-grams may include word-level n-grams and / or character-level n-grams, where the n values for the multiple character-level n-grams may be the same or different from the n values for the multiple word-level n-grams. In one example, the sequence of n-grams includes word-level unigrams and character-level bigrams.
[0118] At block 1408, an embedding layer is used to obtain multiple ordered embedding vectors for the sequence of n-grams. The embedding layer includes a trained model having multiple component vectors.
[0119] In block 1412, a deep network is used to obtain an encoding vector based on the ordered embedding vectors. The deep network includes an attention mechanism (e.g., one or more attention layers). In various embodiments, the deep network may include a trained CNN.
[0120] In block 1416, a classifier is used to obtain a language prediction for the text unit based on the encoding vector. In various embodiments, the classifier can include a feed-forward neural network. In such a case, using the classifier can include applying a softmax function to the output of the final layer of the feed-forward neural network.
[0121] FIG. 15 is a flow chart illustrating a process 1500 of language detection according to an embodiment. The process illustrated in FIG. 15 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective system, hardware, or combination thereof. The software may be stored in a non-transitory storage medium (e.g., memory device). The method illustrated in FIG. 15 and described below is for illustrative purposes and is not intended to be limiting. Although FIG. 15 illustrates various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In an alternative embodiment, steps may be performed in a different order or some steps may also be performed in parallel. In an embodiment, such as the embodiment illustrated in FIGS. 1-3, the process illustrated in FIG. 15 may be performed by a pre-processing subsystem (e.g., language detector 212) to generate extracted information used by one or more other subsystems (e.g., multiple intent subsystems 220 or 310 and / or explicit invocation subsystem 110 or intent classifier 320).
[0122] Blocks 1504, 1512, and 1516 may be implemented as described above for blocks 1404, 1412, and 1416 with reference to FIG. 14. At block 1508, an embedding layer is used to obtain a plurality of embedding vectors ordered for the sequence of n-grams. The embedding layer includes a trained model having a plurality of component vectors. Block 1508 includes blocks 1508a-d that may be executed to obtain a corresponding one of a plurality of embedding vectors ordered for each n-gram in the sequence of n-grams. At block 1508a, a first hash value of the n-gram and a second hash value of the n-gram are obtained. For example, obtaining the first hash value of the n-gram may include applying a hash function having a first seed value to the n-gram, and obtaining the second hash value of the n-gram may include applying a hash function having a second seed value different from the first seed value to the n-gram. At block 1508b, a first component vector is selected from the plurality of component vectors based on the first hash value. At block 1508c, a second component vector is selected from the plurality of component vectors based on the second hash value. For example, the process 1500 may include applying a modulo function to the first hash value to obtain a first index and applying a modulo function to the second hash value to obtain a second index, where the selection of the first component vector may be based on the first index and the selection of the second component vector may be based on the second index. At block 1508d, an n-gram embedding vector based on the first component vector and the second component vector is obtained. For example, the embedding vector may be obtained as a concatenation of the first component vector and the second component vector. Additionally or alternatively, obtaining the n-gram embedding vector may include applying a first weight value to the first component vector to obtain a first weighted vector and applying a second weight value to the second component vector to obtain a second weighted vector, where the embedding vector is based on the first weighted vector and the second weighted vector.
[0123] Explanation System 16 shows a simplified diagram of a distributed system 1600. In the depicted example, the distributed system 1600 includes one or more client computing devices 1602, 1604, 1606, and 1608 coupled to a server 1612 via one or more communications networks 1610. The client computing devices 1602, 1604, 1606, and 1608 may be configured to run one or more applications.
[0124] In various examples, the server 1612 may be adapted to execute one or more services or software applications that enable one or more embodiments described in this disclosure. In certain examples, the server 1612 may also provide other services or software applications, which may include non-virtual and virtual environments. In some examples, these services may be provided as web-based or cloud services, such as under a software-as-a-service (SaaS) model, to users of the client computing devices 1602, 1604, 1606, and / or 1608. Users operating the client computing devices 1602, 1604, 1606, and / or 1608 may in turn utilize one or more client applications to interact with the server 1612 and utilize the services provided by these components.
[0125] In the configuration shown in Figure 16, the server 1612 may include one or more components 1618, 1620, and 1622 that implement the functions performed by the server 1612. These components may include software components that may be executed by one or more processors, hardware components, or combinations thereof. It should be understood that a variety of different system configurations are possible that may differ from the distributed system 1600. Thus, the example shown in Figure 16 is one example of a distributed system for implementing an exemplary system and is not intended to be limiting.
[0126] A user may use client computing devices 1602, 1604, 1606, and / or 1608 to execute one or more applications, models, or chatbots, which may generate one or more events or models that may be implemented or serviced according to the teachings of this disclosure. The client devices may provide an interface that allows a user of the client device to interact with the client device. The client devices may also output information to the user via the interface. Although only four client computing devices are shown in FIG. 16, any number of client computing devices may be supported.
[0127] Client devices may include various types of computing systems, such as portable handheld devices, general purpose computers such as personal computers or laptops, workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-like operating systems, Linux® or Linux-like operating systems such as Google Chrome™ OS), including various mobile operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android™, BlackBerry®, Palm OS®). Portable handheld devices may include mobile phones, smartphones (e.g., iPhone®), tablets (e.g., iPad®), personal digital assistants (PDAs), etc. Wearable devices may include Google Glass® head-mounted displays and other devices. The gaming systems may include various handheld gaming devices, Internet-enabled gaming devices (e.g., Microsoft Xbox® game consoles with or without Kinect® gesture input devices, Sony Play Station® systems, various gaming systems offered by Nintendo®, etc.), etc. The client devices may run a variety of different applications, such as various Internet-related apps, communication applications (e.g., email applications, short message service (SMS) applications), etc., and may use a variety of communication protocols.
[0128] Network 1610 may be any type of network familiar to those skilled in the art and may support data communications using any of a variety of available protocols, including, but not limited to, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internet Packet Exchange), Apple Talk, etc. By way of example only, network 1610 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a network operating under any of the Institute of Electrical and Electronics Engineers (IEEE) 1002.11 protocol suite), Bluetooth, and / or other wireless protocols), and / or any combination of these and / or other networks.
[0129] The servers 1612 may be comprised of one or more general purpose computers, dedicated server computers (including, by way of example, PC (personal computer) servers, UNIX servers, mid-range servers, mainframe computers, rack-mounted servers, etc.), server farms, server clusters, or other suitable arrangements and / or combinations. The servers 1612 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization, such as one or more flexible pools of logical storage that may be virtualized to maintain the virtual storage of the servers. In various examples, the servers 1612 may be adapted to run one or more services or software applications that provide the functionality described in the preceding disclosure.
[0130] The computing system of server 1612 may run one or more operating systems, including any of those mentioned above, as well as any commercially available server operating system. Server 1612 may also run any of a variety of additional server applications and / or mid-tier applications, including a HyperText Transport Protocol (HTTP) server, a File Transfer Protocol (FTP) server, a Common Gateway Interface (CGI) server, a JAVA server, a database server, and the like. Exemplary database servers include, but are not limited to, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), and the like.
[0131] In some implementations, the server 1612 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of the client computing devices 1602, 1604, 1606, and 1608. By way of example, the data feeds and / or event updates may include, but are not limited to, Twitter® feeds, Facebook® updates, or real-time updates received from one or more third party information sources and continuous data streams, which may include real-time events associated with sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, and the like. The server 1612 may also include one or more applications for displaying the data feeds and / or real-time events via one or more display devices of the client computing devices 1602, 1604, 1606, and 1608.
[0132] The distributed system 1600 may also include one or more data repositories 1614, 1616. These data repositories may be used to store data or other information in some examples. For example, one or more of the data repositories 1614, 1616 may be used to store information, such as information related to the performance of a chatbot or models generated, for use by the chatbot used by the server 1612 in performing various functions according to various embodiments. The data repositories 1614, 1616 may reside in a variety of locations. For example, the data repository used by the server 1612 may be local to the server 1612, may be remote from the server 1612, and may communicate with the server 1612 via a network-based or dedicated connection. The data repositories 1614, 1616 may be of different types. In some examples, the data repository used by the server 1612 may be a database, such as a relational database, such as databases provided by Oracle Corporation® or other vendors. One or more of these databases may be adapted to allow for storage, updating, and retrieval of data between the databases in response to SQL-formatted commands.
[0133] In one example, one or more of the data repositories 1614, 1616 may be used by an application to store application data. The data repositories used by the application may be of various types, such as, for example, a key / value store repository, an object store repository, or a general storage repository supported by a file system.
[0134] In an example, the functionality described in this disclosure may be provided as a service via a cloud environment. FIG. 17 is a simplified block diagram of a cloud-based system environment in which various services may be provided as cloud services, according to an example. In the example shown in FIG. 17, a cloud infrastructure system 1702 may provide one or more cloud services that may be requested by users using one or more client computing devices 1704, 1706, and 1708. The cloud infrastructure system 1702 may comprise one or more computers and / or servers, which may include those described above for server 1612. The computers in the cloud infrastructure system 1702 may be organized as general purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0135] The network 1710 can facilitate communication and exchange of data between the clients 1704, 1706, and 1708 and the cloud infrastructure system 1702. The network 1710 can include one or more networks. The networks may be of the same type or different types. The network 1710 can support one or more communication protocols, including wired and / or wireless protocols, to facilitate communication.
[0136] The example shown in Figure 17 is merely one example of a cloud infrastructure system and is not intended to be limiting. It should be understood that in some other examples, cloud infrastructure system 1702 may have more or fewer components than those shown in Figure 17, may combine two or more components, or may have a different configuration or arrangement of components. For example, while Figure 17 shows three client computing devices, any number of client computing devices may be supported in alternative examples.
[0137] The term cloud services is generally used to refer to services provided to users on demand by a service provider's system (e.g., cloud infrastructure system 1702) over a communications network such as the Internet. Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premise servers and systems. The cloud service provider's systems are managed by the cloud service provider. Thus, customers can use cloud services provided by the cloud service provider without separately purchasing licenses, support, or hardware and software resources for the services. For example, the cloud service provider's system can host applications, and users can order and use the applications on demand over the Internet without purchasing infrastructure resources to run the applications. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, several cloud services, such as middleware services, database services, and Java cloud services, are offered by Oracle Corporation of Redwood Shores, California.
[0138] In one example, cloud infrastructure system 1702 can provide one or more cloud services using different models, such as a Software as a Service (SaaS) model, a Platform as a Service (PaaS) model, an Infrastructure as a Service (IaaS) model, and others including hybrid service models. Cloud infrastructure system 1702 can include a set of applications, middleware, databases, and other resources that enable the delivery of various cloud services.
[0139] In the SaaS model, an application or software may be provided to a customer as a service over a communications network such as the Internet without the customer having to purchase the underlying application hardware or software. For example, the SaaS model may be used to provide customers with access to on-demand applications hosted by the cloud infrastructure system 1702. Examples of SaaS services offered by Oracle Corporation® include, but are not limited to, various services for human resource / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytical services, social applications, etc.
[0140] The IaaS model is commonly used to provide infrastructure resources (e.g., servers, storage, hardware and networking resources) as a cloud service to customers to provide elastic computing and storage capabilities. A variety of IaaS services are offered by Oracle Corporation.
[0141] The PaaS model is generally used to provide platform and environment resources as a service that enable customers to develop, run, and manage applications and services without the customer having to procure, build, or maintain such resources. Examples of PaaS services provided by Oracle Corporation include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), data management cloud services, and various application development solution services.
[0142] Cloud services are generally provided on an on-demand self-service basis, subscription-based, elastically scalable, reliable, highly available, and secure manner. For example, a customer may order one or more services provided by cloud infrastructure system 1702 via a subscription order. Cloud infrastructure system 1702 then performs processing to provide the services requested in the customer's subscription order. For example, a user may use utterances to request the cloud infrastructure system to perform certain actions (e.g., intents), as described above, and / or provide services to a chatbot system, as described herein. Cloud infrastructure system 1702 may be configured to provide one or more cloud services.
[0143] Cloud infrastructure system 1702 can provide cloud services through different deployment models. In a public cloud model, cloud infrastructure system 1702 can be owned by a third-party cloud service provider and cloud services are provided to public customers, which can be individuals or businesses. In another example, under a private cloud model, cloud infrastructure system 1702 can be operated within an organization (e.g., within a corporate organization) and services can be provided to customers within the organization. For example, the customers can be various departments of a company, such as human resources, payroll, or individuals within the company. In another example, in a community cloud model, cloud infrastructure system 1702 and the services provided can be shared by multiple organizations within an associated community. Various other models, such as hybrids of the above models, can also be used.
[0144] Client computing devices 1704, 1706, and 1708 may be of different types (e.g., client computing devices 1602, 1604, 1606, and 1608 shown in FIG. 16 ) and may be capable of running one or more client applications. Users may use the client devices to interact with cloud infrastructure system 1702, such as to request services provided by cloud infrastructure system 1702. For example, users may use client devices to request information or actions from a chatbot, as described in this disclosure.
[0145] In some examples, the processing performed by cloud infrastructure system 1702 to provide services may include training and deploying models. This analysis may include using, analyzing, and manipulating data sets to train and deploy one or more models. This analysis may be performed by one or more processors, possibly processing the data in parallel, running simulations using the data, etc. For example, big data analysis may be performed by cloud infrastructure system 1702 to generate and train one or more models for a chatbot system. The data used for this analysis may include structured data (e.g., data stored in a database or structured according to a structured model) and / or unstructured data (e.g., data blobs (binary large objects)).
[0146] 17, cloud infrastructure system 1702 can include infrastructure resources 1730 that are utilized to facilitate the provision of various cloud services provided by cloud infrastructure system 1702. Infrastructure resources 1730 can include, for example, processing resources, storage or memory resources, networking resources, etc. In one example, a storage virtual machine available to provide storage requested by an application may be part of cloud infrastructure system 1702. In other examples, the storage virtual machine may be part of a different system.
[0147] In one example, to facilitate efficient provisioning of these resources to support various cloud services offered by cloud infrastructure system 1702 for various customers, resources may be bundled into resource sets or resource modules (also referred to as "pods"). Each resource module or pod may include a pre-integrated and optimized combination of one or more types of resources. In one example, different pods may be pre-provisioned for different types of cloud services. For example, a first set of pods may be provisioned for database services, a second set of pods may include a different combination of resources than the pods in the first set and may be provisioned for Java services, etc. For some services, resources allocated to provision a service may be shared between services.
[0148] Cloud infrastructure system 1702 may itself use services 1732 internally that are shared by different components of cloud infrastructure system 1702 and that facilitate provisioning of services by cloud infrastructure system 1702. These internal shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelist services, high availability, backup and recovery services, services enabling cloud support, email services, notification services, file transfer services, etc.
[0149] Cloud infrastructure system 1702 may comprise multiple subsystems. These subsystems may be implemented in software or hardware, or a combination thereof. As shown in FIG. 17, the subsystems may include a user interface subsystem 1712 that allows users or customers of cloud infrastructure system 1702 to interact with cloud infrastructure system 1702. User interface subsystem 1712 may include a variety of different interfaces, such as a web interface 1714, an online store interface 1716, and the cloud services offered by cloud infrastructure system 1702 are advertised and available for purchase by consumer and other interfaces 1718. For example, a customer may use a client device to request one or more services offered by cloud infrastructure system 1702 (service request 1734) using one or more of interfaces 1714, 1716, and 1718. For example, a customer may access an online store, browse cloud services offered by cloud infrastructure system 1702, and place a subscription order for one or more services offered by cloud infrastructure system 1702 to which the customer wishes to subscribe. The service request may include information identifying the customer and one or more services to which the customer wishes to subscribe. For example, a customer may place a subscription order for a service provided by cloud infrastructure system 1702. As part of the order, the customer may provide information identifying the chatbot system for which the service is to be provided and, optionally, one or more credentials for the chatbot system.
[0150] In certain examples, such as the example shown in Figure 17, the cloud infrastructure system 1702 may include an order management subsystem (OMS) 1720 configured to process new orders. As part of this process, the OMS 1720 creates an account for the customer if one has not already been created, receives billing and / or accounting information from the customer that is used to bill the customer for providing the requested services to the customer, verifies the customer information, and, once verified, books the order for the customer. Various workflows may be configured to coordinate and prepare the order for provisioning (if one has not already been created).
[0151] Upon proper validation, the OMS 1720 may invoke an order provisioning subsystem (OPS) 1724 configured to provision resources for the order, including processing, memory, and networking resources. Provisioning may include allocating resources for the order and configuring the resources to facilitate the services requested by the customer's order. The manner in which resources are provisioned for the order and the type of resources provisioned may vary depending on the type of cloud service ordered by the customer. For example, according to one workflow, the OPS 1724 may be configured to determine the specific cloud service being requested and identify the number of pods that may be pre-configured for that specific cloud service. The number of pods allocated for the order may vary depending on the size / amount / level / scope of the service being requested. For example, the number of pods allocated may be determined based on the number of users supported by the service, the duration for which the service is being requested, etc. The pods allocated may then be customized for the specific requesting customer to provide the requested service.
[0152] In one example, the setup phase operations may be performed by cloud infrastructure system 1702 as part of a provisioning process, as described above. Cloud infrastructure system 1702 may generate an application ID and select a storage virtual machine for the application from among the storage virtual machines provided by cloud infrastructure system 1702 itself or from among the storage virtual machines provided by other systems other than cloud infrastructure system 1702.
[0153] Cloud infrastructure system 1702 may send a response or notification 1744 to the requesting customer to indicate when the requested service is available for use. In some cases, information (e.g., a link) may be sent to the customer that enables the customer to begin using and taking advantage of the requested service. In one example, for a customer requesting a service, the response may include a chatbot system ID generated by cloud infrastructure system 1702 and information identifying the chatbot system selected by cloud infrastructure system 1702 for the chatbot system corresponding to the chatbot system ID.
[0154] Cloud infrastructure system 1702 may provide services to multiple customers. For each customer, cloud infrastructure system 1702 is responsible for managing information related to one or more subscription orders received from the customer, maintaining customer data related to the orders, and providing the requested services to the customer. Cloud infrastructure system 1702 may also collect usage statistics regarding the customer's use of the subscription services. For example, statistics may be collected such as the amount of storage used, the amount of data transferred, the number of users, system uptime and downtime, etc. This usage information may be used to bill the customer. Billing may occur, for example, on a monthly cycle.
[0155] Cloud infrastructure system 1702 can provide services to multiple customers in parallel. Cloud infrastructure system 1702 can store information for these customers, possibly including proprietary information. In one example, cloud infrastructure system 1702 comprises an identity management subsystem (IMS) 1728 configured to manage customer information and provide separation of management information such that information related to one customer cannot be accessed by another customer. IMS 1728 can be configured to provide various security-related services, such as identity services, such as information access management, authentication and authorization services, and services for managing customer identities and roles and related functions.
[0156] FIG. 18 illustrates an example of a computer system 1800. In some examples, the computer system 1800 may be used to implement either a digital assistant or chatbot system in a distributed environment, as well as the various servers and computer systems described above. As shown in FIG. 18, the computer system 1800 includes various subsystems, including a processing subsystem 1804 that communicates with many other subsystems via a bus subsystem 1802. These other subsystems may include a processing accelerator 1806, an I / O subsystem 1808, a storage subsystem 1818, and a communication subsystem 1824. The storage subsystem 1818 may include non-transitory computer-readable storage media, including a storage medium 1822 and a system memory 1810.
[0157] Bus subsystem 1802 provides a mechanism that allows the various components and subsystems of computer system 1800 to communicate with each other as intended. Although bus subsystem 1802 is shown diagrammatically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1802 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, a local bus using any of a variety of bus architectures, and the like. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus. It may be implemented as a mezzanine bus manufactured in accordance with the IEEE P13156.1 standard, and the like.
[0158] The processing subsystem 1804 controls the operation of the computer system 1800 and may include one or more processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The processors may include single-core or multi-core processors. The processing resources of the computer system 1800 may be organized into one or more processing units 1832, 1834, etc. The processing units may include one or more processors, one or more cores from the same or different processors, combinations of cores and processors, or other combinations of cores and processors. In some examples, the processing subsystem 1804 may include one or more dedicated co-processors, such as a graphics processor, digital signal processor (DSP), etc. In some examples, some or all of the processing units of the processing subsystem 1804 may be implemented using customized circuitry, such as an application specific integrated circuit (ASIC) or field programmable gate array (FPGA).
[0159] In some examples, the processing units in the processing subsystem 1804 can execute instructions stored in the system memory 1810 or the computer readable storage medium 1822. In various examples, the processing units can execute various program or code instructions and maintain multiple simultaneously executing programs or processes. At any time, some or all of the program code being executed can be resident in the system memory 1810 and / or the computer readable storage medium 1822, potentially including one or more storage devices. Through appropriate programming, the processing subsystem 1804 can provide the various functions described above. If the computer system 1800 is running one or more virtual machines, one or more processing units can be assigned to each virtual machine.
[0160] In some examples, a processing accelerator 1806 may optionally be provided to accelerate the overall processing performed by computer system 1800, to perform customized processing, or to offload portions of the processing performed by processing subsystem 1804.
[0161] I / O subsystem 1808 may include devices and mechanisms for inputting information to computer system 1800 and / or outputting information from or through computer system 1800. In general, use of the term input device is intended to include all possible types of devices and mechanisms for inputting information to computer system 1800. User interface input devices may include, for example, keyboards, pointing devices such as mice or trackballs, touchpads or touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, voice input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices, such as Microsoft Kinect® motion sensors that allow a user to control and operate the input device, Microsoft Xbox® 360 game controllers, devices that provide an interface for receiving input using gestures and voice commands, and the like. User interface input devices may also include eye gesture recognition devices, such as a Google Glass® blink detector, that detects a user's eye movements (e.g., "blinking" when taking a photo or selecting a menu) and translates the eye gestures as input to the input device (e.g., Google Glass®). Additionally, user interface input devices may include voice recognition sensing devices that allow a user to interact with a voice recognition system (e.g., the Siri® navigator) through voice commands.
[0162] Other examples of user interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, game pads and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers 3D scanners, 3D printers, laser range finders, eye-tracking devices, etc. Additionally, user interface input devices may include medical imaging input devices such as, for example, computed tomography, magnetic resonance imaging, position emission tomography, and medical ultrasound devices. User interface input devices may also include audio input devices such as, for example, MIDI keyboards, digital musical instruments, etc.
[0163] In general, use of the term output device is intended to include all possible types of devices and mechanisms for outputting information from computer system 1800 to a user or to another computer. User interface output devices may include non-visual displays such as a display subsystem, indicator lights, or audio output devices. Display subsystems may be flat panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs) or plasma displays, projection devices, touch screens, and the like. For example, user interface output devices include, but are not limited to, a variety of display devices that visually convey textual, graphical, and audio / video information, such as monitors, printers, speakers, headphones, automobile navigation systems, plotters, audio output devices, modems, and the like.
[0164] The storage subsystem 1818 provides a repository or data store for storing information and data used by the computer system 1800. The storage subsystem 1818 provides a tangible, non-transitory computer-readable storage medium (e.g., non-transitory computer-readable memory) for storing basic programming and data structures that provide some example functionality. The storage subsystem 1818 can store software (e.g., programs, code modules, instructions) that, when executed by the processing subsystem 1804, provide the functionality described above. The software can be executed by one or more processing units of the processing subsystem 1804. The storage subsystem 1818 can also provide authentication in accordance with the teachings of the present disclosure.
[0165] The storage subsystem 1818 may include one or more non-transitory memory devices, including volatile and non-volatile memory devices. As shown in FIG. 18, the storage subsystem 1818 includes a system memory 1810 and a computer-readable storage medium 1822. The system memory 1810 may include multiple memories, including a volatile main random access memory (RAM) for storing instructions and data during program execution, and a non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some implementations, a basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within the computer system 1800, such as during start-up, may typically be stored in a ROM. The RAM typically contains data and / or program modules currently being operated on and executed by the processing subsystem 1804. In some implementations, the system memory 1810 may include multiple different types of memories, such as static random access memory (SRAM), dynamic random access memory (DRAM), and the like.
[0166] 18, system memory 1810 may load executing application programs 1812, which may include various applications, such as a web browser, a mid-tier application, a relational database management system (RDBMS), program data 1814, and an operating system 1816. By way of example, operating system 1816 may include Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems, such as various versions of iOS, Windows® Phone, Android® OS, BlackBerry® OS, Palm® OS operating systems, etc.
[0167] The computer readable storage medium 1822 can store programming and data structures that provide some example functionality. The computer readable medium 1822 can provide storage of computer readable instructions, data structures, program modules, and other data for the computer system 1800. When executed by the processing subsystem 1804, software (programs, code modules, instructions) that provide the above-mentioned functionality can be stored in the storage subsystem 1818. As an example, the computer readable storage medium 1822 can include non-volatile memory such as a hard disk drive, a magnetic disk drive, a CD-ROM, a DVD, an optical disk drive such as a Blu-Ray® disk, or other optical media. The computer readable storage medium 1822 can include, but is not limited to, a Zip® drive, a flash memory card, a Universal Serial Bus (USB) flash drive, a Secure Digital (SD) card, a DVD disk, a digital video tape, and the like. The computer readable storage medium 822 may also include solid state drives (SSDs) based on non-volatile memory such as flash memory based SSDs, enterprise flash drives, SSDs based on volatile memory such as solid state RAM, dynamic RAM, static RAM, DRAM based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory based SSDs.
[0168] In one example, the storage subsystem 1818 may also include a computer readable storage medium reader 1820, which may be further connected to a computer readable storage medium 1822. The reader 1820 may be configured to receive and read data from a memory device, such as a disk, a flash drive, or the like.
[0169] In an example, computer system 1800 may support virtualization techniques, including but not limited to virtualization of processing and memory resources. For example, computer system 1800 may provide support for running one or more virtual machines. In an example, computer system 1800 may execute a program, such as a hypervisor, that facilitates configuration and management of virtual machines. Each virtual machine may be assigned memory, computing (e.g., processors, cores), I / O, and network resources. Each virtual machine typically runs independently of other virtual machines. A virtual machine typically runs its own operating system, which may be the same or different than the operating systems run by other virtual machines executed by computer system 1800. Thus, multiple operating systems may potentially be run by computer system 1800 simultaneously.
[0170] The communications subsystem 1824 provides an interface to other computer systems and networks. The communications subsystem 1824 serves as an interface for sending and receiving data from the computer system 1800 to and from other systems. For example, the communications subsystem 1824 allows the computer system 1800 to establish a communication channel with one or more client devices over the Internet to send and receive information from the client devices. For example, if the computer system 1800 is used to implement the bot system 120 shown in FIG. 1, the communications subsystem can be used to communicate with a chatbot system selected for the application.
[0171] The communications subsystem 1824 can support both wired and / or wireless communications protocols. In certain examples, the communications subsystem 1824 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular technology, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution)), WiFi (IEEE 1502.XX family of standards, other mobile communications technologies, or a combination thereof), global positioning system (GPS) receiver components, and / or other components. In some examples, the communications subsystem 1824 can provide a wired network connection (e.g., Ethernet) in addition to or instead of a wireless interface.
[0172] The communications subsystem 1824 may send and receive data in a variety of formats. In some examples, in addition to other formats, the communications subsystem 1824 may receive incoming communications in the form of structured and / or unstructured data feeds 1826, event streams 1828, event updates 1830, etc. For example, the communications subsystem 1824 may be configured to receive (or send) data feeds 1826 in real time from users of social media networks and / or other communications services, such as web feeds such as Twitter® feeds, Facebook® updates, rich site summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0173] In one example, the communications subsystem 1824 is configured to receive data in the form of a continuous data stream, which may include an event stream 1828 of real-time events and / or event updates 1830 that has no explicit end and may be continuous or unlimited in nature. Examples of applications that generate continuous data may include sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, and the like.
[0174] The communications subsystem 1824 may also be configured to communicate data from the computer system 1800 to other computer systems or networks. Data may be communicated in a variety of forms, such as structured and / or unstructured data feeds 1826, event streams 1828, event updates 1830, etc., to one or more databases that may be in communication with one or more streaming data source computers coupled to the computer system 1800.
[0175] The computer system 1800 may be one of a variety of types, including a handheld portable device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head-mounted display), a personal computer, a workstation, a mainframe, a kiosk, a server rack, or other data processing system. Due to the ever-changing nature of computers and networks, the description of the computer system 1800 shown in FIG. 18 is intended only as a specific example. Many other configurations are possible having more or fewer components than the system shown in FIG. 18. It should be understood that there are other ways and / or methods for implementing the various examples based on the disclosure and teachings provided herein.
[0176] Although specific examples have been described, various modifications, variations, alternative configurations, and equivalents are possible. The examples are not limited to operating in a specific data processing environment, but may freely operate in multiple data processing environments. Moreover, while certain examples have been described using a particular sequence of transactions and steps, those skilled in the art will appreciate that this is not intended to be limiting. Although some flow charts describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Additionally, the order of operations may be rearranged. A process may include additional steps not included in the figures. Various features and aspects of the above examples may be used individually or in combination.
[0177] Additionally, while certain examples have been described using a particular combination of hardware and software, it should be recognized that other combinations of hardware and software are possible. Certain examples may be implemented exclusively in hardware, exclusively in software, or using a combination thereof. The various processes described herein may be implemented on the same processor or on different processors in any combination.
[0178] Where a device, system, component, or module is described as being configured to perform a certain operation or function, such configuration may be achieved, for example, by designing an electronic circuit to perform the operation, by programming a programmable electronic circuit (e.g., a microprocessor) to perform the operation, such as by executing computer instructions or code, or a processor or core that is programmed to execute code or instructions stored in a non-transitory memory medium, or any combination thereof. Processes may communicate using a variety of techniques, including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0179] Specific details are given in the present disclosure to provide a thorough understanding of the examples. However, the examples may be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary details to avoid obscuring the examples. This description provides only illustrative examples and is not intended to limit the scope, applicability, or configuration of other examples. Rather, the description of the examples above provides those skilled in the art with an effective description for implementing various examples. Various changes may be made in the function and arrangement of elements.
[0180] Accordingly, the specification and drawings are to be regarded in an illustrative, rather than a restrictive, sense. However, it will be apparent that additions, subtractions, deletions, and other modifications and alterations may be made without departing from the broader spirit and scope of the appended claims. Thus, while specific examples have been described, they are not intended to be limiting. Various modifications and equivalents are intended to be within the scope of the following claims.
[0181] In the foregoing specification, aspects of the disclosure have been described with reference to specific examples, but those skilled in the art will appreciate that the disclosure is not limited thereto. Various features and aspects of the above disclosure can be used individually or in combination. Moreover, the examples can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are therefore to be regarded as illustrative rather than restrictive.
[0182] In the preceding description, the methods have been described in a particular order for illustrative purposes. It should be understood that in alternative examples, the methods may be performed in an order different from that described. It should also be understood that the methods described above may be performed by hardware components or embodied in a sequence of machine-executable instructions, which may be used to cause a machine, such as a general-purpose or special-purpose processor or logic circuitry that is programmed with the instructions, to perform the methods. These machine-executable instructions may be stored on one or more machine-readable media, such as a CD-ROM or other type of optical disk, a floppy disk, a ROM, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, or any other type of machine-readable medium suitable for storing electronic instructions. Alternatively, the methods may be performed by a combination of hardware and software.
[0183] Where a component is described as being configured to perform a certain operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or by a combination thereof.
[0184] Although illustrative embodiments of the present application have been described in detail herein, it should be understood that the concepts of the present invention may be embodied and employed in various other forms, and that the appended claims are intended to be construed to include such variations, except insofar as limited by the prior art.
Claims
1. obtaining a sequence of n-grams of text units; using an embedding layer to obtain a plurality of embedded vectors ordered with respect to the sequence of n-grams; using a deep neural network to obtain an encoded vector based on the plurality of ordered embedded vectors; using a classifier to obtain a language prediction of the text unit based on the encoded vector, comprising: the embedding layer includes a trained model having a plurality of component vectors; the deep neural network includes an attention mechanism; obtaining the plurality of ordered embedded vectors using the embedding layer comprises, for each n-gram in the sequence of n-grams, obtaining a first hash value of the n-gram and a second hash value of the n-gram; selecting a first component vector from among the plurality of component vectors based on the first hash value; selecting a second component vector from among the plurality of component vectors based on the second hash value; obtaining an embedded vector of the n-gram based on the first component vector and the second component vector, a language detection method.
2. The sequence of n-grams includes a plurality of character-level n-grams and a plurality of word-level n-grams, the language detection method according to claim 1.
3. The value of n for the plurality of character-level n-grams is different from the value of n for the plurality of word-level n-grams, the language detection method according to claim 2.
4. The deep neural network includes a trained convolutional neural network, the language detection method according to claim 1.
5. For each n-gram in the sequence of n-grams, obtaining the first hash value of the n-gram includes applying a hash function having a first seed value to the n-gram; obtaining the second hash value of the n-gram includes applying the hash function having a second seed value to the n-gram, the second seed value being different from the first seed value, the language detection method according to claim 1.
6. Obtaining the plurality of ordered embedding vectors using the embedding layer includes, for each n-gram in the sequence of n-grams, applying a modulo function to the first hash value to obtain a first index and applying the modulo function to the second hash value to obtain a second index, wherein the selection of the first component vector is based on the first index and the selection of the second component vector is based on the second index. The language detection method according to claim 1.
7. For each n-gram in the sequence of n-grams, obtaining the embedding vector of the n-gram includes concatenating the first component vector and the second component vector. The language detection method according to any one of claims 1 to 6.
8. For each n-gram in the sequence of n-grams, obtaining the embedding vector of the n-gram includes obtaining a first weighted vector by applying a first weight value to the first component vector, and obtaining a second weighted vector by applying a second weight value to the second component vector, wherein the embedding vector is based on the first weighted vector and the second weighted vector. The language detection method according to any one of claims 1 to 6.
9. The classifier includes a feed-forward neural network. The language detection method according to any one of claims 1 to 6.
10. Using the classifier includes applying a softmax function to the output of the final layer of the feed-forward neural network. The language detection method according to claim 9.
11. One or more data processors, and when executed by the one or more data processors, causing the one or more data processors to obtain a sequence of n-grams of text units, use an embedding layer to obtain a plurality of ordered embedding vectors for the sequence of n-grams, use a deep neural network to obtain an encoded vector based on the plurality of ordered embedding vectors, and use a classifier to obtain a language prediction of the text unit based on the encoded vector. One or more computer-readable media storing instructions for executing the process. The embedding layer includes a trained model having a plurality of component vectors, The deep neural network includes an attention mechanism, Obtaining the plurality of ordered embedding vectors using the embedding layer is for each n-gram in the sequence of n-grams, Obtaining a first hash value of the n-gram and a second hash value of the n-gram, Selecting a first component vector from among the plurality of component vectors based on the first hash value, Selecting a second component vector from among the plurality of component vectors based on the second hash value, Obtaining an embedding vector of the n-gram based on the first component vector and the second component vector, a system comprising.
12. For each n-gram in the sequence of n-grams, Obtaining the first hash value of the n-gram includes applying a hash function having a first seed value to the n-gram, Obtaining the second hash value of the n-gram includes applying the hash function having a second seed value to the n-gram, the second seed value being different from the first seed value, the system according to claim 11.
13. Obtaining the plurality of ordered embedding vectors using the embedding layer is for each n-gram in the sequence of n-grams, applying a modulo function to the first hash value to obtain a first index, and applying the modulo function to the second hash value to obtain a second index, the selection of the first component vector being based on the first index, and the selection of the second component vector being based on the second index, the system according to claim 11.
14. For each n-gram in the sequence of n-grams, obtaining the embedding vector of the n-gram is Applying a first weight value to the first component vector to obtain a first weighted vector, Applying a second weight value to the second component vector to obtain a second weighted vector, the embedding vector being based on the first weighted vector and the second weighted vector, the system according to claim 11.
15. The system according to any one of claims 11 to 14, wherein training the deep network includes restricting language prediction according to script information of a corresponding input text unit.
16. A computer program for causing one or more data processors to execute processing, the processing including: obtaining a sequence of n-grams of text units; using an embedding layer to obtain a plurality of ordered embedding vectors for the sequence of n-grams; using a deep network to obtain an encoded vector based on the plurality of ordered embedding vectors; using a classifier to obtain a language prediction of the text unit based on the encoded vector, wherein the embedding layer includes a trained model having a plurality of component vectors; wherein the deep network includes an attention mechanism; obtaining the plurality of ordered embedding vectors using the embedding layer includes, for each n-gram in the sequence of n-grams: obtaining a first hash value and a second hash value of the n-gram; selecting a first component vector from the plurality of component vectors based on the first hash value; selecting a second component vector from the plurality of component vectors based on the second hash value; and obtaining an embedding vector of the n-gram based on the first component vector and the second component vector.
17. For each n-gram in the sequence of n-grams, obtaining the first hash value of the n-gram includes applying a hash function having a first seed value to the n-gram; obtaining the second hash value of the n-gram includes applying the hash function having a second seed value to the n-gram, the second seed value being different from the first seed value. The computer program according to claim 16.
18. Obtaining the plurality of embedded vectors to be ordered using the embedding layer includes, for each n-gram in the sequence of n-grams, applying a modulo function to the first hash value to obtain a first index, and applying the modulo function to the second hash value to obtain a second index, wherein the selection of the first component vector is based on the first index and the selection of the second component vector is based on the second index. The computer program according to claim 16.
19. For each n-gram in the sequence of n-grams, obtaining the embedded vector of the n-gram includes applying a first weight value to the first component vector to obtain a first weighted vector, and applying a second weight value to the second component vector to obtain a second weighted vector, wherein the embedded vector is based on the first weighted vector and the second weighted vector. The computer program according to claim 16.
20. The use of the deep neural network includes applying a first attention weight to a first feature value corresponding to a first n-gram in the sequence of n-grams, and applying a second attention weight different from the first attention weight to a second feature value corresponding to a second n-gram in the sequence of n-grams. The computer program according to any one of claims 16 to 19.