Multi-language intent recognition
Through multilingual training and sharing/language-specific layer combination architecture, data sparsity problems in traditional systems when training end-to-end speech to intent systems are solved, and more efficient intent recognition and adaptability are achieved.
Patent Information
- Application Number
- CN202111323814.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-10
- Filing Date
- 2021-11-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-09
AI Technical Summary
Traditional oral comprehension systems face data sparsity problems when training end-to-end speech to intention systems, especially when a large amount of speech training data is required, resulting in limited system performance.
Multilingual training is used to alleviate data sparseness, collect data from multiple languages and train neural network models, and adopt a combined architecture of a set of shared layers and language-specific layers shared by the language pool.
This method shares features across all languages through the sharing layer, and the language-specific layer learns language-specific construction, thereby better identifying intentions under smaller training data samples, improving the performance and adaptability of the system.
Smart Images

Figure CN114548200B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to intent recognition, and more particularly, to multilingual training for speech-to-intent recognition. Background Art
[0002] Machine learning (ML) is the scientific study of algorithms and statistical models that a computer system uses to perform specific tasks, which do not use explicit instructions but instead rely on patterns and inference. Machine learning is regarded as a subset of artificial intelligence. Machine learning algorithms build mathematical models based on sample data (referred to as training data) in order to make predictions or decisions without being explicitly programmed to perform the task. Machine learning algorithms are used in a variety of applications, such as email filtering and computer vision, where it is difficult or infeasible to develop conventional algorithms for performing the tasks effectively.
[0003] In machine learning, a hyperparameter is a configuration outside the model, and its value cannot be estimated from the data. Hyperparameters are used in the process to help estimate model parameters. Hyperparameters are set before the learning (e.g., training) process begins. In contrast, the values of other parameters are derived via training. Different model training algorithms require different hyperparameters, and some simple algorithms (such as least squares regression) do not require any hyperparameters. Given a set of hyperparameters, the training algorithm learns the parameter values from the instance data. The least absolute shrinkage and selection operator (LASSO) is an algorithm that adds a regularization hyperparameter to least squares regression and needs to be set before the parameters are estimated by the training algorithm. Similar machine learning models may require different hyperparameters (e.g., different constraints, weights, or learning rates) to generalize different data patterns.
[0004] Deep learning is a branch of machine learning based on a set of algorithms that model high-level abstractions of data by using a model architecture that has a complex structure or is otherwise typically composed of multiple non-linear transformations. Deep learning is part of a broader family of machine learning methods for learning representations from data. Observations (e.g., images) can be represented in many ways (such as a vector of intensity values for each pixel) or can be represented more abstractly as a collection of edges, regions of specific shapes, etc. Some representations make it easier to learn tasks from examples (e.g., face recognition or facial expression recognition). Deep learning algorithms typically use a cascade of many non-linear processing unit layers for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. The algorithms can be supervised or unsupervised, and applications include pattern analysis (unsupervised) and classification (supervised). Deep learning models include artificial neural networks (ANNs) inspired by information processing nodes and distributed communication nodes in biological systems. ANNs have various differences from the biological brain.
[0005] A neural network (NN) is a computing system inspired by the biological neural network. The NN is not just an algorithm, but a framework for many different machine learning algorithms to work together and process complex data inputs. Such a system "learns" to perform tasks by considering examples, usually without being programmed with any task-specific rules. For example, in image recognition, the NN learns to recognize images containing cats by analyzing example images correctly labeled as "cat" or "non-cat", and uses the results to identify cats in other images. The NN achieves this without any existing knowledge about cats (e.g., cats have fur, tails, whiskers, and pointed ears). Instead, the NN automatically generates recognition features from the learning materials. The NN is based on a collection of connected units or nodes called artificial neurons, which loosely mimic the neurons in the biological brain. Each connection, like a synapse in the biological brain, can transmit signals from one artificial neuron to another. The artificial neuron receiving the signal can process the signal and then transfer the signal to additional artificial neurons.
[0006] In a typical NN implementation, the signals at the connections between artificial neurons are real numbers, and the output of each artificial neuron is calculated by some non-linear function of the sum of its inputs. The connections between artificial neurons are called "edges". Artificial neurons and edges usually have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal at the connection. An artificial neuron may have a threshold such that the signal is sent only when the aggregated signal crosses the threshold. Typically, artificial neurons are grouped into layers. Different layers can perform different kinds of transformations on their inputs. Signals may travel from the first layer (input layer) to the last layer (output layer) after passing through multiple layers multiple times.
[0007] A convolutional neural network (CNN) is a type of neural network most commonly used for analyzing visual images. The CNN is a regularized version of the multi-layer perceptron (e.g., a fully connected network), where each neuron in one layer is connected to all neurons in the next layer. The CNN exploits the hierarchical pattern of data and uses smaller and simpler patterns to assemble more complex patterns. The CNN breaks an image into small patches (e.g., 5×5 pixel patches), and then moves across the image with a specified stride length. Thus, in terms of the scale of connectivity and complexity, the CNN is at the lower end. Compared with other image classification algorithms, the CNN uses relatively little preprocessing, allowing the network to learn filters that are hand-designed in traditional algorithms.
[0008] An artificial neural network (ANN) is a computational system inspired by the biological neural network. The ANN itself is not an algorithm but a framework for many different machine learning algorithms to work together and process complex data inputs. Such a system "learns" to perform tasks by considering examples, usually without being programmed with any task-specific rules. For example, in image recognition, the ANN learns to recognize images containing cats by analyzing example images correctly labeled as "cat" or "non-cat" and uses the results to identify cats in other images. The ANN achieves this without any prior knowledge about cats (e.g., cats have fur, tails, whiskers, and pointed ears). Instead, the ANN automatically generates recognition features from the learning material. The ANN is based on a collection of connected units or nodes called artificial neurons, which loosely mimic the neurons in the biological brain. Each connection, like a synapse in the biological brain, can transmit signals from one artificial neuron to another. The artificial neuron receiving the signal can process the signal and then transfer the signal to additional artificial neurons.
[0009] In a typical ANN implementation, the signals at the connections between artificial neurons are real numbers, and the output of each artificial neuron is calculated by some non-linear function of the sum of its inputs. The connections between artificial neurons are called "edges". Artificial neurons and edges usually have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal at the connection. An artificial neuron may have a threshold such that the signal is sent only when the aggregated signal crosses the threshold. Typically, artificial neurons are grouped into layers. Different layers can perform different kinds of transformations on their inputs. The signal may travel from the first layer (input layer) to the last layer (output layer) after passing through multiple layers many times.
[0010] A recurrent neural network (RNN) is a class of ANN in which the connections between nodes form a directed graph along a sequence that allows the network to exhibit temporal dynamic behavior for time series. Different from feedforward neural networks, RNNs can use internal states (memories) to process input sequences that allow RNNs to be applicable to tasks such as unsegmented connected handwriting recognition or speech recognition. Long short-term memory (LSTM) units are alternative layer units for recurrent neural networks (RNNs). An RNN composed of LSTM units is called an LSTM network. A common LSTM unit consists of a cell, an input gate, an output gate, and a forget gate. The cell remembers values over arbitrary time intervals, and the gates regulate the flow of information into and out of the cell. For LSTM, the learning rate and network size are the most critical hyperparameters. Summary of the Invention
[0011] According to an aspect of the present invention, there is provided a computer-implemented method. The method includes: accessing one or more intents and associated entities from a limited amount of speech-to-text training data in a single language; using the accessed one or more intents and associated entities to locate speech-to-text training data in one or more other languages different from the single language, thereby locating speech-to-text training data in one or more other languages; and training a neural network based on the limited amount of speech-to-text training data in the single language and the located speech-to-text training data in one or more other languages. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Preferred embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which:
[0013] Figure 1 A block diagram of a computing environment according to an embodiment of the present invention is depicted;
[0014] Figure 2 An example block diagram for training an end-to-end intent classifier according to an embodiment of the present invention is depicted;
[0015] Figure 3 An example block diagram of an end-to-end speech-to-intent classifier adapted to a specific language according to an embodiment of the present invention is depicted;
[0016] Figure 4 An example block diagram of a multi-language end-to-end speech-to-intent classifier for general intent recognition according to an embodiment of the present invention is depicted;
[0017] Figure 5 An example block diagram depicting multi-task training of a multi-language end-to-end speech-to-intent classifier according to an embodiment of the present invention is depicted;
[0018] Figure 6 An operational step for training an end-to-end speech multi-language intent classifier according to an embodiment of the present invention is depicted; and
[0019] Figure 7 is a block diagram of an example system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] Embodiments of the present invention recognize that traditional spoken language understanding systems are typically constructed in two parts: an automatic speech recognition (ASR) system that decodes speech into text, followed by a natural language understanding module for intent recognition, entity extraction, and the like. With current neural network-based architectures, it is now possible to train a single end-to-end system that can directly extract intent and entity information from the speech signal without having to produce an intermediate text representation of the input. Embodiments of the present invention recognize the deficiencies of current neural network-based architectures, namely, that the amount of task-specific training data (e.g., speech data with intent labels) for training these systems is typically limited.
[0021] In these settings, embodiments of the present invention recognize the advantages of traditional systems. Since traditional systems are trained in parts, good performance can be achieved by training each component separately. The automatic speech recognition component can be trained on a large amount of transcribed data collected independently without any intent labels. Then, the subsequent intent classifier can be trained on a relatively smaller amount of data. This is typically the case because the data for training the intent classifier is text-only, thus generally alleviating the data sparsity problem of speech data labeled with intents.
[0022] Embodiments of the present invention recognize the difficulty of training an end-to-end speech-to-intent system due to the need to use a very large amount of speech data. Embodiments of the present invention recognize that the amount of speech training data is often limited. This problem is amplified when considering sub-tasks of automatic speech recognition because the speech data also needs to be transcribed.
[0023] Embodiments of the present invention provide a solution to the limited training data. In other words, embodiments of the present invention provide a solution to alleviate the data training problem via multi-language training. For example, embodiments of the present invention can pool data from various languages in multi-language training to train a neural network model. The trained neural network model includes a set of shared layers shared by the language pool. Then, embodiments of the present invention can use language-specific targets. As will be discussed in more detail later in this specification, the parameters of the network model can be better trained by pooling data from multiple languages.
[0024] Figure 1 is a functional block diagram showing a computing environment (generally labeled as computing environment 100) according to an embodiment of the present invention. Figure 1 Only an illustration of one implementation is provided, and it does not imply any limitation to the environment in which different embodiments can be implemented. Those skilled in the art can make many modifications to the described environment without departing from the scope of the present invention as recited in the claims.
[0025] Computing environment 100 includes client computing device 102 and server computer 108, all of which are interconnected via network 106. Client computing device 102 and server computer 108 can be stand-alone computer devices, management servers, web servers, mobile computing devices, or any other electronic device or computing system capable of receiving, sending, and processing data. In other embodiments, client computing device 102 and server computer 108 can represent server computing systems such as those utilizing multiple computers as server systems in a cloud computing environment. In another embodiment, client computing device 102 and server computer 108 can be laptop computers, tablet computers, netbook computers, personal computers (PCs), desktop computers, personal digital assistants (PDAs), smart phones, or any programmable electronic device capable of communicating with different components and other computing devices (not shown) within computing environment 100. In another embodiment, each of client computing device 102 and server computer 108 represents a computing system that utilizes cluster computers and components (e.g., database server computers, application server computers, etc.) that act as a single seamless resource pool when accessed within computing environment 100. In some embodiments, client computing device 102 and server computer 108 are a single device. Client computing device 102 and server computer 108 can include internal and external hardware components capable of executing machine-readable program instructions, as depicted and described in more detail with respect to Figure 7 More detailed depiction and description.
[0026] In this embodiment, client computing device 102 is a user device associated with a user and includes application 104. Application 104 communicates with server computer 108 to access general intent recognizer 110 (e.g., using TCP / IP) to access content, user information, and database information. Application 104 can further communicate with general intent recognizer 110 to transmit instructions for training a multilingual neural network intent classifier, as discussed in more detail with reference to Figures 2 - 6 More detailed discussion.
[0027] Network 106 can be, for example, a telecommunications network, a local area network (LAN), a wide area network (WAN) (such as the Internet), or a combination of the three, and can include wired, wireless, or fiber optic connections. Network 106 can include one or more wired and / or wireless networks capable of receiving and sending data, voice, and / or video signals (including multimedia signals, which include voice, data, and video information). Generally, network 106 can be any combination of connections and protocols that will support communication between client computing device 102 and server computer 108 and other computing devices (not shown) within computing environment 100.
[0028] Server computer 108 is a digital device that hosts general intent recognizer 110 and database 112. In this embodiment, general intent recognizer 110 resides on server computer 108. In other embodiments, general intent recognizer 110 may have an instance of a program (not shown) stored locally on client computer device 102. In other embodiments, general intent recognizer 110 may be an independent program or system that trains a multilingual neural network intent classifier. In other embodiments, general intent recognizer 110 may be stored on any number of computing devices.
[0029] General intent recognizer 110 trains a multilingual neural network intent classifier, i.e., general intent recognizer 110 can recognize intent from speech or text regardless of the language associated with the received content (e.g., speech or text). In this embodiment, general intent recognizer 110 includes intent classifier 114. Intent classifier 114 classifies the intent based on the received content.
[0030] As used herein, "intent" refers to a state of mind or purpose. For example, an intent can be the purpose, wish, or determination behind an action, thought, or utterance. As a specific example, in the sentence: "I want to book a flight from New York to Boston", the intent is "book a flight". New York and Boston are values corresponding to the entities "departure city" and "arrival city". The set of entities and values in the sentence can also be considered part of the intent of the sentence. There are numerous other such attributes that can be classified or referred to as "intents", such as, for example, part-of-speech tags, dialogue state tags, etc.
[0031] In this embodiment, content refers to the received media. For example, the media may include one or more audio files containing speech. The media may also include received text or text files. In some embodiments, the media may further include video files containing audio (e.g., speech).
[0032] In this embodiment, general intent recognizer 110 uses multilingual speech (e.g., known training data) as the received input and pools the shared parameters trained on all languages to pre-train the neural network, as described in more detail in Figure 2 See more detailed description.
[0033] In this embodiment, the term "shared parameters" is used to denote a set of common layers of a neural network that are trained using multilingual data. A layer of a neural network includes a set of nodes. Each layer is connected to other layers via connections having weights thereon. Nodes are associated with different kinds of non-linearities, bias terms, gated information flows, etc. Depending on the kind of network that can be used, there are several variants of nodes and network connections. LSTM, CNN, RNN, DNN are examples of neural networks. For example, (inFigure 2 The shared parameter 204 (discussed and described in []) is a representation of such a set of common network layers that form part of a multi - language intent classifier.
[0034] In this embodiment, the term "language - specific parameter" is used to denote a set of layers of a neural network that are trained using language - specific data. A layer of a neural network consists of a set of nodes. Each layer is connected to other layers via connections that have weights on them. Nodes are associated with different kinds of non - linearities, bias terms, gated information flow, etc. Depending on the kind of network that can be used, there are several variants of nodes and network connections. LSTM, CNN, RNN, DNN are examples of neural networks. (Shown and described in []) The language - specific parameters 206A - N are representations of language - specific network layers that form part of a multi - language intent classifier. When processing multi - language data from N languages, the component 204 (also referred to as the shared parameter 204) is trained on all the data. On the other hand, the components 206A - N (also referred to as the language - specific parameters 206A - N) are trained on language - specific data for each language. Each component 206A - N is connected to a single shared component 204. Each language also has a language - specific intent prediction layer (also referred to as intent language 208A - N) denoted by 208A - N. Figure 2 Then in this embodiment, the intent classifier 114 can parse the received content into separate language - specific parameters. Consider a booking system that processes flight bookings in Spanish and English. A corpus consisting of Spanish and English speech utterances is used to train the system. The multi - language speech corpus is annotated with various intents such as "book a flight", "cancel a flight", "check flight status", "modify travel booking". Each utterance is also annotated with entities. At test time, when an English utterance passes through the network, the recognized intent and entities will be available at the English - specific output of the network. For a Spanish utterance, the output will be available at the Spanish - specific output layer. Such a network is pre - trained with the network architecture shown in []. The network has layers that are trained on both Spanish data and English data. Then, the multi - language speech representations from these shared network layers are passed to the language - specific layers to produce the desired intent output.
[0035] Using Spanish training data, the training signal passes through the shared layer, through the Spanish - specific layer, and is verified at the Spanish - specific output. Using English data, the training signal similarly passes through the shared layer and then is processed by the English - specific layer, and the output is collected at the English - specific output layer. Figure 2 Shown in []
[0036]
[0037] In different embodiments, the available multilingual data may not have intent labels and may only have transcripts. In such cases, the multilingual network can still be pre-trained. Similar to the pre-training described above, using the available multilingual data, the shared layer and language-specific layers can be trained on the multilingual data.
[0038] In another embodiment, the available data can still be in a single language but correspond to different domains, such as banking, airlines, hotels, etc. In such a setup, data from various domains are pooled together. Similar to the multilingual case, the shared layer is trained on all the available data. However, now the language-specific layers will correspond to domain-specific layers: there will be a set of layers corresponding to the banking domain, a different set of layers corresponding to the airline domain, and so on. These domain-specific layers will all be connected in sequence to a set of shared layers.
[0039] In another embodiment, the available data can still be in a single language but drawn from different data sets (e.g., data sets collected to model dialogue states, data sets of voice commands, etc.). In this setup, data from various data sets will again be pooled together. Similar to the multilingual case, the shared layer is trained on all the available data. However, now the language-specific layers will correspond to data set-specific layers: there will be a set of layers corresponding to the dialogue state data set, a different set of layers corresponding to voice commands, and so on. These domain-specific layers will all be connected in sequence to a set of shared layers.
[0040] In another embodiment, the model can also handle scenarios with language switching. The data can have sentences where a person speaks in English, switches to Spanish, and then returns to English, etc. (also known as code-switching). In such cases, the data is from the same domain but different languages are pooled so that the system can leverage the commonalities in intents and entities. Once trained in such a setup, in this embodiment, when words in two languages are used within the same utterance, the system is able to handle speech with code-switching. This can also be helpful for call center analysis where client data in different languages need to be analyzed together. For example, what is the frequency of people booking flights to Houston this month compared to last month, regardless of the language spoken.
[0041] Then, the intent classifier 114 can identify corresponding intents from each of the languages identified or otherwise recognized. Once the multi - language / multi - domain / multi - corpus network has been pre - trained as described above, it can be refined for the immediate end - intent classification task. Consider the multi - language flight reservation network described above. The network has been trained on English and Spanish data with general flight reservation data and labels. The network is now applicable to a specific airline and its specific data. Depending on the nature of the data, the entire pre - trained network or a portion of the network can be used.
[0042] Case 1: The newly received data is in English and has been labeled with the general labels used for the training data. Due to the nature of data collection across different demographics, the new data contains acoustic characteristics. The shared pre - trained network layers as well as the English - specific layers are used to initialize the new network. Then, the new network is trained with the new data.
[0043] Case 2: The new data has been received in English but has been labeled with a new set of intent labels. In this case, the new network trained on the received data is initialized with the shared multi - language layer and the English - specific layer, but a new intent output layer is used instead of the new set of intent labels. Once initialized, the network is fully trained.
[0044] Case 3: The new data has been received in German and has also been labeled with a new set of intent labels. In this setup, the new network is initialized only with the shared multi - language parameters. Similar use cases can be envisioned in the multi - domain / multi - corpus scenario.
[0045] Then, the intent classifier 114 can receive unknown media in real - time, i.e., unrecognized speech. Then, the intent classifier 114 can access the shared parameters trained on all languages, identify the corresponding language and associated language - specific parameters that match the unknown media, and identify the intent associated with the unknown media. As described in more detail Figure 3 as described in more detail.
[0046] In other embodiments, the general intent recognizer 110 can be modified to predict intents that occur across domains and languages. For example, a multi - language corpus can be represented in two languages (e.g., English and Spanish) for two domains: airlines and hotels.
[0047] In this embodiment, the general intent recognizer 110 can include additional layers trained on top of the language - specific layers, which have learned language - specific constructs, as discussed in more detail Figure 4 as discussed in more detail.
[0048] Consider an embodiment of a prior multilingual airline reservation system. Data available for training the system is available in two parts. In a first data cut, only transcripts of speech utterances in English and Spanish are available. A second data cut is a much smaller part of the corpus having both transcripts and intent labels. Using the first data cut, the network can be trained to identify the key parts of each utterance that convey meaning. For example, the transcripts can be tagged with part-of-speech labels. Words tagged as nouns are tagged with values for entities (such as "destination airport"), while verb phrases (such as "fly to") help identify the intent. However, these constructs occur in different variants across languages. As described in more detail regarding Figure 4 The layer 404 in 404 learns language-agnostic constructs, and the layers 406A-N can be considered layers that learn the POS representations for each language, which are pooled together by the layer in 404. The final layer in 408 maps the various POS representations to intent labels.
[0049] Thus, embodiments of the present invention can provide solutions for parameter sharing / learning from different data sets (multilingual, multi-domain, multiple data sets). Front-end sharing (e.g., shared parameters 404 across all languages) helps, for example, model common speech. However, there can also be sharing that helps model common language structures (such as nouns / verbs). If we have similar domains across multiple languages (e.g., travel), then we can share useful "domain logic". This sharing is captured by a shared layer after the language-specific parameter layers (the second box currently also labeled 404 should be re-numbered as it represents a different level of sharing).
[0050] Embodiments of the present invention also enable sharing of data even when the domains are not similar. By sharing data and pre-training certain parameters, embodiments of the present invention capture some general characteristics of spoken language understanding (SLU) and mitigate data insufficiency in any one particular domain / language combination.
[0051] The multiple embodiments discussed herein provide different levels of organization in terms of how intent labels can be organized. There may be cases where intent labels are standardized / shared and cases where the labels are not standardized / shared. Both cases should benefit from the solutions provided by certain embodiments of the present invention as it allows for shared parameters and language / domain-specific parameters at multiple levels.
[0052] For example, the general intent recognizer 110 can receive content including one or more languages, access shared parameters trained on all languages, identify one or more languages and specific parameters associated with each identified language. Then, the intent classifier 114 can access the shared parameters again to identify language-independent intents. The various layers of the neural network can be considered as transformations applied to the input signal to produce various representations at an abstract level. Figure 5 The shared layers (e.g., the shared parameter 504 across all languages) discussed in Figure 5 are responsible for removing unwanted channels and speaker variability. Then, the language representations in terms of basic sound units etc. produced by the shared layer are transformed to simulate the language characteristics of each language by layers 506A-N. The language characteristics include language-specific grapheme or phonetic representations. For each language, these refined representations can be used to extract actual intents, grapheme symbols to construct entity values simulated by layers 508A-B, and so on.
[0053] In other embodiments, the intent classifier 114 can be adapted to a specific domain within the same language or even a new language without having to retrain from scratch using a multilingual model. In such a case, the final language-specific parameters of the existing model are replaced with new domain and corresponding language-specific layers, thus keeping the shared layers intact. Then, the new model is trained on new data to completion.
[0054] Domains and languages that match (e.g., within a certain threshold percentage) the domains and languages used to train the multilingual network are instances of settings that the trained model can handle without retraining. Using the multilingual travel booking system described previously, after training to produce general travel intents, the multilingual travel booking system can be deployed in English or Spanish. By replacing the final language-specific parameters with a new output layer corresponding to a new set of intent labels as described previously (e.g., regarding how different parts of the network can remain intact or otherwise be replaced).
[0055] In still other embodiments, the intent classifier 114 can be adapted to train the neural network with other related tasks in addition to the primary classification task. In this way, the primary classification task can be improved. As discussed in the examples previously, when a speech utterance is processed by the intent classifier, it not only needs to produce intent labels but also, in many cases, produce values corresponding to various entities. To correctly identify the values, the intent recognition system (e.g., the general intent recognizer 110) should also be able to accurately produce a text transcription. If the primary classification task is intent recognition (e.g., "book a flight"), then a related classification task can be entity recognition (where the departure airport and destination airport are also identified).
[0056] For example, embodiments of the present invention can improve the training of the proposed network for intent recognition and multi-task training with other related tasks (e.g., speech / character recognition, as discussed in more detail with respect to Figure 5 Self-supervision can be used to train these networks after the training data has been appropriately modified (e.g., adding noise, speed / pitch modification).
[0057] Database 112 stores the received information and can represent one or more databases that grant permission for access to the general intent recognizer 110 or publicly available databases. Generally, any non-volatile storage medium known in the art can be used to implement database 112. For example, database 112 can be implemented with a tape library, an optical library, one or more independent hard disk drives, or multiple hard disk drives in a redundant array of independent disks (RAID). In this embodiment, database 112 is stored on server computer 108.
[0058] Figure 2 An example block diagram 200 for training an end-to-end intent classifier according to an embodiment of the present invention is depicted.
[0059] In this example, input 202 is fed into intent classifier 114. Input 202 can include any combination of audio, text, and video. For example, input 202 can include multilingual speech recognized in an audio file. In other embodiments, input 202 can be a live audio stream. In other embodiments, input 202 can be a single language. Then, intent classifier 114 can access shared parameter 204. Shared parameter 204 can be a set of pre-trained data.
[0060] Then, intent classifier 114 can identify language-specific parameters 206A, 206B to 206N and output intent languages 208A, 208B, and 208N respectively. Generally, intent languages 208A, 208B, and 208N are one or more corresponding languages and intents associated with the received input 202.
[0061] Figure 3 An example block diagram 300 of an end-to-end speech-to-intent classifier adapted to a specific language according to an embodiment of the present invention is depicted.
[0062] This example depicts a model that has been initialized from a multilingual model and is specific-task adapted to a specific language. Thus, component 304 is initialized from component 204 (also referred to as shared parameters trained on all languages 304).
[0063] In this example, input 302 is fed into intent classifier 114. Similar to Figure 2The inputs 202, 302 can include any combination of audio, text, and video. In this embodiment, the input 302 can be speech that is unknown or has not been otherwise processed by the intent classifier 114 previously. In other words, the input 302 can be non-training data that is fed through an intent classifier that has been pre-trained. In other embodiments, the input 302 can be a live audio stream. Then, the intent classifier 114 can access shared parameters trained across all languages 304.
[0064] Then, the intent classifier 114 can identify language-specific parameters 306 specific to the received input 302 and, accordingly, identify the intent 308, which is the intent of the received input 302.
[0065] Figure 4 Depicted is an example block diagram 400 of a multilingual end-to-end speech-to-intent classifier for general intent recognition according to an embodiment of the present invention.
[0066] In this example, the input 402 is fed to or otherwise accessed by the intent classifier 114. The input 402 can include any combination of audio, text, and video. In this example, the input 402 is multilingual speech. In other embodiments, the input 202 can be a live audio stream of multilingual speech. Then, the intent classifier 114 can access shared parameters trained across all languages 404.
[0067] Thereafter, then, the intent classifier 114 can identify language-specific parameters 406A, 406B through 406N. Different from Figure 2 the intent recognizer (including the intent classifier 114) discussed in, the intent classifier 114 accesses the shared parameters across all languages 404 a second time. Using the shared parameters across all languages 404, the intent classifier 114 can identify language-agnostic intents 408. As described previously, part-of-speech tags can be considered language-agnostic intents. In Figure 4 this, the layers in 404 learn language-agnostic constructs, and the layers 406A-N can be considered layers that learn the POS representations for each language, which are pooled together by the layers in 404. The final layer in 408 maps the various POS representations to intent labels.
[0068] In other embodiments, an intent recognizer with the intent classifier 114 can be applied to a specific domain within the same language or even a new language without having to re-train from scratch using a multilingual model. In such a case, the final language-specific parameters of the existing model are replaced with new domain and corresponding language-specific layers, thus keeping the shared layers intact. Then, the new model is trained on new data to complete.
[0069] Figure 5 FIG. 500 is an example block diagram depicting multi-task training of a multi-lingual end-to-end speech-to-intent classifier in accordance with an embodiment of the present invention.
[0070] In this example, a general intent recognizer with an intent classifier 114 can be further improved (i.e., better trained) by assigning sub-tasks in addition to the primary task. As used herein, the primary task refers to the main task of identifying an intent from given or accessed speech data. Sub-tasks can be other related tasks (e.g., speech / character recognition). In other embodiments, the general intent recognizer can train these networks with self-supervision after the training data has been appropriately modified (e.g., adding noise, speed / pitch modification, etc.).
[0071] In this example, an input 502 is fed to or otherwise accessed by the intent classifier 114. Similar to Figure 2 input 202, input 502 can include any combination of audio, text, and video. In this example, input 502 is multi-lingual speech. In other embodiments, input 202 can be a live audio stream of multi-lingual speech. Then, the intent classifier 114 can access shared parameters trained on all languages 504.
[0072] Thereafter, the intent classifier 114 can identify language-specific parameters 506A, 506B through 506N and, correspondingly, identify primary outputs 508A, 510A, and 512A and corresponding secondary outputs 508B, 510B, and 512B, respectively. As described previously, the primary task can be intent recognition. The secondary task can be identifying graphemes / phonemes / words in the input, which is also commonly referred to as automatic speech recognition. The sequence of identified graphemes / phonemes / words will be used to construct entity values.
[0073] Figure 6 FIG. 600 is a flowchart depicting operational steps for training an end-to-end speech multi-lingual intent classifier in accordance with an embodiment of the present invention.
[0074] In step 602, a general intent recognizer 110 receives information. In this embodiment, the general intent recognizer 110 receives a request from a client computing device 102. In other embodiments, the general intent recognizer 110 can receive information from one or more other components of the computing environment 100.
[0075] The information received by the general intent recognizer 110 refers to voice or text information. The information received or otherwise accessed by the general intent recognizer 110 can include any combination of audio, text, and video. For example, the information can include one or more languages. The information can also include a limited amount of speech-to-text training data in a single language. For example, a limited amount of speech-to-text training data in a single language can include interpreting intents and associated entities from speech in a single language. In some embodiments, the information can be a live audio stream of multilingual speech. For example, the general intent recognizer 110 can receive information from one or more connected IoT devices.
[0076] In step 604, the general intent recognizer 110 determines an intent from the received information. In a scenario where the general intent recognizer 110 is first trained, the received information can be used to pre-train the general intent recognizer 110. In this embodiment, the general intent recognizer 110 can use a combination of one or more machine learning and artificial intelligence algorithms to determine an intent from the received information. In some embodiments, the general intent recognizer 110 can use an existing neural network-based architecture. Then, the general intent recognizer 110 can store the determined intent in the database 112 as part of the pre-trained data.
[0077] In some embodiments where the general intent recognizer 110 has been pre-trained, the general intent recognizer 110 can access pre-trained data that includes one or more intents and associated entities from a limited amount of speech-to-text training data in a single language.
[0078] In step 606, the general intent recognizer 110 trains a neural network based on the determined intent. For example, the general intent recognizer 110 uses the pre-trained data to train the neural network. In this embodiment, the pre-trained data can include shared parameters across all languages. The shared parameters are the layers of the network that are jointly trained across all languages.
[0079] Then, the general intent recognizer 110 can use the pre-trained data (e.g., the one or more intents and associated entities accessed) to locate speech-to-text training data in one or more other languages different from the single language, thereby locating other speech-to-text training data in one or more other languages. Then, the general intent recognizer 110 can determine language-specific parameters based on the shared parameters across all languages accessed and subsequently store the identified language-specific parameters in a database (e.g., database 112).
[0080] In step 608, the general intent recognizer 110 trains a neural network for natural language processing. In this embodiment, the general intent recognizer 110 trains a neural network for natural language processing by accessing an updated database (e.g., the database 112 updated in the previous step). Thus, the general intent recognizer 110 can identify intents from the received input using the trained intent classifier 114. More specifically, the general intent recognizer 110 can receive multilingual input and identify intents using a relatively small training data sample size, regardless of the language received.
[0081] Thus, the general intent recognizer 110 can be trained by preparing multilingual / multi-domain / multicorpus data with transcripts and intents. Then, the general intent recognizer 110 can be trained with shared layers and language-specific layers. In some embodiments, the general intent recognizer 110 may have prepared or otherwise have domain / language / corpus-specific data. Finally, the general intent recognizer 110 can use data from the shared layers and language-specific layers to tune the general intent recognizer and can initialize as many layers as possible from a pre-trained network.
[0082] Figure 7 Depicts components of a computing system within a Figure 1 computing environment 100 according to an embodiment of the present invention. It should be understood that Figure 7 only an illustration of one implementation is provided and does not imply any limitation as to the environment in which different embodiments may be implemented. Many modifications may be made to the depicted environment.
[0083] The programs described herein are identified based on their applications in which they are implemented in specific embodiments of the present invention. However, it should be understood that any specific program terms herein are used for convenience only, and thus, the present invention should not be limited to use only in any specific application identified and / or implied by such terms.
[0084] The computer system 700 includes a communication structure 702 that provides communication between a cache 716, a memory 706, a permanent storage device 708, a communication unit 712, and one or more input / output (I / O) interfaces 714. The communication structure 702 can be implemented with any architecture designed to transfer data and / or control information between a processor (such as a microprocessor, communication and network processors, etc.), system memory, peripheral devices, and any other hardware components within the system. For example, the communication structure 702 can be implemented with one or more buses or crossbars.
[0085] The memory 706 and the persistent storage device 708 are computer-readable storage media. In this embodiment, the memory 706 includes random access memory (RAM). Generally, the memory 706 can include any suitable volatile or non-volatile computer-readable storage media. The cache 716 is a fast memory that enhances the performance of the (one or more) computer processors 704 by saving recently accessed data from the memory 706 and data near the accessed data.
[0086] The general intent recognizer 110 (not shown) can be stored in the persistent storage device 708 and the memory 706 for execution by one or more of the corresponding computer processors 704 via the cache 716. In an embodiment, the persistent storage device 708 includes a magnetic hard disk drive. Alternatively or in addition to the magnetic hard disk drive, the persistent storage device 708 can also include a solid state drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or any other computer-readable storage media capable of storing program instructions or digital information.
[0087] The medium used by the persistent storage device 708 can also be removable. For example, a removable hard disk drive can be used for the persistent storage device 708. Other examples include optical discs and disks, thumb drives, and smart cards that are inserted into a drive for transfer to another computer-readable storage media that is also part of the persistent storage device 708.
[0088] In these examples, the communication unit 712 provides communication with other data processing systems or devices. In these examples, the communication unit 712 includes one or more network interface cards. The communication unit 712 can provide communication by using physical and / or wireless communication links. The general intent recognizer 110 can be downloaded to the persistent storage device 708 via the communication unit 712.
[0089] (One or more) I / O interfaces 714 allow for the input and output of data with other devices that can be connected to the client computing device and / or the server computer. For example, the I / O interface 714 can provide a connection to external devices 720 such as a keyboard, keypad, touch screen, and / or some other suitable input device. The external device 720 can also include a portable computer-readable storage media such as, for example, a thumb drive, a portable optical disc or disk, and a memory card. The software and data for implementing embodiments of the present invention (e.g., the general intent recognizer 110) can be stored on such portable computer-readable storage media and can be loaded onto the persistent storage device 708 via the (one or more) I / O interfaces 714. The (one or more) I / O interfaces 714 are also connected to the display 722.
[0090] The display 722 provides a mechanism for displaying data to the user and can be, for example, a computer monitor.
[0091] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium (one or more) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0092] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing storage devices. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any appropriate combination of the foregoing devices. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0093] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0094] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider through the Internet). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to perform aspects of the present invention.
[0095] Aspects of the present invention are described herein with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0096] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable storage medium in which the instructions are stored includes an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0097] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a special purpose hardware-based system that performs the specified functions or acts or combinations of special purpose hardware and computer instructions.
[0099] The description of the various embodiments has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or improvements made to the technology found in the marketplace, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method, comprising: accessing one or more intents and associated entities from a limited amount of speech-to-text training data in a single language; using the accessed one or more intents and associated entities to locate speech-to-text training data in one or more other languages different from the single language, thereby locating the speech-to-text training data in the one or more other languages; pooling together shared parameters of the limited amount of speech-to-text training data in the single language and the located speech-to-text training data in the one or more other languages; using the limited amount of speech-to-text training data in the single language, the located speech-to-text training data in the one or more other languages, and the shared parameters to train a neural network; in response to receiving media containing speech related to an unrecognized language, real-time recognizing at least one known language and domain representing the topic discussed in the received media; using the pooled shared parameters to identify language-independent intents from the recognized languages and adding corresponding labels for the recognized languages and language-independent intents; and adapting the trained neural network to a specific domain within the same language by replacing the language-specific parameters of the trained neural network with new domains and corresponding language-specific layers while keeping the shared layers.
2. The computer-implemented method according to claim 1, further comprising: training the neural network for natural language processing based on the limited amount of speech-to-text training data in the single language and the accessed speech-to-text training data in the one or more other languages.
3. The computer-implemented method according to claim 1, further comprising: learning language-specific constructs of each known language through the language-specific layers of the neural network.
4. The computer-implemented method according to claim 1, wherein, the limited amount of speech-to-text training data includes a single language and intents and associated entities interpreted from the speech in the single language.
5. The computer-implemented method according to claim 1, wherein, the limited amount of speech-to-text training data includes a single language derived from different data sets, the different data sets including a data set collected to model a dialogue state and a data set of speech commands.
6. The computer-implemented method according to claim 1, further comprising: enabling language switching by accessing a public domain and pooling different language data sets having shared commonalities in intents and entities.
7. The computer-implemented method according to claim 1, wherein, the speech-to-text training data represents a set of layers of a neural network including a set of nodes, wherein each layer is connected to other layers in the set of layers via corresponding weighted connections.
8. A computer program product, comprising: one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions comprising: Program instructions for accessing one or more intents and associated entities from a limited amount of speech-to-text training data in a single language; Program instructions for using the one or more accessed intents and associated entities to locate speech-to-text training data in one or more other languages different from the single language to locate the speech-to-text training data in the one or more other languages; Program instructions for pooling shared parameters of the limited amount of speech-to-text training data in the single language and the located speech-to-text training data in the one or more other languages; Program instructions for training a neural network using the limited amount of speech-to-text training data in the single language, the located speech-to-text training data in the one or more other languages, and the shared parameters; Program instructions for, in response to receiving media containing speech related to an unrecognized language, real-time recognizing at least one known language and domain representing the topic discussed in the received media; Program instructions for using the pooled shared parameters to identify language-independent intents from the recognized languages and adding corresponding tags for the recognized languages and language-independent intents; and Program instructions for adapting the trained neural network to a specific domain within the same language by replacing the language-specific parameters of the trained neural network with new domains and corresponding language-specific layers while maintaining the shared layers.
9. The computer program product according to claim 8, wherein, the program instructions stored on the one or more computer-readable storage media further include: Program instructions for training the neural network for natural language processing based on the limited amount of speech-to-text training data in the single language and the accessed speech-to-text training data in the one or more other languages.
10. The computer program product according to claim 8, wherein, the program instructions stored on the one or more computer-readable storage media further include: Program instructions for learning language-specific constructs of each known language through the language-specific layers of the neural network.
11. The computer program product according to claim 8, wherein, the limited amount of speech-to-text training data includes a single language and intents and associated entities interpreted from the speech in the single language.
12. The computer program product according to claim 8, wherein, the limited amount of speech-to-text training data includes a single language derived from different data sets, the different data sets including a data set collected for modeling a dialogue state and a data set of speech commands.
13. The computer program product according to claim 8, wherein, the program instructions stored on the one or more computer-readable storage media further include: Program instructions for enabling language switching by accessing the public domain and pooling different language data sets having shared commonalities in intents and entities.
14. The computer program product according to claim 8, wherein, A speech-to-text training data represents a set of layers of a neural network including a set of nodes, wherein each layer is connected to other layers in the set of layers via respective weighted connections.
15. A computer system, comprising: one or more computer processors; one or more computer-readable storage media; and program instructions stored on the one or more computer-readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising: program instructions for accessing one or more intents and associated entities from a limited amount of speech-to-text training data in a single language; program instructions for using the accessed one or more intents and associated entities to locate speech-to-text training data in one or more other languages different from the single language to locate the speech-to-text training data in the one or more other languages; program instructions for pooling together shared parameters of the limited amount of speech-to-text training data in the single language and the located speech-to-text training data in the one or more other languages; program instructions for training a neural network using the limited amount of speech-to-text training data in the single language, the located speech-to-text training data in the one or more other languages, and the shared parameters; program instructions for, in response to receiving media containing speech related to an unrecognized language, real-time recognizing at least one known language and domain representing a topic discussed in the received media; program instructions for using the pooled shared parameters to recognize language-independent intents from the recognized languages and adding corresponding tags for the recognized languages and the language-independent intents; and program instructions for adapting the trained neural network to a specific domain within the same language by replacing language-specific parameters of the trained neural network with new domains and corresponding language-specific layers while maintaining the shared layers.
16. The computer system according to claim 15, wherein the program instructions stored on the one or more computer-readable storage media further comprise: program instructions for training the neural network for natural language processing based on the limited amount of speech-to-text training data in the single language and the accessed speech-to-text training data in the one or more other languages.
17. The computer system according to claim 15, wherein the program instructions stored on the one or more computer-readable storage media further comprise: program instructions for learning language-specific constructs of each known language through language-specific layers of the neural network.
18. The computer system according to claim 15, wherein the limited amount of speech-to-text training data includes a single language and intents and associated entities interpreted from the speech in the single language.
19. The computer system according to claim 15, wherein the limited amount of speech-to-text training data includes a single language derived from different data sets, the different data sets including a data set collected to model a dialogue state and a data set of speech commands.
20. The computer system according to claim 15, wherein, the program instructions stored on the one or more computer-readable storage media further comprise: program instructions for enabling language switching by accessing a common domain and aggregating different language datasets having shared commonalities in intents and entities.
21. A computer system comprising modules for performing the steps of the method according to any one of claims 1 to 7, respectively.