Compression of Neural Networks for Natural Language Understanding

Through the training and iterative optimization process, the large-scale dialogue system model was successfully reduced to a scale suitable for embedded devices, solving the problem of low-efficiency operation of dialogue systems in low-power devices, and achieving efficient natural language understanding tasks.

CN112487783BActive Publication Date: 2025-06-17ORACLE INT CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010907744.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-24
Filing Date
2020-09-02
Publication Date
2025-06-17
Estimated Expiration
2040-09-02

AI Technical Summary

Technical Problem

Large complex models used in existing dialogue systems are slow to execute at runtime and consume a lot of storage resources, making it difficult to embed in low-power and resource-deprived devices.

Method used

By training a smaller compressed neural network model, the tagged model is used to label unlabeled session data, thereby generating tagged data for training the compressed model. The method includes generating the tagged model using a proxy model and gradually reducing the complexity and size of the model through multiple iterative training.

Benefits of technology

It realizes the reduction of neural network models in natural language understanding tasks to a scale suitable for execution on embedded systems, while maintaining high performance and reducing storage and computing requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112487783B_ABST
    Figure CN112487783B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the compression of a neural network for natural language understanding. A model for a natural language understanding task is generated based on data of tokens generated by a token model. The model for the natural language understanding task is smaller than the token model (i.e., has lower computational requirements and memory requirements than the combined model), but has substantially the same performance as the token model. In some cases, the token model can be generated based on a large pre-trained model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit and priority of U.S. Application No. 62 / 899,650, filed on September 12, 2019, entitled "COMPRESSING RECURRENT NEURAL NETWORKS USED IN NATURAL LANGUAGE UNDERSTANDING", and U.S. Application No. 16 / 938,098, filed on July 24, 2020, entitled "COMPRESSING NEURAL NETWORKS FOR NATURAL LANGUAGE UNDERSTANDING", the entire contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field

[0003] The present disclosure generally relates to dialogue systems and machine learning. More specifically but without limitation, the present disclosure describes techniques for using a relatively large model to generate a large amount of tokenized conversation data for training a relatively small model for performing natural language understanding tasks. Background Art

[0004] Now, more and more devices enable users to directly interact with the devices using voice or spoken speech. For example, a user can speak to such a device in natural language, where the user can ask a question or make a statement requesting an action to be performed. In response, the device performs the requested action or uses voice output to respond to the user's question. Since directly interacting using voice is a more natural and intuitive way for humans to communicate with their surrounding environment, the popularity of such voice - based systems is growing at a great rate.

[0005] The ability to interact with a device using spoken speech is facilitated by a dialogue system (sometimes also referred to as a chatbot or digital assistant) that can be in the device. Dialogue systems typically use machine learning models to perform a series of dialogue processing tasks. Dialogue processing models are generally large and complex models. Such complex models typically execute slowly at runtime and can be very large (e.g., pre - trained models may require a large amount of storage resources and processing resources, sometimes to the extent that such models must be accommodated on a supercomputer or multiple servers). This makes it difficult to incorporate such dialogue systems into low - power and resource - constrained devices (e.g., kitchen appliances, lighting fixtures, etc.). Summary of the Invention

[0006] The present disclosure generally relates to natural language understanding. More specifically, techniques for compressing recurrent neural networks for use in natural language understanding tasks are described. Various embodiments are described herein, including methods, systems, non-transitory computer-readable storage media storing programs, code, or instructions executable by one or more processors, and the like.

[0007] In certain embodiments, a method for training a neural network to perform a natural language understanding task includes: obtaining first labeled session data; using the first labeled session data to train a first neural network; obtaining first unlabeled session data; performing the trained first neural network on the first unlabeled session data to label the first unlabeled session data, thereby generating second labeled session data; and using the second labeled session data to train a second neural network for performing the natural language understanding task.

[0008] In some aspects, before training the first neural network, the method further includes: obtaining second unlabeled session data; using the second unlabeled session data to train a third neural network to perform a proxy task; and generating the first neural network based on the third neural network. In some aspects, generating the first neural network based on the third neural network includes constructing the first neural network to include at least a portion of the third neural network. In some aspects, the third neural network is a component of the first neural network. In some aspects, the proxy task includes a prediction language task.

[0009] In some aspects, the natural language understanding task includes one or more of semantic parsing, intent classification, or named entity classification. In some aspects, training the second neural network includes: for a first training input of the second labeled session data, outputting a predicted output by the second neural network; calculating a loss that measures an error between the predicted output and a first label associated with the first training input; based on the loss, calculating an updated value of a first set of parameters of the second neural network; and updating the second neural network by changing the value of the first set of parameters to the updated value.

[0010] Embodiments further include systems and computer-readable memories configured to perform the methods described herein.

[0011] The foregoing, together with other features and embodiments, will become more apparent after referring to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a simplified block diagram showing a dialogue system according to certain embodiments.

[0013] Figure 2 is a simplified block diagram showing a model compression system according to certain embodiments.

[0014] Figure 3 is a flowchart showing a method for compressing a neural network for natural language understanding tasks according to certain embodiments.

[0015] Figure 4 is a flowchart showing additional techniques for compressing a neural network for natural language understanding tasks according to certain embodiments.

[0016] Figure 5 depicts a simplified diagram of a distributed system for implementing embodiments.

[0017] Figure 6 is a simplified block diagram of a cloud-based system environment according to certain embodiments, in which various storage-related services can be provided as cloud services.

[0018] Figure 7 shows an exemplary computer system that can be used to implement certain embodiments. DETAILED DESCRIPTION

[0019] In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of certain embodiments. However, it will be apparent that the various embodiments can be practiced without these specific details. The drawings and description are not intended to be restrictive. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as exemplary is not necessarily to be construed as more preferred or advantageous than other embodiments or designs.

[0020] A voice-enabled system capable of conversing with a user via voice input and voice output can take various forms. For example, such a system can be provided as a stand-alone device, as a digital assistant or virtual assistant, as a service with voice capabilities, etc. In each of these forms, the system is capable of receiving voice input or speech input, understanding the input, generating a response or taking an action in response to the input, and outputting the response using voice output. In certain embodiments, the dialogue function in such a voice-enabled system is provided by a dialogue system or infrastructure ("dialogue system"). Various terms such as chatbot or chatbot system, digital assistant, etc. are used to refer to such a dialogue system.

[0021] The dialogue system can be operated using a combination of processes. Such processes can include natural language understanding (NLU), which can be used to process text utterances to generate meaning representations. Natural language understanding tasks include tasks of converting speech data (e.g., text utterances) into data suitable for computer processing. Natural language understanding tasks include tasks such as semantic parsing, named entity recognition and classification, slot filling, part-of-speech tagging, sentiment analysis, word sense disambiguation, lemmatization, word segmentation, and sentence boundary detection.

[0022] As noted above, machine learning models for dialogue processing tasks are typically large or complex models. Generally, the capabilities of more complex models are stronger than those of simpler models because more complex models can understand more difficult languages and can respond in more complex ways. However, compared with simpler models, complex models are also larger and use more computing resources at runtime. Although it is desirable to perform dialogue processing tasks on small embedded devices, it may be simply infeasible to run complex models on such small devices.

[0023] In some embodiments, the machine learning model used to perform natural language processing tasks in the dialogue system is "compressed" to create a "compressed model" that can be small enough in terms of storage and processing requirements to be executed on an embedded device. The first machine learning model is trained from labeled training data. In some embodiments, this first machine learning model is a neural network ("the first neural network", also referred to as the "labeled model"). Then, this labeled model is used as a teacher model to train a simpler and faster model. The simpler and faster model can be a neural network with a smaller profile (the "second neural network", also referred to herein as the "compressed model") that is ultimately used at runtime. In this way, the compressed model reduces the size of the model and the amount of computing power required to execute the model, as well as the memory required to store the model.

[0024] In some embodiments, deep learning neural network models can be used. Generally, the complexity of such models is characterized by the number of layers in the neural network and the number of nodes in the neural network. These pre-trained models can have 12 layers or 24 layers, and each layer can have millions of nodes in each layer. Simpler models can have 2 layers or 3 layers, with hundreds or thousands of nodes in each layer, and can be an order of magnitude or more smaller than the pre-trained models.

[0025] In some embodiments, labeled training data is used to train a labeling model for making predictions. These predictions can be used to generate a large amount of labeled data that can be used to create a compression model. A complex labeling model can be trained on a relatively small data corpus, whereas a simpler compression model may require more data to achieve the same quality level. Thus, the labeling model can be used to generate a large amount of labeled training data to teach the compression model to perform natural language understanding tasks, effectively transferring the knowledge of the complex labeling model to the smaller compression model. As an example, a dialogue system can access a large number of call logs (e.g., unlabeled data). The labeling model can be used to label the unlabeled call log data and use the call log data to train the compression model.

[0026] The compression model can be compressed to fit within certain constraints. For example, for natural language understanding to be used in an alarm clock, the chip will have certain memory requirements and speed requirements for it. The compression model can be customized according to such hardware requirements. The compression model for natural language understanding tasks is smaller than the labeling model (i.e., has lower computational and memory requirements than the labeling model), but has substantially the same performance as the labeling model.

[0027] In some embodiments, the labeling model itself can be generated based on another model. A model for generating the labeling model can be trained to perform a proxy task related to the ultimate task to be performed. This process is called pre-training. For example, a very large "proxy model" can be trained to predict the next word in a sentence or the next sentence in a text. Then, this model is used to generate a labeling model that is trained to make related predictions (e.g., for semantic parsing or sentiment analysis). This can result in the labeling model having greatly improved performance compared to a basic model generated without any pre-training. Pre-trained models are very useful for tasks such as natural language processing because these models can be trained on vast datasets and learn a very large vocabulary. However, since pre-trained models incorporate many different types of components, such models tend to be very large and use slow components. By using a pre-trained model with a large amount of built-in information, a complex labeling model can be trained on a relatively small corpus of labeled training data. Then, by compressing such a model or its derivatives, a large amount of knowledge can be imparted to a compact compression model.

[0028] Figure 1Shows an example of a dialogue system 100 according to some embodiments. The dialogue system 100 is configured to receive voice input or speech input 104 (also referred to as a speech utterance) from a user 102. Then, the dialogue system 100 can interpret the voice input. The dialogue system 100 can maintain a dialogue with the user 102 and can perform or cause one or more actions to be performed based on the interpretation of the voice input. The dialogue system 100 can prepare an appropriate response and output the response to the user using voice output or speech output.

[0029] In certain embodiments, the processing performed by the dialogue system is implemented by a pipeline of components or subsystems, the pipeline of components or subsystems including a voice input component 105, a wake word detection (WD) subsystem 106, an automatic speech recognition (ASR) subsystem 108, an integrated shared dictionary 109, a natural language understanding (NLU) subsystem 110 including a named entity recognizer (NER) subsystem 112 and a semantic parser subsystem 114, a dialogue manager (DM) subsystem 116, a natural language generator (NLG) subsystem 118, a text-to-speech (TTS) subsystem 120, and a voice output component 124. The subsystems listed above can be implemented only in software (e.g., using code, programs, or instructions executable by one or more processors or cores), in hardware, or in a combination of hardware and software. In certain implementations, one or more of the subsystems can be combined into a single subsystem. Additionally or alternatively, in some implementations, the functions described herein performed by a particular subsystem can be implemented by multiple subsystems.

[0030] The voice input component 105 includes hardware and software configured to receive the voice input 104. In some instances, the voice input component 105 can be part of the dialogue system 100. In some other instances, the voice input component 105 can be separate from the dialogue system 100 and communicatively coupled to the dialogue system 100. The voice input component 105 can include, for example, a microphone coupled to software configured to digitize the voice input and transmit the voice input to the wake word detection subsystem 106.

[0031] The Wake Word Detection (WD) subsystem 106 is configured to listen for and monitor an audio input stream for obtaining an input corresponding to a special sound or word or a set of words (referred to as a wake word). When detecting a wake word configured for the dialogue system 100, the WD subsystem 106 is configured to activate the ASR subsystem 108. In some embodiments, the user may be provided with the ability to activate or deactivate the WD subsystem 106 (e.g., by speaking a wake word for pressing a button). When the WD subsystem 106 is activated (or operating in an activation mode), the WD subsystem 106 is configured to continuously receive the audio input stream and process the audio input stream to identify an audio input or voice input corresponding to the wake word. When an audio input corresponding to the wake word is detected, the WD subsystem 106 activates the ASR subsystem 108.

[0032] As described above, the WD subsystem 106 activates the ASR subsystem 108. In some embodiments of a voice-enabled system, mechanisms other than a wake word may be used to trigger or activate the ASR subsystem 108. For example, in some embodiments, a push button on the device may be used to trigger the ASR subsystem 108 to process without a wake word. In such an embodiment, the WD subsystem 106 may not be provided. When the push button is pressed or activated, the voice input received after the button activation is provided to the ASR subsystem 108 for processing. In some embodiments, the ASR subsystem 108 may be activated when receiving an input to be processed.

[0033] The ASR subsystem 108 is configured to receive and monitor an oral voice input after a trigger signal or wake signal (e.g., a wake signal may be sent by the WD subsystem 106 when detecting a wake word in the voice input, a wake signal may be received when activating a button, etc.) and convert the voice input into text. As part of the processing of the ASR subsystem 108, the ASR subsystem 108 performs speech-to-text conversion. The oral voice input or voice input may be in the form of natural language, and the ASR subsystem 108 is configured to generate a corresponding natural language text in the language of the voice input. Then, the text generated by the ASR subsystem is fed to the NLU subsystem 110 for further processing. The voice input received by the ASR subsystem 108 may include one or more words, phrases, clauses, sentences, questions, etc. The ASR subsystem 108 is configured to generate a text utterance for each oral clause and feed the text utterance to the NLU subsystem 110 for further processing.

[0034] The NLU subsystem 110 receives the text generated by the ASR subsystem 108. The text received by the NLU subsystem 110 from the ASR subsystem 108 may include text utterances corresponding to spoken words, phrases, clauses, etc. The NLU subsystem 110 translates each text utterance (or series of text utterances) into the corresponding logical form of each text utterance. The NLU subsystem 110 may further use the information passed from the ASR subsystem 108 to quickly obtain information from the integrated shared dictionary 109 for use in generating the logical form, as described herein.

[0035] In some embodiments, the NLU subsystem 110 includes a named entity recognizer (NER) subsystem 112 and a semantic parser (SP) subsystem 114. The NER subsystem 112 receives text utterances as input, identifies the named entities in the text utterances, and uses the information associated with the identified named entities to tag the text utterances. Then, the tagged text utterances are fed into the SP subsystem 114, which is configured to generate the logical form of each tagged text utterance (or series of tagged text utterances). The logical form generated for a discourse may identify one or more intents corresponding to the text utterance. The intent of the discourse identifies the purpose of the text utterance. Examples of intents include "order pizza" and "find directions". The intent may identify, for example, the action to be performed as requested. In addition to the intent, the logical form of the text utterance may also identify the slots (also referred to as parameters or arguments) for the identified intent. For example, for the speech input "I’d like to order a large pepperoni pizza with mushrooms and olives", the NLU subsystem 110 may identify the intent to order pizza. The NLU subsystem may also identify and fill the slots, e.g., pizza_size (filled with large) and pizza_toppings (filled with mushrooms and olives). The NLU subsystem may use machine learning-based techniques, rules (which may be domain-specific), or a combination of both to generate the logical form. Then, the logical form generated by the NLU subsystem 110 is fed into the DM subsystem 116 for further processing.

[0036] In some embodiments, the NLU subsystem 110 includes a compressed model 115. The compressed model 115 is a relatively small model for performing natural language understanding tasks, and the compressed model 115 has been compressed using the techniques described herein. In some embodiments, the compressed model 115 is a neural network. The compressed model 115 can be a semantic parser (e.g., part of the semantic parser subsystem 114). Alternatively, or additionally, the compressed model can be a named entity recognizer (e.g., part of the NER subsystem 112). Alternatively, or additionally, the compressed model 115 can be a separate subsystem for performing natural language understanding tasks such as slot filling, part-of-speech tagging, sentiment analysis, word sense disambiguation, etc. The techniques for generating the compressed model 115 are described in further detail below with respect to Figures 2 to 4 The techniques for generating the compressed model 115 are described in further detail below with respect to

[0037] The DM subsystem 116 is configured to manage a conversation with a user based on the logical form received from the NLU subsystem 110. As part of the dialogue management, the DM subsystem 116 is configured to track the dialogue state, initiate the execution of one of a plurality of actions or tasks or perform one of the plurality of actions or tasks itself, and determine how to interact with the user. These actions can include, for example, querying one or more databases, generating execution results, and other actions. For example, the DM subsystem 116 is configured to interpret the intent identified in the logical form received from the NLU subsystem 110. Based on this interpretation, the DM subsystem 116 can initiate one or more actions that the DM subsystem 116 interprets as requested by the voice input provided by the user. In certain embodiments, the DM subsystem 116 performs dialogue state tracking based on the current and past voice inputs and based on a set of rules (e.g., dialogue policies) configured for the DM subsystem 116. These rules can specify different dialogue states, the conditions for transitioning between states, the actions to be performed when in a particular state, etc. These rules can be domain-specific. In certain embodiments, machine learning-based techniques (e.g., machine learning models) can also be used. In some embodiments, a combination of rules and machine learning models can be used. The DM subsystem 116 also generates responses to be passed back to the user involved in the conversation. These responses can be based on the actions initiated by the DM subsystem 116 and the results of the actions. The responses generated by the DM subsystem 116 are fed to the NLG subsystem 118 for further processing.

[0038] The NLG subsystem 118 is configured to generate natural language text corresponding to the response generated by the DM subsystem 116. The text can be generated in a form such that the text can be converted to speech by the TTS subsystem 120. The TTS subsystem 120 receives the text from the NLG subsystem 118 and converts each text in the text to speech or voice audio, and then the speech or voice audio can be output to the user via the audio or voice output component 124 of the dialogue system (e.g., a speaker or a communication channel coupled to an external speaker). In some instances, the voice output component 124 can be part of the dialogue system 100. In some other instances, the voice output component 124 can be separate from the dialogue system 100 and communicatively coupled to the dialogue system 100.

[0039] As described above, the various subsystems of the dialogue system 100 work together to provide the functions that enable the dialogue system 100 to receive voice input 104, respond using voice output 122, and maintain a dialogue with the user using natural language speech. The various subsystems described above can be implemented using a single computer system or using multiple computer systems that work together. For example, for a device implementing a voice-enabled system, the subsystems of the dialogue system 100 described above can be implemented entirely on the device with which the user interacts. In some other embodiments, some components or subsystems of the dialogue system 100 can be implemented on the device with which the user interacts, while other components can be implemented away from the device, possibly on some other computing device, platform, or server.

[0040] As described above, in certain embodiments, the dialogue system 100 can be implemented using a pipeline of subsystems. In some embodiments, one or more of the subsystems can be combined into a single subsystem. In certain embodiments, the functions provided by a particular subsystem can be provided by multiple subsystems. A particular subsystem can also be implemented using multiple subsystems.

[0041] In certain embodiments, machine learning techniques can be used to implement one or more functions of the dialogue system 100. For example, supervised machine learning techniques (e.g., those implemented using neural networks (e.g., deep neural networks)) can be used to implement one or more functions of the dialogue system 100. As an example, a neural network trained to perform the ASR function being performed can be provided, and such a trained model can be used by the ASR subsystem 108 for processing by the ASR subsystem. Such a neural network implementation can take voice input as input and output text utterances to the NLU subsystem. Machine learning-based models can also be used by other subsystems of the dialogue system 100.

[0042] Figure 2Shows an example of a model compression system 200 according to some embodiments. The model compression system 200 is configured to generate a compressed model 115 (e.g., Figure 1 the compressed model 115 of the dialogue system 100) using the labeled model 206 and the unlabeled session data 222 and labeled session data 224 in the training database 220. In some embodiments, the surrogate model 202 is also used to generate the labeled model 206. In some implementations, the model compression system 200 is a subsystem of the dialogue system 100. Alternatively, the model compression subsystem can be separated from the dialogue system and perform offline training and preparation of the compressed model 115, which is pushed to the dialogue system 100 for execution. The subsystems listed above can be implemented only in software (e.g., using code, programs, or instructions executable by one or more processors or cores), in hardware, or in a combination of hardware and software. In certain implementations, one or more of the subsystems can be combined into a single subsystem. Additionally or alternatively, in some implementations, the functions described herein performed by a particular subsystem can be implemented by multiple subsystems.

[0043] In some embodiments, the training database 220 is a storage unit and / or device for storing training data (e.g., a file system, a database, a collection of tables, or other storage mechanism). The training database 220 can include multiple different storage units and / or devices. The training database 220 can be local to the model compression system 200 (e.g., local storage), and / or connected to the model compression system 200 via a network (e.g., cloud storage).

[0044] In some embodiments, the training data stored in the training database 220 is used to train the surrogate model 202, the labeled model 206, and / or the compressed model 115. The training data includes labeled session data 224 and unlabeled session data 222. The unlabeled session data 222 can include session data such as chat logs, phone records, transcripts of speeches or meetings, movie scripts, etc. The unlabeled session data 222 can come from sources where a large amount of data is available. In some cases, the training database 220 can further store additional unlabeled data, such as data obtained from an enterprise database or the Internet (e.g., a set of web pages scraped from the Web, an online encyclopedia, etc.). This can be used to even further enrich the training data. The labeled session data 224 is session data that has been labeled or annotated for training one or more machine learning models (e.g., call log data, etc.). For example, the labeled session data used to train a model to perform named entity classification is annotated such that named entities (e.g., John Goodman and Paris) are labeled with named entity types (e.g., person and city).

[0045] In some embodiments, the tagging model 206 is a relatively large model for performing natural language understanding tasks. In some implementations, the tagging model 206 is a neural network and is referred to as the first neural network. For example, the tagging model 206 can be a recurrent neural network with 12 layers or 24 layers. The tagging model is trained to perform natural language understanding tasks such as semantic parsing or named entity classification. Although the tagging model may not actually be used at runtime, the tagging model can be trained to perform the tasks that are ultimately desired to be performed. The tagging model 206 is trained using all or a subset of the tagged session data 224.

[0046] The tagging model 206 is used to tag all or a subset of the untagged session data 222 in the training database 220 to generate the tagged session data 224. For example, for a training model that performs semantic parsing, the tagging model 206 is executed to parse text clauses in the untagged session data 222 to generate a meaning representation. This meaning representation can be linked to the corresponding part of the untagged data to generate the tagged session data 224.

[0047] In some embodiments, the compression model 115 is a relatively small model for performing natural language understanding tasks. In some implementations, the compression model 115 is a neural network and is referred to as the second neural network. For example, the compression model 115 can be a recurrent neural network with 2 layers or 4 layers. The compression model 115 is trained to perform natural language understanding tasks such as semantic parsing or named entity classification. The compression model 115 can be trained to perform the same tasks as those for which the tagging model 206 is trained to perform. The compression model 115 is trained using the tagged session data 224 generated by the tagging model 206.

[0048] In some implementations, multiple training datasets are used to train multiple machine learning models, and the multiple machine learning models are used to generate the compression model 115 for use at runtime. The training dataset used to train the tagging model 206 is referred to as the first tagged session data in the tagged session data 224. The tagging model 206 tags the first untagged session data in the untagged session data 222. By tagging the first untagged session data, the tagging model generates the second tagged session data in the tagged session data 224, and this second tagged session data is used to train the compression model 115. In some implementations, a third model (the proxy model 202) is also used to generate the tagging model 206. This proxy model 202 can be trained on additional untagged data, which can include the second untagged session data in the untagged session data 222.

[0049] In some embodiments, the proxy model 202 is a large model (e.g., a pre-trained model) for performing natural language tasks. In some implementations, the proxy model 202 is a neural network and is referred to as the third neural network. For example, the proxy model 202 can be a recurrent neural network with 24 layers or more. The proxy model 202 can be trained to perform a proxy task for which a very large amount of data is available. For example, given all the other words in the same sentence or paragraph in which a word appears, the proxy model 202 can be trained to predict a randomly selected word in a sentence from the web. Since the proxy model 202 is trained on a very large amount of different data (e.g., an extremely large vocabulary), the proxy model 202 can be very large.

[0050] In some embodiments, the tagging model 206 is a combined model generated based on the proxy model 202. Using the proxy model 202 to generate the tagging model 206 improves the accuracy of the predictions made by the tagging model 206 because the proxy model 202 enables the tagging model 206 to generalize from task-specific training data in a more reliable manner.

[0051] Specific examples of large models used as the proxy model 202 and / or the tagging model 206 include bidirectional transducer models (e.g., Bidirectional Encoder Representations from Transformers (BERT)) and deep bidirectional language models (e.g., Embeddings from Language Models (ELMo)). BERT is a bidirectional transducer model that learns the contextual relationships between words or parts of words in text. (J. Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” available from https: / / arxiv.org / pdf / 1810.04805.pdf (2019)). ELMo is a deeply contextualized word representation that uses a deep pre-trained neural network trained on a large dataset. The ELMo model characterizes word usage such as syntax and semantics, and how word usage varies across language contexts. (M. Peters et al., “Deep Contextualized Word Representations,” available from https: / / Arxiv.org / abs / 1802.05365 (2017)).

[0052] Figure 3 is a flowchart showing a method 300 for compressing a neural network for natural language understanding tasks according to certain embodiments.Figure 3 The processes depicted in Figure 3 can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, in hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., on a memory device). Figure 3 The methods presented in Figure 3 and described below are intended to be illustrative and not restrictive. Although Figure 3 various processing steps are depicted as occurring in a particular sequence or order, this is not intended to be restrictive. In certain alternative embodiments, the steps can be performed in a different order or some steps can also be performed in parallel. In certain embodiments, Figure 3 the processes depicted in Figure 3 can be performed by Figure 2 the model compression system 200 of Figure 2 .

[0053] At 302, the model compression system obtains first-labeled session data. At an initial time, the labeled session data can be manually or automatically annotated with data specific to a desired task (e.g., intent or named entity). The first-labeled session data can be a small data set or a large data set. The model compression system can obtain the first-labeled session data by downloading the data. Alternatively, or additionally, the first-labeled session data can be generated and stored on the model compression system itself. As described above with respect to Figure 2 Figure 2 , the first-labeled session data can be sourced from sources such as call center logs, chat logs, etc., and can be stored in the training database 220 of the model compression system 200.

[0054] At 304, the model compression system uses the first-labeled session data to train a labeling model. In some implementations, the labeling model is a neural network (e.g., Figure 2 the labeling model 206 of Figure 2 ) for performing natural language understanding tasks. The labeling model is also referred to herein as the first neural network, although it should be understood that in some implementations, other types of machine learning models can be selected as the labeling model.

[0055] The model compression system can use labeled data suitable for the natural language understanding task of interest (e.g., for semantic parsing, named entity classification, etc.) to train the labeling model. For example, session data that has been labeled with a logical form (e.g., a set of utterances labeled with a corresponding logical form) can be used to train the labeling model. As another example, session data that has been labeled with named entity types can be used to train the labeling model such that the labeling model is trained to classify named entities. In some embodiments, training the labeling model includes using a surrogate model 202 that can be used to generate the labeling model, as further described below with respect to Figure 4 Figure 4 .

[0056] In some embodiments, loss minimization techniques are used to train the tagging model. For example, the first tagged session data includes various training inputs. For a first training input of the first tagged session data, the tagging model is executed and a predicted output is produced. As a specific example, the tagging model receives a sample utterance (e.g., "Which country is Paris in?") as an input (e.g., into a first neural network). The tagging model is configured to perform a named entity classification task and output a predicted named entity for the sample utterance (Paris = city). The model compression system calculates a loss that measures the error between the predicted output and a first label associated with the first training input (e.g., tagging the named entity type of Paris as a city). Different types of loss functions can be applicable to different model architectures and end goals. Loss functions include the cross-entropy function (see, e.g., Nielson, "Neural Networks and Deep Learning", chapter 3.2 (last updated July 2020)) and the mean absolute error loss function (see, e.g., mean absolute error in: Sammut C., Webb G.I. (eds.), Encyclopedia of Machine Learning, Springer, Boston, Massachusetts (2011)). The model compression system can use the calculated loss to compute an updated value for a first set of parameters of the tagging model. Then, the tagging model is updated by changing the values of the first set of parameters to the updated values.

[0057] At 306, the dialogue system uses a trained tagging model to tag the first untagged session data to generate second tagged session data. The system can use the trained tagging model to tag untagged session data. The untagged session data can include, for example, call logs, text chat logs, etc. Other examples of untagged session data include emails and other written documents. Using the trained tagging model to tag the untagged data can include using the trained tagging model to perform natural language processing tasks. The input to the tagging model can include specific session data elements (e.g., the utterance "How do I get to the beach?"). The trained tagging model can perform natural language tasks to output a value, and this value can be used to tag the session data. Continuing with the above example, the trained tagging model has been trained to perform intent classification. The output of the trained tagging model for the input utterance "How do I get to the beach" is Directions_To. The model compression system uses Directions_To to tag the utterance "How do I get to the beach?". The result is tagged training data - utterance: "How do I get to the beach?" + intent: "Directions_To". This tagged training data ("second tagged session data") can be saved to the same data set in the training database 220 that includes the first untagged session data used at 302, or to a different data set in the training database 220.

[0058] At 308, the system uses the second tagged session data to train a compression model for natural language understanding tasks. In some embodiments, the compression model is a neural network (e.g., a second neural network). The compression model is a model that is significantly smaller than the tagging model. The compression model is trained to perform natural language understanding tasks, which can be the same type of natural language understanding tasks that the tagging model has been trained to perform (e.g., named entity classification, semantic parsing, etc.). Because the compression model is significantly smaller than the tagging model, the computational requirements and memory requirements of the NLU model are significantly lower than the computational requirements and memory requirements of the compression model.

[0059] In some embodiments, a loss minimization technique is used to train the compression model. This can be performed in a manner similar to the manner described above for training the tagging model at 304. For a first training input of the second tagged session data, the compression model is executed. The compression model outputs a predicted output. The model compression system calculates a loss that measures the error between the predicted output and a first label associated with the first training input. The model compression system can use the calculated loss to calculate an updated value for a first set of parameters of the compression model. Then, the compression model is updated by changing the values of the first set of parameters to the updated values.

[0060] Figure 4 Illustrated for generating based on a surrogate modelFigure 3 Steps of the tagging model. In some embodiments, Figure 4 processing may be initially performed to generate Figure 3 the tagging model. Figure 4 The processing depicted in can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, in hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., on a memory device). Figure 4 The methods presented in and described below are intended to be illustrative and not restrictive. Although Figure 4 depicts various processing steps occurring in a particular sequence or order, this is not intended to be restrictive. In certain alternative embodiments, the steps may be performed in a different order or some steps may be performed in parallel. In certain embodiments, Figure 4 the processing depicted in can be performed by Figure 2 the model compression system 200. Figure 4 The processing can be performed as a preamble to the processing of Figure 3

[0061] At 402, the model compression system obtains second unlabeled session data. The second unlabeled session data can be a large dataset, such as a web crawl, a large group of websites, and / or a large group of chat logs. The model compression system can obtain the second unlabeled session data by downloading the data to the training database 220. Alternatively, or additionally, the model compression system can remotely access the unlabeled session data (e.g., via an API). In some cases, the second unlabeled session data is the same as the first unlabeled session data used in the processing of Figure 3 Alternatively, in some embodiments, different datasets are used (e.g., the first unlabeled session data and the second unlabeled session data are different datasets). For example, if an agent model is to be trained on a very large unstructured dataset, it may be appropriate to use different datasets, while the first unlabeled session data to be tagged by the tagging model is a more moderately sized structured dataset.

[0062] At 404, the model compression system uses the second unlabeled session data to train an agent model to perform an agent task. In some embodiments, the agent model is a neural network (e.g., a third neural network). In some implementations, the agent model is a large pre-trained model such as BERT or ELMo. The agent model is trained to perform an agent task. An agent task is a task that teaches the agent model information that will be useful for training the tagging model and / or the compression model. For example, in Figure 3The token models used at boxes 304 and 306 will be used for semantic parsing. The proxy model is trained to predict the next word in a sentence such that the proxy model can learn to understand language while using a wide-ranging dataset. Thus, the proxy task can include predicting language tasks. In some embodiments, the token model can have the same or very similar architecture as the proxy model. For example, both models can be recurrent neural networks with a relatively large number of layers (e.g., 12 or 24 layers).

[0063] At 406, the model compression system generates a token model based on the proxy model (e.g., the first neural network used in the Figure 3 processing). The model compression system can use the proxy model to build the token model.

[0064] In some embodiments, building a token model based on the proxy model can include building the token model (e.g., the first neural network) to include at least a portion of the proxy model (e.g., the third neural network). For example, the proxy model can be fine-tuned to perform a specific natural language processing task. For example, for a proxy model that is a neural network, the model compression system can replace the input and output layers to adapt to the task to be performed. As a specific example, the proxy model can be adapted to accept an embedding representing an utterance as input and return a value of interest for natural language processing (e.g., logical form, named entity type, sentiment, part-of-speech type, etc.) as output. As described at Figure 3 boxes 302 and 304, the fine-tuning can be done by training the proxy model using labeled training data.

[0065] In some embodiments, the proxy model is a component of the tagging model. For example, the proxy model can be modified to output values of the appropriate type as described above. This model is then combined with one or more additional submodels that are trained to perform equivalent natural language processing tasks. For example, two named entity recognizers can be constructed, or four semantic parser models can be constructed, etc. Then, appropriate techniques (such as ensemble learning (see, e.g., Zhou et al., “Ensembling Neural Networks: Many Could Be Better Than All”, Artificial Intelligence, Vol. 137, pp. 239 - 263 (2002)) or siamese neural networks (see, e.g., Bromley et al., “Signature Verification using a “Siamese””, Time Delay Neural Network (1994))) can be used to combine the models.

[0066] Alternatively, or additionally, the proxy model can be used as a component of the tagging model without modifying the output of the proxy model. This approach can be useful for using the proxy model to perform a different task than the tagging model. For example, the proxy model can be a language model that predicts the next word in the input and works as a component of a tagging model that can be a semantic parser that translates a user request into a logical form. For example, the prediction component can be used to accelerate the semantic parsing performed by the tagging model.

[0067] Using an initial pre - trained initial model can improve the accuracy of the predictions made by the tagging model because the initial model enables the tagging model to generalize from task - specific training data in a more reliable way. For example, the tagging model can generalize appropriately to words that do not appear in the task - specific training data because those words appear in the large amount of training data used to pre - train the initial model. However, because the tagging model incorporates the initial model, and high - performance pre - trained models can be very large and have high computational and memory requirements, the tagging model may also have high computational and memory requirements. Therefore, as described above with respect to Figure 3 the data tagged by the tagging model can be used to generate a smaller NLU model.

[0068] The techniques described herein have several advantages. Natural language understanding models are typically too large such that natural language understanding models require multiple servers to execute. In some cases, it is desirable to embed NLU functionality into a personal device such as a clock radio or a smart watch, in which case the model must be much smaller. The compression techniques described herein can be used to reduce the NLU model to a size suitable for execution on such an embedded system. Even for NLU tasks running in the cloud, it is desirable to reduce the memory usage requirements and the computational requirements, which can be achieved using these techniques. The techniques described herein can start with a very large pre-trained model with a large amount of language knowledge and impart that knowledge onto a small compressed model, thus saving significantly on the memory requirements, storage requirements, and the time for executing the model without sacrificing accuracy.

[0069] The infrastructure described above can be implemented in a variety of different environments including cloud environments (which can be various types of clouds including private cloud environments, public cloud environments, and hybrid cloud environments), on-premises environments, hybrid environments, and the like.

[0070] Figure 5 A simplified diagram of a distributed system 500 for implementing an embodiment is depicted. In the illustrated embodiment, the distributed system 500 includes one or more client computing devices 502, 504, 506, and 508 coupled to a server 512 via one or more communication networks 510. The client computing devices 502, 504, 506, and 508 can be configured to execute one or more applications.

[0071] In various embodiments, the server 512 can be adapted to run one or more services or software applications capable of rationalizing the recognition of an intent from a voice input.

[0072] In certain embodiments, the server 512 can also provide other services or software applications which can include non-virtual environments and virtual environments. In some embodiments, these services can be provided to users of the client computing devices 502, 504, 506, and / or 508 as web-based services or cloud services (e.g., under a software as a service (SaaS) model). The users operating the client computing devices 502, 504, 506, and / or 508 can then utilize one or more client applications to interact with the server 512 to utilize the services provided by these components.

[0073] In Figure 5In the depicted configuration, server 512 may include one or more components 518, 520, and 522 that implement the functions performed by server 512. These components may include software components, hardware components, or combinations thereof that may be executed by one or more processors. It should be understood that a variety of different system configurations different from distributed system 500 are possible. Thus, Figure 5 The illustrated embodiment is an example of a distributed system for implementing an embodiment system and is not intended to be limiting.

[0074] In accordance with the teachings of the present disclosure, a user may use client computing devices 502, 504, 506, and / or 508 to compress a neural network for use in an NLU. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although Figure 5 only four client computing devices are depicted, any number of client computing devices may be supported.

[0075] Client devices may include various types of computing systems, such as portable handheld devices, general-purpose computers (e.g., personal computers and laptop computers), workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, including various mobile operating systems (e.g., Microsoft Windows Windows Android TM , Palm ) and operating systems (e.g., Microsoft Apple or UNIX-like operating systems, Linux or Linux-like operating systems (e.g., Google Chrome TM OS)). Portable handheld devices may include cellular phones, smartphones (e.g., ), tablet computers (e.g., ), personal digital assistants (PDAs), etc. Wearable devices may include Google head-mounted displays and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices (e.g., Microsoft gaming consoles with or without a gesture input device, Sony systems, by The various game systems provided, as well as others). The client device may be capable of executing various different applications, such as various Internet-related applications, communication applications (e.g., email applications, Short Message Service (SMS) applications), and may use various communication protocols.

[0076] One or more networks 510 may be any type of network familiar to those skilled in the art that can support data communication using any one of a variety of available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internetwork Packet Exchange), and so on. By way of example only, one or more networks 510 may be a local area network (LAN), an Ethernet-based network, Token Ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a network operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite, and / or any other wireless protocol) and / or any combination of these networks and / or other networks.

[0077] The server 512 may be composed of: one or more general-purpose computers, dedicated server computers (by way of example including PC (personal computer) servers, servers, midrange servers, mainframes, rack-mounted servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. The server 512 may include one or more virtual machines running a virtual operating system or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain a virtual storage device for the server). In various embodiments, the server 512 may be adapted to run one or more services or software applications that provide the functions described in the foregoing disclosure.

[0078] The computing system in the server 512 may run one or more operating systems, including any of the operating systems discussed above and any commercially available server operating systems. The server 512 may also run any one of a variety of additional server applications and / or middleware applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, servers, database servers, etc. Exemplary database servers include but are not limited to those from Those commercially available database servers such as (International Business Machines Corporation).

[0079] In some embodiments, server 512 may include one or more applications to analyze and combine data feeds and / or event updates received from users of client computing devices 502, 504, 506, and 508. As an example, the data feeds and / or event updates may include, but are not limited to feeds, updates, or real-time updates received from one or more third-party information sources and continuous data streams, which may include real-time events related to sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc. Server 512 may also include one or more applications to display the data feeds and / or real-time events via one or more display devices of client computing devices 502, 504, 506, and 508.

[0080] Distributed system 500 may also include one or more data repositories 514, 516. In certain embodiments, these data repositories may be used to store data and other information. For example, one or more of data repositories 514, 516 may be used to store information for training models for NLU. Data repositories 514, 516 may reside in various locations. For example, a data repository used by server 512 may be local to server 512 or may be remote from server 512 and communicate with server 512 via a network-based or dedicated connection. Data repositories 514, 516 may be of different types. In certain embodiments, a data repository used by server 512 may be a database, such as a relational database (e.g., databases provided by Oracle and other vendors). One or more of these databases may be adapted to implement storage, update, and retrieval of data to and from the database in response to commands in SQL format.

[0081] In certain embodiments, one or more of data repositories 514, 516 may also be used by applications to store application data. A data repository used by an application may be of a different type, such as, for example, a key-value storage repository, an object storage repository, or a general storage repository supported by a file system.

[0082] In certain embodiments, the NLU-related functions described in the present disclosure may be provided as a service via a cloud environment. Figure 6is a simplified block diagram of a cloud-based system environment in which various NLU-related services can be provided as cloud services. In Figure 6 In the depicted embodiment, the cloud infrastructure system 602 can provide one or more cloud services that can be requested by a user using one or more client computing devices 604, 606, and 608. The cloud infrastructure system 602 can include one or more computers and / or servers, which can include those computers and / or servers described above for server 512. The computers in the cloud infrastructure system 602 can be organized as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.

[0083] One or more networks 610 can facilitate data communication and exchange between the clients 604, 606, and 608 and the cloud infrastructure system 602. One or more networks 610 can include one or more networks. The networks can be of the same or different types. One or more networks 610 can support one or more communication protocols (including wired and / or wireless protocols) for facilitating communication.

[0084] Figure 6 The depicted embodiment is only one example of a cloud infrastructure system and is not intended to be limiting. It should be understood that in some other embodiments, the cloud infrastructure system 602 can have more or fewer components than Figure 6 those depicted, can combine two or more components, or can have a different component configuration or arrangement. For example, although Figure 6 three client computing devices are depicted, in alternative embodiments, any number of client computing devices can be supported.

[0085] The term cloud service is generally used to refer to a service that is made available to a user on demand by a service provider's system (e.g., cloud infrastructure system 602) and via a communication network such as the Internet. Generally, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own self-managed servers and systems. The cloud service provider's system is managed by the cloud service provider. Thus, a customer can avail itself of the cloud services provided by the cloud service provider without having to purchase separate licenses, support, or hardware and software resources for the service. For example, the cloud service provider's system can host an application, and a user can order and use the application on demand via the Internet without the user having to purchase the infrastructure resources for executing the application. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Multiple providers offer cloud services. For example, Oracle of Redwood Shores, California offers a variety of cloud services such as middleware services, database services, Java cloud services, and other services.

[0086] In some embodiments, the cloud infrastructure system 602 can provide one or more cloud services using different models (e.g., under a Software as a Service (SaaS) model, a Platform as a Service (PaaS) model, an Infrastructure as a Service (IaaS) model, and other models including hybrid service models). The cloud infrastructure system 602 can include a set of applications, middleware, databases, and other resources that implement the provisioning of various cloud services.

[0087] The SaaS model enables an application or software to be delivered as a service to a customer via a communication network such as the Internet without the customer having to purchase the hardware or software for the underlying application. For example, the SaaS model can be used to provide a customer with access to on-demand applications hosted by the cloud infrastructure system 602. Examples of SaaS services offered by Oracle include, but are not limited to, various services for human resources / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, social applications, and others.

[0088] The IaaS model is generally used to provide infrastructure resources (e.g., servers, storage, hardware, and networking resources) as cloud services to a customer to provide elastic computing and storage capabilities. Oracle offers various IaaS services.

[0089] The PaaS model is generally used to provide platform and environmental resources as a service that enables customers to develop, run, and manage applications and services without the customers having to procure, build, or maintain such resources. Provided by Oracle Examples of PaaS services provided by Oracle include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), Data Management Cloud Service, various application development solution services, and other services.

[0090] Cloud services are typically provided in an on-demand self-service basis, subscription-based, elastically scalable, reliable, highly available, and secure manner. For example, a customer can order one or more services provided by the cloud infrastructure system 602 via a subscription order. Then, the cloud infrastructure system 602 performs processing to provide the services requested in the customer's subscription order. For example, the cloud infrastructure system 602 trains a set of models for NLU-related tasks. The cloud infrastructure system 602 can be configured to provide one or even multiple cloud services.

[0091] The cloud infrastructure system 602 can provide cloud services via different deployment models. In the public cloud model, the cloud infrastructure system 602 can be owned by a third-party cloud service provider, and the cloud services are provided to any general public customers, where the customers can be individuals or enterprises. In some other embodiments, under the private cloud model, the cloud infrastructure system 602 can operate within an organization (e.g., within an enterprise organization) and the services are provided to customers within the organization. For example, the customers can be various departments of an enterprise such as the human resources department, payroll department, etc. or even individuals within the enterprise. In some other embodiments, under the community cloud model, the cloud infrastructure system 602 and the provided services can be shared by multiple organizations in a related community. Various other models can also be used, such as a hybrid of the models mentioned above.

[0092] The client computing devices 604, 606, and 608 can be of different types (e.g., Figure 5 the depicted devices 502, 504, 506, and 508) and can be capable of operating one or more client applications. A user can use the client device to interact with the cloud infrastructure system 602, such as requesting services provided by the cloud infrastructure system 602. For example, a user can use the client device to request the NLU-related services described in this disclosure.

[0093] In some embodiments, the processing for providing NLU-related services performed by the cloud infrastructure system 602 may involve big data analytics. This analysis may involve using, analyzing, and manipulating large datasets to detect and visualize various trends, behaviors, relationships, etc. within the data. This analysis may be performed by one or more processors, enabling parallel processing of data, performing simulations using the data, etc. For example, big data analytics may be performed by the cloud infrastructure system 602 to identify intents based on received voice inputs. The data used for this analysis may include structured data (e.g., data stored in a database or structured according to a structured model) and / or unstructured data (e.g., data blobs (binary large objects)).

[0094] As Figure 6 depicted in the embodiments of [], the cloud infrastructure system 602 may include infrastructure resources 630 that are used to facilitate the provision of various cloud services provided by the cloud infrastructure system 602. The infrastructure resources 630 may include, for example, processing resources, storage or memory resources, networking resources, etc.

[0095] In certain embodiments, to facilitate the efficient provision of these resources to support the various cloud services provided by the cloud infrastructure system 602 for different customers, the resources may be bundled into resource groups or resource modules (also referred to as "pods"). Each resource module or pod may include a pre-integrated and optimized combination of one or more types of resources. In certain embodiments, different pods may be pre-provisioned for different types of cloud services. For example, a first set of pods may be provisioned for database services, a second set of pods may be provisioned for Java services (the second set of pods may include a combination of resources different from those in the first set of pods), etc. For some services, the resources allocated for service provision may be shared among the services.

[0096] The cloud infrastructure system 602 itself may internally use services 632 that are shared by different components of the cloud infrastructure system 602 and facilitate the provision of services by the cloud infrastructure system 602. These internal shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, services for enabling cloud support, email services, notification services, file transfer services, etc.

[0097] The cloud infrastructure system 602 may include multiple subsystems. These subsystems may be implemented in software or hardware or a combination thereof. As Figure 6As depicted, the subsystem may include a user interface subsystem 612 that enables users or customers of the cloud infrastructure system 602 to interact with the cloud infrastructure system 602. The user interface subsystem 612 may include various different interfaces, such as a web interface 614, an online store interface 616 (where cloud services provided by the cloud infrastructure system 602 are advertised and customers can purchase cloud services provided by the cloud infrastructure system 602), and other interfaces 618. For example, a customer may use a client device to request (service request 634) one or more services provided by the cloud infrastructure system 602 using one or more of the interfaces 614, 616, and 618. For example, a customer may access an online store, browse cloud services provided by the cloud infrastructure system 602, and place a subscription order for one or more services provided by the cloud infrastructure system 602 that the customer wishes to subscribe to. The service request may include information identifying the customer and the one or more services the customer expects to subscribe to. For example, a customer may place a subscription order for NLU-related services provided by the cloud infrastructure system 602. As part of the order, the customer may provide a voice input identifying the request.

[0098] In certain embodiments (such as Figure 6 the depicted embodiment), the cloud infrastructure system 602 may include an order management subsystem (OMS) 620 configured to process new orders. As part of this processing, the OMS 620 may be configured to: create an account for the customer (if not already created); receive billing and / or payment information from the customer to be used to bill the customer for providing the requested services; verify the customer information; after verification, book the order for the customer; and orchestrate various workflows to prepare the order for fulfillment.

[0099] Once properly verified, the OMS 620 may invoke an order provisioning subsystem (OPS) 624 configured to provision resources for the order, including processing resources, memory resources, and networking resources. Provisioning may include allocating resources for the order and configuring the resources to facilitate the services requested by the customer order. The manner in which resources are provisioned for an order and the types of resources provisioned may depend on the type of cloud service that the customer has ordered. For example, according to one workflow, the OPS 624 may be configured to determine the specific cloud service being requested and identify the number of clusters that may have been pre-configured for that specific cloud service. The number of clusters allocated for an order may depend on the size / volume / tier / scope of the service requested. For example, the number of clusters to be allocated may be determined based on the number of users to be supported by the service, the duration for which the service is being requested, etc. Then, the allocated clusters may be customized for the specific requesting customer to provide the requested services.

[0100] The cloud infrastructure system 602 can send a response or notification 644 to the requesting customer to indicate when the requested service is now ready for use. In some instances, information (e.g., a link) that enables the customer to start using and leveraging the benefits of the requested service can be sent to the customer. In certain embodiments, for a customer requesting an NLU-related service, the response can include a response generated based on the identified intent.

[0101] The cloud infrastructure system 602 can provide services to multiple customers. For each customer, the cloud infrastructure system 602 is responsible for managing information related to one or more subscription orders received from the customer, maintaining customer data related to the orders, and providing the requested services to the customer. The cloud infrastructure system 602 can also collect usage statistics regarding the customer's use of the subscribed services. For example, statistics can be collected for the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime. This usage information can be used to bill the customer. Billing can be done, for example, on a monthly basis.

[0102] The cloud infrastructure system 602 can provide services to multiple customers in parallel. The cloud infrastructure system 602 can store information for these customers (which may include proprietary information). In certain embodiments, the cloud infrastructure system 602 includes an identity management subsystem (IMS) 628 that is configured to manage customer information and provide separation of the managed information such that information related to one customer cannot be accessed by another customer. The IMS 628 can be configured to provide various security-related services such as identity services (e.g., information access management), authentication and authorization services, services for managing customer identities and roles, and related functions.

[0103] Figure 7 An exemplary computer system 700 that can be used to implement certain embodiments is shown. For example, in some embodiments, the computer system 700 can be used to implement any one of the ASR subsystem, NLU subsystem, and various servers and computer systems described above. As Figure 7 shown, the computer system 700 includes various subsystems that include a processing subsystem 704 that communicates with a number of other subsystems via a bus subsystem 702. These other subsystems can include a processing acceleration unit 706, an I / O subsystem 708, a storage subsystem 718, and a communication subsystem 724. The storage subsystem 718 can include non-transitory computer-readable storage media that includes a storage medium 722 and a system memory 710.

[0104] The bus subsystem 702 provides the mechanism for enabling the various components and subsystems of the computer system 700 to communicate with each other as expected. Although the bus subsystem 702 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 702 can be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, a local bus using any one of a variety of bus architectures. For example, such architectures can include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus (which can be implemented as a Mezzanine bus manufactured to the IEEE P1386.1 standard), etc.

[0105] The processing subsystem 704 controls the operation of the computer system 700 and can include one or more processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The processors can include single-core processors or multi-core processors. The processing resources of the computer system 700 can be organized into one or more processing units 732, 734, etc. The processing units can include one or more processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some embodiments, the processing subsystem 704 can include one or more dedicated coprocessors such as a graphics processor, a digital signal processor (DSP), etc. In some embodiments, some or all of the processing units in the processing subsystem 704 can be implemented using custom circuits such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs).

[0106] In some embodiments, the processing units in the processing subsystem 704 can execute instructions stored in the system memory 710 or on the computer-readable storage medium 722. In various embodiments, the processing units can execute various programs or code instructions and can maintain multiple simultaneously executing programs or processes. At any given time, some or all of the program code to be executed can reside in the system memory 710 and / or on the computer-readable storage medium 722 (including potentially on one or more storage devices). The processing subsystem 704 can provide the various functions described above through suitable programming. In instances where the computer system 700 is executing one or more virtual machines, one or more processing units can be allocated to each virtual machine.

[0107] In some embodiments, a processing acceleration unit 706 may optionally be provided for performing custom processing or for offloading some of the processing performed by the processing subsystem 704 in order to accelerate the overall processing performed by the computer system 700.

[0108] The I / O subsystem 708 may include devices and mechanisms for inputting information to the computer system 700 and / or for outputting information from or via the computer system 700. Generally, the term input device is intended to include all possible types of devices and mechanisms for inputting information to the computer system 700. User interface input devices may include, for example, a keyboard, a pointing device (such as a mouse or trackball), a touchpad or touch screen incorporated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keypad, an audio input device having a voice command recognition system, a microphone, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices, such as Microsoft Motion Sensors, Microsoft Xbox 360 game controllers, devices that provide an interface for receiving input using gestures and voice commands. User interface input devices may also include eye gesture recognition devices, such as Google (e.g., detecting eye movement from the user (e.g., "blinking" when taking a photo and / or making a menu selection) and transforming the eye gesture into an input to an input device such as Google Blink Detector). Additionally, user interface input devices may include voice recognition sensing devices that enable a user to interact with a voice recognition system via voice commands.

[0109] Other examples of user interface input devices include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, and graphics tablets, as well as audio / visual devices (such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye gaze tracking devices). Additionally, user interface input devices may include, for example, medical imaging input devices, such as computed tomography, magnetic resonance imaging, positron emission tomography, and medical ultrasound devices. User interface input devices may also include, for example, audio input devices, such as MIDI keyboards, digital musical instruments, and the like.

[0110] In general, the term output device is intended to include all possible types of devices and mechanisms for outputting information from a computer system 700 to a user or other computer. User interface output devices can include a display subsystem, indicator lights, or non-visual displays such as audio output devices. The display subsystem can be a cathode ray tube (CRT), a flat panel device (e.g., a flat panel device using a liquid crystal display (LCD) or a plasma display), a projection device, a touch screen, etc. For example, user interface output devices can include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automotive navigation systems, plotters, voice output devices, and modems.

[0111] The storage subsystem 718 provides a repository or data storage for storing information and data used by the computer system 700. The storage subsystem 718 provides a tangible non-transitory computer-readable storage medium for storing the basic programming and data constructs that provide the functionality of some embodiments. The storage subsystem 718 can store software (e.g., programs, code modules, instructions) that, when executed by the processing subsystem 704, provides the functionality described above. The software can be executed by one or more processing units of the processing subsystem 704. The storage subsystem 718 can also provide a repository for storing data used in accordance with the teachings of the present disclosure.

[0112] The storage subsystem 718 can include one or more non-transitory memory devices, including volatile memory devices and non-volatile memory devices. As Figure 7 shown, the storage subsystem 718 includes system memory 710 and a computer-readable storage medium 722. The system memory 710 can include multiple memories, including a volatile main random access memory (RAM) for storing instructions and data during program execution and a non-volatile read-only memory (ROM) or flash memory in which fixed instructions are stored. In some embodiments, a basic input / output system (BIOS) that includes basic routines, such as those that help transfer information between elements within the computer system 700 during startup, can typically be stored in the ROM. The RAM typically contains the data and / or program modules that are currently being operated on and executed by the processing subsystem 704. In some embodiments, the system memory 710 can include various different types of memories, such as static random access memory (SRAM), dynamic random access memory (DRAM), etc.

[0113] By way of example and not limitation, as Figure 7As depicted, system memory 710 may load an executing application 712 (which may include various applications such as a web browser, a middleware application, a relational database management system (RDBMS), etc.), program data 714, and an operating system 716. By way of example, the operating system 716 may include various versions of Microsoft Apple and / or Linux operating systems, various commercially available or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google OS, etc.) and / or mobile operating systems (such as iOS, Phone, OS, OS, OS operating systems), and other operating systems.

[0114] Computer-readable storage medium 722 may store the programming and data constructs that provide the functionality of some embodiments. The computer-readable medium 722 may provide storage for the computer system 700 for computer-readable instructions, data structures, program modules, and other data. Software (programs, code modules, instructions) that provides the functionality described above when executed by the processing subsystem 704 may be stored in the storage subsystem 718. By way of example, the computer-readable storage medium 722 may include non-volatile memories such as hard disk drives, disk drives, optical disc drives (such as CD ROM, DVD, disc or other optical media). The computer-readable storage medium 722 may include but is not limited to drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVD discs, digital video tapes, etc. The computer-readable storage medium 722 may also include solid-state drives (SSDs) based on non-volatile memory (such as SSD-based flash memory, enterprise flash drives, solid-state ROM, etc.), SSDs based on volatile memory (such as solid-state RAM, dynamic RAM, static RAM), DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs.

[0115] In certain embodiments, the storage subsystem 718 may also include a computer-readable storage medium reader 720 that may be further connected to the computer-readable storage medium 722. The reader 720 may receive data from a memory device (such as a disc, flash drive, etc.) and is configured to read data from the memory device.

[0116] In some embodiments, computer system 700 may support virtualization technologies, including but not limited to virtualization of processing resources and memory resources. For example, computer system 700 may provide support for executing one or more virtual machines. In some embodiments, computer system 700 may execute a program such as a hypervisor that facilitates the configuration and management of virtual machines. Each virtual machine may be allocated memory, computing (e.g., processors, cores), I / O, and networking resources. Each virtual machine typically runs independently of other virtual machines. A virtual machine typically runs its own operating system, which may be the same as or different from the operating systems executed by other virtual machines of computer system 700. Thus, multiple operating systems may potentially be run simultaneously by computer system 700.

[0117] Communication subsystem 724 provides an interface to other computer systems and networks. Communication subsystem 724 serves as an interface for receiving data from other systems and transmitting data from computer system 700 to other systems. For example, communication subsystem 724 may enable computer system 700 to establish a communication channel to one or more client devices via the Internet for receiving information from and sending information to the client devices. For example, the communication subsystem may be used to communicate with a database to import updated training data for the ASR subsystem and / or the NLU subsystem.

[0118] Communication subsystem 724 may support both wired communication protocols and / or wireless communication protocols. For example, in some embodiments, communication subsystem 724 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular phone technology, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE 802.XX home standards, or other mobile communication technologies, or any combination thereof)), global positioning system (GPS) receiver components, and / or other components. In some embodiments, communication subsystem 724 may provide wired network connectivity (e.g., Ethernet) in addition to or in place of the wireless interface.

[0119] Communication subsystem 724 may receive and transmit various forms of data. For example, in some embodiments, communication subsystem 724 may receive input communications in the form of structured and / or unstructured data feeds 726, event streams 728, event updates 730, etc., among other forms. For example, communication subsystem 724 may be configured to receive (or send) data feeds 726 from users of social media networks and / or other communication services in real time, such as feeds, Updates, web feeds (e.g., Rich Site Summary (RSS) feeds), and / or real-time updates from one or more third-party information sources.

[0120] In some embodiments, the communication subsystem 724 may be configured to receive data in the form of a continuous data stream that can be essentially continuous or unbounded and without a definite end, which may include an event stream 728 of real-time events and / or event updates 730. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc.

[0121] The communication subsystem 724 may also be configured to transfer data from the computer system 700 to other computer systems or networks. The data may be transferred in a variety of different forms (e.g., structured and / or unstructured data feeds 726, event streams 728, event updates 730, etc.) to one or more databases that can communicate with one or more stream data source computers coupled to the computer system 700.

[0122] The computer system 700 may be one of various types, including handheld portable devices (e.g., cellular phones, computing tablet computers, PDAs), wearable devices (e.g., Google head-mounted displays), personal computers, workstations, mainframes, self-service terminals, server racks, or any other data processing system. Due to the ever-changing nature of computers and networks, the description of the Figure 7 depicted computer system 700 is intended to be only a specific example. Many other configurations with more or fewer components than the Figure 7 depicted system are possible. Based on the disclosure and teachings provided herein, those of ordinary skill in the art will understand other ways and / or methods of implementing various embodiments.

[0123] Although specific embodiments have been described, various modifications, changes, alternative constructions, and equivalents are possible. Embodiments are not limited to operating within certain specific data processing environments, but are free to operate within multiple data processing environments. Additionally, although certain embodiments have been described using a specific series of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. Although some flowcharts depict operations as sequential processes, many of the operations can be performed in parallel or simultaneously. Additionally, the order of operations can be rearranged. The process may have additional steps not included in the figures. The various features and aspects of the embodiments described above can be used individually or jointly.

[0124] Further, although certain embodiments have been described using a specific combination of hardware and software, it should be recognized that other combinations of hardware and software are possible. Certain embodiments may be implemented only in hardware, only in software, or using a combination thereof. The various processes described herein may be implemented in any combination on the same processor or different processors.

[0125] In cases where a device, system, component, or module is described as being configured to perform certain operations or functions, such configuration may be accomplished, for example, by designing an electronic circuit to perform the operations, by programming a programmable electronic circuit (such as a microprocessor) to perform the operations (e.g., by executing computer instructions or code, or by a processor or core programmed to execute code or instructions stored on a non-transitory memory medium), or any combination thereof. Processes may communicate using a variety of techniques including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.

[0126] Specific details are given in this disclosure to provide a thorough understanding of the embodiments. However, the embodiments may be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques have been shown without unnecessary detail to avoid obscuring the embodiments. This description provides only example embodiments and is not intended to limit the scope, applicability, or configuration of other embodiments. Rather, the previous description of the embodiments will provide those skilled in the art with an enabling description for implementing the various embodiments. Various changes may be made to the functions and arrangements of the elements.

[0127] Accordingly, the specification and the drawings will be regarded in an illustrative rather than a restrictive sense. However, it will be apparent that additions, deletions, and other modifications and changes may be made without departing from the broader spirit and scope set forth in the claims. Thus, while specific embodiments have been described, these specific embodiments are not intended to be limiting. Various modifications and equivalents are within the scope of the claims.

Claims

1. A computer-implemented method, comprising: Obtain session data with a first tag; Use the session data with the first tag to train a first neural network; Obtain first untagged session data; Use the trained first neural network to tag the first untagged session data, thereby generating second tagged session data; And Use the second tagged session data to train a second neural network for performing natural language understanding tasks.

2. The method according to claim 1, further comprising, before training the first neural network: Obtaining second unlabeled session data; Using the second unlabeled session data to train a third neural network to perform a proxy task; and Generating the first neural network based on the third neural network.

3. The method according to claim 2, wherein, Generating the first neural network based on the third neural network includes constructing the first neural network to include at least a portion of the third neural network.

4. The method according to claim 2, wherein, The third neural network is a component of the first neural network.

5. The method according to claim 2, wherein, The proxy task includes a prediction language task.

6. The method according to claim 1, wherein, The natural language understanding task includes one or more of semantic parsing, intent classification, or named entity classification.

7. The method according to claim 1, wherein, Training the second neural network includes: For a first training input of the second tagged session data, output a predicted output by the second neural network; Calculate a loss that measures an error between the predicted output and a first label associated with the first training input; Based on the loss, calculate an update value for a first set of parameters of the second neural network; and Update the second neural network by changing the values of the first set of parameters to the update value.

8. A non-transitory computer-readable memory storing multiple instructions executable by one or more processors, the multiple instructions including instructions that cause the one or more processors to perform a process including the following operations when executed by the one or more processors: Obtaining first labeled session data; Use the session data with the first tag to train a first neural network; Obtain first untagged session data; Use the trained first neural network to tag the first untagged session data, thereby generating second tagged session data; And Use the second tagged session data to train a second neural network for performing natural language understanding tasks.

9. The non-transitory computer-readable memory according to claim 8, the process further comprising, before training the first neural network: Obtaining second unlabeled session data; Use the second unlabeled session data to train a third neural network to perform a surrogate task; and Generate the first neural network based on the third neural network.

10. The non-transitory computer-readable memory of claim 9, wherein Generating the first neural network based on the third neural network includes constructing the first neural network to include at least a portion of the third neural network.

11. The non-transitory computer-readable memory of claim 9, wherein The third neural network is a component of the first neural network.

12. The non-transitory computer-readable memory of claim 9, wherein The proxy task includes a prediction language task.

13. The non-transitory computer-readable memory of claim 8, wherein The natural language understanding task includes one or more of semantic parsing, intent classification, or entity recognition.

14. The non-transitory computer-readable memory of claim 8, wherein Training the second neural network includes: For a first training input of the second tagged session data, output a predicted output by the second neural network; Calculate a loss that measures an error between the predicted output and a first label associated with the first training input; Based on the loss, calculate an update value for a first set of parameters of the second neural network; and Update the second neural network by changing the values of the first set of parameters to the update value.

15. A computer-implemented system, comprising: One or more processors; A memory coupled to the one or more processors, the memory storing a plurality of instructions executable by the one or more processors, the plurality of instructions including instructions that, when executed by the one or more processors, cause the one or more processors to perform a process including the following operations: Obtain session data with a first tag; Use the session data with the first tag to train a first neural network; Obtain first untagged session data; Use the trained first neural network to tag the first untagged session data, thereby generating second tagged session data; And Use the session data of the second token to train a second neural network for performing natural language understanding tasks.

16. The system of claim 15, wherein the processing further comprises, before training the first neural network: Obtain second unlabeled session data; Use the second unlabeled session data to train a third neural network to perform a surrogate task; and Generate the first neural network based on the third neural network.

17. The system of claim 16, wherein Generating the first neural network based on the third neural network includes constructing the first neural network to include at least a portion of the third neural network.

18. The system of claim 16, wherein The third neural network is a component of the first neural network.

19. The system of claim 15, wherein The natural language understanding tasks include one or more of semantic parsing, intent classification, or named entity classification.

20. The system of claim 15, wherein Training the second neural network includes: For a first training input of the session data of the second token, output a predicted output by the second neural network; Calculate a loss that measures the error between the predicted output and a first label associated with the first training input; Based on the loss, calculate an updated value of a first set of parameters of the second neural network; and Update the second neural network by changing the values of the first set of parameters to the updated values.

21. The system according to claim 16, wherein, The second unlabeled session data is different from the first unlabeled session data.

22. A conversation system, the conversation system using a computer-implemented system according to any one of claims 15 to 21 for natural language understanding.

Citation Information

Patent Citations

  • Machine self-learning construction knowledge atlas training method based on neural network

    CN106875940A

  • Neural network model training method and device

    CN109993300A