Knowledge distillation for speech recognition under domain mismatch
By receiving extra-domain training discourses, generating enhanced discourses, and using teacher ASR models to generate pseudo-labels, training students' ASR models, solving the problem of speech recognition model training under domain mismatch, and achieving efficient speech recognition effect.
Patent Information
- Application Number
- CN202380072458.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-19
- Filing Date
- 2023-10-17
- Publication Date
- 2025-05-23
AI Technical Summary
In the case of domain mismatch, it is difficult for the prior art to effectively train speech recognition models, especially in automatic speech recognition (ASR) applications of long-tail languages, where training data matching the target domain is lacking.
By receiving distillation data including multiple extradomain training discourses, augmented extradomain training discourse is generated and a pseudo-tagged use of the teacher ASR model trained on the target domain, the student ASR model is trained to achieve knowledge distillation.
This method effectively trains students' ASR model in the case of domain mismatch, improving the accuracy and adaptability of speech recognition, especially in the application of long-tail languages.
Smart Images

Figure CN120035858A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to training speech recognition models. Background Art
[0002] Speech recognition systems are increasingly used to transcribe speech into text in many everyday applications. These speech recognition systems can be embedded on user devices such as smart home devices or smartphones, or used in cloud-based services. Summary of the invention
[0003] One aspect of the present disclosure provides a computer-implemented method for training a speech recognition model. The method, when executed by data processing hardware, causes the data processing hardware to perform operations. The operations include receiving distilled data including a plurality of out-of-domain training utterances. For each specific out-of-domain training utterance of the distilled data, the operations include generating a corresponding enhanced out-of-domain training utterance; and using a teacher automatic speech recognition (ASR) model trained on training data corresponding to a target domain to generate a pseudo-label corresponding to the corresponding enhanced out-of-domain training utterance. The operations also include distilling a student ASR model from a teacher ASR model by training the student ASR model using the corresponding enhanced out-of-domain training utterances paired with the corresponding pseudo-labels generated by the teacher ASR model.
[0004] Implementations of the computer-implemented method or system of the present disclosure may include one or more of the following optional features. In some implementations, distilling the student ASR model from the teacher ASR model includes training the student ASR model to recognize utterances corresponding to the target domain. In some examples, before receiving the distillation data, the teacher ASR model is trained on training data including multiple target domain training utterances; and after receiving the distillation data, the training data is unavailable to the teacher ASR model and the student ASR model. In some implementations, the training data includes multiple target domain training utterances, and one or more of the multiple target domain training utterances are not included in the distillation data. In some examples, enhancing the corresponding out-of-domain training utterances includes at least one of adding noise, adding reverberation, or manipulating time.
[0005] In some implementations, the student ASR model is executed in a cloud computing environment, and the teacher ASR model is executed on one or more user devices in communication with the cloud computing environment. The training data may include a plurality of target domain training utterances, each of which is received at a corresponding user device in one or more user devices. In some examples, generating a pseudo-label corresponding to a corresponding enhanced out-of-domain training utterance includes generating a probability distribution of possible speech recognition hypotheses, wherein the pseudo-label includes the N-best speech recognition hypotheses with the highest probability. In some implementations, distilling the student ASR model from the teacher ASR model includes, for each specific enhanced out-of-domain training utterance: using the student ASR model to generate a transcription corresponding to the specific enhanced out-of-domain training utterance; generating a loss based on the pseudo-label corresponding to the specific enhanced out-of-domain training utterance and the transcription corresponding to the specific enhanced out-of-domain training utterance; and using the loss to update the parameters of the student ASR model. The loss may be at least one of a cross entropy loss, a Kulbeck-Leibler (KL) divergence loss, or an L2 loss.
[0006] Another aspect of the present disclosure provides a system comprising data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving distilled data comprising a plurality of out-of-domain training utterances. For each specific out-of-domain training utterance of the distilled data, the operations include generating a corresponding enhanced out-of-domain training utterance; and using a teacher automatic speech recognition (ASR) model trained on training data corresponding to a target domain to generate a pseudo-label corresponding to the corresponding enhanced out-of-domain training utterance. The operations also include distilling a student ASR model from a teacher ASR model by training the student ASR model using the corresponding enhanced out-of-domain training utterances paired with the corresponding pseudo-labels generated by the teacher ASR model.
[0007] Implementations of the computer-implemented method or system of the present disclosure may include one or more of the following optional features. In some implementations, distilling the student ASR model from the teacher ASR model includes training the student ASR model to recognize utterances corresponding to the target domain. In some examples, before receiving the distillation data, the teacher ASR model is trained on training data including multiple target domain training utterances; and after receiving the distillation data, the training data is unavailable to the teacher ASR model and the student ASR model. In some implementations, the training data includes multiple target domain training utterances, and one or more of the multiple target domain training utterances are not included in the distillation data. In some examples, enhancing the corresponding out-of-domain training utterances includes at least one of adding noise, adding reverberation, or manipulating time.
[0008] In some implementations, the student ASR model is executed in a cloud computing environment, and the teacher ASR model is executed on one or more user devices in communication with the cloud computing environment. The training data may include a plurality of target domain training utterances, each of which is received at a corresponding user device in one or more user devices. In some examples, generating a pseudo-label corresponding to a corresponding enhanced out-of-domain training utterance includes generating a probability distribution of possible speech recognition hypotheses, wherein the pseudo-label includes the N-best speech recognition hypotheses with the highest probability. In some implementations, distilling the student ASR model from the teacher ASR model includes, for each specific enhanced out-of-domain training utterance: using the student ASR model to generate a transcription corresponding to the specific enhanced out-of-domain training utterance; generating a loss based on the pseudo-label corresponding to the specific enhanced out-of-domain training utterance and the transcription corresponding to the specific enhanced out-of-domain training utterance; and using the loss to update the parameters of the student ASR model. The loss may be at least one of a cross entropy loss, a Kulbeck-Leibler (KL) divergence loss, or an L2 loss.
[0009] The details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a diagram of an example speech environment using a speech recognition model to transcribe an utterance.
[0011] Figure 2A is a diagram of an example training process for training a speech recognition model using knowledge distillation under domain mismatch.
[0012] Figure 2B is a diagram of another example training process for training a speech recognition model using knowledge distillation under domain mismatch.
[0013] Figure 3 is a flow diagram of an example operational arrangement of a method for training a speech recognition model using knowledge distillation under domain mismatch.
[0014] Figure 4 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0015] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION
[0016] Knowledge distillation can be used to transfer knowledge from one model (i.e., teacher model) to another model (i.e., student model). Conventional knowledge distillation assumes that data in the target domain used to train the teacher model can be used to train the student model. Alternatively, complex and sophisticated methods are required to generate suitable data for training the student model. However, in many cases, training data that matches the target domain is not available for training the student model. For example, for automatic speech recognition (ASR) of long-tail languages (e.g., sub-Saharan African languages), there may not be a sufficient amount of training data available for training the server-side model. Although joint training can be used to train the server-side model, the on-device data that is usually used to train the on-device speech recognition model should not leave the user device for privacy reasons. Therefore, conventional knowledge distillation based on the on-device model (i.e., teacher model) cannot be used to train the server-side model (i.e., student model) because there is no available training data that matches the target domain to which the on-device model was trained. Therefore, when the domain of the training data available for training the student model does not match the domain of the training data used to train the teacher model, a method and system for performing knowledge distillation to train the student model is needed. Although the examples disclosed herein relate to using knowledge distillation to train a student ASR model under domain mismatch, the disclosed examples can be used to train other types of models using knowledge distillation under domain mismatch, that is, when the training data used to train the teacher model is from a domain that does not match (i.e., is different from) the domain of the training data that can be used to train the student model.
[0017] refer to Figure 1In some implementations, the speech environment 100 includes a user 104 who interacts with a speech-enabled device 10 (also referred to as a user device 10) using a spoken utterance 106. Here, the spoken utterance 106 corresponds to a dictation or query, such as for transcription, which is used to request a response from the user device 10 or to cause the user device 10 to perform a task specified by the query. In this sense, the user 104 can interact with the user device 10 in a conversation-like manner to perform computing activities or find answers to questions. In the example shown, the system 102 includes an ASR model 170 of a speech recognition system 150 for generating a transcription 172 of the utterance 106. The transcription 172 can then be processed by the digital assistant 20 to generate a response to the utterance 106 or to perform a task specified by the utterance 106. In some implementations, the digital assistant 20 includes a natural language processing / understanding (NLP / NLU) module executed on the user device 10 or a remote computing system 70 to process the transcription 172 to understand the utterance 106. The digital assistant 20 may provide a response in the form of text in a user interface 22 on a display 16 c of the user device 10 or in the form of an audio signal 108 output by a speaker 16 b of the user device 10. In some examples, the digital assistant 20 generates text representing the response, and a text-to-speech (TTS) system converts the text into an audio signal 108 as synthesized speech. In the example shown, the user 104 speaks an utterance 106 asking “Who taught Alexander the Great”, and the digital assistant 20 responds with audio data 108 representing the response “Aristotle”.
[0018] The user device 10 may correspond to any computing device associated with the user 104 and capable of capturing audio data 162 and providing text or auditory output. Some examples of the user device 10 include, but are not limited to, mobile devices (e.g., mobile phones, tablet computers, laptop computers, etc.), computers, wearable devices (e.g., smart watches, smart glasses, smart goggles, AR headsets, VR headsets, etc.), smart home appliances, Internet of Things (IoT) devices, in-vehicle infotainment systems, smart displays, smart speakers, etc. The user device 10 includes data processing hardware 12 and memory hardware 14 that communicates with the data processing hardware 12 and stores instructions that, when executed by the data processing hardware 12, cause the data processing hardware 12 to perform one or more operations. The user device 10 further includes one or more input / output devices 16 (16a-c), such as an audio capture device 16 (16a) (e.g., a microphone) for capturing the spoken utterance 106 and converting it into electrical signals, an audio output device 16 (16b) (e.g., a speaker) for transmitting auditory audio signals (e.g., as output audio data from the user device 10), and a display 16 (16c) for displaying visual content. Of course, any number and / or type of other input / output devices 16 may be used. The input / output devices 16 may reside on the user device 10 or communicate with the user device.
[0019] The speech recognition system 150 is executed on the user device 10 of the user 104 and / or a remote computing system 70 (e.g., one or more remote servers of a distributed system executed in a cloud computing environment) that communicates with the user device 10 via a network 40. The speech recognition system 150 includes an input subsystem 160 that is configured to receive utterances 106 spoken by the user 104 and captured by the audio capture device 16a, and convert each utterance 106 into a corresponding digital format associated with an input acoustic frame 162 (also commonly referred to as audio data 162) that can be processed by the speech recognition system 150. Thereafter, the ASR model 170 of the speech recognition system 150 receives the audio data 162 corresponding to the current utterance 106 as input, and generates / predicts a corresponding transcription 172 (e.g., a recognition result / hypothesis) of the utterance 106 as output.
[0020] Remote computing system 70 includes data processing hardware 72 and memory hardware 74 in communication with data processing hardware 72. Memory hardware 74 stores instructions that, when executed by data processing hardware 72, cause data processing hardware 72 to perform one or more operations, such as those disclosed herein.
[0021] The ASR model 170 can have different types of neural network architectures. In some examples, the ASR model 170 includes a conformer-based encoder and a listen-attention-spelling (LAS) decoder. In some implementations, the conformer-based encoder can include 600 million trainable parameters, and the LAS decoder can include 200 million trainable parameters. In other examples, the ASR model 170 includes an RNN-T model with an encoder-decoder architecture. Here, the decoder architecture of the RNN-T model includes a joint / prediction network. The audio encoder of the ASR model 170 can include a cascaded encoder architecture.
[0022] Figure 2A An example training process 200a for training a student ASR model 210 from an ASR model 170 (i.e., teacher ASR model 170) using knowledge distillation under domain mismatch is depicted. That is, a domain mismatch occurs when the training data used to train the teacher ASR model 170 is from a target domain that does not match (i.e., is different from) the domain of the training data that can be used to train the student ASR model 210. The training process 200a includes a knowledge distillation training process in which the teacher ASR model 170 distills its knowledge to the student ASR model 210. Here, the teacher ASR model 170 distills its knowledge to the student ASR model 210 by training the student ASR model 210 with the distilled data 220 including a plurality of out-of-domain training utterances 222 (222a-n). Notably, the out-of-domain training utterances 222 are from a different domain than the target domain of the training data (i.e., target domain training utterances) used to train the teacher ASR model 170 prior to distilling the knowledge to the student ASR model 210. Here, the student ASR model 210 learns to predict transcriptions 212 (212a-n) that match pseudo labels 240 (240a-n) generated by the teacher ASR model 170 in order to train the student ASR model 210 to recognize utterances corresponding to the target domain. In the federated learning example, the student ASR model 210 is executed in a cloud computing environment (e.g., a server-side model executed on a remote computing system 70), and the teacher ASR model 170 is executed on a user device 10 that is in communication with the cloud computing environment. Here, the target domain training utterances used to train the teacher ASR model 170 are received at the user device 10 from the user 104 of the user device 10 and are not available for training the student ASR model 210. Therefore, for the purpose of training the teacher ASR model 170, the target domain training utterances are only processed by the teacher ASR model 170, whereby neither the target domain training utterances nor sensitive information derived from the target domain training utterances are sent / transmitted to the cloud computing environment.
[0023] Figure 2AThe use of a single teacher ASR model 170 executed on a user device 10 to train the student ASR model 210 is depicted. However, the student ASR model 210 may be trained using one or more teacher ASR models 170 executed on one or more user devices 10, wherein the target domain training utterances used to train each teacher ASR model 170 are received from the corresponding user at the corresponding user device in the one or more user devices 10. Notably, each teacher ASR model 170 may be associated with its own target domain training utterances, which are not available for training other teacher ASR models 170 or the student ASR model 210. The architecture of the teacher ASR model 170 may be different from the architecture of the student ASR model 210. Additionally or alternatively, the teacher ASR model 170 and the student ASR model 210 may have different sizes. Notably, it has been found that for ASR of long-tail languages, the teacher ASR model 170 may be as small as 1 / 32 of the student ASR model 210 and still achieve good knowledge distillation.
[0024] During the training process 200a, the enhancer 230 receives the out-of-domain training utterances 222, and for each specific out-of-domain training utterance 222, generates a corresponding enhanced out-of-domain training utterance 232 (232a-n). The enhancer 230 generates the enhanced out-of-domain training utterance 232 by, for example, adding noise to the corresponding out-of-domain training utterance 222, adding reverberation, or manipulating its time. In some implementations, the enhancer 230 uses Gaussian noise injection (GNI) to generate the enhanced out-of-domain training utterance 232. For example, for the out-of-domain training utterance 222 represented by Log-Mel input features, the enhancer 230 can enhance the out-of-domain training utterance 222 by randomly changing the value of one or more input features of the out-of-domain training utterance 222 by adding a corresponding random value extracted from a Gaussian distribution to the value of each input feature in one or more changed input features. Here, the enhancer 230 may not change all input features in the input features. The mean and variance of the Gaussian distribution can be calculated from the corresponding frequency channels in the input features. Alternatively, one or more input features may be replaced by corresponding random values drawn from a Gaussian distribution. Here, GNI may be tuned using a single hyperparameter representing a probability ranging from zero to one hundred percent indicating whether a particular input feature is changed / replaced, such that GNI-based enhancements may be easily adjusted. Notably, after receiving the distilled data 220, target domain training utterances for training the teacher ASR model 170 are not available, and the distilled data 220 need not include any such target domain training utterances. In the example shown, the enhancer 230 generates a single enhanced training utterance 232 for each training utterance 222. However, the enhancer 230 may generate multiple enhanced training utterances 232 for each training utterance 222.
[0025] For each specific enhanced out-of-domain training utterance 232, the teacher ASR model 170 generates a pseudo-label 240 (240a-n) corresponding to the specific enhanced out-of-domain training utterance 232. In some implementations, the teacher ASR model 170 generates the corresponding pseudo-label 240 by generating a probability distribution of possible speech recognition hypotheses for the specific enhanced out-of-domain training utterance 232, and the pseudo-label 240 is the speech recognition hypothesis with the highest probability. It is worth noting that the teacher ASR model 170 can generate a pseudo-label 240 for a specific enhanced out-of-domain training utterance 232 that is different from the pseudo-label generated by the teacher ASR model 170 for the corresponding out-of-domain training utterance 222. Therefore, the enhancer 230 effectively and dynamically creates a large unlabeled dataset at almost zero cost to increase the abundance and coverage of the distilled data 220, which provides the teacher ASR model 170 with more opportunities to distill knowledge in the target domain to the student ASR model 210.
[0026] For each specific enhanced out-of-domain training utterance 232, the student ASR model 210 generates a corresponding predicted transcription 212 corresponding to the specific enhanced out-of-domain training utterance 232. In some implementations, the student ASR model 210 generates the corresponding predicted transcription 212 by generating a probability distribution of possible speech recognition hypotheses for the specific enhanced out-of-domain training utterance 232, and the corresponding predicted transcription 212 is the speech recognition hypothesis with the highest probability.
[0027] Thereafter, the training process 200a distills knowledge from the teacher ASR model 170 to the student ASR model 210 by training the student ASR model 210 using the corresponding augmented out-of-domain training utterances 232 paired with the corresponding pseudo-labels 240 generated by the teacher ASR model 170. Specifically, for each particular augmented out-of-domain training utterance 232, the loss term module 250 receives the corresponding pseudo-labels 240 generated by the teacher ASR model 170 and the corresponding transcriptions 212 predicted by the student ASR model 210. Thereafter, the loss term module 250 generates corresponding losses 252 (252a-n) based on the corresponding pseudo-labels 240 and the corresponding predicted transcriptions 212, and uses the losses 252 to update the parameters of the student ASR model 210. In some implementations, the corresponding losses 252 are represented as:
[0028]
[0029] in, is the corresponding augmented training utterance 232 of the particular training utterance x 222, and is from The pseudo labels 240 generated by the teacher ASR model inferred in . Function P s and P t are the output logits from the student ASR model 210 and the teacher ASR model 170, respectively, and have the same shape. The loss function L dist It can be any distillation loss, such as cross entropy loss, Kulbeck–Leibler (KL) divergence loss, or L2 loss.
[0030] Alternatively, the corresponding loss 252 can be expressed as:
[0031]
[0032] For each specific enhanced out-of-domain training utterance 232, utilizing the plurality of pseudo labels 240 generated by the teacher ASR model 170 from the augmented out-of-domain training utterances 232 in a beam-search. In some implementations, the teacher ASR model 170 performs a beam-search to generate pseudo labels 240 for each specific augmented out-of-domain training utterance 232 generates a probability distribution over possible speech recognition hypotheses to generate corresponding pseudo labels 240, wherein the plurality of pseudo labels 240 represent the N-best speech recognition hypotheses with the highest probabilities. Similarly, the student ASR model 210 can be trained by selecting each specific augmented out-of-domain training utterance in a beam search. 232 generates a probability distribution over possible speech recognition hypotheses to generate corresponding predicted transcriptions 212, wherein the plurality of predicted transcriptions 212 represent the N-best speech recognition hypotheses with the highest probabilities. Here, each particular hypothesis The loss of EQN(1) is multiplied by the specific assumption The posterior probability The weighting operation can be implemented as a weighted sum or random sampling. In general, compared with the loss of EQN(1), the loss of EQN(2) can provide better word coverage given a specific augmented out-of-domain training utterance 232.
[0033] Figure 2B Another example training process 200b for training a student ASR model 210 from an ASR model 170 (ie, a teacher ASR model 170) using knowledge distillation under domain mismatch is depicted. Figure 2A Compared with the training process 200a, Figure 2BThe training process 200b includes an additional enhancer 260 for transforming the enhanced out-of-domain training utterances 232 to form student training utterances 262, 262a-n for input to the student ASR model 210. Here, the teacher ASR model 170 still processes the enhanced out-of-domain training utterances 232, and thus the pseudo-labels 240 are not affected by the enhancer 260. In some examples, the enhancer 260 adds noise (e.g., using GNI) or performs some other type of spectral enhancement. Thus, the student ASR model 210 learns from a noisy student learning framework, such that the student ASR model 210 learns that small differences in the input utterances may produce the same pseudo-labels 240, and thus provides an additional degree of robustness to the training process 200b. Here, the enhancer 260 can create more than one student training utterance 262 for each enhanced out-of-domain training utterance 232.
[0034] Figure 3 4 is a flowchart of an exemplary arrangement of operations for a computer-implemented method 300 for training an ASR model using knowledge distillation under domain mismatch. These operations may be performed by data processing hardware 410 ( Figure 4 ) (e.g., data processing hardware 12 of user device 10 or data processing hardware 72 of remote computing system 70) is executed based on execution instructions stored on memory hardware 420 (e.g., memory hardware 14 of user device 10 or memory hardware 74 of remote computing system 70).
[0035] At operation 302, the method 300 includes receiving distilled data 220 including a plurality of out-of-domain training utterances 222. For each particular out-of-domain training utterance 222 of the distilled data 220, the method 300 includes: generating a corresponding augmented out-of-domain training utterance 232 at operation 304; and generating a pseudo-label 240 corresponding to the corresponding augmented out-of-domain training utterance 232 using a teacher ASR model 170 trained on training data corresponding to a target domain at operation 306. At operation 308, the method 300 includes distilling a student ASR model 210 from a teacher ASR model 170 by training the student ASR model 210 using the corresponding augmented out-of-domain training utterance 232 paired with the corresponding pseudo-label 240 generated by the teacher ASR model 170.
[0036] Figure 44 is a schematic diagram of an example computing device 400 that can be used to implement the systems and methods described in this document. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only, and are not intended to limit implementations of the inventions described and / or claimed in this document.
[0037] The computing device 400 includes a processor 410 (i.e., data processing hardware) that can be used to implement the data processing hardware 12 and / or 72, a memory 420 (i.e., memory hardware) that can be used to implement the memory hardware 14 and / or 74, a storage device 430 (i.e., memory hardware) that can be used to implement the memory hardware 14 and / or 74, a high-speed interface / controller 440 connected to the memory 420 and a high-speed expansion port 450, and a low-speed interface / controller 460 connected to a low-speed bus 470 and the storage device 430 that can be used to store the distillation data 220. Each of the components 410, 420, 430, 440, 450, and 460 are interconnected using various buses and can be mounted on a common motherboard or otherwise installed as appropriate. The processor 410 can process instructions for execution within the computing device 400, including instructions stored in the memory 420 or on the storage device 430, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 480 coupled to the high-speed interface 440. In other implementations, multiple processors and / or multiple buses and multiple memories and types of memories may be used as appropriate. In addition, multiple computing devices 400 may be connected, each providing a portion of the necessary operations (e.g., as a server group, a group of blade servers, or a multi-processor system).
[0038] Memory 420 stores information non-temporarily within computing device 400. Memory 420 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-temporary memory 420 may be a physical device for temporarily or permanently storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 400. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0039] Storage device 430 is capable of providing mass storage for computing device 400. In some implementations, storage device 430 is a computer-readable medium. In various implementations, storage device 430 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as memory 420, storage device 430, or memory on processor 410.
[0040] The high-speed controller 440 manages bandwidth-intensive operations of the computing device 400, while the low-speed controller 460 manages less bandwidth-intensive operations. Such a division of responsibilities is exemplary only. In some implementations, the high-speed controller 440 is coupled to the memory 420, the display 480 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 450 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 460 is coupled to the storage device 430 and the low-speed expansion port 490. The low-speed expansion port 490, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device, such as a switch or a router, for example, through a network adapter.
[0041] As shown, the computing device 400 may be implemented in a variety of different forms. For example, the computing device may be implemented as a standard server 400a or multiple times in a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.
[0042] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor that can be either special purpose or general purpose and can be coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0043] A software application (i.e., software resource) may refer to computer software that enables a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0044] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0045] The processes and logic flows described in this specification can be performed by one or more programmable processors (also referred to as data processing hardware), which execute one or more computer programs to perform functions by operating on input data and generating outputs. The processes and logic flows can also be performed by a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). For example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as a magnetic disk, a magneto-optical disk, or an optical disk, or be operably coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer does not have to have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0046] To provide interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen) for displaying information to the user, and possibly a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, speech, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.
[0047] Unless expressly stated to the contrary, the phrase “at least one of A, B, or C” is intended to refer to any combination or subset of A, B, and C, such as: (1) only at least one A; (2) only at least one B; (3) only at least one C; (4) at least one A and at least one B; (5) at least one A and at least one C; (6) at least one B and at least one C; and (7) at least one A and at least one B and at least one C. Furthermore, unless expressly stated to the contrary, the phrase “at least one of A, B, and C” is intended to refer to any combination or subset of A, B, and C, such as: (1) only at least one A; (2) only at least one B; (3) only at least one C; (4) at least one A and at least one B; (5) at least one A and at least one C; (6) at least one B and at least one C; and (7) at least one A and at least one B and at least one C. Furthermore, unless expressly stated to the contrary, “A or B” is intended to refer to any combination of A and B, such as: (1) only A; (2) only B; and (3) A and B.
[0048] A variety of implementations have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Therefore, other implementations are within the scope of the following claims.
Claims
1. A computer-implemented method (300) executed by data processing hardware (510), the computer-implemented method causing the data processing hardware (510) to perform an operation, the operation include: receiving distilled data (220) including a plurality of out-of-domain training utterances (222); For each specific out-of-domain training utterance (222) of the distilled data (220): generating corresponding enhanced out-of-domain training utterances (232); as well as generating pseudo labels (240) corresponding to the corresponding augmented out-of-domain training utterances (232) using a teacher automatic speech recognition (ASR) model (170) trained on training data corresponding to the target domain; as well as The student ASR model (210) is distilled from the teacher ASR model (170) by training the student ASR model (210) using the corresponding augmented out-of-domain training utterances (232) paired with the corresponding pseudo labels (240) generated by the teacher ASR model (170).
2. The computer-implemented method (300) of claim 1, in, Distilling the student ASR model (210) from the teacher ASR model (170) includes training the student ASR model (210) to recognize utterances corresponding to the target domain.
3. The computer-implemented method (300) of claim 1 or 2, in: Prior to receiving the distilled data (220), the teacher ASR model (170) is trained on the training data comprising a plurality of target domain training utterances; and After receiving the distilled data (220), the training data is unavailable to the teacher ASR model (170) and the student ASR model (210).
4. The computer-implemented method (300) of any one of claims 1 to 3, in: The training data includes a plurality of target domain training utterances; and One or more target domain training utterances among the plurality of target domain training utterances are not included in the distilled data (220).
5. The computer-implemented method (300) of any one of claims 1 to 4, in, Enhancing the corresponding out-of-domain training utterance (222) includes at least one of adding noise, adding reverberation, or manipulating time.
6. The computer-implemented method (300) of any one of claims 1 to 5, in: The student ASR model (210) is executed in a cloud computing environment (70); and The teacher ASR model (170) is executed on one or more user devices (10) in communication with the cloud computing environment (70).
7. The computer-implemented method (300) of claim 6, in, The training data includes a plurality of target domain training utterances, each target domain training utterance being received at a corresponding user device among the one or more user devices (10).
8. The computer-implemented method (300) of any one of claims 1 to 7, in, Generating the pseudo-label (240) corresponding to the corresponding enhanced out-of-domain training utterance (232) includes: Generate a probability distribution over possible speech recognition hypotheses, Wherein, the pseudo labels (240) include the N-best speech recognition hypotheses with the highest probabilities.
9. The computer-implemented method (300) of any one of claims 1 to 8, in, Distilling the student ASR model (210) from the teacher ASR model (170) includes, for each specific augmented out-of-domain training utterance (232): generating a transcription (212) corresponding to the specific augmented out-of-domain training utterance (232) using the student ASR model (210); generating a loss (252) based on the pseudo-label (240) corresponding to the specific augmented out-of-domain training utterance (232) and the transcription (212) corresponding to the specific augmented out-of-domain training utterance (232); and The loss (252) is used to update the parameters of the student ASR model (210).
10. The computer-implemented method (300) of claim 9, in, The loss (252) includes at least one of a cross entropy loss, a Kulbeck-Leibler (KL) divergence loss, or an L2 loss.
11. A system (400), include: Data processing hardware (410); as well as Memory hardware (420) in communication with the data processing hardware (410), the memory hardware (420) storing instructions that, when executed on the data processing hardware (410), cause the data processing hardware (410) to perform operations, the operations comprising: receiving distilled data (220) including a plurality of out-of-domain training utterances (222); For each specific out-of-domain training utterance (222) of the distilled data (220): generating corresponding augmented out-of-domain training utterances (232); and generating pseudo labels (240) corresponding to the corresponding augmented out-of-domain training utterances (232) using a teacher automatic speech recognition (ASR) model (170) trained on training data corresponding to the target domain; and The student ASR model (210) is distilled from the teacher ASR model (170) by training the student ASR model (210) using the corresponding augmented out-of-domain training utterances (232) paired with the corresponding pseudo labels (240) generated by the teacher ASR model (170).
12. The system (400) of claim 11, in, Distilling the student ASR model (210) from the teacher ASR model (170) includes training the student ASR model (210) to recognize utterances corresponding to the target domain.
13. The system (400) according to claim 11 or 12, in: Prior to receiving the distilled data (220), the teacher ASR model (170) is trained on the training data comprising a plurality of target domain training utterances; and After receiving the distilled data (220), the training data is unavailable to the teacher ASR model (170) and the student ASR model (210).
14. The system (400) according to any one of claims 11 to 13, in: The training data includes a plurality of target domain training utterances; and One or more target domain training utterances among the plurality of target domain training utterances are not included in the distilled data (220).
15. The system (400) according to any one of claims 11 to 14, in, Enhancing the corresponding out-of-domain training utterance (222) includes at least one of adding noise, adding reverberation, or manipulating time.
16. The system (400) according to any one of claims 11 to 15, in: The student ASR model (210) is executed in a cloud computing environment (70); and The teacher ASR model (170) is executed on one or more user devices (10) in communication with the cloud computing environment (70).
17. The system (400) of claim 16, in, The training data includes a plurality of target domain training utterances, each target domain training utterance being received at a corresponding user device among the one or more user devices (10).
18. The system (400) according to any one of claims 11 to 17, in, Generating the pseudo-label (240) corresponding to the corresponding enhanced out-of-domain training utterance (232) includes: Generate a probability distribution over possible speech recognition hypotheses, Wherein, the pseudo labels (240) include the N-best speech recognition hypotheses with the highest probabilities.
19. The system (400) according to any one of claims 11 to 18, in, Distilling the student ASR model (210) from the teacher ASR model (170) includes, for each specific augmented out-of-domain training utterance (232): generating a transcription (212) corresponding to the specific augmented out-of-domain training utterance (232) using the student ASR model (210); generating a loss (252) based on the pseudo-label (240) corresponding to the specific augmented out-of-domain training utterance (232) and the transcription (212) corresponding to the specific augmented out-of-domain training utterance (232); and The loss (252) is used to update the parameters of the student ASR model (210).
20. The system (400) of claim 19, in, The loss (252) includes at least one of a cross entropy loss, a Kulbeck-Leibler (KL) divergence loss, or an L2 loss.