Self-adaptive distillation
The method of distilling teacher ASR models into multilingual student models with adjustable loss weights addresses performance degradations in ASR models, enhancing accuracy and robustness across languages, particularly in low-resource scenarios.
Patent Information
- Application Number
- JP2023558805
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-26
- Filing Date
- 2021-12-07
- Publication Date
- 2025-09-08
- Estimated Expiration
- 2041-12-07
AI Technical Summary
Automatic speech recognition (ASR) models face performance degradations due to resource imbalances, affecting user experience, particularly in low-resource languages, where there is a lack of training data.
A method and system for distilling trained teacher ASR models into multilingual student models using adjustable distillation loss weights, allowing the student models to learn from both teacher and student training examples, with weights that can decrease over time, especially for RNN-T architectures.
Enhances the performance of multilingual ASR models by balancing learning processes, improving accuracy and robustness across multiple languages, especially in low-resource scenarios.
Smart Images

Figure 0007735419000006 
Figure 0007735419000007 
Figure 0007735419000008
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to self-adaptive distillation. [Background technology]
[0002] In recent years, automatic speech recognition (ASR) has gained popularity and has been more widely applied to languages around the world. However, some languages have limitations that affect the quality or robustness of ASR models. For example, languages may vary from those with high resources to those with low resources. Resources refer to the resources that an ASR model utilizes for training and to improve its accuracy and robustness. Due to resource imbalances, ASR models may encounter various performance degradations, which inevitably affect the user experience of applications or programs that employ ASR models. Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method for distilling one or more trained teacher ASR (automatic speech recognition) models into multilingual student models. When executed on data processing hardware, the computer-implemented method causes the data processing hardware to perform operations including receiving a plurality of teacher training examples and a plurality of student training examples. The operations also include training one or more teacher ASR models using the plurality of teacher training examples. Each teacher ASR model is configured to output a respective text representation of a respective speech input. The operations also include generating multilingual student ASR models by training the multilingual student ASR models using the plurality of student training examples and distilling the trained one or more teacher ASR models into multilingual student ASR models using adjustable distillation loss weights. Each student ASR model is configured to receive a speech input and output a text representation corresponding to the received speech input.
[0004] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, the one or more teacher ASR models are configured to collectively recognize fewer languages than the multilingual student ASR model; the adjustable distillation loss weights may include constant values; and in some additional implementations, training the multilingual student model is performed over n training steps, and the adjustable distillation loss weights include a decreasing function that decreases based on the n training steps.
[0005] In some examples, each of the one or more teacher ASR models and the multilingual student ASR models includes an RNN-T (recurrent neural network-transducer) architecture. In these examples, the adjustable distillation loss weights may include a reduction function based on an RNN-T loss corresponding to the one or more teacher ASR models. Alternatively, the adjustable distillation loss weights in these examples may include a reduction function based on a first RNN-T loss corresponding to the one or more teacher ASR models and a second RNN-T loss corresponding to the multilingual student ASR model. Here, the reduction function may decrease the first RNN-T loss corresponding to the one or more teacher ASR models over time instances and increase the second RNN-T loss corresponding to the multilingual student ASR model over time instances.
[0006] Each teacher ASR model of the one or more teacher ASR models may correspond to a monolingual ASR model, or the one or more teacher ASR models may correspond to a single multilingual ASR model.
[0007] Another aspect of the present disclosure provides a system for distilling one or more trained teacher ASR (automatic speech recognition) models into multilingual student models. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations including receiving a plurality of teacher training examples and a plurality of student training examples. The operations also include training one or more teacher ASR models using the plurality of teacher training examples. Each teacher ASR model is configured to output a respective text representation of a respective speech input. The operations also include generating multilingual student ASR models by training the multilingual student ASR models using the plurality of student training examples and distilling the trained one or more teacher ASR models into multilingual student ASR models using adjustable distillation loss weights. Each student ASR model is configured to receive a speech input and output a text representation corresponding to the received speech input.
[0008] This aspect may include one or more of the following optional features: In some implementations, the one or more teacher ASR models are configured to collectively recognize fewer languages than the multilingual student ASR model; the adjustable distillation loss weights may comprise constant values; and in some additional implementations, training the multilingual student model is performed over n training steps, and the adjustable distillation loss weights comprise a decreasing function that decreases based on the n training steps.
[0009] In some examples, each of the one or more teacher ASR models and the multilingual student ASR models includes an RNN-T (recurrent neural network-transducer) architecture. In these examples, the adjustable distillation loss weights may include a reduction function based on an RNN-T loss corresponding to the one or more teacher ASR models. Alternatively, the adjustable distillation loss weights in these examples may include a reduction function based on a first RNN-T loss corresponding to the one or more teacher ASR models and a second RNN-T loss corresponding to the multilingual student ASR model. Here, the reduction function may decrease the first RNN-T loss corresponding to the one or more teacher ASR models over time instances and increase the second RNN-T loss corresponding to the multilingual student ASR model over time instances.
[0010] Each teacher ASR model of the one or more teacher ASR models may correspond to a monolingual ASR model, or the one or more teacher ASR models may correspond to a single multilingual ASR model.
[0011] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0012] [Figure 1A] 1 is a schematic diagram of an exemplary audio environment using an adaptive automatic speech recognition model. [Figure 1B] 1 is a schematic diagram of an exemplary audio environment using an adaptive automatic speech recognition model. [Figure 2A] Schematic diagram of an exemplary adaptive model formed from two or more monolingual teacher models. [Figure 2B] Schematic diagram of an exemplary adaptive model formed from a single multilingual teacher model. [Figure 3] FIG. 2C is a schematic diagram of an exemplary model architecture for the adaptive model of FIGS. 1A-2B. [Figure 4]1 is a flowchart of an exemplary configuration of operations for a method of generating an adaptive model. [Figure 5] 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] Like reference symbols in the various drawings indicate like elements. 1A and 1B , in some implementations, a system 100 includes a user 10 communicating spoken utterances 12 to a voice-enabled device 110 (also referred to as a device 110 or a user device 110). The user 10 (i.e., the speaker of the utterance 12) may speak the utterances 12 as queries or commands to solicit a response from the device 110 or to cause the device 110 to perform a task specified by the query. The device 110 is configured to capture sounds from one or more users 10 within the audio environment of the user device 110. As used herein, audio sounds may refer to utterances 12 spoken by the user 10 that function as audible queries, commands to the device 110, or audible communications captured by the device 110. A voice-enabled system (e.g., a digital assistant interface) of or associated with the device 110 may address the command or query by answering the query and / or causing the command to be executed.
[0014] Here, device 110 captures voice data 14 corresponding to utterances 12 spoken by user 10. Device 110 may correspond to any computing device associated with user 10 and capable of receiving voice data 14. Some examples of user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, e-readers, etc.), computers, wearable devices (e.g., smart watches), music players, casting devices, smart appliances (e.g., smart TVs) and Internet of Things (IoT) devices, remote controls, smart speakers, etc. Device 110 includes data processing hardware 112 and memory hardware 114. Memory hardware 114 communicates with data processing hardware 112 and stores instructions that, when executed by data processing hardware 112, cause data processing hardware 112 to perform one or more operations related to voice and / or text processing. In some examples, device 110 includes one or more applications (i.e., software applications). Each application may utilize one or more audio processing systems 140, 200 associated with the device 110 to perform various functions within the application.
[0015] Device 110 further includes an audio subsystem having an audio capture device (e.g., a microphone) 116 for capturing and converting audio data 14 in the audio environment into electrical signals, and an audio output device (e.g., a speaker) 118 for communicating an audible audio signal (e.g., a synthesized playback signal 154 from device 110). Although device 110 implements a single audio capture device 116 in the illustrated example, device 110 may implement a series of audio capture devices 116 to communicate with the audio subsystem (e.g., peripherals of device 110) without one or more of the series of audio capture devices 116 physically residing on device 110 without departing from the scope of this disclosure. For example, device 110 may correspond to a vehicle infotainment system that utilizes a series of microphones located throughout the vehicle.
[0016] Additionally, the device 110 is configured to communicate with a remote system 130 over the network 120. The remote system 130 may include remote resources 132, such as remote data processing hardware 134 (e.g., a remote server or CPU) and / or remote memory hardware 136 (e.g., a remote database or other storage hardware). The device 110 may utilize the remote resources 132 to perform various functions related to voice processing. For example, the device 110 is configured to perform voice recognition using a voice recognition system 140. These systems 140, 200 may reside on the device 110 (referred to as on-device systems) or may reside remotely (e.g., on the remote system 130) while communicating with the device 110. In some embodiments, some of these systems 140, 200 reside locally or on-device, and others reside remotely. In other words, these systems 140, 200 may be local or remote in any combination. For example, if the size or processing requirements of the systems 140, 200 are significant, the systems 140, 200 may reside in the remote system 130. However, if the device 110 can support the size or processing requirements of one or more of the systems 140, 200, one or more of the systems 140, 200 may reside on the device 110 using the data processing hardware 112 and / or memory hardware 114. Alternatively, one or more of the systems 140, 200 may reside locally or both on-device and remotely. For example, one or more of the systems 140, 200 may run by default on the remote system 130 if a connection to the network 120 between the device 110 and the remote system 130 is available, but if the connection is lost or the network 120 is unavailable, the systems 140, 200 instead run locally on the device 110.
[0017] The speech recognition system 140 receives speech data 14 as input and uses an adaptive ASR (automatic speech recognition) model 200 (also referred to as adaptive model 200) to transcribe the speech signal into a transcription 142 as output. Generally, by converting the speech data 14 into a transcription 142, the speech recognition system 140 enables the device 110 to recognize when a spoken utterance 12 from the user 10 corresponds to a query, a command, or some other form of voice communication. The transcription 142 refers to a sequence of text that the device 110 may then use to generate a response to the query or command. For example, if the user 10 asks the device 110 the question, "What's the weather going to be like today?", the device 110 passes speech data 14 corresponding to the question, "What's the weather going to be like today?" to the speech recognition system 140. The speech recognition system 140 converts the speech data 14 into a transcript that includes the text, "What's the weather going to be like today?" Device 110 may then use the text or portions of the text to determine a response to the query. For example, to determine the weather for the current day (i.e., today), device 110 passes the text (e.g., "What's the weather like today?") or identifying portions of the text (e.g., "weather" and "today") to a search engine. The search engine may then return one or more search results that device 110 interprets to generate a response for user 10.
[0018] 1B, the adapted model 200 of the speech recognition system 140 may be a multilingual speech recognition model. A multilingual speech recognition model is a model that can generate transcriptions 142 in more than one language (i.e., multiple languages). For example, FIG. 1B illustrates how the speech recognition system 140 receives speech data 14 and how the multilingual adapted model 200 converts the speech data 14 corresponding to the utterance 12, "What's the weather going to be like today?" into three different transcriptions 142, 142a-c. Here, the first transcription 142a is a Spanish (denoted SP) translation of "What's the weather going to be like today?",
[0019] [Table 1]
[0020] The second transcription 142b is a Swedish (denoted SW) translation of "What's the weather going to be like today?"
[0021] [Table 2]
[0022] The third transcription, 142c, is a German (denoted DE) translation of "What's the weather going to be like today?"
[0023] [Table 3]
[0024] Multilingual speech recognition models can be advantageous for multilingual users who can speak different languages, or to improve the performance of speech recognition models for low-resource languages by learning shared representations from available data from other languages (i.e., high-resource languages). For example, there may be a large amount of training data for a language like U.S. English, but a small amount of training data for a language like Zulu. Here, if adapted model 200 is a multilingual model, it can leverage a large amount of U.S. English training examples to compensate for the lack of training data for Zulu.
[0025] 2A and 2B illustrate an example process for generating an adaptive model 200 (also referred to as a student model 200). The adaptive model 200 may be formed from one or more teacher models 210, such that the adaptive model 200 may be referred to as a student model. That is, the teacher model 210 has a neural network that is distilled into a student model (e.g., the adaptive model 200) to form or in some way influence the neural network of the student model. Distillation generally refers to the process of training a neural network using a pre-trained network. Using distillation, neurons of the pre-trained network that are less important to the desired output (e.g., similar to deadweight) may be reduced to form a more streamlined neural network (i.e., a distilled neural network). Distillation may enable the distilled neural network to be more accurate and / or more compact in size when compared to the pre-trained network. In other words, when a pre-trained network is formed, the pre-trained network may have formed neurons that ultimately have less influence on the desired output once training of the pre-trained network is complete. Therefore, the pre-trained network contains neurons that may be removed or modified to reduce the detrimental effects from these neurons or to remove unnecessary neurons. For ASR models, distillation can be advantageous in low-resource situations where a student model may learn behavior in low-resource situations from a teacher model that learned in high-resource situations.
[0026] However, distillation presents challenges in imparting knowledge to the student model 200. For example, one difficulty when performing knowledge distillation on a student model 200 is how to balance the learning processes 220. That is, the student model 200 may be taught by both the distillation process 220, 220a and its own training process 220, 220b. Because multiple learning processes 220 are involved in generating the student model 200, the performance of the trained student model 200 may vary based on the balance between these processes 220. During the learning process 220, one or more teacher models 210 are first trained to establish a neural network for the distillation process 220a. During the training process for one or more teacher models 210, the teacher model 210 receives multiple teacher training samples 152, 152a-n (e.g., from the training sample database 150) and trains using the teacher training samples 152 to teach each teacher model 210 to predict, as output, a textual representation of a respective speech input. In this regard, the training samples (e.g., teacher training samples 152 or student training examples 154) allow the model to learn ground truth because the training samples 152, 154 include audio samples and corresponding transcriptions (i.e., textual representations) of the audio samples. Once one or more teacher models 210 are trained, the trained one or more teacher models 210 may then distill their knowledge into the student model 200.
[0027] In addition to the distillation process 220a from one or more teacher models 210, the student model 200 also learns from training processes 220, 220b. In the training process 220b, much like the teacher training process, the student model 200 learns from the student training samples 154 so that it can predict text expressions. Using both the distillation process 220a and the training process 220b, the student model 200 is configured to balance how much knowledge it gains from these processes 220a, 220b by using weights 222, 222a-b. That is, each process 220a, 220b is a series of training steps. In each training step, the loss of each process 220 is calculated and used to influence the next training step. For example, generally speaking, the student model 200 wants to minimize the loss of a given process 220 to approach a neural network that can accurately predict text expressions for a given input speech. Because each process 220 has an associated loss, the entire learning process can be represented by a total loss as a combination of the distillation loss for the distillation process 220a and the training loss (e.g., RNN-T loss) for the training process 220b. Thus, to determine how the student model 200 balances these processes 220a, 220b, the student model 200 uses adjustable weights 222 that are applied to either process loss. In some examples, the adjustable weights 222 are referred to as adjustable distillation weights 222a because they are applied to the distillation loss.
[0028] In some configurations, the adjustable distillation weights 222a are configured as constant values. However, in other configurations, the adjustable distillation weights 222a may be a decreasing function that decreases as the number of training steps increases. That is, the student model 200 becomes less concerned about distillation process losses over time. When the one or more teacher models 210 and the student model 200 have an RNN-T model architecture (e.g., in an end-to-end streaming application), the adjustable distillation weights 222a may be a decreasing function based on the RNN-T losses corresponding to the one or more teacher models 210. Furthermore, with the RNN-T architecture of both models 200, 210, the adjustable distillation weights 222a may also compensate for the RNN-T losses of the student model 200. Here, the adjustable distillation weights 222a may be a decreasing function based on a first RNN-T loss from the one or more teacher models 210 and an increasing function based on a second RNN-T loss from the student model 200, thereby compensating for the RNN-T losses of the student model 200.
[0029] 2A, the one or more teacher models 210 may correspond to multiple monolingual teacher models 210 that each distill their knowledge into the student model 200 to form the multilingual student model 200. Alternatively, FIG. 2B illustrates the one or more teacher models 210 as a single multilingual teacher model 210 that distills its knowledge into the multilingual student model 200. In either of these cases, the one or more teacher models 210 may collectively recognize fewer languages than the resulting multilingual student model 200. For example, because the student model 200 has its own training process 220a, the student model 200 can expand its language base to include more languages than the teacher model 210 distilled into the student model 200.
[0030] Referring to FIG. 3, the adaptive distillation process may be applicable to different types of speech recognition models. Generally, when the adapted model 200 is a specific type of speech recognition model, both the teacher model 210 and the student model (i.e., the adapted model 200) are the same type of model for distillation. One model that is gaining popularity is a sequence-to-sequence model known as RNN-T (Recurrent Neural Network Transducer). RNN-T does not employ an attention mechanism. Unlike other sequence-to-sequence models that generally require processing an entire sequence (e.g., an audio waveform) to generate an output (e.g., a sentence), RNN-T processes input samples continuously and streams output symbols. This feature is particularly useful for real-time communications. For example, speech recognition using an RNN-T may output spoken characters one by one. Here, the RNN-T uses a feedback loop to feed back the symbols predicted by the model to itself to predict the next symbol. Because decoding an RNN-T involves a beam search through a single neural network instead of a large decoder graph, an RNN-T can be compacted to a fraction of the size of a server-based speech recognition model. Due to its reduced size, an RNN-T can be deployed entirely on-device and run offline (i.e., without a network connection), thereby avoiding unreliability issues associated with communication networks.
[0031] When the adapted model 200 is an RNN-T model, it is a neural network model corresponding to an encoder-decoder framework that can be trained end-to-end to map an input sequence (e.g., an input speech signal) to a target sequence (e.g., words or characters spoken in the speech signal). In other words, given an input sequence (e.g., a real-valued vector), the RNN-T model attempts to predict a target sequence of labels. Here, the input sequence may be a raw feature vector, such as a log-mel filterbank energy feature or other neural network coded feature.
[0032] Continuing to refer to FIG. 3, the adapted model 200 includes an encoder network 302 and a decoder network 304. The encoder network 302 includes an encoder 310. The encoder 310 receives a d-dimensional feature vector x=(x1, x2, ..., x T ) sequence, where
[0033]
number
[0034] , and at each time step, generates a high-order feature representation, also called an encoder embedding e. The decoder network 304 receives the high-order feature representation e and decodes the high-order feature representation e using the joint layer 320 and the prediction network 330. The joint layer 320 in combination with the prediction network 330 can be viewed as a feedforward neural network, where the joint layer 320 computes logits that are fed to the prediction network 330. In other words, the joint layer 230 computes the decoder output y r To generate the high-dimensional feature representation e output by the encoder network 302, we apply it to the previous prediction y r-1 The decoder output is the same as the previous unit {y i-1,...,y0} and given input x, the current subword unit y i Probability distribution for
[0035]
number
[0036] Although not shown, the decoder network 304 may be configured to receive the output y r The output of the softmax layer is then used in a beam search process to select the orthogonal elements. The softmax layer may be integrated with the decoder network 304 or may be separate, depending on the configuration of the model 200.
[0037] 4 is a flowchart of an example configuration of operations for a method 400 of distilling one or more trained teacher ASR (automatic speech recognition) models 210 into a multilingual student model 200. At operation 402, the method 400 includes receiving a plurality of teacher training examples and a plurality of student training examples.
[0038] At operation 404, the method 400 includes using the plurality of teacher training examples to train one or more teacher ASR models 210. Each teacher ASR model 210 is configured to output a respective text representation of a respective speech input.
[0039] At operation 406, the method includes generating a multilingual student ASR model 200 by performing suboperations 406a and 406b. Suboperation 406a includes training the multilingual student ASR model 200 using a plurality of student training examples. The student ASR model 200 is configured to receive speech input and output a text representation corresponding to the received speech input. Suboperation 406b includes distilling one or more trained teacher ASR models 210 into the multilingual student ASR model 200 using adjustable distillation loss weights.
[0040] 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described herein. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the implementation of the invention as described and / or claimed herein.
[0041] Computing device 500 includes a processor 510 (e.g., data processing hardware 134), a memory 520 (e.g., memory hardware 136), a storage device 530, a high-speed interface / controller 540 that connects to memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 that connects to a low-speed bus 570 and storage device 530. Each of components 510, 520, 530, 540, 550, and 560 are interconnected using various buses and may be implemented on a common motherboard or in other manners as desired. Processor 510 can process instructions for execution within computing device 500, including instructions stored in memory 520 or storage device 530 to display graphical information for a GUI (graphical user interface) on an external input / output device, such as a display 580 connected to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used as desired, along with multiple memories and memory types. Also, multiple computing devices 500 may be connected, each providing a portion of the required operations, for example as a bank of servers, a group of blade servers, or a multi-processor system.
[0042] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and ROM (Read Only Memory) / PROM (Programmable Read Only Memory) / EPROM (Erasable Programmable Read Only Memory) / EEPROM (Electronically Erasable Programmable Read Only Memory) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, RAM (Random Access Memory), DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), PCM (Phase Change Memory), and disk or tape.
[0043] The storage device 530 is capable of providing mass storage for the computing device 500. In some implementations, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 may be a floppy disk device, a hard disk device, an optical disk device, a tape device, a flash memory or other similar solid-state memory device, or an array of devices including devices in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied as an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 520, the storage device 530, or memory on the processor 510.
[0044] The high-speed controller 540 manages bandwidth-intensive operations for the computing device 500, while the low-speed controller 560 manages less bandwidth-intensive operations. This duty allocation is merely exemplary. In some implementations, the high-speed controller 540 is connected to the memory 520, the display 680 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 550, which may accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is connected to the storage device 530 and a low-speed expansion port 590. The low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be connected to one or more input / output devices, such as a keyboard, pointing device, scanner, or network device, such as a switch or router, for example, through a network adapter.
[0045] Computing device 500 may be implemented in a number of different forms, as shown in the figure. For example, computing device 500 may be implemented as a laptop computer 500b, as part of a rack server system 500c, or as a standard server 500a or multiple times in a group of such servers 500a.
[0046] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or translatable on a programmable system with one or more special-purpose or general-purpose programmable processors coupled to send data and instructions to, and receive data and instructions from, a storage device, one or more input devices, and one or more output devices.
[0047] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, PLD (Programmable Logic Device)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0048] The processes and logic flows described herein may be implemented by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from, transfer data to, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0049] To provide for user interaction, one or more aspects of the present disclosure may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be captured in any form, including acoustic input, speech input, or tactile input. Furthermore, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0050] While multiple implementations have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (400) that, when executed by data processing hardware (134), causes the data processing hardware (134) to perform an operation, comprising: The operation is receiving a plurality of teacher training examples (152) and a plurality of student training examples (154); training one or more teacher automatic speech recognition (ASR) models (210) using the plurality of teacher training examples (152), each teacher ASR model (210) configured to output a respective text representation of a respective speech input (14); generating a multilingual student ASR model (200), a student model training step of training the multilingual student ASR model (200) using the plurality of student training examples (154), the multilingual student ASR model (200) being configured to receive a speech input (14) and output a corresponding text representation (142) of the received speech input (14); distilling the one or more trained teacher ASR models (210) into the multilingual student ASR model (200); and generating the multilingual student ASR model (200) by A method (400) in which a distillation loss weight (222a) applied to a distillation loss in the distillation step and a training weight (222b) applied to a training loss in the student model training step are determined so as to minimize the total loss of the distillation loss and the training loss.
2. 10. The method of claim 1, wherein the one or more teacher ASR models are configured to collectively recognize fewer languages than the multilingual student ASR model.
3. The method (400) of claim 1 or 2, wherein the distillation loss weight (222a) comprises a constant value.
4. training the multilingual student ASR model (200) over n training steps; 3. The method of claim 1, wherein the distillation loss weight comprises a decreasing function that decreases based on the n training steps.
5. 5. The method (400) of any one of claims 1 to 4, wherein each of the one or more teacher ASR models (210) and the multilingual student ASR model (200) comprises a recurrent neural network-transducer (RNN-T) architecture.
6. 6. The method of claim 5, wherein the distilled loss weights include a decreasing function based on RNN-T losses corresponding to the one or more teacher ASR models.
7. 7. The method of claim 5 or 6, wherein the distillation loss weights include a function that decreases based on a first RNN-T loss corresponding to the one or more teacher ASR models and increases based on a second RNN-T loss corresponding to the multilingual student ASR model.
8. The method (400) of any one of claims 1 to 7, wherein each teacher ASR model (210) of the one or more teacher ASR models (210) corresponds to a monolingual teacher ASR model (210).
9. The method (400) of any one of claims 1 to 8, wherein the one or more teacher ASR models (210) correspond to a single multilingual ASR model.
10. A system (100), comprising: data processing hardware (134); memory hardware (136) in communication with the data processing hardware (135), the memory hardware (136) storing instructions that, when executed on the data processing hardware (134), cause the data processing hardware (134) to perform operations; The operation is receiving a plurality of teacher training examples (152) and a plurality of student training examples (154); training one or more teacher automatic speech recognition (ASR) models (210) using the plurality of teacher training examples (152), each teacher ASR model (210) configured to output a respective text representation of a respective speech input (14); generating a multilingual student ASR model (200), a student model training step of training the multilingual student ASR model (200) using the plurality of student training examples (154), the multilingual student ASR model (200) being configured to receive a speech input (14) and output a corresponding text representation (142) of the received speech input (14); distilling the one or more trained teacher ASR models (210) into the multilingual student ASR model (200); and generating the multilingual student ASR model (200) by A distillation loss weight (222a) applied to the distillation loss in the distillation step and a training weight (222b) applied to the training loss in the student model training step are determined so as to minimize the total loss of the distillation loss and the training loss, in the system (100).
11. The system of claim 10 , wherein the one or more teacher ASR models are configured to collectively recognize fewer languages than the multilingual student ASR model.
12. 12. The system (100) of claim 10 or 11, wherein the distillation loss weight (222a) comprises a constant value.
13. training the multilingual student ASR model (200) over n training steps; 12. The system (100) of claim 10 or 11, wherein the distillation loss weight (222a) comprises a decreasing function that decreases based on the n training steps.
14. 14. The system (100) of any one of claims 10 to 13, wherein each of the one or more teacher ASR models (210) and the multilingual student ASR model comprises a recurrent neural network-transducer (RNN-T) architecture.
15. The system of claim 14, wherein the distilled loss weights include a decreasing function based on RNN-T losses corresponding to the one or more teacher ASR models.
16. 16. The system of claim 14, wherein the distillation loss weights include a function that decreases based on a first RNN-T loss corresponding to the one or more teacher ASR models and increases based on a second RNN-T loss corresponding to the multilingual student ASR model.
17. 17. The system (100) of any one of claims 10 to 16, wherein each teacher ASR model (210) of the one or more teacher ASR models (210) corresponds to a monolingual teacher ASR model (210).
18. 18. The system (100) of any one of claims 10 to 17, wherein the one or more teacher ASR models (210) correspond to a single multilingual ASR model.
Citation Information
Patent Citations
Method for training a multilingual speech recognition network, speech recognition system, and multilingual speech recognition system
JP2020537765A
Large-scale multilingual speech recognition with a streaming end-to-end model
WO2020242580A1