Knowledge distillation from non-streaming encoder to streaming encoder

By applying knowledge distillation and dedicated loss functions in the encoder part of the ASR model, the migration difficulty of non-streaming ASR models to streaming ASR models is solved, achieving faster training and better on-device real-time ASR performance.

CN120752696APending Publication Date: 2025-10-03QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480014101.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-19
Filing Date
2024-02-16
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively transfer knowledge from non-streaming ASR models to streaming ASR models, resulting in the need to label all data during training and misaligned output, which affects the performance of real-time ASR on the device.

Method used

By applying knowledge distillation (KD) only to the encoder part of the ASR model during training and using a dedicated loss function and transformer attention mask to avoid labeled data, the performance of the streaming ASR model is improved.

Benefits of technology

This enables faster training and processing, avoids output misalignment issues, and improves on-device real-time ASR performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752696A_ABST
    Figure CN120752696A_ABST
Patent Text Reader

Abstract

An example apparatus includes a memory configured to store a speech signal representing speech and a streaming model. The streaming model includes an on-device real-time streaming model. An apparatus includes one or more processors implemented in circuitry coupled to a memory. The one or more processors are configured to determine one or more words in the speech signal based on one or more migrations of learned knowledge from a non-streaming model to a streaming model. The one or more processors are further configured to take an action based on the determined one or more words.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Application No. 18 / 355,055, filed on July 19, 2023, and U.S. Provisional Application No. 63 / 487,449, filed on February 28, 2023, the entire contents of which are incorporated herein by reference. U.S. Application No. 18 / 355,055, filed on July 19, 2023, claims the benefit of U.S. Provisional Application No. 63 / 487,449, filed on February 28, 2023. Technical Field

[0002] The present disclosure relates to non-streaming model encoders and streaming model encoders. Background Art

[0003] Automatic speech recognition (ASR) models may be used to automatically recognize speech, such as words, thereby facilitating processing of such speech into text, commands, queries, etc. ASR models may be part of popular digital assistants, which may include standalone virtual assistant devices, smartphone applications, etc. Summary of the Invention

[0004] The present disclosure generally relates to techniques and devices for speech-related streaming models (such as ASR models), and to training techniques for such models. Various aspects of the techniques of the present disclosure can provide improved streaming model performance. Although the techniques of the present disclosure are generally discussed with respect to ASR models, these techniques are applicable to any speech-related model that can be classified as non-streaming or streaming.

[0005] There may be an information gap between non-streaming ASR models and streaming ASR models, where non-streaming ASR models generally perform better than streaming ASR models. However, non-streaming ASR models may also have issues. The processing associated with streaming ASR models generally has much lower latency because there is no need to wait for the speech (e.g., utterance) to end before starting to process the captured utterance. Therefore, non-streaming ASR models may not be desirable for real-time ASR on devices (e.g., in non-cloud computing environments) because the latency properties of streaming ASR models may be more suitable for real-time ASR on devices.

[0006] Knowledge distillation (KD) is a technique that can be used to transfer learned knowledge from one model (sometimes called a teacher) to another model (sometimes called a student). For example, KD can be used to distill or compress one or more large models to train smaller models. KD techniques can be used to improve the performance of streaming ASR models while maintaining the latency benefits of streaming ASR models. For example, a streaming ASR model student can be trained by applying KD techniques from a non-streaming ASR model teacher. In this case, the streaming ASR model (streaming ASR model student) can imitate the behavior of the non-streaming ASR model teacher.

[0007] However, applying KD from a non-streaming ASR model teacher to a streaming ASR model student can be difficult. Because the non-streaming ASR model and the streaming ASR model may use very different contexts, the streaming ASR model student may often fail to follow the non-streaming ASR model teacher. Previous KD studies have applied distillation to the final output probabilities, but this approach may have problems: (1) all data may need to be annotated (e.g., text transcription of captured audio data); and (2) the output data alignment between the non-streaming ASR model teacher and the streaming ASR model student often does not match (e.g., the output data are misaligned).

[0008] Therefore, it may be desirable to overcome these difficulties in applying KD from a non-streaming model teacher to a streaming model student. According to the techniques of the present disclosure, a system may apply KD only to the system's encoder, which may be a portion (e.g., not the entire model). The techniques of the present disclosure may result in faster training and / or processing, may not require labeling all data, and may not result in misalignment of outputs between a non-streaming model teacher and a streaming model student.

[0009] In one example, various aspects of these techniques relate to a device comprising: a memory configured to store a speech signal representing speech and a streaming model, the streaming model comprising an on-device real-time streaming model; one or more processors implemented in a circuit coupled to the memory, the one or more processors configured to: determine one or more words in the speech signal based on one or more transfers of learned knowledge from a non-streaming model to the streaming model; and take an action based on the determined one or more words.

[0010] In another example, various aspects of the techniques relate to a method comprising: determining one or more words in a speech signal based on one or more transfers of learned knowledge from a non-streaming model to a streaming model, the streaming model comprising a real-time on-device streaming model; and taking an action based on the determined one or more words.

[0011] In another example, various aspects of these techniques relate to a method that includes migrating learned knowledge from a non-streaming model to an on-device real-time streaming model.

[0012] In another example, various aspects of the techniques relate to a non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine one or more words in a speech signal based on one or more transitions of learned knowledge from a non-streaming model to a streaming model, the streaming model including a real-time on-device streaming model; and take an action based on the determined one or more words.

[0013] In another example, various aspects of the techniques relate to a device comprising: components for determining one or more words in a speech signal based on one or more transitions of learned knowledge from a non-streaming model to a streaming model, the streaming model comprising a real-time streaming model on a device; and components for taking an action based on the determined one or more words.

[0014] The details of one or more examples of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of various aspects of the technology will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a block diagram of an example system for automatic speech recognition in accordance with the techniques of this disclosure.

[0016] Figure 2 is a block diagram illustrating a specific implementation of a system for training a streaming model according to the techniques of this disclosure.

[0017] Figure 3 is a block diagram illustrating an example application of KD from a non-streaming ASR model encoder teacher to a streaming ASR model encoder student in accordance with the techniques of this disclosure.

[0018] Figure 4 is a conceptual diagram illustrating an example use of a KD loss function according to the techniques of this disclosure.

[0019] Figure 5 is a conceptual diagram illustrating an example transformer attention mask according to techniques of this disclosure.

[0020] Figure 6 is a chart illustrating the test results of the streaming ASR model encoder.

[0021] Figure 7 is a block diagram illustrating an example of a device according to the technology of this disclosure.

[0022] Figure 8 is a flow chart illustrating an example technique for KD from a non-streaming encoder to a streaming encoder in accordance with one or more aspects of the present disclosure. DETAILED DESCRIPTION

[0023] ASR models can be classified into two groups: non-streaming ASR models or streaming ASR models. Non-streaming ASR models can use the entire captured audio signal to transcribe captured utterances (e.g., phrases, sentences, commands, queries, etc.). An utterance can be a continuous speech segment or an uninterrupted chain of spoken language that can be started and / or ended with pauses. For example, an utterance can be a word, a sentence, or a sentence fragment (e.g., one or more words). The entire captured audio signal may include the captured utterance. Therefore, a non-streaming ASR model may not make inferences about the captured audio signal (e.g., what words are in the spoken utterance) until the entire utterance is captured. A non-streaming ASR model can be implemented as a neural network with multiple non-streaming layers. Each layer of a non-streaming ASR model can perform a specific task. Such layers may include an input layer, an output layer, and one or more hidden layers.

[0024] A streaming ASR model may use only past context and therefore may not need to use the entire captured utterance. For example, a streaming ASR model may make inferences in real time based on a portion of the utterance captured up to that point in time. In some examples, a streaming ASR model may update inferences as more utterances are captured. A streaming ASR model may be implemented as a neural network with multiple streaming layers. Speech-related models other than ASR models may also be classified as non-streaming or streaming. Each layer of a streaming ASR model may perform a specific task. Such layers may include an input layer, an output layer, and one or more hidden layers.

[0025] Because there is an information gap between non-streaming ASR models and streaming ASR models, non-streaming ASR models typically perform better than streaming ASR models (which typically perform worse). However, the processing associated with streaming ASR models typically has much lower latency because there is no need to wait for the speech (e.g., utterance) to end before starting to process the captured utterance. Therefore, the latency properties of streaming ASR models are desirable for real-time ASR on devices.

[0026] KD techniques can be used to transfer learned knowledge from a non-streaming ASR model to a streaming ASR model in an attempt to improve the performance of the streaming ASR model while maintaining the latency benefits of the streaming ASR model. However, there may be difficulties in applying KD from a non-streaming ASR model teacher to a streaming ASR model student. Because the non-streaming ASR model and the streaming ASR model may use very different contexts, the streaming ASR model student may often be unable to follow the non-streaming ASR model teacher. Additionally, if KD is applied for the final output probability, all data may need to be labeled (e.g., text transcription of the captured audio data); and (2) the output data will most likely be misaligned between the non-streaming ASR model teacher and the streaming ASR model student, which may adversely affect further training. Therefore, it may be desirable to overcome these difficulties in applying KD from a non-streaming ASR model teacher to a streaming ASR model student.

[0027] The present disclosure relates to systems, devices, and techniques for applying KD only to an encoder of an ASR model and the resulting systems, devices, and techniques. Such an encoder can be part of (e.g., not the entire) ASR model. The techniques of the present disclosure can result in faster training and / or processing, may not require all data to be labeled, and may not result in misalignment of output between a non-streaming ASR model teacher and a streaming ASR model student.

[0028] Such techniques may include using auxiliary non-streaming layers during training. Additionally or alternatively, the system may include a dedicated loss function for KD from a non-streaming ASR model teacher to a streaming ASR model student. Compared to other techniques, the techniques of the present disclosure may achieve significant improvement margins and may not require labeled data, thereby fundamentally removing the heavy data labeling costs. The techniques of the present disclosure may not provide the additional overhead of the inference stage (e.g., the streaming ASR model makes inferences after training).

[0029] The disclosed techniques can improve the performance of streaming ASR models for various speech-related tasks, such as keyword detection, voice assistance, speaker verification, etc.

[0030] Disclosed are systems, devices, and methods for performing automatic speech recognition. The disclosed system is configured to automatically recognize speech and take actions based on the recognized speech (e.g., elements of the recognized speech, such as words of an utterance). The system can be integrated into devices such as mobile devices, smart speaker systems (e.g., speakers in a user's home that can play audio, receive verbal user commands, and perform actions based on the user commands), vehicles, robots, and the like.

[0031] To illustrate, a user may be using utensils in the kitchen and utter the command "Turn on the kitchen lights." The system may receive audio data corresponding to the user's speech (e.g., "Turn on the kitchen lights.") The system may identify words within the utterance "Turn on the kitchen lights" and respond by taking an action to turn on the kitchen lights.

[0032] Figure 1 is a block diagram of an example system for automatic speech recognition according to the techniques of this disclosure. System 100 includes a processor 102, a memory 104 coupled to processor 102, a microphone 110, a transmitter 140, and a receiver 142. Transmitter 140 and receiver 142 can be configured to facilitate interaction between system 100 and a second device 144 (e.g., a kitchen light switch or another device). System 100 may optionally include an interface device 106, a display device 108, a camera 116, and an orientation sensor 118. In some examples, system 100 is implemented in a smart speaker system (e.g., a wireless speaker and voice command device integrated with a virtual assistant). In other examples, system 100 is implemented in a mobile device such as a mobile phone (e.g., a smartphone), a laptop computer, a tablet computer, a computerized watch, etc. In other examples, system 100 is implemented in one or more Internet of Things (IoT) devices such as smart appliances, etc. In other examples, system 100 is implemented in a vehicle such as a car or an autonomous vehicle, etc. In other examples, system 100 is implemented in a robot. The system 100 is configured to perform automatic speech recognition and take actions based on the recognized speech.

[0033] The processor 102 is configured to automatically recognize speech, such as spoken utterances, and perform one or more tasks based on the recognized speech. Such tasks may include processing the recognized speech into text, responding to commands, responding to queries (such as requesting information from the internet by retrieving information from the internet), and the like. For example, the processor 102 may be configured to identify a specific spoken word or words of an utterance and take an action based on the identity of the word. The memory 104 is configured to store a streaming ASR model 120 including an encoder 121. Although the streaming ASR model 120 is illustrated as being stored in the memory 104 of the system 100 (e.g., an on-device model), in other implementations, the streaming ASR model 120 (or a portion thereof) may be stored remotely in a network-based storage device (e.g., the "cloud"). The encoder 121 may include multiple streaming layers. In some examples, the encoder 121 does not include any non-streaming layers. For example, when the encoder 121 has been trained, the encoder 121 may not include any non-streaming layers.

[0034] Microphone 110 (which may be one or more microphones) is configured to capture audio input 112 (e.g., speech) and generate input audio data 114 (e.g., a speech signal) based on the audio input 112. The audio input 112 may include utterances (e.g., speech) from a speaker (e.g., a person).

[0035] The processor 102 is configured to automatically recognize the input audio data 114. For example, the processor 102 may execute the streaming ASR model 120 to automatically recognize the input audio data 114. For example, as part of automatically identifying the word represented by the input audio data 114, the processor 102 may compare the input audio data 114 (or a portion thereof) to known models for different words. Such models may be learned by the streaming ASR model 120 as described in the present disclosure.

[0036] The processor 102 may take an action based on the recognized input audio data 114. For example, the processor 102 may process the recognized speech into text, respond to a command, respond to a query, etc. In some examples, the action may include converting the recognized input audio data 114 into a text string 150. For example, the processor 102 may be configured to perform speech-to-text conversion on the input audio data 114 to convert the input audio data 114, or a portion thereof that includes speech, into a text string 150. The text string 150 may include a text representation of the speech included in the input audio data 114.

[0037] Figure 2 is a block diagram illustrating a specific implementation of a system for training a streaming model according to the techniques of this disclosure. In some examples, system 200 may include or correspond to system 100 or portions thereof. For example, elements of system 200 may include or correspond to hardware within processor 102. Streaming ASR model 210 may be an example of streaming ASR model 120. Each of the elements of system 200 may be represented by hardware, such as via an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA), or the operations described with reference to the elements may be performed by one or more processors executing computer-readable instructions.

[0038] System 200 includes a streaming ASR model 210. Streaming ASR model 210 includes an encoder 202 and an inferer 204. Encoder 202 may include multiple streaming layers and may be configured to receive input audio data, such as input audio data 114, and encode the input audio data 114 for use by inferer 204. Inferer 204 may use the encoded input audio data to perform inference. Inference may include operations to identify words within speech of the audio input.

[0039] The streaming ASR model 210 can be trained using an already trained non-streaming ASR model 220 (e.g., using a KD technique). For example, the encoder 214 of the non-streaming ASR model 220, which includes multiple non-streaming layers, can be used to train the encoder 202. In some examples, the training of the encoder 202 by the encoder 214 can utilize multiple auxiliary non-streaming layers 230.

[0040] Once the encoder 202 is trained by the encoder 214, the non-streaming ASR model 220 and the auxiliary non-streaming layer 230 can be removed from the system 200 so that when the streaming ASR model 210 receives input audio data, the streaming ASR model 210 can automatically recognize speech in the input audio data, and the system 200 can take action 208 based on the recognized speech. This automatic speech recognition can be performed by the streaming ASR model 210 without any additional overhead (e.g., processing power, memory usage, latency, etc.) that may be associated with the auxiliary non-streaming layer 230 or the non-streaming ASR model 220. In some examples, the streaming ASR model 210 can perform automatic speech recognition on a device (e.g., on the system 100) and in real time.

[0041] Figure 3 is a block diagram illustrating an example application of KD from a non-streaming ASR model encoder teacher to a streaming ASR model encoder student according to techniques of this disclosure. System 300 includes a non-streaming ASR model 310, an auxiliary non-streaming layer 304, and a streaming ASR model 312. System 300 may represent an example of system 200 during training, wherein non-streaming ASR model 310 represents non-streaming ASR model 220, auxiliary non-streaming layer 304 represents auxiliary non-streaming layer 230, and streaming ASR model 212 represents streaming ASR model 210. Non-streaming ASR model 310 may include encoder 308, which may represent encoder 214. Streaming ASR model 312 may include encoder 302, which may represent encoder 202.

[0042] The encoder 302 can be trained by the encoder 308 via the auxiliary non-streaming layer 304 using a KD technique. For example, the encoder 302 can be trained by layer-by-layer distillation of selected layers (indicated by the layers pointed to by the dashed lines). The same input data may include labeled data (e.g., spoken words with annotations of the identified words) and / or unlabeled data (e.g., spoken words). In some examples, according to the techniques of this disclosure, the KD technique can be utilized without including any labeled data as input data for training the encoder 302.

[0043] During KD training of encoder 302 by encoder 308, auxiliary non-streaming layer 304 can be inserted between encoder 308 and encoder 302. In some examples, auxiliary non-streaming layer 304 can be inserted between encoder 302 and encoder 308 on streaming ASR model 312. In some examples, auxiliary non-streaming layer 304 can be inserted between encoder 302 and encoder 308 on non-streaming ASR model 310. In some examples, auxiliary non-streaming layer 304 can be inserted between encoder 302 and encoder 308 in a location other than on non-streaming ASR model 310 and streaming ASR model 312. Once encoder 302 is trained, auxiliary non-streaming layer 304 can be removed so as not to impact the overhead of operating streaming ASR model 312 when operating to automatically recognize speech (e.g., to make inferences).

[0044] When using the KD technique to train the encoder 302, the trained encoder 308 of the non-streaming ASR model 310 can be used as a teacher, and the encoder 302 can be used as a student. As shown, the layers of the encoder 302 can include only streaming layers.

[0045] In some examples, system 350 can use at least one of two losses: an ASR loss and a KD loss during the training of encoder 302 by encoder 308. The ASR loss can be used to train streaming ASR model 312 to accurately transcribe speech, and the KD loss can be used to influence encoder 302 (e.g., student) to better follow the behavior of encoder 308 (e.g., teacher).

[0046] If the input data is labeled, the system 350 may use both the ASR loss and the KD loss. If the input data is not labeled, the system 350 may use only the KD loss instead of both the ASR loss and the KD loss. KD ) can be applied to selected layers of the encoder 308 involved in training. The KD loss is shown as being applied between the selected layers of the encoder 308 and the auxiliary non-streaming layers 304.

[0047] Figure 4 is a conceptual diagram illustrating an example use of a KD loss function by a system according to the techniques of this disclosure. System 400 may be Figure 3 An example of system 300 is provided. System 400 may use a dedicated KD loss for streaming ASR training. For example, the KD loss function may be determined or calculated as a weighted sum of up to three losses. The three losses may include a distance (DIS) loss, a Kullback–Leibler divergence (KLD) loss, and an autoregressive predictive coding (APC) loss.

[0048] These three losses can be applied at different points. For example, the KLD loss can be applied between the non-streaming layer of the encoder 308 and the associated non-streaming layer of the auxiliary non-streaming layer 304. The DIS loss can be applied before the N-step shift performed by the encoder 308 and the unidirectional long short-term memory (LSTM) of the encoder 302. The APC loss can be applied after the N-step shift 402 performed by the encoder 308 and the unidirectional LSTM layer 404 of the encoder 302.

[0049] DIS loss can be determined as:

[0050]

[0051] where h is the output feature sequence of the teacher / student layer, D is the feature dimension, t is the index of each element in the sequence, and T is the number of elements in the sequence.

[0052] KLD loss can be determined as:

[0053]

[0054] Where A is the query / key / value matrix inside the self-attention layer of the transformer, h is the self-attention head index, H is the number of heads, and d h is the head feature dimension.

[0055] APC loss can be determined as:

[0056]

[0057] where K is a distance indicating how far away the target for prediction is located, and y is an output feature sequence of the LSTM layer (e.g., the unidirectional LSTM layer 404).

[0058] For example, the three losses can be reinterpreted as performing the following functions: (1) DIS loss: reducing the gap between the extracted features of each frame extracted by encoder 308 and encoder 302; (2) KLD loss: matching frame-to-frame relationships between all frames between encoder 308 and encoder 302; (3) APC loss: predicting future frames by using only past context. In some examples, the KLD loss can be a combination of three losses: KLD query loss, KLD key loss, and KLD value loss. For example, KLD query / key / value can correspond to the above KLD loss equation, where A is the query / key / value matrix, respectively.

[0059] The weighted sum of the three losses can be expressed as:

[0060]

[0061] Among them L KD is the knowledge distribution loss, LDIS is the distance loss, is the KLD query loss, is the KLD bond loss, is the KLD value loss, L APC is the APC loss, and α, β, and γ are weights, such as Figure 4 shown.

[0062] Figure 5 is a conceptual diagram illustrating an example transformer attention mask according to the techniques of this disclosure. Figure 4 The three losses discussed may be particularly helpful for streaming ASR models, where the model does not have access to future information when making inferences.

[0063] exist Figure 5 In the non-streaming attention mask 500, the X-axis represents the current frame index and the Y-axis represents the target frame. The non-streaming attention mask 500 represents the attention mask of the non-streaming ASR model. Using the non-streaming attention mask 500, during the 12 frames represented, each individual current frame will be able to access all 12 frames, whether they are past, present or future frames (relative to the current frame). Using Figure 5 The “block-by-block” streaming attention mask 502 of FIG. 5 , given a current frame, can only access frames within a “block”. For example, for a 2-frame block, the current frame can access the current frame and the nearest past frame or the nearest future frame that together with the current frame constitute the 2-frame block.

[0064] To improve Figure 4 In order to adjust the APC loss of the auxiliary non-streaming layer 304, the system 300 may modify the transformer attention mask. For example, instead of applying the non-streaming attention mask 500 or the block-by-block streaming attention mask 502, the system 400 may apply non-streaming attention to the APC loss mask 504. In the example of applying non-streaming attention to the APC loss mask 504, the current frame may not have access to the next four frames. With such application of the modified transformer attention mask during training, for example to the auxiliary non-streaming layer 304, the encoder 302 cannot directly obtain the APC target frame from the self-attention mechanism, thereby preventing the encoder 302 from "cheating". For example, the encoder 302 cannot infer the next frame by simply performing a "cut and paste" operation.

[0065] Figure 6is a chart illustrating test results for a streaming ASR model encoder. A streaming ASR model 312 having an encoder trained according to the techniques of the present disclosure (e.g., encoder 302) may exhibit improved ASR performance when compared to streaming ASR models trained in other ways. The streaming ASR models are trained using different techniques. A baseline model is tested without using KD. Another model is tested after being trained using a prior art technique that performs KD on output word unit probabilities. A third model is tested after training the encoder using the DIS loss as the KD loss. A fourth model is tested after training the encoder using the DIS and KLD losses to determine the KD loss. A fifth model is tested after training the encoder using the DIS, KLD, and APC losses to determine the KD loss. Figure 6 The metrics shown in are word error rates (WER) expressed as a percentage, where lower percentage WER is better than higher percentage WER. The dataset used for testing is LibriSpeech (dev-clean subset, dev-other subset). Figure 6 The numbers stated in are those compared using the same settings (e.g., the same epoch). Figure 6 It can be seen that the techniques of this disclosure result in better WER of the streaming ASR model compared to the prior art.

[0066] Figure 7 is a block diagram illustrating an example of a device according to the techniques of the present disclosure. In some examples, device 700 may be a wireless communication device, such as a smartphone. In some examples, device 700 may have a Figure 7 The illustrated components may be more or less than the illustrated components. In illustrative aspects, the device 700 may perform the following operations with reference to Figures 1 to 6 The discussed technique describes one or more operations.

[0067] In a particular implementation, the device 700 includes a processor 710, such as a central processing unit (CPU) or a digital signal processor (DSP), coupled to a memory 732. The memory 732 may include instructions 768 (e.g., executable instructions) such as computer-readable instructions or processor-readable instructions. The instructions 768 may include one or more instructions that can be executed by a computer, such as the processor 710. The memory 732 also includes instructions such as those described in reference 768. Figure 1 The streaming ASR model 120 is described.

[0068] The device 700 also includes a display controller 726 coupled to the processor 710 and to a display 728. A codec / decoder (CODEC) 734 may also be coupled to the processor 710. A speaker 736 and a microphone 738 may be coupled to the CODEC 734.

[0069] Figure 7 It is also illustrated that a wireless interface 740 (such as a wireless controller) and a transceiver 746 can be coupled to the processor 710 and the antenna 742 so that wireless data received via the antenna 742, the transceiver 746, and the wireless interface 740 can be provided to the processor 710. In some examples, the processor 710, the display controller 726, the memory 732, the codec 734, the wireless interface 740, and the transceiver 746 are included in a system-in-package or system-on-chip device 722. In some examples, an input device 730 and a power supply 744 are coupled to the system-on-chip device 722. Furthermore, in certain examples, as Figure 7 As illustrated, the display 728, input device 730, speaker 736, microphone 738, antenna 742, and power supply 744 are external to the system-on-chip device 722. In particular implementations, each of the display 728, input device 730, speaker 736, microphone 738, antenna 742, and power supply 744 can be coupled to a component of the system-on-chip device 722, such as an interface or controller.

[0070] In some examples, memory 732 includes or stores instructions 768 (e.g., executable instructions), such as computer-readable instructions or processor-readable instructions. For example, memory 732 may include or correspond to a non-transitory computer-readable medium storing instructions 768. Instructions 768 may include one or more instructions executable by a computer, such as processor 710.

[0071] In some examples, the device 700 includes a non-transitory computer-readable medium (e.g., memory 732) storing instructions (e.g., instructions 768) that, when executed by one or more processors (e.g., processor 710), may cause the one or more processors to perform operations including: determining one or more words in a speech signal (e.g., input audio data 114) based on learned knowledge from a non-streaming model (e.g., non-streaming ASR model 220) to a streaming model (e.g., streaming ASR model 210), including an on-device real-time streaming model (e.g., streaming ASR model 120). The instructions may also cause the one or more processors to take an action (e.g., action 208) based on the determined one or more words.

[0072] Device 700 may include a wireless telephone, a mobile communication device, a mobile device, a mobile phone, a smartphone, a cellular phone, a laptop computer, a desktop computer, a computer, a tablet computer, a set-top box, a personal digital assistant (PDA), a display device, a television, a game console, an augmented reality (AR) device, a virtual reality (VR) device, a music player, a radio, a video player, an entertainment unit, a communication device, a fixed location data unit, a personal media player, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a decoder system, an encoder system, a vehicle, a component of a vehicle, or any combination thereof.

[0073] The present disclosure may state that one or more words in the speech signal may be determined based on one or more transfers of learned knowledge from a non-streaming model to a streaming model, and that the streaming model is a real-time streaming model on the device. This language is intended to include the following examples. Example 1: The transfer of learned knowledge may occur while the streaming model is on the device (e.g., device 700), for example, while the streaming model resides in memory 732. Example 2: The transfer of learned knowledge may occur before the streaming model is installed on the device, where the streaming model is intended to perform real-time processing on the device, such as ASR. In other words, the streaming ASR model 120 may be trained (e.g., migrating learned knowledge from the non-streaming ASR model 220 to the streaming ASR model 210) while on a device different from the device 700 (e.g., in a lab, in a cloud computing environment, on a server, etc.) before being loaded onto the device 700. Example 3: The transfer of learned knowledge may occur partially before the streaming model is on the device and partially while the streaming model is on the device. Thus, transfer of learned knowledge may occur in one or more transfers (eg, in a single transfer or in multiple transfers).When residing on the device 700, the streaming ASR model 120 may be referred to as an on-device real-time streaming model.

[0074] It should be noted that various functions performed by one or more components of the system described with reference to FIG and device 700 are described as being performed by certain components or circuits. This division of components and circuits is for illustration only. In alternative aspects, the functions performed by a particular component may be divided among multiple components. Furthermore, in alternative aspects, reference to FIG Figures 1 to 7 Two or more components described may be integrated into a single component. Figures 1 to 7 Each component described may be implemented using hardware (e.g., field programmable gate array (FPGA) devices, application specific integrated circuits (ASICs), DSPs, controllers, etc.), software (e.g., instructions executable by a processor), or any combination thereof.

[0075] In conjunction with the described aspects, an apparatus or device may include a component for storing one or more category labels associated with one or more categories of a natural language processing library. The component for storing may include or correspond to Figure 1 Memory 104, Figure 7 a memory 732, one or more other structures or circuits configured to store one or more category labels associated with one or more categories of a natural language processing library, or any combination thereof.

[0076] The apparatus or device may further include means for processing. The means for processing may include means for determining one or more words in the speech signal (e.g., input audio data 114) based on one or more transfers of learned knowledge from a non-streaming model (e.g., non-streaming ASR model 220) to a streaming model (e.g., streaming ASR model 210), the streaming model including an on-device real-time streaming model (e.g., streaming ASR model 120). The means for processing may also include means for taking an action (e.g., action 208) based on the determined one or more words.

[0077] One or more of the disclosed aspects may be implemented in a system or device, such as device 700, which may include a communication device, a fixed location data unit, a mobile location data unit, a mobile phone, a cellular phone, a satellite phone, a computer, a tablet computer, a portable computer, a display device, a media player, or a desktop computer. Alternatively or additionally, device 700 may include a set-top box, an entertainment unit, a navigation device, a personal digital assistant (PDA), a monitor, a computer monitor, a television, a tuner, a radio, a satellite radio, a music player, a digital music player, a portable music player, a video player, a digital video player, a digital video disc (DVD) player, a portable digital video player, a satellite, a vehicle, a component integrated within a vehicle, any other device that includes a processor or stores or retrieves data or computer instructions, or any combination thereof. As another illustrative, non-limiting example, device 700 may include a remote unit, such as a handheld personal communication system (PCS) unit, a portable data unit (such as a device that supports a global positioning system (GPS)), a meter reading device, or any other device that includes a processor or stores or retrieves data or computer instructions, or any combination thereof.

[0078] Although Figure 7 The wireless communication device including the processor configured to perform automatic speech recognition is exemplified, but the processor configured to perform automatic speech recognition may be included in various other electronic devices. For example, the processor configured to perform the automatic speech recognition as described in reference Figures 1 to 7 The described automatic speech recognition processor may be included in one or more components of a base station.

[0079] A base station may be part of a wireless communication system. A wireless communication system may include multiple base stations and multiple wireless devices. The wireless communication system may be a Long Term Evolution (LTE) system, a Code Division Multiple Access (CDMA) system, a Global System for Mobile Communications (GSM) system, a Wireless Local Area Network (WLAN) system, or some other wireless system. A CDMA system may implement Wideband CDMA (WCDMA), CDMA 1X, Evolution Data Optimized (EVDO), Time Division Synchronous CDMA (TD-SCDMA), or some other version of CDMA.

[0080] Various functions may be performed by one or more components of the base station, such as sending and receiving messages and data (e.g., audio data). One or more components of the base station may include a processor (e.g., a CPU), a transcoder, a memory, a network connection, a media gateway, a demodulator, a transmit data processor, a receiver data processor, a transmit multiple-input multiple-output (MIMO) processor, a transmitter and a receiver (e.g., a transceiver), an antenna array, or a combination thereof. The base station or one or more of the components of the base station may include a processor configured to perform adaptive audio analysis, as described above with reference to Figures 1 to 7 described.

[0081] During operation of a base station, one or more antennas of the base station may receive a data stream from a wireless device. A transceiver may receive the data stream from the one or more antennas and may provide the data stream to a demodulator. The demodulator may demodulate the modulated signal of the data stream and provide the demodulated data to a receiver data processor. The receiver data processor may extract audio data from the demodulated data and provide the extracted audio data to a processor.

[0082] The processor may provide the audio data to a transcoder for transcoding. The decoder of the transcoder may decode the audio data from a first format into decoded audio data, and the encoder may encode the decoded audio data into a second format. In some embodiments, the encoder may encode the audio data using a data rate (e.g., up-conversion) or a data rate (e.g., down-conversion) higher than the data rate received from the wireless device. In other embodiments, the audio data may not be transcoded. The transcoding operation (e.g., decoding and encoding) may be performed by multiple components of the base station. For example, decoding may be performed by a receiver data processor, and encoding may be performed by a transmit data processor. In other embodiments, the processor may provide the audio data to a media gateway for conversion to another transmission protocol, a decoding scheme, or both. The media gateway may provide the converted data to another base station or a core network via a network connection.

[0083] Figure 81 is a flow chart illustrating an example technique for KD from a non-streaming encoder to a streaming encoder according to one or more aspects of the present disclosure. Processor 102 may determine one or more words in a speech signal based on one or more transfers of learned knowledge from a non-streaming model to a streaming model (800). For example, audio input 112 may include speech. Microphone 110 may capture speech in audio input 112 and convert audio input 112 into a signal, such as input audio data 114. Input audio data 114 may include a speech signal.

[0084] The processor 102 may execute a streaming ASR model 120. The streaming ASR model 120 may be an example of the streaming ASR model 210 and / or the streaming ASR model 312. The streaming ASR model 210 may be trained by transferring learned knowledge from the non-streaming ASR model 220 one or more times. The streaming ASR model 312 may be trained by transferring learned knowledge from the non-streaming ASR model 310 one or more times. This training may be utilized by the processor 102 executing the streaming ASR model 120 to determine one or more words in the speech signal. The streaming ASR model 120 may be an on-device, real-time streaming model. In some examples, the streaming ASR model 120 may be trained while resident in the memory 104. In some examples, the streaming ASR model 120 may be trained before the streaming ASR model 120 becomes resident in the memory 104. For example, the streaming ASR model 120 may be trained while the streaming ASR model 120 resides on a different device (eg, trained before being stored in the memory 104 ).

[0085] The processor 102 may take an action based on the determined one or more words (802). For example, the processor 102 may process the one or more words into text, respond to a command within the one or more words (e.g., turn on the kitchen lights in response to determining that the one or more words correspond to "turn on the kitchen lights"), respond to a query (e.g., retrieve and audibly present a weather report from the Internet in response to determining that the one or more words correspond to "what's the current weather like?"), etc.

[0086] In some examples, the non-streaming model includes a trained non-streaming model, and one or more transfers of learned knowledge from the non-streaming model to the streaming model include using the non-streaming model to train the streaming model. In some examples, one or more transfers of learned knowledge are based on an encoder configured to encode speech (e.g., encoder 121, encoder 202, encoder 214, encoder 302, and / or encoder 308). In some examples, the encoder includes multiple layers. In some examples, the encoder transfers knowledge at a selected layer among the multiple layers.

[0087] In some examples, one or more transfers of learned knowledge are from an encoder of a non-streaming model (e.g., encoder 214 and / or encoder 308) to an encoder of a streaming model (e.g., encoder 202 and / or encoder 302). In some examples, the streaming model includes a streaming ASR model (e.g., streaming ASR model 120, streaming ASR model 210, and / or streaming ASR model 312), and the non-streaming model includes a non-streaming ASR model (e.g., non-streaming ASR model 220 and / or non-streaming ASR model 310). In some examples, one or more transfers of learned knowledge are based on a plurality of auxiliary non-streaming layers (e.g., auxiliary non-streaming layer 230 and / or auxiliary non-streaming layer 304) between the streaming model and the non-streaming model. In some examples, one or more transfers of learned knowledge are based on associated with the plurality of auxiliary non-streaming layers (e.g., Figure 5 C) Modified attention mask.

[0088] In some examples, one or more transfers of learned knowledge are based on KD. In some examples, one or more transfers of learned knowledge are based on a KD loss function. In some examples, the KD loss function includes at least one of a distance loss, a KLD loss, or an autoregressive predictive decoding loss. In some examples, the KD loss function includes at least two of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive decoding loss. In some examples, the KD loss function includes a weighted sum of a distance loss, a Kullback-Leibler divergence loss, and an autoregressive predictive decoding loss. In some examples, the KD loss function includes:

[0089]

[0090] Among them L KD is the knowledge distribution loss, L DIS is the distance loss, is the KLD query loss, is the KLD bond loss, is the KLD value loss, L APC is the APC loss, and α, β, and γ are weights.

[0091] In some examples, speech includes an utterance. In some examples, an utterance includes one or more words.

[0092] In some examples, at least one of the one or more transfers of learned knowledge occurs before the streaming model is located on the device. For example, the transfer of learned knowledge can occur when the streaming ASR model 120 is located on the server. In some examples, at least one of the one or more transfers of learned knowledge occurs after the streaming model is located on the device. For example, the transfer of learned knowledge can occur when the streaming ASR model 120 is located on the device 700. In some examples, the one or more transfers of learned knowledge can occur partially before the streaming model is located on the device and partially after the streaming model is located on the device.

[0093] although Figures 1 to 8 One or more figures in the drawings may illustrate systems, devices and / or methods according to the teachings of the present disclosure, but the present disclosure is not limited to these illustrated systems, devices and / or methods. Figures 1 to 8 One or more functions or components in any of the diagrams may be used with Figures 1 to 8 Therefore, no single embodiment described herein should be interpreted as limiting, and the embodiments of the present disclosure may be appropriately combined without departing from the teachings of the present disclosure. As an example, refer to Figures 1 to 7 One or more of the operations described may be optional, may be performed at least partially concurrently, and / or may be performed in a different order than illustrated or described.

[0094] This disclosure includes the following non-limiting terms.

[0095] Item 1A. A device configured to automatically recognize speech, the device comprising: a memory configured to store a speech signal representing speech and a real-time streaming model on the device; one or more processors implemented in a circuit coupled to the memory, the one or more processors configured to: determine one or more words in the speech signal based on the migration of learned knowledge from a non-streaming model to the streaming model; and take action based on the determined one or more words.

[0096] Clause 2A. The apparatus of Clause 1A, wherein the transfer of learned knowledge is based on an encoder configured to encode the speech.

[0097] Clause 3A. The apparatus of clause 2A, wherein the encoder comprises a plurality of layers.

[0098] Clause 4A. The apparatus of clause 3A, wherein the encoder transfers knowledge at a selected layer of the plurality of layers.

[0099] Clause 5A. The apparatus of any of Clauses 2A to 4A, wherein the transfer of learned knowledge is from an encoder of the non-streaming model to an encoder of the streaming model.

[0100] Clause 6A. The apparatus of any of clauses 1A to 5A, wherein the streaming model comprises a streaming automatic speech recognition (ASR) model and the non-streaming model comprises a non-streaming ASR model.

[0101] Clause 7A. The apparatus of any of Clauses 1A to 6A, wherein the transfer of learned knowledge is based on a plurality of auxiliary non-streaming layers between the streaming model and the non-streaming model.

[0102] Clause 8A. The apparatus of Clause 7A, wherein the transfer of learned knowledge is based on modified attention masks associated with the plurality of auxiliary non-streaming layers.

[0103] Clause 9A. The apparatus of any of clauses 1A to 8A, wherein the transfer of learned knowledge is based on a knowledge distribution (KD).

[0104] Clause 10A. The apparatus of clause 9A, wherein the transfer of learned knowledge is based on a KD loss function.

[0105] Clause 11A. The apparatus of clause 10A, wherein the KD loss function comprises at least one of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

[0106] Clause 12A. The apparatus of Clause 11A, wherein the KD loss function comprises at least two of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

[0107] Clause 13A. The apparatus of Clause 12A, wherein the KD loss function comprises a weighted sum of a distance loss, a Kullback-Leibler divergence loss, and an autoregressive predictive coding loss.

[0108] Clause 14A. The apparatus of clause 13A, wherein the KD loss function comprises:

[0109]

[0110] Among them L KD is the knowledge distribution loss, L DIS is the distance loss, is the KLD query loss, is the KLD bond loss, is the KLD value loss, LAPC is the APC loss, and α, β, and γ are weights.

[0111] Clause 15A. The apparatus of any of clauses 1A to 14A, wherein the speech comprises speech.

[0112] Clause 16A. The apparatus of clause 15A, wherein the utterance comprises the one or more words.

[0113] Clause 17A. The device of any of clauses 1A to 16A, further comprising one or more microphones configured to capture the speech signal.

[0114] Clause 18A. The apparatus of any of Clauses 1A to 17A, wherein the transfer of learned knowledge occurs before the streaming model is located on the apparatus.

[0115] Clause 19A. The apparatus of any of Clauses 1A to 17A, wherein the transfer of learned knowledge occurs after the streaming model is on the apparatus.

[0116] Clause 20A. The apparatus of any of clauses 1A to 19A, wherein the action comprises processing speech into text, responding to a command, or responding to a query.

[0117] Clause 21A. A method comprising: determining one or more words in a speech signal based on the transfer of learned knowledge from a non-streaming model to a streaming model, wherein the streaming model is a real-time streaming model on a device; and taking an action based on the determined one or more words.

[0118] Clause 22A. The method of Clause 21A, wherein the transfer of learned knowledge is based on an encoder configured to encode the speech.

[0119] Clause 23A. The method of Clause 22A, wherein the encoder comprises a plurality of layers.

[0120] Clause 24A. The method of clause 23A, wherein the encoder transfers knowledge at a selected layer of the plurality of layers.

[0121] Clause 25A. The method of any one of Clauses 21A to 24A, wherein the transfer of learned knowledge is from an encoder of the non-streaming model to an encoder of the streaming model.

[0122] Clause 26A. The method of any one of clauses 21A to 25A, wherein the streaming model comprises a streaming automatic speech recognition (ASR) model, and the non-streaming model comprises a non-streaming ASR model.

[0123] Clause 27A. The method of any one of Clauses 21A to 26A, wherein the transfer of learned knowledge is based on a plurality of auxiliary non-streaming layers between the streaming model and the non-streaming model.

[0124] Clause 28A. The method of Clause 27A, wherein the transfer of learned knowledge is based on modified attention masks associated with the plurality of auxiliary non-streaming layers.

[0125] Clause 29A. The method of any one of Clauses 21A to 28A, wherein the transfer of learned knowledge is based on a knowledge distribution (KD).

[0126] Clause 30A. The method of Clause 29A, wherein the transfer of learned knowledge is based on a KD loss function.

[0127] Clause 31A. The method of clause 30A, wherein the KD loss function comprises at least one of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

[0128] Clause 32A. The method of Clause 31A, wherein the KD loss function comprises at least two of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

[0129] Clause 33A. The method of Clause 32A, wherein the KD loss function comprises a weighted sum of a distance loss, a Kullback-Leibler divergence loss, and an autoregressive predictive coding loss.

[0130] Clause 34A. The method of clause 33A, wherein the KD loss function comprises:

[0131]

[0132] Among them L KD is the knowledge distribution loss, L DIS is the distance loss, is the KLD query loss, is the KLD bond loss, is the KLD value loss, L APC is the APC loss, and α, β, and γ are weights.

[0133] Clause 35A. The method of any one of clauses 21A to 34A, wherein the speech comprises speech.

[0134] Clause 36A. The method of clause 35A, wherein the utterance comprises the one or more words.

[0135] Clause 37A. A method of training a streaming model, the method comprising: transferring learned knowledge from a non-streaming model to an on-device real-time streaming model.

[0136] Item 38A. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine one or more words in a speech signal based on the migration of learned knowledge from a non-streaming model to a streaming model, the streaming model including a real-time streaming model on a device; and take action based on the determined one or more words.

[0137] Item 39A. A device comprising: a component for determining one or more words in a speech signal based on the transfer of learned knowledge from a non-streaming model to a streaming model, the streaming model including a real-time streaming model on the device; and a component for taking an action based on the determined one or more words.

[0138] Item 1B. A device configured to automatically recognize speech, the device comprising: a memory configured to store a speech signal representing speech and a streaming model, the streaming model comprising an on-device real-time streaming model; one or more processors implemented in a circuit coupled to the memory, the one or more processors configured to: determine one or more words in the speech signal based on one or more transitions from a non-streaming model to the streaming model based on learned knowledge; and take action based on the determined one or more words.

[0139] Clause 2B. The apparatus of clause 1B, wherein the non-streaming model comprises a trained non-streaming model, and wherein the one or more transfers of learned knowledge from the non-streaming model to the streaming model comprise training the streaming model using the non-streaming model.

[0140] Clause 3B. The apparatus of Clause 1B or Clause 2B, wherein the one or more transfers of learned knowledge are based on an encoder configured to encode the speech.

[0141] Clause 4B. The apparatus of clause 3B, wherein the encoder comprises a plurality of layers.

[0142] Clause 5B. The apparatus of clause 4B, wherein the encoder transfers knowledge at a selected layer of the plurality of layers.

[0143] Clause 6B. The apparatus of any of clauses 3B to 5B, wherein the one or more transfers of learned knowledge are from an encoder of the non-streaming model to an encoder of the streaming model.

[0144] Clause 7B. The apparatus of any of clauses 1B to 6B, wherein the streaming model comprises a streaming automatic speech recognition (ASR) model and the non-streaming model comprises a non-streaming ASR model.

[0145] Clause 8B. The apparatus of any of Clauses 1B to 7B, wherein the one or more transfers of learned knowledge are based on a plurality of auxiliary non-streaming layers between the streaming model and the non-streaming model.

[0146] Clause 9B. The apparatus of clause 8B, wherein the one or more transfers of learned knowledge are based on modified attention masks associated with the plurality of auxiliary non-streaming layers.

[0147] Clause 10B. The apparatus of any of clauses 1B to 9B, wherein the one or more transfers of learned knowledge are based on a KD loss function.

[0148] Clause 11B. The apparatus of clause 10B, wherein the KD loss function comprises at least one of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

[0149] Clause 12B. The apparatus of clause 11B, wherein the KD loss function comprises at least two of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

[0150] Clause 13B. The apparatus of clause 12B, wherein the KD loss function comprises a weighted sum of a distance loss, a Kullback-Leibler divergence loss, and an autoregressive predictive coding loss.

[0151] Clause 14B. The apparatus of clause 13B, wherein the KD loss function comprises:

[0152]

[0153] Among them L KD is the knowledge distribution loss, L DIS is the distance loss, is the Kullback-Leibler divergence (KLD) query loss, is the KLD bond loss, is the KLD value loss, L APC is the autoregressive predictive coding (APC) loss, and α, β, and γ are weights.

[0154] Clause 15B. The apparatus of any of clauses 1B to 14B, wherein the speech comprises an utterance comprising the one or more words.

[0155] Clause 16B. The device of any of clauses 1B to 15B, further comprising one or more microphones configured to capture the speech signal.

[0156] Clause 17B. The apparatus of any of Clauses 1B to 16B, wherein at least one of the one or more transfers of learned knowledge occurs before the streaming model is located on the apparatus.

[0157] Clause 18B. The apparatus of any of Clauses 1B to 16B, wherein at least one of the one or more transfers of learned knowledge occurs after the streaming model is on the apparatus.

[0158] Clause 19B. The apparatus of any of clauses 1B to 18B, wherein the action comprises at least one of processing speech into text, responding to a command, or responding to a query.

[0159] Clause 20B. A method comprising: determining one or more words in a speech signal based on one or more transfers of learned knowledge from a non-streaming model to a streaming model, the streaming model including a real-time streaming model on a device; and taking an action based on the determined one or more words.

[0160] Clause 21B. The method of clause 20B, wherein the non-streaming model comprises a trained non-streaming model, and wherein the one or more transfers of learned knowledge from the non-streaming model to the streaming model comprise using the non-streaming model to train the streaming model.

[0161] Clause 22B. The method of clause 20B or clause 21B, wherein the one or more transfers of learned knowledge are based on an encoder configured to encode the speech.

[0162] Clause 23B. The method of clause 22B, wherein the encoder comprises a plurality of layers.

[0163] Clause 24B. The method of clause 23B, wherein the encoder transfers knowledge at a selected layer of the plurality of layers.

[0164] Clause 25B. The method of any of Clauses 22B to 24B, wherein the one or more transfers of learned knowledge are from an encoder of the non-streaming model to an encoder of the streaming model.

[0165] Clause 26B. The method of any one of clauses 20B to 25B, wherein the streaming model comprises a streaming automatic speech recognition (ASR) model, and the non-streaming model comprises a non-streaming ASR model.

[0166] Clause 27B. The method of any of Clauses 20B to 26B, wherein the one or more transfers of learned knowledge are based on a plurality of auxiliary non-streaming layers between the streaming model and the non-streaming model.

[0167] Clause 28B. The method of any of clauses 20B to 27B, wherein the action comprises at least one of processing speech into text, responding to a command, or responding to a query.

[0168] Item 29B. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine one or more words in a speech signal based on one or more transitions of learned knowledge from a non-streaming model to a streaming model, the streaming model including a real-time streaming model on a device; and take action based on the determined one or more words.

[0169] Item 30B. A device comprising: a component for determining one or more words in a speech signal based on one or more transitions from a non-streaming model to a streaming model based on learned knowledge; the streaming model comprising a real-time streaming model on a device; and a component for taking an action based on the determined one or more words.

[0170] In one or more examples, the functionality described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or codes on a computer-readable medium or sent via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, which includes, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0171] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is appropriately referred to as a computer-readable medium. For example, if the instruction is sent from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave), the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals or other transient media, but are instead directed to non-transient tangible storage media. As used herein, disks and optical disks include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, wherein disks typically reproduce data magnetically, while optical discs utilize lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0172] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Therefore, the term "processor" as used herein may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Additionally, these techniques may be fully implemented in one or more circuits or logic elements.

[0173] The techniques of this disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require implementation by different hardware units. Specifically, as described above, the various units can be combined in a codec hardware unit, or the various units can be provided by a collection of interoperable hardware units (including one or more processors as described above) in combination with appropriate software and / or firmware.

[0174] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A device configured to automatically recognize speech, the device comprising: a memory configured to store a speech signal representing speech and a streaming model, the streaming model including an on-device real-time streaming model; one or more processors implemented in circuitry coupled to the memory, the one or more processors configured to: determining one or more words in the speech signal based on one or more transfers of learned knowledge from a non-streaming model to the streaming model; as well as An action is taken based on the determined word or words.

2. The apparatus of claim 1 , wherein the non-streaming model comprises a trained non-streaming model, and wherein the one or more transfers of learned knowledge from the non-streaming model to the streaming model comprise training the streaming model using the non-streaming model. 3 . The apparatus of claim 1 , wherein the one or more transfers of learned knowledge are based on an encoder configured to encode the speech. The apparatus of claim 3 , wherein the encoder comprises a plurality of layers.

5. The apparatus of claim 4, wherein the encoder transfers knowledge at a selected layer among the plurality of layers.

6. The apparatus of claim 3, wherein the one or more transfers of learned knowledge are from an encoder of the non-streaming model to an encoder of the streaming model.

7. The apparatus of claim 1, wherein the streaming model comprises a streaming automatic speech recognition (ASR) model, and the non-streaming model comprises a non-streaming ASR model. 8 . The apparatus of claim 1 , wherein the one or more transfers of learned knowledge are based on a plurality of auxiliary non-streaming layers between the streaming model and the non-streaming model.

9. The apparatus of claim 8, wherein the one or more transfers of learned knowledge are based on modified attention masks associated with the plurality of auxiliary non-streaming layers.

10. The apparatus of claim 1, wherein the one or more transfers of learned knowledge are based on a KD loss function.

11. The apparatus of claim 10, wherein the KD loss function comprises at least one of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

12. The apparatus of claim 11, wherein the KD loss function comprises at least two of a distance loss, a Kullback-Leibler divergence loss, or an autoregressive predictive coding loss.

13. The apparatus of claim 12, wherein the KD loss function comprises a weighted sum of the distance loss, the Kullback-Leibler divergence loss, and the autoregressive predictive coding loss.

14. The apparatus according to claim 13, wherein the KD loss function comprises: Among them L KD is the knowledge distribution loss, L DIS is the distance loss, is the Kullback-Leibler divergence (KLD) query loss, is the KLD bond loss, is the KLD value loss, L APC is the autoregressive predictive coding (APC) loss, and α, β, and γ are weights.

15. The apparatus of claim 1, wherein the speech comprises an utterance comprising the one or more words.

16. The device of claim 1, further comprising one or more microphones configured to capture the voice signal.

17. The device of claim 1, wherein at least one of the one or more transfers of learned knowledge occurs before the streaming model is located on the device.

18. The device of claim 1, wherein at least one of the one or more transfers of learned knowledge occurs after the streaming model is on the device.

19. The device of claim 1, wherein the action comprises at least one of processing speech into text, responding to a command, or responding to a query.

20. A method comprising: determining one or more words in the speech signal based on one or more transfers of learned knowledge from a non-streaming model to a streaming model, the streaming model including an on-device real-time streaming model; and An action is taken based on the determined word or words.

21. The method of claim 20, wherein the non-streaming model comprises a trained non-streaming model, and wherein the one or more transfers of learned knowledge from the non-streaming model to the streaming model comprise training the streaming model using the non-streaming model.

22. The method of claim 20, wherein the one or more transfers of learned knowledge are based on an encoder configured to encode the speech signal.

23. The method of claim 22, wherein the encoder comprises a plurality of layers.

24. The method of claim 23, wherein the encoder transfers knowledge at selected layers among the plurality of layers.

25. The method of claim 22, wherein the one or more transfers of learned knowledge are from an encoder of the non-streaming model to an encoder of the streaming model.

26. The method of claim 20, wherein the streaming model comprises a streaming automatic speech recognition (ASR) model and the non-streaming model comprises a non-streaming ASR model.

27. The method of claim 20, wherein the one or more transfers of learned knowledge are based on a plurality of auxiliary non-streaming layers between the streaming model and the non-streaming model.

28. The method of claim 20, wherein the action comprises at least one of processing speech into text, responding to a command, or responding to a query.

29. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: determining one or more words in the speech signal based on one or more transfers of learned knowledge from a non-streaming model to a streaming model, the streaming model including an on-device real-time streaming model; and An action is taken based on the determined word or words.

30. A device comprising: means for determining one or more words in a speech signal based on one or more transfers of learned knowledge from a non-streaming model to a streaming model, the streaming model including an on-device real-time streaming model; and A component for taking action based on the determined word or words.